7 Best ETL Tools in 2024

Choosing the right ETL tools is pivotal for seamless data integration and effective business analytics. Are you overwhelmed by the options or unsure which ETL tool suits your company’s scale and needs? Join us as we review the top ETL tools of 2024, comparing their strengths in handling diverse data volumes, compliance standards, and integration […]

Best Practices for Quality ETL Testing

You’ve likely heard horror stories of businesses making costly decisions based on bad data. JPMorgan Chase experienced one such debacle when a data error led to a $6 billion trading loss. The infamous “London Whale” incident in 2012 was partly due to values being incorrectly recorded by an automated system. This costly error might’ve been […]

How To Create A Data Pipeline Automation

The essence of data pipeline automation lies in its ability to streamline the flow from data extraction to the delivery of actionable insights. With automation, you can ensure that data quality stays maintained as it travels through the stages of your pipeline.  Moreover, it allows for continuous data ingestion and processing, which means that your […]

How AI Agents Work: A Practical Guide for Marketing and Product Leaders

Despite the prevalence and demonstrated superiority of ML and AI tools, many companies are still making decisions — even relatively trivial decisions, like determining messaging frequency — manually. What’s worse, each manual decision affects more than a single individual; even simple decisions typically require time and attention from multiple teams, which leads to significant bottlenecks […]

Scalable Event-Based Clustering for User Segmentation

There are generally limited options for clustering large numbers of records. A typical Aampe customer has up to 400 app events (“product viewed”, “purchase completed”, etc.) instrumented, for sometimes as many as 100 million end users. Clustering 300 features is simple, even with 100 million records – something like mini-batch k-means can do the trick. […]

You don’t know Jacc(ard)

I’ve been thinking lately about one of my go-to data science tools, something we use quite a bit at Aampe: the Jaccard index.  It’s a similarity metric that you compute by taking the size of the intersection of two sets and dividing it by the size of the union of two sets.  In essence, it’s […]

Aampe + Braze

Braze has gained significant traction as one of the most commonly used platforms in the customer engagement space. Offering most of the features and functionality that CRM teams have come to expect, such as email and push notification campaigns, user segmentation, and basic campaign analytics, Braze is recognized for its utilitarian approach to customer engagement. […]

A 30,000,000,000 row join! And how we reduced runtime of a query by >99%

Last week we at Aampe faced a simple yet interesting scale problem when using BigQuery. The culprit: A simple inner join. We wrote a simple query that joins the users who are eligible to be messaged with their corresponding messages CMS table. The output needed: for each user, all messages that they are eligible to […]

How we handle messy data

As our platform has to work with various data providers, CDPs, and many different companies with very different data lakes and schemas, it’s actually more common for us to encounter messy data than anything else. Here’s our approach to cleaning up this data so it can be usable and actionable: What is “messy” data? “Messy […]

What is the Best Time to Send SMS Marketing in eCommerce?

SMS marketing can get incredibly expensive, so it’s important that every message is adding value. To that end, we took all the “best practices” we could find online for ‘the best times and days to send SMS marketing messages for an e-commerce app,’ and compared it to actual data for an actual app with over […]