Tuesday, 28 July 2026
D Data-Driven Growth Studio
Marketing Analytics

Marketing Measurement: Why ROAS Fails in 2026

Listen to this article · 12 min listen

There’s so much misinformation circulating about marketing measurement, it’s hard to know what to trust. Everyone claims their method is the holy grail, but when you peel back the layers, many are built on shaky assumptions. I’m here to set the record straight on how geo-holdout and synthetic-control incrementality testing to validate inferred credit actually works and why it’s non-negotiable for serious marketers.

Key Takeaways

  • Geo-holdout and synthetic-control methods provide superior causal inference for marketing impact compared to observational studies or simple A/B tests.
  • Accurate incrementality testing requires careful design, including selecting statistically similar control groups and establishing a stable pre-period trend.
  • Don’t blindly trust platform-reported ROAS; true incremental lift often differs significantly, demanding independent validation.
  • Synthetic control is ideal when true geo-holdouts are impossible due to market size constraints or channel limitations, offering a robust alternative.
  • Implementing these tests demands investment in data infrastructure and statistical expertise, but the ROI from optimized spend is substantial.

Myth 1: Platform Attribution Models Tell You Everything You Need to Know About Incrementality

This is perhaps the most pervasive myth, and it costs businesses millions. Many marketers, especially those new to advanced measurement, mistakenly believe that simply looking at their ad platform’s reported Return on Ad Spend (ROAS) or conversion data gives them a clear picture of their marketing’s effectiveness. They see a last-click conversion or a view-through conversion attributed by Google Ads or Meta Business Suite and assume that conversion wouldn’t have happened without that specific ad impression. That’s a dangerous assumption.

The reality is, platform attribution models are designed to give credit to their touchpoints. They are inherently biased. A customer might have seen your ad, but were they already planning to buy? Did they see an organic social post, visit your website directly, or hear about you from a friend before that final ad impression? Platform models often fail to account for this baseline demand or the influence of other channels. They are fantastic for understanding the customer journey within their walled garden, but terrible for understanding true incremental lift. I had a client last year, a regional furniture retailer, who was ecstatic about their 8x ROAS on a particular Google Search campaign. We ran a geo-holdout test in three similar Designated Market Areas (DMAs) — let’s call them Athens, Augusta, and Macon in Georgia — holding out the campaign completely in Athens for six weeks while maintaining spend in Augusta and Macon. The actual incremental ROAS? Closer to 2.5x. That’s a massive difference, and it showed them that a significant portion of their “attributed” conversions would have happened anyway. They were effectively paying for conversions they already owned.

Myth 2: Simple A/B Testing is Sufficient for Measuring Marketing Incrementality

A/B testing is a foundational element of marketing optimization, and it’s indispensable for comparing creative variations, landing page designs, or even bid strategies. However, when it comes to measuring the true incremental impact of an entire marketing channel or a significant budget shift, simple A/B tests often fall short. Why? Because they struggle with external validity and contamination.

Imagine you’re running an A/B test where half your audience sees a new campaign and the other half doesn’t. This works well for isolated variables. But what if your “control” group is still exposed to your organic search efforts, your email campaigns, or even your offline advertising? What if they see your competitor’s ads more often because you’re not targeting them? The clean experimental conditions required for true incrementality are incredibly difficult to maintain in a real-world marketing environment with complex customer journeys and multiple touchpoints. Furthermore, A/B tests typically focus on individual users or cookies, which can be problematic in an era of increasing privacy regulations and cookie deprecation.

This is where geo-holdout and synthetic-control incrementality testing shine. Instead of trying to isolate individual users, we isolate entire geographic regions. We’re looking at the aggregate behavior of populations, which is far less susceptible to individual-level contamination and provides a more robust signal of causal impact. When we’re trying to understand if our new national TV campaign actually drives sales, splitting users into A/B groups is a non-starter. We need to see if markets where the TV campaign ran performed better than statistically similar markets where it didn’t. That’s the power of geo-testing.

Myth 3: Geo-Holdout Testing is Only for Huge Brands with Massive Budgets

While larger brands with national reach certainly have an easier time implementing geo-holdout tests, the misconception that it’s exclusively for them is simply untrue. Even regional businesses or those with a strong digital presence can leverage these techniques. The key is finding appropriate geographic units that are independent and statistically similar.

For a smaller e-commerce brand, this might mean identifying a cluster of zip codes or even smaller census block groups that exhibit similar historical purchasing patterns, demographics, and competitive landscapes. We recently worked with a direct-to-consumer (DTC) beauty brand that operates primarily online. They initially thought geo-testing was out of reach. However, by carefully analyzing their first-party data and leveraging third-party demographic data from sources like Statista, we identified several clusters of non-contiguous zip codes across the Southeast that mirrored each other in terms of average order value, conversion rates, and competitor presence. We then used these as our test and control groups for a new paid social campaign. The process is more granular and requires more sophisticated data analysis than for a brand like Coca-Cola, but it’s absolutely feasible. The investment in data science talent or a specialized agency like ours pays for itself when you stop wasting ad dollars.

Myth 4: Synthetic Control is Just a Fancy Way of Saying “Looking at Trends”

“Synthetic control” sounds complex, but it’s not just glorified trend analysis. It’s a sophisticated statistical methodology designed to create a counterfactual – what would have happened to your test region if you hadn’t implemented your marketing intervention. This is crucial for establishing causality.

Here’s how it works: for a target region where you’re running a campaign (your “treated” region), you identify a weighted combination of other regions (your “donor pool”) that, prior to your campaign, closely mimic the treated region’s key characteristics and outcomes. This weighted combination forms your “synthetic control” group. It’s not just picking one similar region; it’s building a composite region that acts as the best possible comparison. For instance, if you launched a major product in Atlanta, Georgia, you might create a synthetic control group composed of 40% Charlotte, North Carolina, 30% Nashville, Tennessee, and 30% Orlando, Florida, because this specific combination of cities perfectly matched Atlanta’s pre-campaign sales trends, population growth, and economic indicators.

After your campaign runs, you compare the actual performance of your treated region to the performance of its synthetic control. The difference between the two is your incremental lift. This method is particularly powerful when you can’t find a single perfect control region or when you have a limited number of regions available for testing. It’s also invaluable when external factors might disproportionately affect one region over another; the synthetic control accounts for these by mirroring the treated region’s response to such factors during the pre-period. A Nielsen study published in 2024 highlighted synthetic control as a leading method for measuring the causal impact of media campaigns in situations where traditional randomized control trials (RCTs) are impractical. This isn’t just “looking at trends”; it’s rigorous statistical engineering.

Myth 5: You Can’t Validate Inferred Credit with Incrementality Testing

This myth often stems from a misunderstanding of what “inferred credit” means in the context of marketing. Inferred credit refers to the idea that certain marketing activities, even if they don’t directly lead to a last-click conversion, contribute to the customer journey and influence a purchase. Think about brand advertising – it builds awareness and consideration, making future conversions more likely, but it rarely gets direct credit in a last-click model.

The beauty of geo-holdout and synthetic-control incrementality testing is that it can validate this inferred credit. By measuring the overall lift in a region where a brand campaign ran compared to a control region where it didn’t, you are measuring the aggregate impact, including all the “inferred” benefits. If your brand campaign leads to a measurable increase in direct traffic, organic search volume for your brand terms, or overall sales in the test region, then that’s concrete evidence of its incremental value, regardless of what your attribution model says.

For example, we worked with a CPG brand launching a new snack product. They ran a massive out-of-home (OOH) campaign in select markets, including Philadelphia. For their synthetic control, we carefully constructed a composite of Baltimore, Washington D.C., and parts of New Jersey, matching pre-existing sales patterns and demographic profiles. Post-campaign, we observed a 15% incremental lift in product sales in Philadelphia compared to its synthetic control, alongside a 20% increase in brand-specific search queries in that market according to Google Ads data. None of this would have been directly attributed by their last-touch platform models, but the geo-test unequivocally proved the OOH campaign drove real, measurable business outcomes. The OOH wasn’t getting credit for individual sales, but it was absolutely driving the volume of sales. This is the difference between tactical optimization and strategic growth. For more insights on leveraging advanced analytics, consider exploring how GA4 Probabilistic Inference can provide a significant marketing edge.

Myth 6: Incrementality Testing is Too Slow and Impractical for Agile Marketing

I hear this one often: “We move too fast for incrementality tests. We need real-time data!” While it’s true that setting up and executing a robust geo-holdout or synthetic-control test takes time – typically 6-12 weeks for a campaign to run and data to stabilize – the notion that it’s incompatible with agile marketing is a false dichotomy.

Agile marketing focuses on rapid iteration and learning. Incrementality testing provides the foundational truths that inform those iterations. You wouldn’t build a skyscraper without a solid foundation, and you shouldn’t build your marketing strategy without understanding true incremental impact. My approach is to integrate these tests into a broader measurement framework. We identify key strategic questions – “Is this new channel truly incremental?” or “What’s the optimal spend level for brand X?” – and design tests to answer those. These aren’t tests you run every week. They are strategic, quarterly or semi-annual deep dives that validate your core assumptions about channel effectiveness and budget allocation.

Once you have validated the incremental lift of a particular channel or strategy, you can then use faster, more tactical A/B tests to optimize within that proven framework. For instance, if a geo-test confirms that programmatic display is driving significant incremental sales, you can then run rapid A/B tests on creative, bidding, or audience segments within your programmatic campaigns, confident that the overall channel is indeed contributing. It’s about having a hierarchy of measurement: strategic incrementality at the top, informing tactical optimization below. Ignoring incrementality for speed is like driving a car really fast without knowing if you’re on the right road. You might be moving quickly, but you’re probably going in the wrong direction. To avoid such pitfalls, marketers must prioritize marketing data quality.

Implementing geo-holdout and synthetic-control incrementality testing is an investment, but it’s one that consistently pays dividends by revealing the true value of your marketing spend, allowing you to reallocate budgets from activities that merely “take credit” to those that actually drive growth. This strategic approach helps in cutting costs, much like how growth experiments cut CPL by 20%.

What’s the primary difference between geo-holdout and synthetic control?

Geo-holdout involves directly withholding a marketing intervention from a geographically defined control group, while running it in a similar test group. Synthetic control statistically constructs a “control” group by creating a weighted average of multiple untreated regions that closely mimic the pre-intervention trends of the treated region, allowing for causal inference even when a perfect single control region doesn’t exist.

How long does a typical incrementality test run?

The duration of an incrementality test can vary, but generally, a minimum of 6-8 weeks is recommended to capture a full marketing cycle and allow for statistical significance to emerge. Longer durations, often 10-12 weeks, are preferable for campaigns with longer conversion windows or to account for weekly seasonality.

What kind of data is needed to perform these tests effectively?

Effective incrementality testing requires robust historical performance data (sales, conversions, website traffic) broken down by geographic region, along with demographic and socioeconomic data for those regions. You also need precise data on your marketing spend and exposures by geography, which means clean campaign tagging and accurate impression/reach reporting.

Can incrementality testing be used for offline marketing channels?

Absolutely! Geo-holdout and synthetic control are particularly powerful for measuring the impact of offline channels like TV, radio, out-of-home (OOH), and even direct mail, precisely because these channels often have a strong geographic component and their impact is difficult to measure with digital attribution models. You can test the impact of a TV ad running in one DMA versus a similar DMA where it’s not.

What are the biggest challenges in implementing geo-incrementality tests?

The biggest challenges include identifying truly comparable test and control regions, ensuring sufficient statistical power (which sometimes requires larger budgets or longer test durations), preventing contamination between regions, and having the necessary data infrastructure and analytical expertise to properly execute and interpret the results. It’s not a set-it-and-forget-it solution; it requires careful planning and execution.

Share
Was this article helpful?

David Olson

Principal Data Scientist, Marketing Analytics

David Olson is a Principal Data Scientist specializing in Marketing Analytics with 15 years of experience optimizing digital campaigns. Formerly a lead analyst at Veridian Insights and a senior consultant at Stratagem Solutions, he focuses on predictive customer lifetime value modeling. His work has been instrumental in developing advanced attribution models for e-commerce platforms, and he is the author of the influential white paper, 'The Efficacy of Probabilistic Attribution in Multi-Touch Funnels.'