The short answer

The five questions that separate a rigorously validated Media Mix Modeling provider from one that only looks convincing are: how they demonstrate their results reflect causation rather than correlation, how their model accounts for the delay between an ad exposure and a conversion, how granular their model can reliably go, how they model diminishing returns to recommend an optimal budget rather than reporting a historical average, and how they account for the halo effect where spend on one channel lifts sales on another. A provider who can answer all five with evidence, not just reassurance, is one whose numbers are worth acting on.

Media Mix Modeling has become the default measurement layer for retail eCommerce brands since privacy changes eroded user-level tracking, and the vendor market has grown just as fast. That growth has made it harder to tell a rigorously validated model from one that simply presents well. Before you sign, these five questions help you see whether a provider's model has been tested against evidence, or whether it's still resting on assurance alone.

Why vendor diligence matters more than the pitch deck

A Media Mix Model (MMM) is a statistical model that estimates how much each marketing channel, and non-marketing factors like seasonality or pricing, contributed to sales, using aggregated data rather than individual user-level tracking. Most providers will show you a dashboard, a media plan, and a story about how their model works. What determines whether that model is worth acting on is whether the methodology underneath holds together.

A study by Harvard Business Review Analytic Services, sponsored by Google and surveying 547 marketing leaders, found that 87% call MMM important to their measurement strategy, but only around 28% describe their organization as very effective at turning MMM output into timely action. The research traces that gap mostly to data quality and siloed inputs, slow internal processes, and organizational silos between teams, more than to the modeling itself. Measurement that's hard to interrogate is one contributor among several, and it's the one these five questions are aimed at closing.

Here are the five, in the order worth asking them.

1. How do you ensure your results reflect causation, not just correlation?

A provider worth working with treats in-sample fit and out-of-sample accuracy as two different questions, and answers both. Fit statistics like R-squared describe the past; they don't guarantee a relationship will hold going forward. Get this wrong and you risk cutting a channel that was driving growth, or continuing to fund one that was riding a seasonal trend.

What good looks like:

  • Validates performance out-of-sample, using back-testing against future periods, geo holdout tests, or in-platform lift tests, tracked over time rather than reported as a one-off
  • Explains the assumptions and statistical controls used to separate media effects from confounders like seasonality, pricing and promotions
  • Can trace an output back through the model and explain why a channel received the contribution it did

Red flags:

  • Fit statistics (R², in-sample accuracy) offered as the only evidence, with no back-testing or out-of-sample check
  • Vague assurances of causal accuracy with no explanation of the assumptions or controls behind them
  • Reluctance to explain why a particular channel received the contribution it did

2. How does your model account for the time delay between customers seeing ads and converting?

Marketing impact is rarely instant; a video ad seen today can influence a purchase a week or two later. A model that gets this wrong tends to under-credit the upper-funnel channels doing the work of building demand, and over-credit whichever channel sits closest to conversion, the same bottom-of-funnel bias that shows up whenever measurement leans too heavily on last-click data.

What good looks like:

  • Applies different decay curves across channels, typically modeled through adstock, running longer for awareness formats like video and social and shorter for demand-capture channels like paid search
  • Estimates decay from your own data rather than choosing it arbitrarily, and validates it against evidence
  • Can show how changing the decay assumption shifts the resulting channel contributions

Red flags:

  • One fixed lag or decay curve applied to every channel with no clear rationale
  • No explanation for how lag parameters were estimated or constrained

3. What is the most granular level an MMM can reliably model, and how do you ensure robustness at that level?

Granularity, whether a model resolves to channel, campaign, or individual ad, requires proportionally more data and variation at each step down to stay robust. A decision made on a number that looks specific but isn't statistically sound can be worse than no number at all.

What good looks like:

  • Gives a clear answer on the level of granularity the data can  support, and why: channel-level estimates typically hold up with far less data than campaign or ad-set level ones
  • Uses techniques suited to low-signal levels, such as hierarchical modeling or regularization, or triangulates with incrementality tests rather than stretching one regression too far
  • Presents estimates with uncertainty ranges rather than as precise facts

Red flags:

  • Campaign- or ad-level numbers presented as reliable without stating how much spend and conversion history backs them
  • Single-point estimates with no indication of uncertainty
  • Same confidence level applied at every tier, from channel down to ad, with no distinction in how far each one should be trusted

4. How do you model diminishing returns, and can you show the optimal budget for a channel rather than just its historical return?

An average historical ROI tells you how a channel performed at the spend level you've already tested, not what to expect if you double that budget next quarter. Scale a channel that looks efficient on paper without seeing the underlying curve, and you can just as easily push it past the point where it stops paying off.

What good looks like:

  • Models a full saturation curve, showing how each additional dollar spent tends to produce a smaller incremental return as spend increases, rather than relying on historical average ROI alone
  • Uses that curve to recommend a budget range instead of a single average figure
  • Refreshes that curve as spend and market conditions shift, rather than setting it once and leaving it
  • Makes the assumptions and uncertainty explicit, especially when forecasting beyond historically observed spend

Red flags:

  • Budget recommendations based only on historical average ROI, with no saturation curve behind them
  • A single number handed over for spend well beyond what's  been tested, with no flag that it's an extrapolation
  • Uncertainty disappears the further the recommendation moves from tested spend levels; the number is simply presented as fact

Worth knowing that this is roughly how, Fospha's incremental forecasting tool, approaches it: a saturation curve for each channel, used to suggest a budget range rather than lean on a single historical average, and refreshed as spend and market conditions move.

5. How do you model the halo effect, where spend on one channel drives sales on another?

A TikTok ad that drives a branded search a week later won't show up as TikTok revenue unless the model is designed to catch it. Skip this check, and channels that quietly support the rest of the mix can look far less valuable alone than they really are, making the case for scaling them harder to win internally than it should be.

What good looks like:

  • Reports a metric that captures cross-channel spillover, with a concrete example of how it's separated from a channel's direct, first-touch contribution
  • Controls for other explanations of the same pattern, like seasonality or promotions
  • Incorporates marketplaces such as Amazon and TikTok Shop where the data and methodology support it

Red flags:

  • Measurement that automatically excludes marketplace sales without explaining why
  • Cross-channel effects inferred purely from timing or correlation
  • Every channel credited only for directly attributed sales, with no consideration of potential spillover

It's this kind of gap that led Fospha to build Halo, its total commerce measurement across DTC, Amazon, and TikTok Shop.

How do you know if the answers hold up?

The thread running through all five questions is really one question: can the provider show their work? That's what Fospha is built on: a system designed to be interrogated, not just trusted. Every model output is shown at every stage, credit shifts are compared against Last Click, forecasts are back-tested against what actually happened, and results from incrementality tests like GeoLift calibrate forecasting and budget planning.

Whichever provider you land on, that's the standard worth holding them to.

Common questions

Q: Can incrementality testing replace MMM as your core measurement approach?

No, the two answer different questions. A well-designed incrementality test gives you a causal read of a specific channel or campaign at a single point in time, while a Daily MMM provides continuous, day-to-day signal across the full channel mix. The strongest setups treat incrementality tests as calibration inputs that sharpen the model, rather than treating either one as a replacement for the other.

Q: Is a high R-squared enough to prove an MMM is accurate?

Not on its own, though it isn't a red flag either. R-squared is an in-sample fit statistic: it shows how much of your historical sales variation the model can explain, which is useful, but it describes the past rather than the future. The complementary check is back-testing, which holds back recent periods and checks whether the model's relationships still hold on data it hasn't seen. A provider who can show both a credible R-squared and a stable back-tested error over time has given you the fuller picture. Fospha has published a detailed breakdown of how these signals work together, if you want to go deeper on this one.

Q: How much historical data does an MMM typically need?

Google's Meridian documentation, one of the more widely referenced open-source MMM frameworks, recommends at least two years of weekly data for geo-level models and three years for models run at a national level. Below that, there often isn't enough variation in spend and outcomes for the model to separate one channel's effect from another's with any confidence.

Q: What's the difference between channel-level and ad-level MMM output?

Channel-level output estimates the contribution of an entire channel, like Meta or paid search, and typically holds up reliably with less data. Ad-level output estimates the contribution of an individual ad or ad set, which needs far more consistent spend and conversion history to stay statistically robust. A good provider is transparent about which level their model is confidently resolving for your data volume, rather than reporting ad-level numbers that look precise but rest on too little signal underneath.

Q: What should I expect during the first few weeks after signing with an MMM provider?

Expect a data validation and reconciliation phase before treating the model's outputs as decision-ready. A provider that produces highly confident-looking numbers immediately, without checking data quality, definitions and historical consistency, is worth questioning. Setting up the model properly takes time because the quality of the underlying data and assumptions directly affects the usefulness of the output.

Related reading

See how Fospha works in practice. 30-minute walkthrough on your data, your channels, this quarter.

Book a demo