
Why how often a channel appears doesn't measure its contribution
A channel appearing in 620 of 1,000 paths looks indispensable. The number that matters is how many paths it closed on its own, and that's a different calculation.

Every attribution debate eventually arrives at the same question: if we stopped spending on this channel, how many orders would we lose?
Models divide credit for orders that happened. This question is about something that didn't happen, which is a different claim and needs a different method.
A holdout means you stop a channel for a defined slice: one region, one audience segment, one set of days. You hold everything else constant, then compare total orders in that slice against a comparable slice where the channel kept running.
Note the wording: total orders, not attributed orders. You want to see what happened to the business, not how credit got assigned.
Everything else is inference from observational data: first touch, last touch, multi-touch, path analysis, solo rates. Inference can't separate a channel that creates demand from a channel standing where demand already was.
Pick a slice you can genuinely isolate. Geography is usually cleanest. Audience segments leak, because platforms overlap audiences and people move between them.
Run it for at least one full buying cycle. If your customers take three weeks to decide, a one-week holdout measures nothing but the tail of journeys already in flight. See buying cycle.
Hold everything else constant. No creative changes, no budget shifts elsewhere, no promotion that hits one slice and not the other. The test is fragile and contaminates easily.
Decide the success threshold before you start. Without one, the result gets interpreted by whoever wanted a particular answer.
The number you judge the test on is your store's total orders split by region and by source, which is live in Flowfy over a 90-day window. Running the holdout itself happens inside the ad platforms.
Test openers: broad prospecting, awareness, creator content. Their contribution is genuinely unsettled under every model, and the gap between their first-touch and last-touch numbers is exactly the uncertainty a holdout resolves.
Don't bother with closers in the same way. Pausing retargeting produces an immediate visible drop, and that drop tells you almost nothing about causation. The demand was already there, and the question is whether it would have converted anyway. A retargeting holdout needs a much longer window to mean anything, because the customers it would have reached may still buy later.
A holdout also costs real money. You're deliberately not spending somewhere spending would probably have produced orders. That's why you run it on the channel whose answer moves the most budget, and why you run it rarely. A holdout on a channel taking 3% of your spend is an expensive way to learn something that won't change a decision.
When a holdout isn't justified, two observational proxies narrow the uncertainty, though neither proves causation:
Solo conversion rate. What share of paths containing this channel converted with it as the only touch. See frequency is not contribution.
First-touch share. A channel that appears first in most of its paths is doing opening work, and that's the role a holdout usually confirms as incremental.
Can't I just use the platform's own lift test? You can, and it beats nothing. Read the result against your own total orders as well, since a lift test measures the platform's conversions.
How often should I run one? Once or twice a year, on the channel whose answer moves the most budget.
What if the holdout shows almost no effect? That's a real and valuable finding, and it's worth re-running once before acting on it, because a single test on a noisy slice can mislead in either direction.
Does a holdout replace attribution? No. Attribution runs continuously and guides your day-to-day allocation. A holdout answers one question, occasionally, at a cost.
Pick one opening channel, one geographic slice you can isolate, and a full buying cycle as the duration. Write down the success threshold before you turn anything off, and compare total orders across the two slices when it ends.

A channel appearing in 620 of 1,000 paths looks indispensable. The number that matters is how many paths it closed on its own, and that's a different calculation.

The gap between first touch and purchase is a fact about your product that no benchmark supplies. It sets your attribution window, when you start spending before a season, and how long you wait before judging a campaign.

Most arguments about channel performance aren't arguments about data. Identify the type first, because only one of the three is settled by opening a record.