
How to compare two attribution models on the same orders
Most attribution tools calculate credit at collection time. That means switching models needs new data and a month of waiting, and you need the comparison now.

All the engineering in attribution tooling goes into computing credit correctly. The question that comes up in every meeting isn't about correctness at all.
On the dashboard, an 890 SAR order sits next to a channel nobody spent on this month.
The merchant asks the agency. The agency sends a screenshot of the same dashboard. Asked again, the answer arrives as a sentence: "the model attributes it there." A sentence, not evidence.
After that exchange the number is dead. It stays on the screen and stops informing any decision, and every figure from the same source carries the same doubt. One unexplained number is enough to make a merchant suspicious of the whole tool.
All three have to be visible without asking anyone:
1. The sequence. Every touch on this order, in order, with dates and sources. Not a summary, the actual list.
2. The weight and the model's name. What each touch received under the model currently running, and what that model is. A figure with no model named isn't an explanation, it's a second unexplained figure.
3. The identity evidence. Which identifiers tied these touches to one person. This is the part almost nobody provides, and it's the one that answers the hardest follow-up: are you sure that was the same customer?
With all three, the conversation moves from "is this right" to "is this model right for us", and the second is a disagreement a group can settle.
In Flowfy, every order stores its full touchpoint sequence with timestamps and sources, so you can open an order and read the path that produced it, and the resolved identity behind it is auditable. Showing per-touch weight across models side by side on one screen is on the roadmap and hasn't shipped. Today the sequence and the identity are visible, and the weighting is inferred from the model you selected.
A more sophisticated model produces a figure that's harder to explain, not easier. Data-driven attribution is the clearest example: it may well be more accurate, and it can rarely tell a merchant why one specific order was credited to one specific channel.
That trade is worth making after order-level explainability exists, not before. A model nobody can interrogate gets trusted while the numbers look plausible and abandoned the first time they don't, which is exactly the moment you need it.
Storing journeys instead of pre-computed credit is more work per query. That's the trade that makes both explanation and model comparison possible, and it's covered in comparing models on the same orders.
There's a version of this that has nothing to do with tooling.
When a number can't be explained, the decision defaults to whoever is most senior or most confident. The person with the data has a figure, the person disagreeing has a conviction, and the conviction wins because the figure can't defend itself.
So explanation isn't a nicety for analysts. It's the mechanism that lets evidence outrank seniority in a meeting. Without it, your investment in measurement produces reports that lose arguments.
Nobody opens a single order very often, and that's fine. The value is that it can be opened when it's questioned. And if opening it reveals the attribution was wrong, you've found a real problem, usually caused upstream of the model by a bad identity merge or a touch that never arrived.
Open one order from your dashboard now and try to explain it: the sequence, the model, the identity evidence. If you can't get all three without asking anyone, that's the first gap to close, well before you think about a more accurate model.

Most attribution tools calculate credit at collection time. That means switching models needs new data and a month of waiting, and you need the comparison now.

Most arguments about channel performance aren't arguments about data. Identify the type first, because only one of the three is settled by opening a record.

No attribution model answers this question. The only method is a holdout, and it costs real money. Here's how to design one and which channel deserves it.