Skip to content
Proofbell

Attribution models

A customer clicks a Google ad, comes back through organic search a week later, then types your address in and phones. Which channel gets the credit?

There is no single correct answer, which is exactly why Proofbell computes several and shows you all of them. Credit is calculated when you ask for it and never stored, so switching model is instant and changes nothing about your data.

Where the models disagree is the useful information. If a campaign looks strong under last-click and weak under first-click, its apparent performance is partly a modelling choice. Anyone showing you one number as the truth is hiding that.

The six rule-based models

Using our example journey: Paid Search → Organic Search → Direct → phone call.

ModelCredit goes toUse it when
Last click Direct — 100% Almost never on its own. It over-credits whatever happened last, which is usually Direct or brand search — channels that captured demand rather than created it.
Last non-direct Organic Search — 100% A sensible default, and ours. Direct usually means "already knew you", so crediting the last real marketing touch is closer to useful.
First click Paid Search — 100% Judging what creates demand. Answers "what introduced this customer to us" and ignores everything that closed them.
Linear 33% each A neutral starting point when you have no view. Simple and defensible; treats a passing Display impression as equal to the search that converted.
Time decay Weighted towards Direct Short sales cycles, where recent touches genuinely matter more. Credit halves every 7 days by default.
Position based Paid 40%, Direct 40%, Organic 20% The most common practical choice: the introduction and the close matter most, and the middle still counts for something.

Fractional conversions are correct, not a rounding error

Under linear, one conversion becomes 0.33 for each of three channels. So a report may say 2.4 conversions from Paid Search. That is the only honest way to express shared credit — rounding it to 2 or 3 would mean picking a winner, which is what these models exist to avoid.

Offline sources are whole deals, not shares

If you import deals from a CRM with a lead source column, the Pipeline screen shows a second table: what each offline source — a trade show, a webinar, outbound calling, a walk-in, a lead form, an inbound sales call — actually closed. It sits beside the channel table rather than as extra rows inside it, and the reason is the section above.

A channel row is a share of a deal, because credit is split across every touchpoint the buyer had. An offline source row is a whole deal, because “this lead came from the NEC show” is something you told us rather than something we modelled. Put both in one table and the column adds up to more than you closed: a trade-show lead who also clicked an ad is counted once in full and again in slices. So the two are kept apart, and the screen says so under the table.

The offline table is also the only place a deal with no tracked marketing appears at all. Somebody who walked up to your stand and never visited the site has no touchpoint to credit, so no channel report can show them — ours included. The screen tells you how many deals that is, because otherwise the two tables look like they disagree.

“No source stated” is not a source. It covers a file with no lead source column, and a value we would not place — if your file says “Partner referral” we will not file it as the nearest match. Deals synced directly from HubSpot or Salesforce land here too: your own lead source field is your vocabulary, not ours, and mapping it is not something we will guess at.

The lookback window

Touchpoints older than 90 days before the conversion are excluded by default. That is deliberately generous: a short window is a systematic bias, not a neutral simplification — it quietly moves credit away from the channels that create demand early towards whatever happened most recently.

Every report tells you how many conversions had touchpoints only outside the window. If that number is large, your window is too short for your sales cycle, which is a setting to change rather than missing data.

Data-driven attribution

The five models above apply a rule someone chose. The data-driven model derives credit from your own data instead: it builds a map of the journeys your customers actually take, then asks of each channel "how many conversions would we lose if this channel disappeared?"

That is the removal effect, and it is the closest thing to a measurement rather than an opinion.

It refuses when it cannot answer

The model needs roughly 100 customer journeys and 20 conversions. Below that, Proofbell tells you it is not ready and recommends a rule-based model.

That refusal is the feature. A data-driven model fitted on thirty journeys will contradict Google Ads, and you would have no way to defend it when it did.

Reading the two numbers it gives you

ColumnAnswers
Credit What share of conversions would be lost without this channel. Total contribution — and it is volume-weighted, so a wide-reach channel scores highly simply by touching more journeys.
Efficiency Credit relative to reach, where 1.0 is average. This is the budget number. Above 1.0 a channel earns its exposure; below 1.0 it does not.

Do not move budget on credit alone. A channel touching every journey will always show high credit — it is on all the winning paths. That answers "what contributes", not "where should the next pound go". We show both because using the first as the second moves money the wrong way.

Confidence intervals, and ties

Each channel shows a 95% range as well as a figure. A channel on 30% credit with a range of 5–55% is a very different claim from 30% with a range of 28–32%.

Where two adjacent channels' ranges overlap, we mark them as tied rather than ranking them. A ranking is only as trustworthy as the gaps in it, and rows that reshuffle each month for no reason are how people stop believing a tool.

An honest limit

A channel with genuinely no effect still tends to attract a few percent of credit, because it appears alongside channels that do work and the model cannot fully separate correlation from cause. Treat small shares as noise. The reports say so too, rather than implying a precision that is not there.

Reconciliation against your own revenue

Every report carries a reconciliation summary, and the difference in it must be zero. It proves that credit spread across channels still totals exactly the revenue you imported — nothing invented, nothing lost.

Check it against your own CRM total. If it does not tie out, tell us: that is a bug in our arithmetic, not a modelling nuance, and we would want to know immediately.

Reconciliation against the ad platforms

A different question, and the one that comes up more often: why does Google Ads say 40 conversions when Proofbell says 52? The panel above is our own arithmetic and must come to zero. This one never comes to zero, and it is not supposed to — the point is that the difference has reasons, and you can see each of them.

At the bottom of the Attribution page, per connected platform, is a chain that starts with every conversion we recorded and subtracts a line for each reason one did not reach that platform:

  • No click id. Usually the biggest line and the least understood. A caller who found you through organic search, dialled from a business card, or whose click id the browser dropped has no join key — so no platform can match them. These are real conversions that cannot appear in any ad platform's count, ever.
  • Rejected by the platform. They told us the conversion was unusable, and the response code is shown. An expired click id is the commonest cause.
  • Sent in dry run. The connection is still in test mode, so these reached nobody. They look accepted everywhere else, which is exactly why they get their own line.
  • Skipped, failed, or still queued. Ours to fix rather than yours to explain. A number that stays in "failed" needs looking at, and telling us is the fastest route.

Underneath, what each platform said about the clicks. The surprising one is expired: a platform's click lookback is finite, so a call four months after the click carries an id the platform itself will no longer match. Your conversion is real, our count is right, and theirs cannot include it.

What this page cannot tell you. It compares what we recorded against what each platform told us when we sent it. It does not read their reports, so a figure in Google Ads today can still differ for reasons on their side: their attribution window may credit a conversion to a different day, their own modelling adds conversions we never sent, and a conversion action set to a different attribution model counts the same click differently.

Those differences are theirs to explain. This page exists so ours are not a mystery.

So which should I use?

  • Starting out: last non-direct. It is the default and it is defensible.
  • Reporting to a client or a board: position based. It reflects how people intuitively think about introduction and close.
  • Deciding where money goes: data-driven, on the efficiency column, once you have the volume for it.
  • When a number looks surprising: compare all of them. If the models agree, trust it. If they disagree wildly, the honest answer is that your data cannot yet settle the question.
Next Keyword data and its limits