Most agency reviews answer the easy question: is the dashboard green? The question that should decide a renewal is harder. Did contribution margin rise by more than the agency cost, compared with what the business would have done without the agency? Formal reviews are nothing new. An ANA survey fielded in 2009 found that 82% of marketers already ran formal agency performance evaluations, and those reviews have traditionally weighed service quality and relationship at least as heavily as money. This guide covers how to evaluate marketing agency performance on incremental revenue and profit: setting a decision rule, separating agency effect from baseline, reconciling agency reports with your own systems, choosing KPIs by business model, running the CAC and margin arithmetic, reading fee incentives, and turning the result into a retain, restructure or replace decision.

GPI's position runs against the usual review. Most evaluations score the relationship and accept the numbers; the relationship matters, but the numbers are what deserve interrogation. Attributed revenue is a claim, produced with settings the agency chose, about a business whose revenue would have moved anyway through seasonality, pricing and sales effort. The test that matters is whether the agency can show, with a method you can independently check, that contribution margin is higher than it would have been without them. When that evidence is weak, the first fix is usually measurement design and contract terms rather than a new agency. Switching partners without changing how you judge one restarts the same uncertainty with a new logo on the deck.

TL;DR: How to evaluate marketing agency performance in five checks

If you want the answer to how to evaluate marketing agency performance in one pass, run these five checks in order:

  • Define the decision first. Decide whether the review ends in retain, restructure or replace, and write down the contribution-margin threshold that would justify each outcome before anyone opens a report.
  • Measure against a baseline. Compare revenue trends before and after agency activity, or use a holdout, before accepting an attributed figure. Investopedia's ROI guidance makes the same point: ROI only means something against prior-month sales comparisons.
  • Reconcile before you debate. Match agency-reported spend, conversions and revenue to your platform exports, CRM and finance ledger for the same window.
  • Replace vanity metrics with the two or three indicators that fit your model: pipeline and LTV for long sales cycles, contribution margin per order for DTC. ActiveCollab's KPI guide describes how marketers get caught up in vanity metrics instead of outcome numbers.
  • Read the fee model as an incentive map. Pricing structures vary widely, per NetSuite's summary of Promethean data, and whatever the agency is paid on is what it will optimise.

Score the agency on evidence quality, not on the size of the numbers in the deck. Note which of these checks your current process skips.

Start with the decision, not the dashboard

What 'impact on revenue and profit' actually means

Four numbers get called revenue in agency conversations, and only two of them can settle a contract decision.

  • Gross revenue is everything the business booked. It moves with pricing, sales headcount, seasonality and product launches, none of which the agency controls.
  • Attributed revenue is the share a reporting system assigns to agency-managed channels. It depends on attribution windows, view-through settings and de-duplication rules chosen by whoever configured the report.
  • Incremental revenue is the revenue that would not have happened without the agency's work. It requires a counterfactual, which means a baseline, a holdout or a model.
  • Contribution margin after media and fees is incremental revenue less cost of goods, less media spend, less agency fees, production and tooling. This is the figure that tells you whether the engagement made money.

The retain or replace question is answered by the last two. Everything else is context or a hypothesis about them.

Set the threshold before you open the report

Write the decision rule in advance, in one sentence, and have finance sign it. A workable form: incremental contribution margin over the evaluation window must exceed total agency cost, fees plus media, by an agreed multiple. The multiple is a business choice that reflects your cost of capital and risk appetite; the discipline is in fixing it before the numbers arrive, so the threshold cannot drift toward whatever the report shows.

The same 2009 ANA survey that found 82% of marketers conducting formal evaluations described a practice built around relationship and service scoring. That survey is old and says nothing about today's tooling, but it illustrates a durable pattern: evaluation has long been routine, while financial rigour has not.

Set the window relative to your sales cycle. Ninety days is meaningless for enterprise B2B, where an opportunity created in month one may close in month nine. The same 90 days is generous for a DTC replenishment product with a three-week repurchase cycle. Agree the window with the agency at contract signing, not at renewal, so neither side can pick a convenient period later.

Separate incremental revenue from the baseline the agency did not create

Revenue moves for reasons that have nothing to do with the agency. Organic growth compounds, seasonality repeats, a price increase lifts average order value, and two new sales hires close deals that marketing did not source. An honest evaluation nets these out before crediting anyone. Investopedia's guidance on campaign ROI is blunt on this point: without comparisons to the business line's sales in the months before launch, the ROI figure has no real meaning. Four methods do that work, with very different costs and confidence levels. What follows is a practical framework, not a set of study findings.

Pre/post baseline comparison

Start here because every company can do it. Pull monthly revenue for the twelve months before the engagement began, fit a simple trend that captures growth and seasonality, and project it across the engagement period. The gap between projected baseline and actual revenue is the maximum plausible incremental claim. If agency-attributed revenue is larger than that gap, the report is counting baseline revenue as its own. State the confidence honestly: a trend projection cannot distinguish agency effect from a concurrent price change or a competitor's exit, so treat it as a ceiling, not a measurement.

Holdouts and geo tests

A holdout switches agency-managed activity off for a matched group of users or regions and compares outcomes. Geo tests do the same at market level, which sidesteps user-level tracking entirely. These are the most direct incrementality evidence available. They have limits: you need enough spend and enough regional volume to detect a difference, the test costs revenue in the dark cells, and the agency has to cooperate with a design that may show its work produced less than the dashboard implied. Cooperation, or resistance, is itself informative.

Marketing mix modelling for larger budgets

MMM uses aggregate time-series data to estimate each channel's contribution to revenue after controlling for price, seasonality and other drivers. It works at budget level and needs a couple of years of weekly data with genuine variation in spend. Treat it as a periodic check, quarterly or annual, that validates whether the mix the agency runs is worth its cost. It is too slow and too coarse to serve as a monthly scorecard.

Platform attribution: useful for optimisation, weak for judgement

Each ad platform claims conversions using its own window and its own view of the user. Search, social and retail media can each claim the same purchase, so summing platform-reported conversions produces a number larger than the orders that actually happened. Platform attribution is useful for deciding which ad set to scale this week. It is weak evidence for whether the agency created value, and it should never be summed into a revenue claim.

Choosing the method your data can actually support

Separate incremental revenue from the baseline the agency did not create
MethodWhat it isolatesData and budget requiredMain blind spotSurvives loss of user-level tracking?
Pre/post baselineCeiling on total incremental effect12+ months of monthly revenue; no extra spendConcurrent changes in price, sales, competitionYes
Holdout or geo testCausal effect of a channel or the whole programmeEnough spend and regions to detect a difference; agency cooperationSpillover between test and control; short test windowsYes
Marketing mix modellingChannel-level contribution at budget levelTwo or more years of weekly data with spend variation; analyst timeSlow; struggles with new channels and small budgetsYes
Platform attributionWhich ads and audiences convert relative to each otherFree with the platformDouble-counting across channels; window chosen by the platformNo

The first three methods do not depend on cookies or user-level identifiers, which is why they matter for measurement durability. Ask the agency which of them it can run on your account today. This week, build the twelve-month pre-engagement trendline described above and set it next to the agency's attributed revenue. The size of that gap tells you how much of the report is baseline, following the comparison logic Investopedia describes.

Pricing and RGM Insider explains how marketing mix modeling separates baseline revenue generated without ad spend from true channel-driven incremental sales.
Comparison matrix showing pre/post baseline, holdout tests, and platform attribution against what they isolate and their blind spots.
Comparing incrementality measurement methods helps determine which approach your budget and data can support. A pre/post baseline provides a fast estimate, while holdouts offer stronger causal proof.GPI original conceptual framework

Audit the agency's numbers against your own systems

An agency dashboard is derived data. Someone chose the date range, the attribution window, the conversion events that count, and whether to de-duplicate across platforms. None of those choices is necessarily dishonest, but each one is a decision made by the party being evaluated. Reconciliation replaces trust in the output with a check on the inputs.

Three sources of truth: ad platforms, CRM, finance

Three systems hold the raw material. The ad platforms record spend and their own view of conversions. The CRM records leads, opportunities and closed-won deals with source fields. The finance ledger records invoiced or recognised revenue. Agency reports sit downstream of the first, sometimes the second, and almost never the third. ActiveCollab's KPI guide notes how easily marketers get caught up in vanity metrics instead of outcome numbers; the same drift shows up when a report's line items are chosen for availability rather than traceability.

A reconciliation routine that takes an afternoon

  1. Fix the window. Pick one closed period, last quarter is ideal, and use the same start and end dates in every system.
  2. Export raw spend from each ad platform for that window, with your own login, not a screenshot from the agency.
  3. Export platform conversions and note the attribution window each platform applied.
  4. Pull CRM closed-won for the same window, with source and campaign fields, and pull first-touch and last-touch views separately.
  5. Pull finance revenue for the same window, split by new versus existing customers where possible.
  6. Line up the agency report against each export and record every variance above an agreed tolerance.
  7. Ask for the mapping. For each unresolved variance, request the query logic or filter definitions that produced the agency's figure.

The headline test: the agency's reported revenue for the quarter should trace to closed-won in the CRM and then to the ledger, with each step's difference explained.

Common causes of discrepancy and which ones matter

Audit the agency's numbers against your own systems
Report line itemSource of truthReconciliation checkTypical cause of varianceAction if unresolved
Media spendAd platform billing exportSum of platform invoices equals reported spendMarkup, currency conversion, date cutoffRequest invoice-level breakdown and markup disclosure
Conversions or leadsPlatform export plus CRM lead recordsPlatform count versus CRM records with matching sourceDouble-counting across platforms, view-through, duplicate submissionsAgree one counting rule and one window in the contract
Attributed revenueCRM closed-won with source fieldAgency revenue versus closed-won for the windowAttribution window, model choice, deals sourced by salesRequire the attribution settings document and a last-touch cross-check
Return on ad spendPlatform spend divided by finance revenueRecompute with ledger revenue rather than platform revenuePlatform-reported revenue includes refunds, cancellations, unattributed ordersReport ROAS on ledger revenue only
Period-over-period growthFinance ledgerCompare the same period definitions in both systemsSelective periods, missing months, weekday alignmentFix reporting calendar by contract

The table above is a framework for classifying variances, not a set of measured error rates. Discrepancies fall into three groups. Definitional variances, such as a seven-day click window against a thirty-day one, are expected and become non-issues once written down. Technical variances, such as tag loss after a site release or consent-driven under-reporting, are worth investigating because they degrade the agency's optimisation as much as your judgement. Presentational variances, such as a report that quietly starts after a weak month, are the ones that signal a problem with the relationship.

What a trustworthy report looks like

A report you can trust shows its settings on the page: window, model, de-duplication rule, and the date the definitions last changed. The agency grants read-only access to every ad account and shares the query logic behind the summary figures, not just the summary. It reports revenue on your finance number and treats platform revenue as a leading indicator. Before the next review, request read-only access to every ad account and the CRM report the agency uses, then reconcile last quarter's headline revenue figure to closed-won and to the ledger.

Flowchart showing the reconciliation process from agency report line to platform export, CRM closed-won, and finally finance ledger.
Reconciling agency-reported figures replaces trust in the dashboard with a check on the inputs. Follow the data from the platform to the finance ledger to identify definitional or technical variances.GPI original conceptual framework

True KPIs versus vanity metrics, by business model

Why vanity metrics are a financial risk, not just noise

A vanity metric is any figure that can rise while contribution margin falls. Impressions, reach, clicks and follower counts all qualify, because each can be bought more cheaply by shifting spend toward low-intent inventory. If a team is judged on cost per click, the fastest way to improve is to buy cheaper clicks, and cheaper clicks convert at lower rates and produce lower-value orders. Spend migrates to inventory that flatters the metric and starves the margin. That is the financial risk behind the vanity metrics critique ActiveCollab raises: the numbers can look better every month while the business gets poorer. Which metrics count as true KPIs depends on how you sell.

B2B and long sales cycles

For B2B, the KPIs that map to money are qualified pipeline created, opportunity velocity, win rate on agency-sourced deals, and eventually cohort LTV. The problem is lag. A deal created this quarter may not close for two or three, so a monthly scorecard cannot show revenue. Bridge the gap with leading indicators that have known conversion rates: if your historical close rate from sales-qualified opportunity to closed-won is stable, pipeline created weighted by that rate is a defensible interim measure. Review the weighting quarterly against actual closes so the proxy stays honest.

DTC and ecommerce

For DTC, blended ROAS hides too much. Measure contribution margin per order after discounts, returns and media; the split between new and returning customer revenue, because retargeting existing buyers inflates ROAS without acquiring anyone; and cohort repeat rate, which tells you whether the customers the agency acquires come back.

True KPIs versus vanity metrics, by business model
Metric the agency reportsBusiness model where it misleads mostWhat to measure insteadData source
Impressions and reachBothQualified pipeline (B2B) or new-customer orders (DTC)CRM; order system
Cost per clickBothFully loaded cost per qualified opportunity or per new customerPlatform spend plus CRM or orders
Lead volumeB2BSales-qualified opportunities and win rate on agency-sourced dealsCRM stage history
Blended ROASDTCContribution margin per order, new versus returning splitFinance and order data
Engagement rateBothCohort repeat rate or opportunity velocityOrder cohorts; CRM

The table restates the mapping above as a checklist; none of its rows is a benchmark.

Brand metrics: valuable, but rarely available in time to act

Brand metrics are real, and for some engagements they are the point. Adobe's survey of more than 400 U.S. marketing professionals found brand awareness metrics ranked as most valuable, yet only about 1 in 10 had access to real-time brand data. The survey's publication date is not verified, so treat it as an illustration of a structural problem rather than a current benchmark. The practical consequence: brand measurement belongs in a quarterly review with a defined method, such as branded search volume or a tracked survey, not in a monthly scorecard where it will be filled with impressions. Rewrite the agency's KPI sheet so every line maps to pipeline, margin per order, or retained revenue. Anything that cannot map is a diagnostic, useful for the agency's optimisation, but not a KPI you renew on.

CAC, LTV and contribution margin: the arithmetic that decides the contract

Fully loaded CAC includes the agency fee

The standard customer acquisition cost formula, as Improvado states it, is (total marketing spend + total sales spend) / new customers acquired. The formula is uncontroversial; the argument is over what goes in the numerator. Agency reports routinely show media spend divided by new customers and call it CAC. Insist that the numerator also carry the agency retainer, production costs, tooling and the internal sales cost of converting the leads the agency generates. A CAC that excludes the people you pay to acquire customers is not a cost of acquisition.

Blended CAC and channel CAC answer different questions. An agency can honestly report a low CAC on its channel while blended CAC rises, because the channel is capturing customers who would have arrived through organic search or direct anyway. That cannibalisation is invisible in channel reporting and visible only when you compare blended CAC before and after the engagement, which is the baseline discipline from earlier.

LTV on a cohort basis, not a forecast

Lifetime value should be computed from what customers actually did. Take the customers acquired in a given month, track their gross margin over the following months, and report the observed value at 3, 6 and 12 months. Forecast LTV built from an assumed churn curve is a hypothesis dressed as a result, and agencies have every incentive to use a generous curve. For cash-constrained businesses the more useful number is payback period: how many months of observed contribution margin it takes to recover fully loaded CAC.

Contribution margin after media and fees

The template below shows where each cost enters. It uses placeholders, not measured figures.

  • Fully loaded CAC = (media spend + agency fees + production + tooling + internal sales cost) / new customers acquired
  • Cohort contribution margin at month N = (cohort revenue at month N less cost of goods, discounts and returns)
  • Payback period = first month N where cohort contribution margin per customer equals fully loaded CAC
  • Engagement contribution = incremental contribution margin over the window less (agency fees + media)

Compare the last line with the decision rule you wrote before the review. A hypothetical illustration, not a measured case: if the threshold was that incremental contribution margin must be at least twice total agency cost, and the recomputed figure with fees in the numerator comes in at 1.4 times, the agency has not met the bar regardless of what the ROAS slide says.

Benchmarks: what external comparisons can and cannot tell you

The research behind this article contains no reliable CAC or ROAS benchmarks, and none are offered here. Published figures vary by category, margin structure, price point and, critically, by the attribution method used to produce them, so a benchmark computed on platform-attributed revenue cannot be compared with your ledger-based CAC. The defensible comparison is internal: your pre-engagement baseline and your written margin threshold. Recompute last quarter's CAC with the agency fee and internal sales cost in the numerator using the Improvado formula as the base, then compare payback with the threshold you set.

Read the fee model as an incentive map

Percentage of spend, retainer, performance fee and hybrids

Agency pricing is rarely one thing. NetSuite's summary of Promethean Research's 2025 Digital Agency Industry Report notes that only 4% of surveyed agencies relied on a single pricing model, with most tailoring fees to project, client and risk. That figure reaches you secondhand, so treat the precise number cautiously, but the pattern matches what buyers see: a retainer for strategy, a percentage of media for management, and sometimes a performance component on top.

What each model quietly rewards

Every fee model is an instruction about what to optimise, whether or not anyone wrote it down.

Read the fee model as an incentive map
Fee modelWhat it rewardsRisk to your marginContract clause that offsets it
Percentage of media spendScale over efficiencyBudget growth without margin growthFee cap or declining percentage above a spend tier
Fixed retainerTenure and low varianceUnder-staffing once the account is stableNamed staffing commitment and scope review each quarter
Performance feeWhatever metric the contract namesOptimising a proxy while margin fallsTie the bonus to ledger-verified contribution margin, not platform metrics
HybridA blend of the aboveComplexity hides which lever is being pulledRequire a disclosed breakdown of every revenue line from your account

The table is a framework for reading incentives and does not summarise study findings. The performance-fee row matters most: if the bonus metric is not the true KPI from the earlier sections, the contract is paying for a vanity metric.

Agency economics and what they mean for your account

Agency margins provide context, not verdicts. Promethean's 2026 State of Digital Services, reporting self-reported 2025 figures, found that agencies that narrowed their offerings averaged 30% net margins against a 13% industry average. Read cautiously, this suggests a thin-margin generalist has pressure to under-staff or push scope, while a specialist may be more capable but less flexible about work outside its lane. Margin does not predict client results, and nothing here should be read as saying it does. GPI's guide to evaluating an agency staffing plan covers the staffing side in depth. For this review, list every way the agency earns money from your account, including media markups and any platform rebates, and check whether any of those lines rises when a vanity metric rises.

Bar chart comparing 30 percent net margin for specialized agencies against a 13 percent industry average.
Agencies that narrowed their service offerings reported net margins more than double the industry average in 2025. Data from Promethean Research.Sources: www.hausadvisors.com · hausadvisors.com

Transparency, cadence and whether the measurement will still work next year

Access, not just reports

The transparency standard is short: you hold read-only access to every ad account, the attribution settings are documented, and a change log records every time a definition moves. Reports are a courtesy; access is the requirement. An agency should also be able to explain why each metric is on the page, which is the practical form of the vanity metrics critique ActiveCollab makes. A metric with no stated link to pipeline, margin or retention should be defended or removed.

A cadence tied to decision points

Generic monthly reporting fits no decision in particular. Tie cadence to what you decide at each interval: weekly operational signals for budget and creative moves; a monthly reconciliation of agency figures to platform, CRM and ledger, following the routine described earlier; and a quarterly incrementality and margin review using whichever method from the baseline section your data supports, aligned to the decision window you wrote at the start.

Questions that test measurement durability

This article does not track specific platform cookie timelines, so the questions below are framed generally. Five questions for the quarterly review:

  1. Which reported headline metrics depend on third-party cookies or user-level tracking?
  2. What first-party data and consent instrumentation has the agency built on our properties?
  3. Which reported figures would change if user-level tracking disappeared, and what replaces them?
  4. Can the agency run a holdout, geo test or aggregate model on this account today?
  5. When was the last time the agency volunteered bad news or flagged a data break before we noticed?

That last question is a competence proxy. An agency that finds and reports its own tracking failures is more likely to be optimising on real data.

Turn the evaluation into a retain, restructure or replace decision

Score evidence quality before scoring results

A result you cannot verify cannot carry a renewal decision. Score each dimension twice: once for the quality of evidence available, once for the outcome against the threshold you set. The two inputs that most often move a strong performer into the restructure column are the baseline comparison, in the spirit of Investopedia's prior-month comparison, and a CAC recomputed on the full marketing-plus-sales numerator.

Turn the evaluation into a retain, restructure or replace decision
DimensionEvidence availableResult versus thresholdScoreImplied action
Incremental contribution marginBaseline, holdout or MMM; or platform attribution onlyAbove, near or below the written thresholdEvidence 0-2, result 0-2Strong evidence and below threshold points to replace; weak evidence points to restructure measurement
Reconciliation varianceFull access and traced figures; or report onlyVariances explained or unexplainedEvidence 0-2, result 0-2Unexplained variance with full access is a result problem; no access is an evidence problem
KPI alignmentContract KPIs mapped to pipeline, margin or retentionProportion of reported metrics that mapEvidence 0-2, result 0-2Rewrite KPI sheet before judging results
Incentive alignmentDisclosed fee lines and markupsAny fee line rises with a vanity metricEvidence 0-2, result 0-2Move to hybrid fees tied to ledger margin
Measurement durabilityAnswers to the five durability questionsHeadline metrics survive loss of user-level trackingEvidence 0-2, result 0-2Add holdout or aggregate method requirements

The scorecard is a framework and the scoring scale is a suggestion, not a validated instrument. Its value is in forcing the evidence column to be filled before the result column.

Restructure before you replace

When evidence scores are low and results are unclear, the problem is measurement, and a new agency inherits it. Restructure first: change contract KPIs to true KPIs, move to hybrid fees tied to ledger-verified margin, add a holdout or geo test requirement, and narrow scope to the work the agency has demonstrably done well. Give the restructured arrangement one full decision window before judging it.

When replacement is the right call

Replace when evidence is strong and results are weak: the baseline, reconciliation and margin arithmetic are all in place, and the agency still falls short of the threshold. Also replace when an agency refuses the access and methods that would produce strong evidence, because that refusal makes evaluation impossible by design. Before you search, define what the next candidate must show, so the same gap does not recur; GPI's marketing agency due diligence checklist covers the pre-signing side. Complete the scorecard with your own data and share it with the agency before the renewal conversation, inviting them to contest the evidence rather than the conclusion.

How GPI's evidence standard supports your next agency evaluation

The through-line of this article is a single standard: judge an agency on the quality of its evidence, the method behind it, and the limitations it acknowledges, before you judge the size of its numbers. That is the same standard GPI applies when documenting agencies in its directory, and it is why the baseline principle from Investopedia's ROI guidance is a good first question for any candidate: how would you show that revenue on our account is higher than it would have been without you?

The reconciliation habits and incrementality methods described here are not wasted if you decide to switch. They become the brief for the replacement search. A candidate agency that can explain how it would evidence incrementality on your account, what access it would grant, and which of its fee lines could conflict with your margin has already passed the tests this article describes.

For examples of how GPI documents agency evidence, see profiles such as AB Marketing Group and Amazing Agency. These profiles show the documented-evidence format; they are not claims about any agency's client results, and the scorecard you completed here is the right tool to carry into conversations with any of them.

Frequently asked questions

How long should an agency have before you judge revenue impact?

Long enough for one full sales cycle to complete, plus the ramp period the agency needed to build data. For a DTC brand with a short repurchase cycle, two quarters usually suffice to see acquisition and first repeat behaviour. For enterprise B2B, a fair first judgement on closed revenue may take twelve months or more, so judge interim performance on pipeline weighted by historical close rates and defer the revenue verdict to the window you agreed at signing.

What if the agency refuses to share platform access?

Treat refusal as a transparency failure, not a negotiating position. Read-only access costs the agency nothing and is the only way to reconcile spend and conversions independently. If the accounts are owned by the agency rather than you, the refusal also creates a switching cost that should have been priced into the contract. Make access a written condition of any renewal.

Can you evaluate incrementality without a holdout test?

Yes, with lower confidence. A projected pre-engagement baseline, following the prior-month comparison logic Investopedia describes, gives you a ceiling on the incremental claim. Marketing mix modelling gives channel-level estimates if you have enough history. Neither proves causation the way a well-designed holdout does, so state the method and its limits in the review rather than presenting the result as measured incrementality.

Should the agency fee count in CAC?

Yes. The standard formula divides total marketing plus total sales spend by new customers, and the retainer is marketing spend. Excluding it understates acquisition cost and flatters payback. Include production, tooling and the internal sales cost of converting agency-sourced leads as well.

How do you evaluate an agency that mostly runs brand campaigns?

On a slower cadence and with different evidence. Adobe's survey of more than 400 U.S. marketers found awareness metrics ranked most valuable while only about 1 in 10 had real-time brand data, which explains why monthly brand scorecards fill up with impressions. Agree a quarterly method in advance, such as branded search trend, a tracked awareness survey or a geo test on brand spend, and judge the agency against that method rather than reach.

What variance between agency and finance numbers is acceptable?

There is no fixed percentage. Acceptable variance is whatever the two sides have defined and can explain: a documented attribution window, a known refund lag, a stated de-duplication rule. Unexplained variance of any size is the problem, because it means one of the numbers has an unknown source. Write the tolerance and the explanation process into the contract so the question is settled before the next review.