Most agency reviews answer the easy question: is the dashboard green? The question that should decide a renewal is harder. Did contribution margin rise by more than the agency cost, compared with what the business would have done without the agency? Formal reviews are nothing new. An ANA survey fielded in 2009 found that 82% of marketers already ran formal agency performance evaluations, and those reviews have traditionally weighed service quality and relationship at least as heavily as money. This guide covers how to evaluate marketing agency performance on incremental revenue and profit: setting a decision rule, separating agency effect from baseline, reconciling agency reports with your own systems, choosing KPIs by business model, running the CAC and margin arithmetic, reading fee incentives, and turning the result into a retain, restructure or replace decision.
GPI's position runs against the usual review. Most evaluations score the relationship and accept the numbers; the relationship matters, but the numbers are what deserve interrogation. Attributed revenue is a claim, produced with settings the agency chose, about a business whose revenue would have moved anyway through seasonality, pricing and sales effort. The test that matters is whether the agency can show, with a method you can independently check, that contribution margin is higher than it would have been without them. When that evidence is weak, the first fix is usually measurement design and contract terms rather than a new agency. Switching partners without changing how you judge one restarts the same uncertainty with a new logo on the deck.
TL;DR: How to evaluate marketing agency performance in five checks
If you want the answer to how to evaluate marketing agency performance in one pass, run these five checks in order:
- Define the decision first. Decide whether the review ends in retain, restructure or replace, and write down the contribution-margin threshold that would justify each outcome before anyone opens a report.
- Measure against a baseline. Compare revenue trends before and after agency activity, or use a holdout, before accepting an attributed figure. Investopedia's ROI guidance makes the same point: ROI only means something against prior-month sales comparisons.
- Reconcile before you debate. Match agency-reported spend, conversions and revenue to your platform exports, CRM and finance ledger for the same window.
- Replace vanity metrics with the two or three indicators that fit your model: pipeline and LTV for long sales cycles, contribution margin per order for DTC. ActiveCollab's KPI guide describes how marketers get caught up in vanity metrics instead of outcome numbers.
- Read the fee model as an incentive map. Pricing structures vary widely, per NetSuite's summary of Promethean data, and whatever the agency is paid on is what it will optimise.
Score the agency on evidence quality, not on the size of the numbers in the deck. Note which of these checks your current process skips.
Start with the decision, not the dashboard
What 'impact on revenue and profit' actually means
Four numbers get called revenue in agency conversations, and only two of them can settle a contract decision.
- Gross revenue is everything the business booked. It moves with pricing, sales headcount, seasonality and product launches, none of which the agency controls.
- Attributed revenue is the share a reporting system assigns to agency-managed channels. It depends on attribution windows, view-through settings and de-duplication rules chosen by whoever configured the report.
- Incremental revenue is the revenue that would not have happened without the agency's work. It requires a counterfactual, which means a baseline, a holdout or a model.
- Contribution margin after media and fees is incremental revenue less cost of goods, less media spend, less agency fees, production and tooling. This is the figure that tells you whether the engagement made money.
The retain or replace question is answered by the last two. Everything else is context or a hypothesis about them.
Set the threshold before you open the report
Write the decision rule in advance, in one sentence, and have finance sign it. A workable form: incremental contribution margin over the evaluation window must exceed total agency cost, fees plus media, by an agreed multiple. The multiple is a business choice that reflects your cost of capital and risk appetite; the discipline is in fixing it before the numbers arrive, so the threshold cannot drift toward whatever the report shows.
The same 2009 ANA survey that found 82% of marketers conducting formal evaluations described a practice built around relationship and service scoring. That survey is old and says nothing about today's tooling, but it illustrates a durable pattern: evaluation has long been routine, while financial rigour has not.
Set the window relative to your sales cycle. Ninety days is meaningless for enterprise B2B, where an opportunity created in month one may close in month nine. The same 90 days is generous for a DTC replenishment product with a three-week repurchase cycle. Agree the window with the agency at contract signing, not at renewal, so neither side can pick a convenient period later.
Separate incremental revenue from the baseline the agency did not create
Revenue moves for reasons that have nothing to do with the agency. Organic growth compounds, seasonality repeats, a price increase lifts average order value, and two new sales hires close deals that marketing did not source. An honest evaluation nets these out before crediting anyone. Investopedia's guidance on campaign ROI is blunt on this point: without comparisons to the business line's sales in the months before launch, the ROI figure has no real meaning. Four methods do that work, with very different costs and confidence levels. What follows is a practical framework, not a set of study findings.
Pre/post baseline comparison
Start here because every company can do it. Pull monthly revenue for the twelve months before the engagement began, fit a simple trend that captures growth and seasonality, and project it across the engagement period. The gap between projected baseline and actual revenue is the maximum plausible incremental claim. If agency-attributed revenue is larger than that gap, the report is counting baseline revenue as its own. State the confidence honestly: a trend projection cannot distinguish agency effect from a concurrent price change or a competitor's exit, so treat it as a ceiling, not a measurement.
Holdouts and geo tests
A holdout switches agency-managed activity off for a matched group of users or regions and compares outcomes. Geo tests do the same at market level, which sidesteps user-level tracking entirely. These are the most direct incrementality evidence available. They have limits: you need enough spend and enough regional volume to detect a difference, the test costs revenue in the dark cells, and the agency has to cooperate with a design that may show its work produced less than the dashboard implied. Cooperation, or resistance, is itself informative.
Marketing mix modelling for larger budgets
MMM uses aggregate time-series data to estimate each channel's contribution to revenue after controlling for price, seasonality and other drivers. It works at budget level and needs a couple of years of weekly data with genuine variation in spend. Treat it as a periodic check, quarterly or annual, that validates whether the mix the agency runs is worth its cost. It is too slow and too coarse to serve as a monthly scorecard.
Platform attribution: useful for optimisation, weak for judgement
Each ad platform claims conversions using its own window and its own view of the user. Search, social and retail media can each claim the same purchase, so summing platform-reported conversions produces a number larger than the orders that actually happened. Platform attribution is useful for deciding which ad set to scale this week. It is weak evidence for whether the agency created value, and it should never be summed into a revenue claim.
Choosing the method your data can actually support
| Method | What it isolates | Data and budget required | Main blind spot | Survives loss of user-level tracking? |
|---|---|---|---|---|
| Pre/post baseline | Ceiling on total incremental effect | 12+ months of monthly revenue; no extra spend | Concurrent changes in price, sales, competition | Yes |
| Holdout or geo test | Causal effect of a channel or the whole programme | Enough spend and regions to detect a difference; agency cooperation | Spillover between test and control; short test windows | Yes |
| Marketing mix modelling | Channel-level contribution at budget level | Two or more years of weekly data with spend variation; analyst time | Slow; struggles with new channels and small budgets | Yes |
| Platform attribution | Which ads and audiences convert relative to each other | Free with the platform | Double-counting across channels; window chosen by the platform | No |
The first three methods do not depend on cookies or user-level identifiers, which is why they matter for measurement durability. Ask the agency which of them it can run on your account today. This week, build the twelve-month pre-engagement trendline described above and set it next to the agency's attributed revenue. The size of that gap tells you how much of the report is baseline, following the comparison logic Investopedia describes.

Audit the agency's numbers against your own systems
An agency dashboard is derived data. Someone chose the date range, the attribution window, the conversion events that count, and whether to de-duplicate across platforms. None of those choices is necessarily dishonest, but each one is a decision made by the party being evaluated. Reconciliation replaces trust in the output with a check on the inputs.
Three sources of truth: ad platforms, CRM, finance
Three systems hold the raw material. The ad platforms record spend and their own view of conversions. The CRM records leads, opportunities and closed-won deals with source fields. The finance ledger records invoiced or recognised revenue. Agency reports sit downstream of the first, sometimes the second, and almost never the third. ActiveCollab's KPI guide notes how easily marketers get caught up in vanity metrics instead of outcome numbers; the same drift shows up when a report's line items are chosen for availability rather than traceability.
A reconciliation routine that takes an afternoon
- Fix the window. Pick one closed period, last quarter is ideal, and use the same start and end dates in every system.
- Export raw spend from each ad platform for that window, with your own login, not a screenshot from the agency.
- Export platform conversions and note the attribution window each platform applied.
- Pull CRM closed-won for the same window, with source and campaign fields, and pull first-touch and last-touch views separately.
- Pull finance revenue for the same window, split by new versus existing customers where possible.
- Line up the agency report against each export and record every variance above an agreed tolerance.
- Ask for the mapping. For each unresolved variance, request the query logic or filter definitions that produced the agency's figure.
The headline test: the agency's reported revenue for the quarter should trace to closed-won in the CRM and then to the ledger, with each step's difference explained.
Common causes of discrepancy and which ones matter
| Report line item | Source of truth | Reconciliation check | Typical cause of variance | Action if unresolved |
|---|---|---|---|---|
| Media spend | Ad platform billing export | Sum of platform invoices equals reported spend | Markup, currency conversion, date cutoff | Request invoice-level breakdown and markup disclosure |
| Conversions or leads | Platform export plus CRM lead records | Platform count versus CRM records with matching source | Double-counting across platforms, view-through, duplicate submissions | Agree one counting rule and one window in the contract |
| Attributed revenue | CRM closed-won with source field | Agency revenue versus closed-won for the window | Attribution window, model choice, deals sourced by sales | Require the attribution settings document and a last-touch cross-check |
| Return on ad spend | Platform spend divided by finance revenue | Recompute with ledger revenue rather than platform revenue | Platform-reported revenue includes refunds, cancellations, unattributed orders | Report ROAS on ledger revenue only |
| Period-over-period growth | Finance ledger | Compare the same period definitions in both systems | Selective periods, missing months, weekday alignment | Fix reporting calendar by contract |
The table above is a framework for classifying variances, not a set of measured error rates. Discrepancies fall into three groups. Definitional variances, such as a seven-day click window against a thirty-day one, are expected and become non-issues once written down. Technical variances, such as tag loss after a site release or consent-driven under-reporting, are worth investigating because they degrade the agency's optimisation as much as your judgement. Presentational variances, such as a report that quietly starts after a weak month, are the ones that signal a problem with the relationship.
What a trustworthy report looks like
A report you can trust shows its settings on the page: window, model, de-duplication rule, and the date the definitions last changed. The agency grants read-only access to every ad account and shares the query logic behind the summary figures, not just the summary. It reports revenue on your finance number and treats platform revenue as a leading indicator. Before the next review, request read-only access to every ad account and the CRM report the agency uses, then reconcile last quarter's headline revenue figure to closed-won and to the ledger.

True KPIs versus vanity metrics, by business model
Why vanity metrics are a financial risk, not just noise
A vanity metric is any figure that can rise while contribution margin falls. Impressions, reach, clicks and follower counts all qualify, because each can be bought more cheaply by shifting spend toward low-intent inventory. If a team is judged on cost per click, the fastest way to improve is to buy cheaper clicks, and cheaper clicks convert at lower rates and produce lower-value orders. Spend migrates to inventory that flatters the metric and starves the margin. That is the financial risk behind the vanity metrics critique ActiveCollab raises: the numbers can look better every month while the business gets poorer. Which metrics count as true KPIs depends on how you sell.
B2B and long sales cycles
For B2B, the KPIs that map to money are qualified pipeline created, opportunity velocity, win rate on agency-sourced deals, and eventually cohort LTV. The problem is lag. A deal created this quarter may not close for two or three, so a monthly scorecard cannot show revenue. Bridge the gap with leading indicators that have known conversion rates: if your historical close rate from sales-qualified opportunity to closed-won is stable, pipeline created weighted by that rate is a defensible interim measure. Review the weighting quarterly against actual closes so the proxy stays honest.
DTC and ecommerce
For DTC, blended ROAS hides too much. Measure contribution margin per order after discounts, returns and media; the split between new and returning customer revenue, because retargeting existing buyers inflates ROAS without acquiring anyone; and cohort repeat rate, which tells you whether the customers the agency acquires come back.
| Metric the agency reports | Business model where it misleads most | What to measure instead | Data source |
|---|---|---|---|
| Impressions and reach | Both | Qualified pipeline (B2B) or new-customer orders (DTC) | CRM; order system |
| Cost per click | Both | Fully loaded cost per qualified opportunity or per new customer | Platform spend plus CRM or orders |
| Lead volume | B2B | Sales-qualified opportunities and win rate on agency-sourced deals | CRM stage history |
| Blended ROAS | DTC | Contribution margin per order, new versus returning split | Finance and order data |
| Engagement rate | Both | Cohort repeat rate or opportunity velocity | Order cohorts; CRM |
The table restates the mapping above as a checklist; none of its rows is a benchmark.
Brand metrics: valuable, but rarely available in time to act
Brand metrics are real, and for some engagements they are the point. Adobe's survey of more than 400 U.S. marketing professionals found brand awareness metrics ranked as most valuable, yet only about 1 in 10 had access to real-time brand data. The survey's publication date is not verified, so treat it as an illustration of a structural problem rather than a current benchmark. The practical consequence: brand measurement belongs in a quarterly review with a defined method, such as branded search volume or a tracked survey, not in a monthly scorecard where it will be filled with impressions. Rewrite the agency's KPI sheet so every line maps to pipeline, margin per order, or retained revenue. Anything that cannot map is a diagnostic, useful for the agency's optimisation, but not a KPI you renew on.
CAC, LTV and contribution margin: the arithmetic that decides the contract
Fully loaded CAC includes the agency fee
The standard customer acquisition cost formula, as Improvado states it, is (total marketing spend + total sales spend) / new customers acquired. The formula is uncontroversial; the argument is over what goes in the numerator. Agency reports routinely show media spend divided by new customers and call it CAC. Insist that the numerator also carry the agency retainer, production costs, tooling and the internal sales cost of converting the leads the agency generates. A CAC that excludes the people you pay to acquire customers is not a cost of acquisition.
Blended CAC and channel CAC answer different questions. An agency can honestly report a low CAC on its channel while blended CAC rises, because the channel is capturing customers who would have arrived through organic search or direct anyway. That cannibalisation is invisible in channel reporting and visible only when you compare blended CAC before and after the engagement, which is the baseline discipline from earlier.
LTV on a cohort basis, not a forecast
Lifetime value should be computed from what customers actually did. Take the customers acquired in a given month, track their gross margin over the following months, and report the observed value at 3, 6 and 12 months. Forecast LTV built from an assumed churn curve is a hypothesis dressed as a result, and agencies have every incentive to use a generous curve. For cash-constrained businesses the more useful number is payback period: how many months of observed contribution margin it takes to recover fully loaded CAC.
Contribution margin after media and fees
The template below shows where each cost enters. It uses placeholders, not measured figures.
- Fully loaded CAC = (media spend + agency fees + production + tooling + internal sales cost) / new customers acquired
- Cohort contribution margin at month N = (cohort revenue at month N less cost of goods, discounts and returns)
- Payback period = first month N where cohort contribution margin per customer equals fully loaded CAC
- Engagement contribution = incremental contribution margin over the window less (agency fees + media)
Compare the last line with the decision rule you wrote before the review. A hypothetical illustration, not a measured case: if the threshold was that incremental contribution margin must be at least twice total agency cost, and the recomputed figure with fees in the numerator comes in at 1.4 times, the agency has not met the bar regardless of what the ROAS slide says.
Benchmarks: what external comparisons can and cannot tell you
The research behind this article contains no reliable CAC or ROAS benchmarks, and none are offered here. Published figures vary by category, margin structure, price point and, critically, by the attribution method used to produce them, so a benchmark computed on platform-attributed revenue cannot be compared with your ledger-based CAC. The defensible comparison is internal: your pre-engagement baseline and your written margin threshold. Recompute last quarter's CAC with the agency fee and internal sales cost in the numerator using the Improvado formula as the base, then compare payback with the threshold you set.
Read the fee model as an incentive map
Percentage of spend, retainer, performance fee and hybrids
Agency pricing is rarely one thing. NetSuite's summary of Promethean Research's 2025 Digital Agency Industry Report notes that only 4% of surveyed agencies relied on a single pricing model, with most tailoring fees to project, client and risk. That figure reaches you secondhand, so treat the precise number cautiously, but the pattern matches what buyers see: a retainer for strategy, a percentage of media for management, and sometimes a performance component on top.
What each model quietly rewards
Every fee model is an instruction about what to optimise, whether or not anyone wrote it down.
| Fee model | What it rewards | Risk to your margin | Contract clause that offsets it |
|---|---|---|---|
| Percentage of media spend | Scale over efficiency | Budget growth without margin growth | Fee cap or declining percentage above a spend tier |
| Fixed retainer | Tenure and low variance | Under-staffing once the account is stable | Named staffing commitment and scope review each quarter |
| Performance fee | Whatever metric the contract names | Optimising a proxy while margin falls | Tie the bonus to ledger-verified contribution margin, not platform metrics |
| Hybrid | A blend of the above | Complexity hides which lever is being pulled | Require a disclosed breakdown of every revenue line from your account |
The table is a framework for reading incentives and does not summarise study findings. The performance-fee row matters most: if the bonus metric is not the true KPI from the earlier sections, the contract is paying for a vanity metric.
Agency economics and what they mean for your account
Agency margins provide context, not verdicts. Promethean's 2026 State of Digital Services, reporting self-reported 2025 figures, found that agencies that narrowed their offerings averaged 30% net margins against a 13% industry average. Read cautiously, this suggests a thin-margin generalist has pressure to under-staff or push scope, while a specialist may be more capable but less flexible about work outside its lane. Margin does not predict client results, and nothing here should be read as saying it does. GPI's guide to evaluating an agency staffing plan covers the staffing side in depth. For this review, list every way the agency earns money from your account, including media markups and any platform rebates, and check whether any of those lines rises when a vanity metric rises.

Transparency, cadence and whether the measurement will still work next year
Access, not just reports
The transparency standard is short: you hold read-only access to every ad account, the attribution settings are documented, and a change log records every time a definition moves. Reports are a courtesy; access is the requirement. An agency should also be able to explain why each metric is on the page, which is the practical form of the vanity metrics critique ActiveCollab makes. A metric with no stated link to pipeline, margin or retention should be defended or removed.
A cadence tied to decision points
Generic monthly reporting fits no decision in particular. Tie cadence to what you decide at each interval: weekly operational signals for budget and creative moves; a monthly reconciliation of agency figures to platform, CRM and ledger, following the routine described earlier; and a quarterly incrementality and margin review using whichever method from the baseline section your data supports, aligned to the decision window you wrote at the start.
Questions that test measurement durability
This article does not track specific platform cookie timelines, so the questions below are framed generally. Five questions for the quarterly review:
- Which reported headline metrics depend on third-party cookies or user-level tracking?
- What first-party data and consent instrumentation has the agency built on our properties?
- Which reported figures would change if user-level tracking disappeared, and what replaces them?
- Can the agency run a holdout, geo test or aggregate model on this account today?
- When was the last time the agency volunteered bad news or flagged a data break before we noticed?
That last question is a competence proxy. An agency that finds and reports its own tracking failures is more likely to be optimising on real data.
Turn the evaluation into a retain, restructure or replace decision
Score evidence quality before scoring results
A result you cannot verify cannot carry a renewal decision. Score each dimension twice: once for the quality of evidence available, once for the outcome against the threshold you set. The two inputs that most often move a strong performer into the restructure column are the baseline comparison, in the spirit of Investopedia's prior-month comparison, and a CAC recomputed on the full marketing-plus-sales numerator.
| Dimension | Evidence available | Result versus threshold | Score | Implied action |
|---|---|---|---|---|
| Incremental contribution margin | Baseline, holdout or MMM; or platform attribution only | Above, near or below the written threshold | Evidence 0-2, result 0-2 | Strong evidence and below threshold points to replace; weak evidence points to restructure measurement |
| Reconciliation variance | Full access and traced figures; or report only | Variances explained or unexplained | Evidence 0-2, result 0-2 | Unexplained variance with full access is a result problem; no access is an evidence problem |
| KPI alignment | Contract KPIs mapped to pipeline, margin or retention | Proportion of reported metrics that map | Evidence 0-2, result 0-2 | Rewrite KPI sheet before judging results |
| Incentive alignment | Disclosed fee lines and markups | Any fee line rises with a vanity metric | Evidence 0-2, result 0-2 | Move to hybrid fees tied to ledger margin |
| Measurement durability | Answers to the five durability questions | Headline metrics survive loss of user-level tracking | Evidence 0-2, result 0-2 | Add holdout or aggregate method requirements |
The scorecard is a framework and the scoring scale is a suggestion, not a validated instrument. Its value is in forcing the evidence column to be filled before the result column.
Restructure before you replace
When evidence scores are low and results are unclear, the problem is measurement, and a new agency inherits it. Restructure first: change contract KPIs to true KPIs, move to hybrid fees tied to ledger-verified margin, add a holdout or geo test requirement, and narrow scope to the work the agency has demonstrably done well. Give the restructured arrangement one full decision window before judging it.
When replacement is the right call
Replace when evidence is strong and results are weak: the baseline, reconciliation and margin arithmetic are all in place, and the agency still falls short of the threshold. Also replace when an agency refuses the access and methods that would produce strong evidence, because that refusal makes evaluation impossible by design. Before you search, define what the next candidate must show, so the same gap does not recur; GPI's marketing agency due diligence checklist covers the pre-signing side. Complete the scorecard with your own data and share it with the agency before the renewal conversation, inviting them to contest the evidence rather than the conclusion.
How GPI's evidence standard supports your next agency evaluation
The through-line of this article is a single standard: judge an agency on the quality of its evidence, the method behind it, and the limitations it acknowledges, before you judge the size of its numbers. That is the same standard GPI applies when documenting agencies in its directory, and it is why the baseline principle from Investopedia's ROI guidance is a good first question for any candidate: how would you show that revenue on our account is higher than it would have been without you?
The reconciliation habits and incrementality methods described here are not wasted if you decide to switch. They become the brief for the replacement search. A candidate agency that can explain how it would evidence incrementality on your account, what access it would grant, and which of its fee lines could conflict with your margin has already passed the tests this article describes.
For examples of how GPI documents agency evidence, see profiles such as AB Marketing Group and Amazing Agency. These profiles show the documented-evidence format; they are not claims about any agency's client results, and the scorecard you completed here is the right tool to carry into conversations with any of them.
Frequently asked questions
How long should an agency have before you judge revenue impact?
Long enough for one full sales cycle to complete, plus the ramp period the agency needed to build data. For a DTC brand with a short repurchase cycle, two quarters usually suffice to see acquisition and first repeat behaviour. For enterprise B2B, a fair first judgement on closed revenue may take twelve months or more, so judge interim performance on pipeline weighted by historical close rates and defer the revenue verdict to the window you agreed at signing.
What if the agency refuses to share platform access?
Treat refusal as a transparency failure, not a negotiating position. Read-only access costs the agency nothing and is the only way to reconcile spend and conversions independently. If the accounts are owned by the agency rather than you, the refusal also creates a switching cost that should have been priced into the contract. Make access a written condition of any renewal.
Can you evaluate incrementality without a holdout test?
Yes, with lower confidence. A projected pre-engagement baseline, following the prior-month comparison logic Investopedia describes, gives you a ceiling on the incremental claim. Marketing mix modelling gives channel-level estimates if you have enough history. Neither proves causation the way a well-designed holdout does, so state the method and its limits in the review rather than presenting the result as measured incrementality.
Should the agency fee count in CAC?
Yes. The standard formula divides total marketing plus total sales spend by new customers, and the retainer is marketing spend. Excluding it understates acquisition cost and flatters payback. Include production, tooling and the internal sales cost of converting agency-sourced leads as well.
How do you evaluate an agency that mostly runs brand campaigns?
On a slower cadence and with different evidence. Adobe's survey of more than 400 U.S. marketers found awareness metrics ranked most valuable while only about 1 in 10 had real-time brand data, which explains why monthly brand scorecards fill up with impressions. Agree a quarterly method in advance, such as branded search trend, a tracked awareness survey or a geo test on brand spend, and judge the agency against that method rather than reach.
What variance between agency and finance numbers is acceptable?
There is no fixed percentage. Acceptable variance is whatever the two sides have defined and can explain: a documented attribution window, a known refund lag, a stated de-duplication rule. Unexplained variance of any size is the problem, because it means one of the numbers has an unknown source. Write the tolerance and the explanation process into the contract so the question is settled before the next review.

