A media agency pitch is designed to be persuasive. The slides are polished, the case studies show strong returns, and the senior people in the room are articulate. None of that tells you whether the agency can isolate what its media actually caused, whether your data will still be yours when the contract ends, or whether the people presenting will be the people running the account. The cost of finding out later is rising. Gartner's 2025 CMO Spend Survey of 402 CMOs, most at large enterprises, found marketing budgets flat at 7.7% of company revenue, which leaves little room to absorb a wasted year and a re-pitch. This guide sets out media agency evaluation criteria for the RFP and pitch phase, before spend is committed: what to ask, what to request in writing, and how to score the answers.

Most agency evaluations reward the strongest presentation and then discover the measurement gaps twelve months in. GPI's position is that the pitch is the cheapest moment to demand evidence. Every performance claim should arrive with its method, its comparison group and its stated limitation. A ROAS figure without those is an attribution output, and attribution assigns credit rather than proving cause. Buyers should also settle the business question the agency must answer before hearing any plan, then judge each proposal on how directly it answers that question. Culture matters once the work begins. Before that, the criteria that predict a good year are methodological rigor, technical readiness, fair terms and a named team.

What media agency evaluation criteria should test before a contract exists

The criteria in this guide apply to one moment: the period between shortlisting agencies and signing a contract. Post-hire questions, such as reconciling platform-reported results against your finance ledger or deciding whether to retain an incumbent, are covered in GPI's guide to evaluating marketing agency performance by profit. Here the object under evaluation is a proposal and the reasoning behind it, not a live campaign.

Why flat budgets raise the cost of a wrong hire

The Gartner figure in the introduction matters for a specific reason. When budgets hold steady as a share of revenue, a misjudged agency hire cannot be offset by incremental spend the following year. The money lost to a weak first year is simply gone, and the second year starts with a new pitch process, new onboarding and a delay before any readable measurement exists. Agency selection is therefore a capital allocation decision, and it deserves the same evidence standard you would apply to any other.

The stakes concentrate in digital. Nielsen's 2024 Annual Marketing Report found global marketers planning to put more than 63% of media spend into digital channels, based on a survey of marketers managing budgets of $1 million or more. Digital is also where measurement is most contested, because each platform reports on its own performance and signal loss is degrading the deterministic data those reports rely on.

Strategy claims versus deliverable lists

Many proposals describe deliverables: channels to be run, reporting cadence, meeting rhythm, dashboard access. A deliverables list tells you what the agency will do. A strategy tells you why doing it should produce a business outcome, and how the agency would know if it did not. The difference is a causal argument. Proposals without one should be scored as capability descriptions, not strategies.

Four dimensions this guide scores

The sections that follow test four verifiable dimensions: the incrementality logic behind the plan, the technical capability to measure under signal loss, the contract terms that govern fees, data and exit, and the named team that will run the work. These map closely to the categories in GPI's public methodology, which scores documented evidence across Proof of Measurable Outcomes, Creative-Media Integration and Operating Maturity. Most published selection checklists give heavy weight to chemistry and cultural fit and comparatively little to measurement rigor, data ownership or privacy readiness. Those gaps are where the expensive surprises tend to appear.

Before the next pitch, write down the single business decision the media investment must inform, for example whether paid media can grow contribution margin at a set acquisition cost. Then score every proposal on how directly it answers that decision.

How to test an agency's incrementality and ROI logic during the pitch

Put one question to every agency in the room: if we paused this plan for a month, how would you know what we lost? The answer tells you more about measurement maturity than most of the deck.

A two-panel causal impact chart plotting actual sales against a counterfactual baseline and isolating the pointwise impact margin.
A credible measurement case study isolates the causal effect of a campaign from baseline demand by plotting observed outcomes against a predicted counterfactual.Source: 11  Time series methods: Measuring impact without a control group – Everyday causal inference · www.everydaycausal.com · everydaycausal.com

Ask how they would prove the plan added revenue you would not have earned anyway

A credible answer names a design. Holdout groups, geographic tests, matched-market comparisons and media mix modeling are all legitimate approaches, and each has a cost, a minimum scale and a set of conditions under which it fails. An equally credible answer is an honest one: that your budget or sales volume is too small for some of these designs, and here is what the agency would do instead. What you are listening for is whether the agency has a method for separating media impact from baseline demand. A confident promise of a number, unaccompanied by a method, is the weak answer.

This is the question behind a common buyer worry: that the agency will take credit for sales that would have happened anyway. The criterion is whether the agency can describe how it would isolate its own contribution, not whether it forecasts a return.

Separate platform-reported ROAS from a causal claim

Platform ROAS is an input to the argument, never the argument itself. Attribution assigns credit to touchpoints according to rules the platform sets, and each platform grades its own work. A conversion that would have happened anyway can still carry a full attribution credit. So when a deck shows a ROAS figure, the evaluation question is what comparison group sits behind it. If the answer is none, the number describes what the platform reported, which is a different thing from what the media caused.

Request a measurement plan, not a results promise

During the RFP, ask each agency to supply the following in writing:

  1. A proposed measurement plan for the first six months, naming the test design it would use to demonstrate incremental impact.
  2. The baseline it would measure against and how that baseline would be established.
  3. The minimum spend or duration required before the test produces a readable result.
  4. The decision it would recommend if the test detects no lift.

Make this a mandatory RFP field and score blank or promise-only responses at zero. The fourth item is the most revealing. An agency that can say what it would recommend if its own plan shows no lift has thought past the sale. One that cannot answer has not.

Complivex's agency evaluation checklist includes the observation that agencies scoring below the midpoint tend to underperform projections because gaps in attribution rigor, team transparency or cost detail compound across a contract year. This is a qualitative practitioner observation without a published sample, not a statistical finding, but it captures the mechanism well: methodology weaknesses that look minor at pitch stage become structural once the engagement is running.

Read case studies for method, sample and limitation

Apply the same standard to case studies that you would apply to a research paper. Does the study name the method used to establish the result? Does it state the observation period? Is there a comparison group or a stated baseline? Does it acknowledge a limitation? A case study that offers a ROAS figure and nothing else is marketing material. It may still be true, but it cannot be scored as evidence, and it should be rated accordingly in the template later in this guide.

How to test an agency's incrementality and ROI logic during the pitch
Question to ask in the pitchStrong answerWeak answerDocument to request
If we paused this plan for a month, how would you know what we lost?Names a test design, its minimum scale and its failure conditionsCites platform ROAS or historical growth without a comparisonProposed measurement plan with test design
What baseline would you measure against?Describes how baseline demand would be established and updatedAssumes all attributed revenue is incrementalBaseline methodology note
How much spend or time is needed before results are readable?States a minimum with reasoning, or says the budget is too small for some designsPromises results within weeks regardless of scaleMinimum budget and duration statement
What would you recommend if the test shows no lift?Names a specific reallocation or scope changeDeflects or says lift is certainWritten contingency recommendation
What method sits behind this case study?Names method, period, comparison group and limitationROAS figure with no methodCase study methodology appendix

The table structure follows the evidence standard described above; the compounding-gap observation is drawn from Complivex's checklist.

How to assess technical and data capabilities under signal loss

Measurement capability used to be an implementation detail settled after signing. Signal loss has made it a selection criterion. LiveRamp's summary of a 2024 EMARKETER report states that roughly half of advertisers are prioritizing authenticated channels with logged-in audiences as open-web targeting and attribution degrade. That is a secondary reference to a 2024 report rather than a primary reading, but the direction is consistent with what buyers see in their own reporting: deterministic tracking covers less of the customer journey each year, and an agency's plan has to account for that.

Ask each finalist whether it has implemented server-side event delivery and consent management within a stack like yours, meaning your CMS, your CRM and your consent platform. Then ask for two things: a documented implementation plan for your environment and a named technical owner who will be accountable for it. A generic slide about first-party data strategy does not satisfy this criterion. A one-page architecture diagram for your specific stack does.

Diagram of a site instrumented using a server-side tagging container.
A verified server-side tagging architecture routes data from the client to a server container you control, rather than relying on legacy third-party tags.Source: An introduction to server-side tagging  |  Google Tag Manager - Server-side  |  Google for Developers · developers.google.com · developers.google.com

This is also where scope boundaries need to be written down. Define where agency responsibility ends on CRM data passing and where your team's begins. Typically the client owns the CRM, its data quality and any offline conversion definitions; the agency owns the connection to ad platforms and the documentation of what is sent. Creative production follows a similar logic: the proposal should state who produces assets, at what velocity, and what happens when creative volume becomes the constraint on performance. Unresolved boundaries here tend to surface as friction in the first quarter.

Who owns the accounts, pixels and data pipelines

Ad accounts, pixels, conversion APIs, audience lists, data warehouse connections and reporting dashboards should be created in, or transferred to, your ownership at setup. The agency works inside them with administrative access that ends with the contract. In the RFP, specify the access levels you require: client as owner or admin on every ad account and business manager, client as owner of any conversion API configuration, client as owner of the warehouse and dashboard instances. Reject any proposal in which the agency retains ownership of ad accounts or conversion data. The contract section below covers the clause that enforces this; here the point is that the technical setup must make the clause meaningful.

Modeling readiness: what they can do when the pixel goes dark

Ask what the agency does when deterministic tracking degrades further. Credible answers include aggregated conversion modeling, preparing inputs for media mix modeling, improving first-party match rates through authenticated touchpoints, and using the test designs discussed in the previous section. Then ask what data the agency needs from your finance and CRM teams to do this. An agency that has done this work will have a specific list; one that has not will describe it in the abstract. A plan that depends solely on last-click pixel data should be scored as fragile regardless of how good the projected numbers look.

AI tooling: ask what it changes about measurement, not whether they use it

Most proposals now list AI tools. The list itself is not a differentiator and tells you little. What matters is whether the agency has thought about how AI assistants change discovery. Google's Gemini app passed 750 million monthly active users, according to Alphabet's fourth-quarter 2025 earnings release. That is a company-reported figure without third-party audit, and it says nothing measured about advertising impact. It does indicate that a large share of information seeking now happens in interfaces that do not leave a click. Ask the agency how its measurement plan accounts for discovery that never generates a trackable event. An answer that connects back to modeling and testing is a good sign.

How to assess technical and data capabilities under signal loss
CapabilityEvidence to request in the RFPMinimum acceptableRed flag
Server-side event deliveryArchitecture diagram for your stack and a named technical ownerDocumented plan with owner and timelineGeneric first-party data slide with no implementation detail
Consent managementDescription of how consent state governs data sent to platformsConsent respected in every pipeline, documentedConsent treated as a legal team problem
Account and data ownershipList of every asset with owner and access levelClient owns all accounts, pixels, APIs, warehouses and dashboardsAgency retains ownership of any ad account or conversion data
Modeling readinessList of data required from finance and CRM, and the modeling approachSpecific inputs named, method describedReliance on last-click pixel data alone
AI and clickless discoveryExplanation of how the plan measures discovery without clicksTies back to testing or modelingList of AI tools with no measurement implication

The authenticated-channel context is from LiveRamp's summary of the 2024 EMARKETER report; the Gemini adoption figure is from TechCrunch's coverage of Alphabet's Q4 2025 earnings.

How to stress-test the proposed channel strategy against your business model

Once you know how an agency measures, test whether its channel logic fits your business rather than its own habits.

Does the plan reflect your sales cycle or the agency's habits?

A B2B pipeline business with a 90-day sales cycle and a DTC brand converting in three days need different attribution windows, different primary KPIs and different channel logic. Ask each agency to explain how its plan would change if your sales cycle were 90 days instead of three, or the reverse. An agency that understands your category will describe changes to conversion definitions, to the role of upper-funnel channels, and to the timing of any readable result. One that does not will reuse the same plan with different labels.

A practical variant: include one deliberately awkward constraint in the RFP brief, such as a channel you cannot use for regulatory reasons or a sales cycle that outlasts the platform attribution window, and score how each plan adapts to it.

Channel shifts the plan should be able to explain

Channel economics move quickly. EMARKETER estimates that TikTok Shop grew US sales 407.0% in 2024 and 108.0% in 2025, reaching a forecast $15.82 billion. Those are forecast estimates subject to market and regulatory change, and they are not a reason to include the channel in your plan. They are a reason to expect the agency to argue for or against it based on your category, margin structure and audience, rather than on the trend. The same applies to any fast-growing channel: the credible plan explains inclusion and exclusion with reference to your constraints. With more than 63% of planned media going to digital in Nielsen's 2024 survey, the digital mix is where this strategic argument mostly lives.

What a specific strategy looks like versus a templated one

Hand two finalists the same brief and compare. If the recommendations differ according to your constraints, you are looking at strategy. If they differ only by each agency's preferred platforms, you are looking at templates. Score specificity: named audiences, stated assumptions, explicit tradeoffs and a clear statement of what the agency would cut first under a budget reduction.

Scaling assumptions deserve the same treatment. Ask at what spend level the agency expects returns to flatten and how it would detect that point. GPI's guide on how to scale paid media when increasing spend stops increasing growth covers the mechanics; at pitch stage the question is simply whether the agency has a method for spotting diminishing returns before you fund them.

How to review contract terms, fee transparency and data ownership

Contract terms are usually treated as a negotiation detail after selection. They are better treated as evaluation criteria, because an agency's willingness to disclose costs and transfer ownership tells you a good deal about how the engagement will run. This is not legal advice; have counsel review any final terms.

Fee structure: what is billed, what is marked up, what is rebated

Request a full fee schedule: retainer or percentage of media, technology and platform fees, markups on third-party costs, and any rebates, volume incentives or principal-based media the agency earns on your spend. Ask directly whether these will be disclosed and how. The depth of cost detail in the pitch is itself a signal. Complivex's checklist notes that gaps in cost detail compound over a contract year alongside gaps in attribution rigor; it is a practitioner observation rather than a measured finding, but it matches how undisclosed margins tend to surface. Resistance to disclosing markups should be disqualifying.

Data and account ownership on day one and on exit

The clause should state that the client owns ad accounts, pixels, audiences, creative assets, raw performance data and dashboards, and that the agency's administrative access ends with the contract. The technical audit in the earlier section establishes whether the setup makes this true in practice; the clause makes it enforceable.

Termination, notice and transition obligations

Describe the exit scenario in writing before you sign. The contract should specify the notice period, the handover of all credentials, the export of historical data in a usable format, and a transition period with named support. The point is to make an exit boring rather than adversarial. An agency confident in its work has little reason to resist these terms.

Performance clauses that do not distort behavior

Avoid performance clauses tied to platform-reported ROAS. They reward attribution manipulation, because the easiest way to lift a platform-reported number is to shift spend toward audiences that would have converted anyway. Prefer clauses tied to agreed test outcomes, operational commitments such as staffing and reporting, or jointly defined business metrics that your finance team can verify. Also define the scope change process and how out-of-scope creative or data work is priced, so that the first request outside the original brief does not become a dispute.

How to review contract terms, fee transparency and data ownership
ClauseWhat to requireCommon wording to rejectWhy it matters at exit
Fee disclosureFull schedule of fees, markups, rebates and principal mediaFees described as bundled or proprietaryUndisclosed margins are unrecoverable once paid
Account and data ownershipClient owns accounts, pixels, audiences, assets, raw data and dashboardsAgency-owned accounts with client accessAgency-owned assets can be withheld or lost at termination
Termination and noticeDefined notice period and clear termination rightsAutomatic renewal with long lock-inPrevents a hostage period while you re-pitch
Transition obligationsCredential handover, data export and named support for a set periodHandover at agency discretionEnsures continuity of campaigns and history
Performance clausesTied to agreed tests, operational commitments or finance-verified metricsBonuses tied to platform-reported ROASAvoids incentives to game attribution
Scope changeWritten process and pricing for out-of-scope workChange requests handled case by caseReduces disputes over creative and data work

The compounding-gap observation referenced above is from Complivex's evaluation checklist.

Before final selection, send each finalist a one-page term sheet covering fee disclosure, ownership, exit and scope change, and score their redlines. Once a contract is signed, the evaluation shifts from proposed terms to delivered profit, which is the subject of the post-hire performance guide linked in the first section.

How to validate team structure and staffing before the switch happens

The people who present rarely run the account, and most buyers know this. The criterion is whether the agency will commit in writing to who does.

Name the day-to-day team in the proposal

Require the proposal to name the account lead, the media planner, the analyst and the technical owner, with a short description of each person's role in your engagement. Pair this with a written commitment on replacement: notice before any named person is moved, and client approval of the replacement. These commitments belong in the contract draft, not in a verbal assurance at the final presentation.

Check tenure, capacity and account load

Ask for team capacity in numbers: how many accounts each named person supports, and what staff tenure looks like across the agency. This guide cites no benchmark for either figure, so treat the answers as inputs to your own judgment rather than as pass or fail thresholds. What you can score is the willingness to answer. Refusal to share capacity or tenure is a signal about how the engagement will be staffed.

Senior involvement defined in writing

Proposals often list agency leadership under a heading like strategic oversight. Ask for the concrete version: hours per month, cadence of involvement, and the specific decisions the senior person owns. Oversight with no hours or decisions attached rarely turns into involvement after award. This is the staffing element of GPI's Operating Maturity category, where documented processes, staffing stability and escalation paths are observable before signing.

Test the team, not the presenters

Request a working session with the named delivery team, not the pitch team, on a real problem from your brief. Score the quality of their questions rather than the polish of their answers. A delivery team that asks about your conversion definitions, your CRM data quality and your creative approval process is showing you how the first month will go.

A hypothetical illustrates the risk. Imagine a pitch attended by eight agency staff in which nobody can say who will own weekly optimization once the account is live. However well the meeting goes, that is a staffing risk until the agency commits names in writing. The scenario is constructed for this guide and does not describe any specific agency.

Add a named-team schedule to the contract draft with a replacement clause requiring notice and client approval, and hold the delivery-team working session before award rather than after.

A scoring template for comparing media agency pitches

The earlier sections supply the questions. This section supplies the instrument for comparing answers across finalists.

Field-by-field template

For each of the four dimensions, define three to five criteria. Each criterion has a description of the evidence required, a score from 0 to 3, and a notes field recording the specific document the agency supplied to support its score. The notes field is what makes identical-looking decks diverge: once every claim must point to a document, presentation quality stops carrying the evaluation.

A scoring template for comparing media agency pitches
DimensionCriterionEvidence requiredScore 0 to 3Gating criterion?
Incrementality logicTest design proposed for first six monthsWritten measurement plan with design, baseline and minimum scale0 to 3No
Incrementality logicContingency if no lift detectedWritten recommendation0 to 3No
Incrementality logicCase study evidence qualityMethod, period, comparison group and limitation stated0 to 3No
Technical capabilityTracking architecture for your stackOne-page diagram with named technical owner0 to 3No
Technical capabilityAccount and data ownership at setupAsset list with owners and access levels0 to 3No
Technical capabilityModeling readiness under signal lossData requirements and method described0 to 3No
Contract termsFee disclosureFull fee schedule including markups and rebates0 to 3Yes
Contract termsOwnership, exit and transition clausesRedlined term sheet accepted0 to 3Yes
Contract termsPerformance clause designClauses tied to tests or operational commitments0 to 3No
TeamNamed delivery team with replacement clauseTeam schedule in contract draft0 to 3Yes
TeamCapacity and tenure disclosedWritten answers0 to 3No
TeamDelivery-team working sessionSession held and scored0 to 3Yes

Scoring anchors: 0 means no evidence supplied, 1 means a verbal or slide-level claim, 2 means a document with gaps, 3 means a document that names method, sample or owner and states a limitation.

Weighting the dimensions

For enterprise media buyers, weight incrementality logic and technical capability most heavily, since they determine whether you will ever know what the spend achieved. Treat contract terms and team as gating criteria: a score of 0 on any gated line disqualifies the agency regardless of total. Complivex's observation that below-midpoint agencies tend to underperform projections, offered as practitioner experience rather than a statistical result, is the reason for gates rather than averages. A weak contract or an unnamed team is not offset by a strong deck.

Reading the result and handling ties

A high total supported by slide-level claims is a presentation score, not a capability score. Require at least one primary document per dimension before accepting any total, and treat a midpoint score as a sign of compounding gaps rather than an acceptable average. The template maps to GPI's public methodology categories: incrementality logic and case study quality correspond to Proof of Measurable Outcomes, channel and creative reasoning to Creative-Media Integration, and team, process and contract discipline to Operating Maturity, which lets you compare pitch scores with a finalist's documented directory evidence.

Run the process as follows:

  1. Two evaluators score each finalist independently against the supplied documents.
  2. Reconcile discrepancies by returning to the document, not to memory of the pitch.
  3. Apply the gates, then compare weighted totals.
  4. Break ties on evidence quality: the finalist with more scores of 3 backed by primary documents advances.
  5. Only then schedule the final presentation.

How GPI's methodology fits your agency shortlist

GPI's public 100-point methodology scores agencies on documented evidence across Proof of Measurable Outcomes, Creative-Media Integration and Operating Maturity. It takes the same stance this guide does: a claim counts when it arrives with the method, sample and limitation behind it, and it counts less when it does not. The ownership and editorial disclosures behind the index are published on the about page.

A directory score has limits. It assesses the evidence an agency has documented, not whether that agency fits your brief, your sales cycle or your data environment. Inclusion in a directory does not predict results for any specific client, and GPI does not run campaigns. What a documented score can do is give you a shortlist whose claims have already been checked against evidence, so that your pitch evaluation starts from proof rather than from reputation.

Build the shortlist from documented evidence. Score pitches with the template, applying the gates before the totals. Once a contract is signed, move to the post-hire performance guide referenced in the first section to set the first-year measurement plan against profit rather than platform reporting.

Hire the agency whose claims come with method, sample and limitation attached. To see how documented evidence is scored across the three categories, review GPI's public methodology linked above.

FAQ

How much budget do we need before an agency can credibly propose an incrementality test, and what should it propose if we are below that?

There is no single threshold; it depends on your conversion volume, geography and the test design. Ask the agency to state the minimum scale for the design it proposes and to explain how it would proceed if you fall below it, for example with aggregated modeling or a phased test. Score the honesty of that answer, not the size of the promise.

Should performance clauses in a media agency contract be tied to ROAS, and if not, what should they reference?

No. Platform-reported ROAS rewards shifting spend toward audiences that would have converted anyway. Tie clauses instead to agreed test outcomes, operational commitments such as named staffing and reporting cadence, or finance-verified business metrics. Draft the clause with the measurement plan from the RFP so that both documents describe the same success condition.

What is a reasonable notice and transition period to require if we terminate a media agency?

This guide's evidence does not establish an industry norm, so set the period by your own operational needs: long enough to re-pitch and onboard a replacement, with credential handover and data export completed early in the window. Specify named support during transition. Ask each finalist to redline your proposed term and score their response.

How do we evaluate a proposed tracking setup if our own team lacks technical depth?

Require the one-page architecture diagram with a named technical owner and have an independent analytics consultant or your IT team review it against your stack. Focus your own review on ownership: confirm every account, pixel and pipeline is created under your control. Ownership is checkable without engineering expertise; implementation quality can be reviewed by a third party.

How should we score an agency that admits a measurement limitation versus one that promises a number?

Score the admission higher. A stated limitation shows the agency understands the boundaries of its method, which is the same standard applied to case studies in this guide. A promised number without method scores 1 at most on the template. Ask the promising agency for the design behind the number before adjusting its score.

Can a directory score such as GPI's replace our own pitch evaluation?

No. A directory score assesses documented evidence an agency has published or submitted; it cannot assess fit with your brief, data environment or contract terms. Use it to build a shortlist whose claims have been checked, then run the template on the finalists.