Search for an agency performance evaluation template and most of what comes back is an employee review form with the word agency swapped in. Those forms rate competencies and growth. An advertising agency controls your media budget, bills fees against a scope and reports its own results, so the instrument that governs it has to be commercial. You need two scorecards, not one. A quarterly check governs workflow, service levels and scope. An annual review governs renewal, fees and contract terms. Every number on either sheet is reconciled against your own platform, CRM and finance records before it earns a score. The same discipline GPI recommends at selection, an itemized resource plan that justifies the fee, carries over into the review year (see the RFP template guide).
GPI's position is that a score only matters when it maps to a decision someone in the room is authorized to make. A quarterly review that cannot change scope or process is a status meeting. An annual review that cannot touch fees or the contract is a courtesy. Separate instruments keep each decision attached to the evidence that supports it. The second principle follows: a contribution figure belongs on the scorecard only with its method and limits written beside it. Platform attribution describes correlation inside one model, and an agency's deck is a claim to test rather than a verdict. The template below is built to hold claims to their evidence.
Why HR review templates fail for an advertising agency
What an employee review measures that an agency scorecard cannot
The templates that dominate this search are written for managers reviewing staff. Asana's performance evaluation template, Smartsheet's employee review templates, Indeed's review templates and monday.com's employee review templates are all built around an individual's competencies, goals and development. The Management Center's evaluation form sits in the same family, and the academic literature on measurement, such as this systematic review, is about employees too. Even the word agency is ambiguous in search results; the NCBI hosts a template for agency reports that concerns government agencies in a research program. None of these carry the fields a buyer needs when a vendor spends the media budget.
A competency rating, a growth plan and a 1 to 5 star rating assume a manager who can coach, promote or exit the person. A buyer's levers are different: retain, re-scope, reprice or tender.
What an agency scorecard must add: scope, fees, spend and data ownership
An agency evaluation is a procurement instrument. It has to record contracted scope against delivered work, the fee model and what was actually billed, stewardship of media spend, who owns platform and analytics access, and whether reported numbers survive contact with your own records. Agency-branded templates exist, for example from Demand Metric and Visme, and they can supply a layout, but the commercial fields and their definitions have to come from your contract.
The one thing worth borrowing: defined rating anchors
HR practice does get one thing right: a score needs a written definition at each level. A 1 to 5 rating for strategic thinking with no anchor measures how much the reviewer likes the account lead. A criterion that reads "delivered a written test plan with hypotheses and success thresholds before the quarter's spend was released" can be scored from documents. Everything that follows is written for a buyer with authority over retention, compensation or scope, and every field carries an anchor.
Two cadences, two decisions: quarterly check-ins versus annual reviews
What the quarterly review is allowed to decide
The quarterly review answers operational questions. Is contracted work shipping? Are service levels met? Is the agreed test plan running and producing readouts? Where are the bottlenecks, and are they on the agency's side or yours? Its outputs are corrective actions, process changes and scope clarifications. Marketing leads it.
What the annual review is allowed to decide
The annual review answers commercial questions. Did verified contribution justify fees plus media spend? Should the contract be renewed as is, re-scoped, repriced or put to tender? Marketing, finance and procurement own it jointly, because each holds part of the evidence. GPI's guide to evaluating agency performance by profit covers the commercial logic in more depth.
Why mixing the two cadences corrupts both
Three months is too short to judge contribution. Attribution is noisy, seasonality dominates a single quarter, and a test started in month one may not read out by month three. Twelve months is too long to catch execution drift; by the time a slow creative pipeline shows up in annual revenue, you have paid for four quarters of it. Quarterly scores feed the annual review as inputs, but they never substitute for commercial evidence.
A conceptual example: a Q2 review finds creative refresh turnaround slipped from 10 to 18 working days (hypothetical figures). That is a quarterly finding with an owner and a due date. It is not a renewal argument. The annual review later asks a different question: did that slippage cost verified pipeline?
| Dimension | Quarterly operational check-in | Annual commercial review |
|---|---|---|
| Question answered | Is the work shipping to standard? | Did verified contribution justify the cost? |
| Decision authority | Marketing lead and agency account lead | Marketing, finance and procurement jointly |
| Data required | Delivery logs, SLA records, test readouts, that quarter's reconciliation | Reconciled contribution, fee and staffing records, four quarterly totals |
| Time horizon | Past 90 days | Past 12 months |
| Attendees | Day-to-day teams on both sides | Senior sponsors, finance, procurement, agency leadership |
| Outputs | Action log with owners and dates | Renewal decision with documented evidence |
| What it may not decide | Fees, contract term, renewal | Day-to-day process fixes better handled in-quarter |
Write a one-line decision statement for each cadence and send it to the agency before the first review, so both sides know what is on the table.

The quarterly scorecard template, field by field
Copy the five sections below into a shared document, assign a source of truth to each field, and score the current quarter retrospectively as a calibration run before the first live review. Scores use a 0 to 3 scale with written anchors; the numbers in this section are placeholders for you to replace.
Section A: Scope and deliverable compliance
Record deliverables contracted versus delivered, change requests raised, and out-of-scope work either absorbed or billed. Source of truth is the scope schedule plus the delivery tracker. The marketing owner fills it. A 3 means all contracted items delivered and every scope change documented; a 0 means material items missing with no change request.
Section B: Service levels and responsiveness
Record turnaround on briefs, response time on escalations and unplanned account team changes. Source of truth is the ticketing or email log and the staffing roster. No independent benchmark for these norms is cited here, so the standard is whatever the contract states; if the contract is silent, agree a target in the first quarter and score against it from the second.
Section C: Test plan and learning velocity
Record the number of pre-registered tests launched, tests concluded with a written readout, and decisions that changed as a result. Source of truth is the test register agreed at the start of the quarter. A hypothetical worked row: "Tests concluded with written readout" scored 1 of 3 because two of five planned tests launched and neither produced a readout. Action logged: agency delivers both readouts by week 2 of next quarter.
Section D: Reporting integrity and data hygiene
Record whether the agency's reported spend and results matched your platform billing, CRM and finance records that quarter, and how large any unexplained gap was. The full reconciliation procedure is described later in this article; here you record only its result and who signed it off.
Section E: Blockers attributed to the buyer
Record late approvals, missing assets, delayed data access and internal decisions that stalled work. Marketing fills this honestly because the annual review depends on it: without this section, a year of client friction reads as agency underperformance.
Scoring anchors and the quarter's action log
| Field | What is recorded | Source of truth | Owner | Score anchor (0–3) |
|---|---|---|---|---|
| Deliverables delivered vs contracted | Count and list of variances | Scope schedule, delivery tracker | Marketing | 3 all delivered; 2 minor slips documented; 1 material gap with change request; 0 material gap, no request |
| Brief turnaround | Working days from brief to first draft | Ticket or email log | Marketing | 3 within contract target; 2 within tolerance; 1 repeated misses; 0 no tracking possible |
| Account team changes | Unplanned departures or swaps | Staffing roster | Marketing | 3 none; 2 one with handover; 1 one without handover; 0 multiple |
| Tests concluded with readout | Count vs plan | Test register | Marketing | 3 all planned tests read out; 2 most; 1 minority; 0 none |
| Reported vs reconciled spend | Variance after classification | Platform billing, invoices | Finance | 3 zero unexplained; 2 within tolerance; 1 above tolerance, explained late; 0 unexplained |
| Buyer-side blockers | Count and days lost | Approval log | Marketing | Recorded as context, not scored against agency |
All anchors above are illustrative placeholders to be tightened to your contract. The required output of every quarterly review is an action log with an owner and a due date for each item. A quarterly review that produces no actions was a status meeting.
The annual scorecard template, field by field
Define the weight of each of the five sections and the renewal thresholds before the review year begins, and record them in the contract or an addendum. The annual sheet uses seven columns so that every commercial figure appears twice: as the agency reported it and as you reconciled it.
Section 1: Verified commercial contribution
Record the revenue, pipeline or margin that your own finance and CRM data attributes to agency-managed activity. The attribution or counterfactual method, and what it cannot establish, is written on the scorecard itself; the next section covers how to choose that method. Where no verified figure exists, the field reads "unverified" and scores accordingly.
Section 2: Media and fee efficiency
Record total fees plus media spend against verified contribution. Decompose fees using the itemized resource plan agreed at selection. GPI's RFP guidance describes TrinityP3's requirement for a plan listing "names, roles, hours and rates" so that fee totals are justified and commercial models can be read side by side; the annual scorecard reuses that same plan as its fee-efficiency baseline, so you can see where senior time actually went.
Section 3: Strategic alignment and planning quality
Record whether the agency's annual plan connected to the business decisions you actually made, and whether its recommendations changed spend allocation. Evidence is the plan document, recommendation memos and the budget decisions that followed.
Section 4: Operating maturity and team stability
Record account team turnover, documentation quality, incident handling, and security and data access practices. GPI's published methodology describes how it weighs documented measurable outcomes and operating maturity, the same two categories the annual template scores in Sections 1 and 4. That is a statement of GPI's criteria, not a claim that GPI has scored your agency.
Section 5: Twelve-month trend of quarterly scores
Plot the four quarterly totals alongside the buyer-side blocker scores from each quarter. Improvement or decline is read in context: a falling agency score in a quarter where your team lost three weeks to approvals is a different finding from the same fall in a clean quarter.
| Field | Metric definition | Buyer-side data source | Agency-reported figure | Reconciled figure | Weight | Score |
|---|---|---|---|---|---|---|
| Verified contribution | Margin or pipeline attributed by stated method | Finance ledger, CRM with source fields | Placeholder | Placeholder | Set in advance | 0–3 |
| Fee and media efficiency | Fees plus spend divided by reconciled contribution | Invoices, platform billing | Placeholder | Placeholder | Set in advance | 0–3 |
| Resource plan adherence | Senior hours delivered vs planned | Staffing record vs proposal plan | Placeholder | Placeholder | Set in advance | 0–3 |
| Strategic alignment | Recommendations adopted into budget decisions | Plan document, budget minutes | Placeholder | Placeholder | Set in advance | 0–3 |
| Operating maturity | Turnover, documentation, incidents, access hygiene | Roster, incident log, access audit | Placeholder | Placeholder | Set in advance | 0–3 |
| Quarterly trend | Four quarterly totals with blocker context | Quarterly scorecards | Not applicable | Four totals | Set in advance | 0–3 |
All rows are template placeholders; the resource plan mechanism is drawn from GPI's RFP template guidance referenced in the opening.
Renewal decision matrix
Before the year starts, tie each of four outcomes to weighted score thresholds: renew as is, renew with re-scope, renew with repricing, or tender. Record the thresholds where the agency can see them. A hypothetical illustration: a buyer might set renewal at a weighted total above 2.2 with contribution scoring at least 2, re-scope or repricing for totals between 1.5 and 2.2, and tender below 1.5. The numbers are yours to set; the point is that they exist before the results do.
High-stakes fields: proving contribution without trusting attribution
Pick the highest-rigor counterfactual your budget allows for the coming year and write it into the annual scorecard's method field now, so the test runs before the review rather than being reconstructed after it.

Why platform-reported ROAS is an input, not a verdict
An attribution report describes correlation within one platform's model: conversions that followed an exposure, credited by rules the platform chose. It cannot tell you what would have happened without the spend. That is the question the annual contribution field has to answer, so the scorecard must state which method produced its figure and what that method cannot establish.
Counterfactual options ranked by rigor and cost
No marketing-specific incrementality benchmarks support this article, so the options below are described conceptually, without numbers.
| Counterfactual method | What it can establish | What it cannot | Cost and duration | When to use |
|---|---|---|---|---|
| Trend against pre-engagement baseline | Whether outcomes moved after the agency started | Whether the move was caused by the agency rather than season, price or market | Low; needs a clean baseline period | First year, or when experiments are not feasible |
| Matched-market or geo holdout | Difference between exposed and unexposed regions | Effects where markets differ in ways you did not match | Moderate; several weeks to months | Regional media with enough comparable markets |
| Incrementality or conversion-lift test | Lift within a platform's randomized exposure | Cross-platform or long-horizon effects | Moderate; platform dependent | Channel-level contribution questions |
| Controlled experiment | Isolated effect of the intervention on a defined outcome | Anything outside the tested condition | Highest; needs design discipline | High-stakes fee or renewal decisions |
The design principle behind the last row, the controlled experiment, comes from outside advertising. In a 2022 GitHub study, developers using Copilot completed a defined JavaScript task 55% faster than a control group who did not. The study is historical and from software development; it says nothing about advertising outcomes. Its value for agency buyers is the shape: a treated group and an untreated group on the same task, and an explicit limit, since the result applied to that one task. That is the pair of statements the scorecard requires of any contribution claim.
Documenting the method and its limits on the scorecard
The method field records: the method used, the sample or spend it covered, the observation window, confounders acknowledged (seasonality, pricing changes, competitor activity, tracking changes) and the confidence you place in the resulting figure. Where no counterfactual was feasible, the field reads "unverified" and the contribution score reflects that, rather than accepting the reported number as a stand-in.
How to evaluate the agency's own contribution claims
Apply the same fields to the agency's deck. If a slide claims incremental revenue, ask for the design, the control condition and the period. A claim that arrives with those answers can be scored. One that arrives without them is a hypothesis, and it is fair to say so in the review.
Reconciling agency-reported numbers against your own data
Schedule a reconciliation in the two weeks before each quarterly review, owned by someone outside the day-to-day agency relationship, and add its result as a mandatory scorecard field. The precondition is buyer-owned access: ad accounts, analytics and tracking must sit under your organization's login, or the reconciliation cannot be independent.
- Freeze the reporting definitions. Write down, before the first quarter, what counts as spend (gross or net, fees included or excluded), what counts as a lead, a qualified opportunity and a conversion, the conversion window, the currency and the date basis. Disagreements about definitions are cheap to settle in advance and expensive to settle in a review.
- Pull buyer-side records independently. Export ad platform billing under your own login, invoices from the finance ledger, CRM pipeline with source fields, and web analytics under your ownership. Do not ask the agency for these pulls; the point is a second, independent read.
- Reconcile spend, then volume, then value. Spend should match invoices exactly; there is no tolerance for money. Volume (leads, orders, qualified opportunities) is reconciled within a tolerance you set. Value (revenue or margin) is where attribution differences are expected, and every gap must be explained rather than eliminated.
- Classify discrepancies. Sort each gap into one of five classes: definition mismatch, timing mismatch, tracking gap, attribution model difference, or unexplained. Only the last two count against the agency's reporting integrity score; the first three are shared housekeeping.
- Record the reconciliation result as a scorecard field. Store the reconciled figures, the tolerance used, the classification and the sign-off name in the quarterly scorecard. Four clean quarterly reconciliations are what make the annual contribution figure credible.
A hypothetical illustration: an agency dashboard reports 1,240 leads for the quarter; the CRM shows 1,010 carrying the agency's source tag. Classification finds 150 duplicates removed by CRM rules, a definition mismatch, and 80 records with no source tag, a tracking gap. Only the 80 count against reporting integrity, and the fix, a tagging repair with an owner and a date, goes into the action log. The exercise takes a few hours once definitions are frozen, and it removes the single largest source of argument from the annual review: whether the numbers being scored are real.

Auditing fees, resource plans and hidden senior hours
Before the annual review, request the current-year staffing record by name, role and hours, and lay it beside the resource plan from the original proposal.
Reading the itemized resource plan against actual staffing
The resource plan agreed at selection is the baseline. GPI's RFP guidance describes TrinityP3's approach of tying commercial transparency to the scope document through a plan of names, roles, hours and rates, and that same document is what makes the annual fee audit possible. Compare it with who actually worked the account and at what seniority. A conceptual application: an account pitched with 30% strategy-director time whose staffing record shows 8% is a scored finding, not an anecdote. The percentages are hypothetical; the pattern is what you are checking for.

Fee model checks by structure: retainer, percentage of spend, performance fee
| Fee model | What to audit annually | Typical drift risk | Evidence to request |
|---|---|---|---|
| Retainer | Hours delivered by role against hours retained | Retainer never reconciles to time; senior hours pitched, junior hours delivered | Timesheets by name and role, staffing roster changes |
| Percentage of spend | Fee growth against contribution growth | Fee rises with budget whether or not verified contribution does | Spend history, reconciled contribution by period |
| Performance fee | Trigger metric and its counterfactual method | Fee paid on attributed results rather than incremental ones | Method field from the annual scorecard, test design and readout |
No benchmark hourly rates or margin norms exist in the evidence for this article, so the table lists patterns to check rather than thresholds. Present them as questions, not accusations; most drift is unmanaged rather than deliberate.
A performance fee deserves particular care. Its trigger should be the verified contribution field, with the counterfactual method described earlier. A performance fee triggered by platform-reported ROAS pays for attribution, not results.
Scope creep and absorbed work as a two-way ledger
Record work billed beyond scope, and record work the agency absorbed without billing. A fair audit runs in both directions, and an agency that has quietly carried extra work deserves to have it appear in the review. The output of this section is a fee-efficiency score for the annual sheet plus a short written recommendation to procurement: hold, reprice or restructure, with the staffing comparison attached.
Roles, weighting and scoring rules that keep the scorecard honest
Publish a one-page RACI for the scorecard and lock the annual weights in writing before the next quarter opens.
Who scores what: marketing, finance, procurement
Marketing owns the operational and strategic fields, finance owns reconciliation and contribution, procurement owns fees and contract compliance. The rule that makes this work: no field is scored by the person whose own performance depends on the answer. The marketing lead who hired the agency should not be the sole scorer of its contribution.
| Scorecard section | Scoring owner | Evidence required | Reviewer |
|---|---|---|---|
| Quarterly A, B, C, E | Marketing | Delivery tracker, logs, test register, approval log | Agency account lead sees and may respond |
| Quarterly D and annual Section 1 | Finance | Reconciliation record, ledger and CRM exports | Marketing |
| Annual Section 2 and fee audit | Procurement | Invoices, staffing record, resource plan | Finance |
| Annual Sections 3 and 4 | Marketing | Plan documents, roster, incident log | Procurement |
| Renewal decision | All three jointly | Completed annual sheet | Executive sponsor |
Setting weights before the year, not after the results
Weights are fixed at the start of the review year and recorded. Adjusting them after seeing results is the most common way a scorecard loses credibility, because it converts a measurement into a justification.
Score anchors, evidence attachments and dissent notes
Every score carries an attached artifact: an export, an invoice, a readout. Every field also carries a dissent note, so that when finance and marketing disagree, the disagreement is visible instead of being averaged into a number nobody believes. Keep the set of metrics escalated to senior leadership small. An adjacent-domain illustration, labelled historical: Cisco's 2021 privacy benchmark study found 93% of respondent organizations reported at least one privacy metric to their board, while 14% reported five or more. It concerns privacy programs, not agencies, and the Future of Privacy Forum has written on measuring such programs, but the discipline transfers: escalate a defined handful of measures, not a dashboard.
Sharing the scorecard with the agency
Share the template, the anchors and the weights at contract start. An evaluation the agency has never seen measures surprise, not performance.
Using the GPI scorecard to decide renewal, renegotiation or a new search
Record this year's renewal decision and its evidence in the scorecard file, and reuse the annual template unchanged as the evaluation sheet for any agency search that follows.
Reading the annual result against thresholds you set in advance
Map the weighted result to the four outcomes fixed at the start of the year, and document which threshold was crossed and by what evidence. A conceptual case: a score below the renewal threshold on verified contribution but above it on operating maturity points to a re-scope or repricing conversation, not an immediate tender. A stable, well-run team that is pointed at the wrong work is a scope problem.
Turning the scorecard into a shortlist brief when you tender
If the outcome is tender, the annual scorecard becomes the evaluation baseline for prospective agencies: same fields, same anchors, same reconciliation demands. Incumbents and challengers are then compared on one instrument rather than on a pitch scorecard that measures presentation. Generic pitch evaluation templates, such as Template.net's, can supply a layout, but the criteria should be yours. GPI's published methodology emphasizes documented evidence of measurable outcomes and operating maturity, the same categories your annual sheet scores, and the Growth Partner Index directory lists scored paid, creator, Amazon and creative agencies as a starting point for a shortlist. GPI's ownership disclosure is published alongside it; treat directory scores as one input to be verified with your own fields, not as an endorsement.
What the scorecard cannot tell you
The scorecard measures what was contracted and verified. It does not capture market shifts, category headwinds, or opportunities nobody briefed. It cannot rescue a relationship where the buyer never froze definitions or owned its data. Within those limits, the argument stands: two cadences, two decisions, one reconciled data foundation.
Frequently asked questions
How do we score a quarter where our own team caused most of the delays?
Score the agency's fields on what it controlled and record the delays in Section E with days lost and causes. Do not inflate agency scores to compensate; the blocker record does that work at the annual review, where the trend is read against your own friction. Assign actions to your side in the same log.
Can one scorecard template evaluate media, creative and Amazon agencies on the same roster?
The structure transfers: cadences, anchors, reconciliation and the reconciled-versus-reported pair are common. Field definitions differ. A creative agency's test plan concerns concepts and formats; an Amazon agency's spend reconciliation runs against retail media invoices and marketplace sales data. Keep the sections and weights consistent across the roster so cross-agency comparison is fair, and vary the metric definitions inside them.
How long should an agency be in place before the first annual commercial review is fair?
A full four quarters, so that seasonality and at least one completed counterfactual are in the record. A shorter tenure can have an interim commercial review, but treat its contribution field as provisional and do not attach renewal or repricing decisions to it.
What do we do when finance and marketing disagree on the reconciled contribution figure?
Use the dissent note. Record both figures, the method behind each and the reason for the gap, then classify it using the five discrepancy classes. If it is an attribution model difference, the scorecard carries the more conservative figure with the alternative noted. Averaging the two hides the question that matters.
Should the agency see its scores, and should it be allowed to submit a rebuttal?
Yes to both. The agency should have seen the template and weights at contract start and should receive its scores with the evidence attached. A written rebuttal, stored with the scorecard, gives the annual review a fuller record and often surfaces buyer-side blockers that Section E missed.
How should a performance-fee clause reference the scorecard without paying for attribution noise?
Tie the trigger to the annual verified contribution field and name the counterfactual method in the clause itself. Specify that a field scored "unverified" triggers no performance payment, and that platform-reported figures alone do not qualify. That aligns the fee with the evidence rather than with whichever model reports the largest number.

