Search for an agency performance evaluation template and most of what comes back is an employee review form with the word agency swapped in. Those forms rate competencies and growth. An advertising agency controls your media budget, bills fees against a scope and reports its own results, so the instrument that governs it has to be commercial. You need two scorecards, not one. A quarterly check governs workflow, service levels and scope. An annual review governs renewal, fees and contract terms. Every number on either sheet is reconciled against your own platform, CRM and finance records before it earns a score. The same discipline GPI recommends at selection, an itemized resource plan that justifies the fee, carries over into the review year (see the RFP template guide).

GPI's position is that a score only matters when it maps to a decision someone in the room is authorized to make. A quarterly review that cannot change scope or process is a status meeting. An annual review that cannot touch fees or the contract is a courtesy. Separate instruments keep each decision attached to the evidence that supports it. The second principle follows: a contribution figure belongs on the scorecard only with its method and limits written beside it. Platform attribution describes correlation inside one model, and an agency's deck is a claim to test rather than a verdict. The template below is built to hold claims to their evidence.

Why HR review templates fail for an advertising agency

What an employee review measures that an agency scorecard cannot

The templates that dominate this search are written for managers reviewing staff. Asana's performance evaluation template, Smartsheet's employee review templates, Indeed's review templates and monday.com's employee review templates are all built around an individual's competencies, goals and development. The Management Center's evaluation form sits in the same family, and the academic literature on measurement, such as this systematic review, is about employees too. Even the word agency is ambiguous in search results; the NCBI hosts a template for agency reports that concerns government agencies in a research program. None of these carry the fields a buyer needs when a vendor spends the media budget.

A competency rating, a growth plan and a 1 to 5 star rating assume a manager who can coach, promote or exit the person. A buyer's levers are different: retain, re-scope, reprice or tender.

What an agency scorecard must add: scope, fees, spend and data ownership

An agency evaluation is a procurement instrument. It has to record contracted scope against delivered work, the fee model and what was actually billed, stewardship of media spend, who owns platform and analytics access, and whether reported numbers survive contact with your own records. Agency-branded templates exist, for example from Demand Metric and Visme, and they can supply a layout, but the commercial fields and their definitions have to come from your contract.

The one thing worth borrowing: defined rating anchors

HR practice does get one thing right: a score needs a written definition at each level. A 1 to 5 rating for strategic thinking with no anchor measures how much the reviewer likes the account lead. A criterion that reads "delivered a written test plan with hypotheses and success thresholds before the quarter's spend was released" can be scored from documents. Everything that follows is written for a buyer with authority over retention, compensation or scope, and every field carries an anchor.

Two cadences, two decisions: quarterly check-ins versus annual reviews

What the quarterly review is allowed to decide

The quarterly review answers operational questions. Is contracted work shipping? Are service levels met? Is the agreed test plan running and producing readouts? Where are the bottlenecks, and are they on the agency's side or yours? Its outputs are corrective actions, process changes and scope clarifications. Marketing leads it.

What the annual review is allowed to decide

The annual review answers commercial questions. Did verified contribution justify fees plus media spend? Should the contract be renewed as is, re-scoped, repriced or put to tender? Marketing, finance and procurement own it jointly, because each holds part of the evidence. GPI's guide to evaluating agency performance by profit covers the commercial logic in more depth.

Why mixing the two cadences corrupts both

Three months is too short to judge contribution. Attribution is noisy, seasonality dominates a single quarter, and a test started in month one may not read out by month three. Twelve months is too long to catch execution drift; by the time a slow creative pipeline shows up in annual revenue, you have paid for four quarters of it. Quarterly scores feed the annual review as inputs, but they never substitute for commercial evidence.

A conceptual example: a Q2 review finds creative refresh turnaround slipped from 10 to 18 working days (hypothetical figures). That is a quarterly finding with an owner and a due date. It is not a renewal argument. The annual review later asks a different question: did that slippage cost verified pipeline?

Two cadences, two decisions: quarterly check-ins versus annual reviews
DimensionQuarterly operational check-inAnnual commercial review
Question answeredIs the work shipping to standard?Did verified contribution justify the cost?
Decision authorityMarketing lead and agency account leadMarketing, finance and procurement jointly
Data requiredDelivery logs, SLA records, test readouts, that quarter's reconciliationReconciled contribution, fee and staffing records, four quarterly totals
Time horizonPast 90 daysPast 12 months
AttendeesDay-to-day teams on both sidesSenior sponsors, finance, procurement, agency leadership
OutputsAction log with owners and datesRenewal decision with documented evidence
What it may not decideFees, contract term, renewalDay-to-day process fixes better handled in-quarter

Write a one-line decision statement for each cadence and send it to the agency before the first review, so both sides know what is on the table.

Comparison matrix showing quarterly check-ins focused on workflow over a 90-day horizon, versus annual reviews focused on commercial contribution over a 12-month horizon.
Quarterly check-ins fix delivery bottlenecks; annual reviews judge commercial contribution against cost. Merging them dilutes the evidence needed to change fees or scope.GPI original conceptual framework

The quarterly scorecard template, field by field

Copy the five sections below into a shared document, assign a source of truth to each field, and score the current quarter retrospectively as a calibration run before the first live review. Scores use a 0 to 3 scale with written anchors; the numbers in this section are placeholders for you to replace.

Section A: Scope and deliverable compliance

Record deliverables contracted versus delivered, change requests raised, and out-of-scope work either absorbed or billed. Source of truth is the scope schedule plus the delivery tracker. The marketing owner fills it. A 3 means all contracted items delivered and every scope change documented; a 0 means material items missing with no change request.

Section B: Service levels and responsiveness

Record turnaround on briefs, response time on escalations and unplanned account team changes. Source of truth is the ticketing or email log and the staffing roster. No independent benchmark for these norms is cited here, so the standard is whatever the contract states; if the contract is silent, agree a target in the first quarter and score against it from the second.

Section C: Test plan and learning velocity

Record the number of pre-registered tests launched, tests concluded with a written readout, and decisions that changed as a result. Source of truth is the test register agreed at the start of the quarter. A hypothetical worked row: "Tests concluded with written readout" scored 1 of 3 because two of five planned tests launched and neither produced a readout. Action logged: agency delivers both readouts by week 2 of next quarter.

Section D: Reporting integrity and data hygiene

Record whether the agency's reported spend and results matched your platform billing, CRM and finance records that quarter, and how large any unexplained gap was. The full reconciliation procedure is described later in this article; here you record only its result and who signed it off.

Section E: Blockers attributed to the buyer

Record late approvals, missing assets, delayed data access and internal decisions that stalled work. Marketing fills this honestly because the annual review depends on it: without this section, a year of client friction reads as agency underperformance.

Scoring anchors and the quarter's action log

The quarterly scorecard template, field by field
FieldWhat is recordedSource of truthOwnerScore anchor (0–3)
Deliverables delivered vs contractedCount and list of variancesScope schedule, delivery trackerMarketing3 all delivered; 2 minor slips documented; 1 material gap with change request; 0 material gap, no request
Brief turnaroundWorking days from brief to first draftTicket or email logMarketing3 within contract target; 2 within tolerance; 1 repeated misses; 0 no tracking possible
Account team changesUnplanned departures or swapsStaffing rosterMarketing3 none; 2 one with handover; 1 one without handover; 0 multiple
Tests concluded with readoutCount vs planTest registerMarketing3 all planned tests read out; 2 most; 1 minority; 0 none
Reported vs reconciled spendVariance after classificationPlatform billing, invoicesFinance3 zero unexplained; 2 within tolerance; 1 above tolerance, explained late; 0 unexplained
Buyer-side blockersCount and days lostApproval logMarketingRecorded as context, not scored against agency

All anchors above are illustrative placeholders to be tightened to your contract. The required output of every quarterly review is an action log with an owner and a due date for each item. A quarterly review that produces no actions was a status meeting.

The annual scorecard template, field by field

Define the weight of each of the five sections and the renewal thresholds before the review year begins, and record them in the contract or an addendum. The annual sheet uses seven columns so that every commercial figure appears twice: as the agency reported it and as you reconciled it.

Section 1: Verified commercial contribution

Record the revenue, pipeline or margin that your own finance and CRM data attributes to agency-managed activity. The attribution or counterfactual method, and what it cannot establish, is written on the scorecard itself; the next section covers how to choose that method. Where no verified figure exists, the field reads "unverified" and scores accordingly.

Section 2: Media and fee efficiency

Record total fees plus media spend against verified contribution. Decompose fees using the itemized resource plan agreed at selection. GPI's RFP guidance describes TrinityP3's requirement for a plan listing "names, roles, hours and rates" so that fee totals are justified and commercial models can be read side by side; the annual scorecard reuses that same plan as its fee-efficiency baseline, so you can see where senior time actually went.

Section 3: Strategic alignment and planning quality

Record whether the agency's annual plan connected to the business decisions you actually made, and whether its recommendations changed spend allocation. Evidence is the plan document, recommendation memos and the budget decisions that followed.

Section 4: Operating maturity and team stability

Record account team turnover, documentation quality, incident handling, and security and data access practices. GPI's published methodology describes how it weighs documented measurable outcomes and operating maturity, the same two categories the annual template scores in Sections 1 and 4. That is a statement of GPI's criteria, not a claim that GPI has scored your agency.

Section 5: Twelve-month trend of quarterly scores

Plot the four quarterly totals alongside the buyer-side blocker scores from each quarter. Improvement or decline is read in context: a falling agency score in a quarter where your team lost three weeks to approvals is a different finding from the same fall in a clean quarter.

The annual scorecard template, field by field
FieldMetric definitionBuyer-side data sourceAgency-reported figureReconciled figureWeightScore
Verified contributionMargin or pipeline attributed by stated methodFinance ledger, CRM with source fieldsPlaceholderPlaceholderSet in advance0–3
Fee and media efficiencyFees plus spend divided by reconciled contributionInvoices, platform billingPlaceholderPlaceholderSet in advance0–3
Resource plan adherenceSenior hours delivered vs plannedStaffing record vs proposal planPlaceholderPlaceholderSet in advance0–3
Strategic alignmentRecommendations adopted into budget decisionsPlan document, budget minutesPlaceholderPlaceholderSet in advance0–3
Operating maturityTurnover, documentation, incidents, access hygieneRoster, incident log, access auditPlaceholderPlaceholderSet in advance0–3
Quarterly trendFour quarterly totals with blocker contextQuarterly scorecardsNot applicableFour totalsSet in advance0–3

All rows are template placeholders; the resource plan mechanism is drawn from GPI's RFP template guidance referenced in the opening.

Renewal decision matrix

Before the year starts, tie each of four outcomes to weighted score thresholds: renew as is, renew with re-scope, renew with repricing, or tender. Record the thresholds where the agency can see them. A hypothetical illustration: a buyer might set renewal at a weighted total above 2.2 with contribution scoring at least 2, re-scope or repricing for totals between 1.5 and 2.2, and tender below 1.5. The numbers are yours to set; the point is that they exist before the results do.

High-stakes fields: proving contribution without trusting attribution

Pick the highest-rigor counterfactual your budget allows for the coming year and write it into the annual scorecard's method field now, so the test runs before the review rather than being reconstructed after it.

Infographic summarizing a controlled experiment where developers using GitHub Copilot completed a task 55% faster than a control group.
A controlled experiment isolates the effect of an intervention by comparing treated and untreated groups on an identical task. Agency claims of incremental revenue should provide a similar control comparison. Source: GitHub Blog.Source: Research: quantifying GitHub Copilot’s impact on developer productivity and happiness - The GitHub Blog LinkedIn icon Instagram icon YouTube icon X icon TikTok icon Twitch icon GitHub icon · github.blog · github.blog

Why platform-reported ROAS is an input, not a verdict

An attribution report describes correlation within one platform's model: conversions that followed an exposure, credited by rules the platform chose. It cannot tell you what would have happened without the spend. That is the question the annual contribution field has to answer, so the scorecard must state which method produced its figure and what that method cannot establish.

Counterfactual options ranked by rigor and cost

No marketing-specific incrementality benchmarks support this article, so the options below are described conceptually, without numbers.

High-stakes fields: proving contribution without trusting attribution
Counterfactual methodWhat it can establishWhat it cannotCost and durationWhen to use
Trend against pre-engagement baselineWhether outcomes moved after the agency startedWhether the move was caused by the agency rather than season, price or marketLow; needs a clean baseline periodFirst year, or when experiments are not feasible
Matched-market or geo holdoutDifference between exposed and unexposed regionsEffects where markets differ in ways you did not matchModerate; several weeks to monthsRegional media with enough comparable markets
Incrementality or conversion-lift testLift within a platform's randomized exposureCross-platform or long-horizon effectsModerate; platform dependentChannel-level contribution questions
Controlled experimentIsolated effect of the intervention on a defined outcomeAnything outside the tested conditionHighest; needs design disciplineHigh-stakes fee or renewal decisions

The design principle behind the last row, the controlled experiment, comes from outside advertising. In a 2022 GitHub study, developers using Copilot completed a defined JavaScript task 55% faster than a control group who did not. The study is historical and from software development; it says nothing about advertising outcomes. Its value for agency buyers is the shape: a treated group and an untreated group on the same task, and an explicit limit, since the result applied to that one task. That is the pair of statements the scorecard requires of any contribution claim.

Documenting the method and its limits on the scorecard

The method field records: the method used, the sample or spend it covered, the observation window, confounders acknowledged (seasonality, pricing changes, competitor activity, tracking changes) and the confidence you place in the resulting figure. Where no counterfactual was feasible, the field reads "unverified" and the contribution score reflects that, rather than accepting the reported number as a stand-in.

How to evaluate the agency's own contribution claims

Apply the same fields to the agency's deck. If a slide claims incremental revenue, ask for the design, the control condition and the period. A claim that arrives with those answers can be scored. One that arrives without them is a hypothesis, and it is fair to say so in the review.

Reconciling agency-reported numbers against your own data

Schedule a reconciliation in the two weeks before each quarterly review, owned by someone outside the day-to-day agency relationship, and add its result as a mandatory scorecard field. The precondition is buyer-owned access: ad accounts, analytics and tracking must sit under your organization's login, or the reconciliation cannot be independent.

  1. Freeze the reporting definitions. Write down, before the first quarter, what counts as spend (gross or net, fees included or excluded), what counts as a lead, a qualified opportunity and a conversion, the conversion window, the currency and the date basis. Disagreements about definitions are cheap to settle in advance and expensive to settle in a review.
  2. Pull buyer-side records independently. Export ad platform billing under your own login, invoices from the finance ledger, CRM pipeline with source fields, and web analytics under your ownership. Do not ask the agency for these pulls; the point is a second, independent read.
  3. Reconcile spend, then volume, then value. Spend should match invoices exactly; there is no tolerance for money. Volume (leads, orders, qualified opportunities) is reconciled within a tolerance you set. Value (revenue or margin) is where attribution differences are expected, and every gap must be explained rather than eliminated.
  4. Classify discrepancies. Sort each gap into one of five classes: definition mismatch, timing mismatch, tracking gap, attribution model difference, or unexplained. Only the last two count against the agency's reporting integrity score; the first three are shared housekeeping.
  5. Record the reconciliation result as a scorecard field. Store the reconciled figures, the tolerance used, the classification and the sign-off name in the quarterly scorecard. Four clean quarterly reconciliations are what make the annual contribution figure credible.

A hypothetical illustration: an agency dashboard reports 1,240 leads for the quarter; the CRM shows 1,010 carrying the agency's source tag. Classification finds 150 duplicates removed by CRM rules, a definition mismatch, and 80 records with no source tag, a tracking gap. Only the 80 count against reporting integrity, and the fix, a tagging repair with an owner and a date, goes into the action log. The exercise takes a few hours once definitions are frozen, and it removes the single largest source of argument from the annual review: whether the numbers being scored are real.

Flowchart showing five steps: freeze definitions, pull buyer records independently, reconcile spend/volume/value, classify discrepancies, and record result.
Reconcile spend, volume, and value independently before scoring the agency. Only track definition and timing gaps to find the true reporting variance.GPI original conceptual framework

Auditing fees, resource plans and hidden senior hours

Before the annual review, request the current-year staffing record by name, role and hours, and lay it beside the resource plan from the original proposal.

Reading the itemized resource plan against actual staffing

The resource plan agreed at selection is the baseline. GPI's RFP guidance describes TrinityP3's approach of tying commercial transparency to the scope document through a plan of names, roles, hours and rates, and that same document is what makes the annual fee audit possible. Compare it with who actually worked the account and at what seniority. A conceptual application: an account pitched with 30% strategy-director time whose staffing record shows 8% is a scored finding, not an anecdote. The percentages are hypothetical; the pattern is what you are checking for.

Screenshot of an agency resource planning template tracking roles, FTE, and planned assignments.
An itemized resource plan tracks the specific names and roles allocated to an account, allowing buyers to verify if senior strategy hours were delivered as contracted. Source: BreezeLeave.Source: Agency Resource Planning Template: Roles, FTE, PTO, and Planned Slots - BreezeLeave Blog · breezeleave.com · breezeleave.com

Fee model checks by structure: retainer, percentage of spend, performance fee

Auditing fees, resource plans and hidden senior hours
Fee modelWhat to audit annuallyTypical drift riskEvidence to request
RetainerHours delivered by role against hours retainedRetainer never reconciles to time; senior hours pitched, junior hours deliveredTimesheets by name and role, staffing roster changes
Percentage of spendFee growth against contribution growthFee rises with budget whether or not verified contribution doesSpend history, reconciled contribution by period
Performance feeTrigger metric and its counterfactual methodFee paid on attributed results rather than incremental onesMethod field from the annual scorecard, test design and readout

No benchmark hourly rates or margin norms exist in the evidence for this article, so the table lists patterns to check rather than thresholds. Present them as questions, not accusations; most drift is unmanaged rather than deliberate.

A performance fee deserves particular care. Its trigger should be the verified contribution field, with the counterfactual method described earlier. A performance fee triggered by platform-reported ROAS pays for attribution, not results.

Scope creep and absorbed work as a two-way ledger

Record work billed beyond scope, and record work the agency absorbed without billing. A fair audit runs in both directions, and an agency that has quietly carried extra work deserves to have it appear in the review. The output of this section is a fee-efficiency score for the annual sheet plus a short written recommendation to procurement: hold, reprice or restructure, with the staffing comparison attached.

Roles, weighting and scoring rules that keep the scorecard honest

Publish a one-page RACI for the scorecard and lock the annual weights in writing before the next quarter opens.

Who scores what: marketing, finance, procurement

Marketing owns the operational and strategic fields, finance owns reconciliation and contribution, procurement owns fees and contract compliance. The rule that makes this work: no field is scored by the person whose own performance depends on the answer. The marketing lead who hired the agency should not be the sole scorer of its contribution.

Roles, weighting and scoring rules that keep the scorecard honest
Scorecard sectionScoring ownerEvidence requiredReviewer
Quarterly A, B, C, EMarketingDelivery tracker, logs, test register, approval logAgency account lead sees and may respond
Quarterly D and annual Section 1FinanceReconciliation record, ledger and CRM exportsMarketing
Annual Section 2 and fee auditProcurementInvoices, staffing record, resource planFinance
Annual Sections 3 and 4MarketingPlan documents, roster, incident logProcurement
Renewal decisionAll three jointlyCompleted annual sheetExecutive sponsor

Setting weights before the year, not after the results

Weights are fixed at the start of the review year and recorded. Adjusting them after seeing results is the most common way a scorecard loses credibility, because it converts a measurement into a justification.

Score anchors, evidence attachments and dissent notes

Every score carries an attached artifact: an export, an invoice, a readout. Every field also carries a dissent note, so that when finance and marketing disagree, the disagreement is visible instead of being averaged into a number nobody believes. Keep the set of metrics escalated to senior leadership small. An adjacent-domain illustration, labelled historical: Cisco's 2021 privacy benchmark study found 93% of respondent organizations reported at least one privacy metric to their board, while 14% reported five or more. It concerns privacy programs, not agencies, and the Future of Privacy Forum has written on measuring such programs, but the discipline transfers: escalate a defined handful of measures, not a dashboard.

Sharing the scorecard with the agency

Share the template, the anchors and the weights at contract start. An evaluation the agency has never seen measures surprise, not performance.

Using the GPI scorecard to decide renewal, renegotiation or a new search

Record this year's renewal decision and its evidence in the scorecard file, and reuse the annual template unchanged as the evaluation sheet for any agency search that follows.

Reading the annual result against thresholds you set in advance

Map the weighted result to the four outcomes fixed at the start of the year, and document which threshold was crossed and by what evidence. A conceptual case: a score below the renewal threshold on verified contribution but above it on operating maturity points to a re-scope or repricing conversation, not an immediate tender. A stable, well-run team that is pointed at the wrong work is a scope problem.

Turning the scorecard into a shortlist brief when you tender

If the outcome is tender, the annual scorecard becomes the evaluation baseline for prospective agencies: same fields, same anchors, same reconciliation demands. Incumbents and challengers are then compared on one instrument rather than on a pitch scorecard that measures presentation. Generic pitch evaluation templates, such as Template.net's, can supply a layout, but the criteria should be yours. GPI's published methodology emphasizes documented evidence of measurable outcomes and operating maturity, the same categories your annual sheet scores, and the Growth Partner Index directory lists scored paid, creator, Amazon and creative agencies as a starting point for a shortlist. GPI's ownership disclosure is published alongside it; treat directory scores as one input to be verified with your own fields, not as an endorsement.

What the scorecard cannot tell you

The scorecard measures what was contracted and verified. It does not capture market shifts, category headwinds, or opportunities nobody briefed. It cannot rescue a relationship where the buyer never froze definitions or owned its data. Within those limits, the argument stands: two cadences, two decisions, one reconciled data foundation.

Frequently asked questions

How do we score a quarter where our own team caused most of the delays?

Score the agency's fields on what it controlled and record the delays in Section E with days lost and causes. Do not inflate agency scores to compensate; the blocker record does that work at the annual review, where the trend is read against your own friction. Assign actions to your side in the same log.

Can one scorecard template evaluate media, creative and Amazon agencies on the same roster?

The structure transfers: cadences, anchors, reconciliation and the reconciled-versus-reported pair are common. Field definitions differ. A creative agency's test plan concerns concepts and formats; an Amazon agency's spend reconciliation runs against retail media invoices and marketplace sales data. Keep the sections and weights consistent across the roster so cross-agency comparison is fair, and vary the metric definitions inside them.

How long should an agency be in place before the first annual commercial review is fair?

A full four quarters, so that seasonality and at least one completed counterfactual are in the record. A shorter tenure can have an interim commercial review, but treat its contribution field as provisional and do not attach renewal or repricing decisions to it.

What do we do when finance and marketing disagree on the reconciled contribution figure?

Use the dissent note. Record both figures, the method behind each and the reason for the gap, then classify it using the five discrepancy classes. If it is an attribution model difference, the scorecard carries the more conservative figure with the alternative noted. Averaging the two hides the question that matters.

Should the agency see its scores, and should it be allowed to submit a rebuttal?

Yes to both. The agency should have seen the template and weights at contract start and should receive its scores with the evidence attached. A written rebuttal, stored with the scorecard, gives the annual review a fuller record and often surfaces buyer-side blockers that Section E missed.

How should a performance-fee clause reference the scorecard without paying for attribution noise?

Tie the trigger to the annual verified contribution field and name the counterfactual method in the clause itself. Specify that a field scored "unverified" triggers no performance payment, and that platform-reported figures alone do not qualify. That aligns the fee with the evidence rather than with whichever model reports the largest number.