Agency reference check questions rarely get the preparation the pitch does. The U.S. Office of Personnel Management's guidance on reference checking notes that structure improves the validity of reference checks, and the logic holds when the candidate is an agency. Use the same written script for every shortlisted partner. This guide explains how to check agency references: which clients to call, what to ask about staffing, scope, communication and results, and how to compare their answers. Before drafting your question list, it helps to see how documented agency evidence is laid out on a GPI profile, for example ACE Agency on Growth Partner Index.

GPI's position is that an agency claim is only as strong as the evidence, methodology and limitations behind it, and that standard should not stop at the public record. A reference call is where a buyer gets to test the case study against the client's own numbers and to ask how a result was measured before deciding what it proves. It is also where attribution stories meet operational reality: who actually ran the account, what slipped, and how the agency behaved when a test failed. The purpose is not to collect endorsements. It is to answer a concrete decision, which is whether this team, working on your category and budget, is likely to behave the way its pitch promised.

Why curated agency references still leak useful signal

An agency chooses its references; you choose the questions to ask agency references. Even a satisfied client can describe account-lead changes, undelivered scope and discrepancies between reported results and finance records.

References are a behavioral interview, not a satisfaction survey

OPM's hiring guidance invokes behavioral consistency: past behaviour informs future expectations. Apply that cautiously to comparable agency accounts. 'Who ran your weekly call in month six?' yields a checkable fact; 'Were you happy?' yields an opinion.

The analogy with individual hiring has limits, and one figure from that world shows why unstructured vetting carries risk there: in Checkr's 2025 survey of hiring managers, 31% said they had personally interviewed a candidate using a fake identity. That is a statistic about job candidates and says nothing about agencies. It does illustrate what happens when a check rests on impression rather than verification.

Structure is what separates signal from politeness

For a marketing agency reference check, adapt GoodHire's hiring recommendation: use a shared question template. Ask every reference the same questions in the same order to make answers comparable.

This article is written for buyers of advertising and media services. The current search results for this keyword are dominated by templates for hiring individuals, such as Rutgers' sample reference questions, and that guidance does not transfer cleanly to a contract with a team. Employment-law and background-screening compliance questions are out of scope here.

Set up the reference call before you ask a single question

Send the agency a written reference request before any call is scheduled. Specify the category match you need, the spend band, at least one former or friction-period client, and the names and roles you expect to speak with. The quality of the call is mostly decided at this stage.

Ask for references that match your category and budget band

GPI's Amazon ads agency buyer guide distinguishes CPG, durables and apparel playbooks. Apply the same relevance test to references: compare category, business model and spend band before weighing an endorsement.

  • Category
  • Business model
  • Spend band
  • Engagement type

When two of the four miss, ask for a substitute before you accept the call.

Insist on at least one reference from a hard period

Crosschq recommends a challenging-period reference for executive hiring. Adapt that request to agencies: seek a former client or one whose account had a rough quarter. Record any refusal as a limit on the evidence available.

Agree the format: 30 minutes, one caller, one note-taker, same script

  1. Confirm the reference's role, account dates and firsthand involvement.
  2. Explain how notes will be used and who can see them; clarify confidentiality expectations.
  3. Explain that all references receive the same questions.
  4. Have the future relationship owner ask questions and a colleague take notes.
  5. Reserve five minutes for the rehire question and follow-up.

Decide what you will do with a lukewarm answer before you hear one

Write down, before the first call, which answers are disqualifiers and which trigger a follow-up question to the agency. For example:

  • A reference who learned about their lead strategist's departure after the fact might trigger a follow-up.
  • A fee increase the client discovered on an invoice might be a disqualifier.

Teams that skip this step tend to rationalise warnings once they have already pictured working with a particular agency.

Flow diagram showing reference call setup steps: establish criteria, agree on format, use one script, score within an hour, and reconcile against proposal.
Establishing structure and consistency before calling past clients ensures that every agency's references are evaluated on an identical rubric.Sources: www.goodhire.com, www.opm.gov · goodhire.com

Questions that reveal who actually works on your account

Ask every reference to name the people they worked with most and how long each of them stayed on the account, then reconcile those names against the roster and named roles in the agency's proposal. This block addresses a familiar complaint in agency relationships: the people who won the pitch were rarely seen again after signing.

Pitch team versus delivery team

Ask who pitched, who remained after 90 days, who ran weekly calls and how often the named strategist appeared. Names and approximate dates can be checked against the proposal. Follow up on 'the whole team was great' with a request for specific people.

Turnover and handoffs during the engagement

'How many account leads did you have over the engagement?' 'How were handoffs handled?' 'Did you learn about a staffing change before or after it happened?' The third question is the sharpest, because it tests whether the agency manages its own churn as a client-facing event or as an internal matter the client discovers later.

Seniority and decision rights of the day-to-day lead

'When you needed a strategic call made, who made it and how long did it take?' 'Could your day-to-day lead reallocate budget or change a test plan without escalating?' These questions test whether the seniority in the proposal matched the seniority on the account.

Questions that reveal who actually works on your account
QuestionWhat you are testingStrong answer patternWarning signal
Who presented in the pitch, and how many were still on your account 90 days later?Pitch-to-delivery continuityNames two or three people from the pitch who stayed in named rolesCannot recall the pitch team, or says none of them stayed
Who ran your weekly call?Identity and stability of the day-to-day leadOne named person for most of the engagement, or a planned handoff explained in advanceSeveral names, uncertain order, or the client ran the call themselves
How many account leads did you have?Turnover rate on the accountOne or two, with reasons for any changeThree or more in a year, or the reference has to count
Did you learn about staffing changes before or after they happened?Whether churn is managed as a client eventAdvance notice with a named successor and overlap periodFound out from an out-of-office reply or a new face on a call
How were handoffs handled?Knowledge transfer disciplineWritten account history, joint calls, no repeated briefingClient had to re-explain the business to the new lead
Who made strategic calls, and how long did it take?Decision rights of the account teamDay-to-day lead had authority, or escalation took days rather than weeksEvery decision went to someone the client never met
How often did you see the named strategist?Whether senior time was real or decorativeRegular cadence, present at planning and quarterly reviewsAppeared at the pitch and the renewal
Did the team change when your spend changed?Whether staffing tracks revenue rather than needTeam stayed stable or changes were discussed openlySenior people disappeared after a budget cut without comment

The question structure follows the OPM guidance on structured reference checking and the Crosschq recommendation to include a reference from a challenging period. The answer patterns are the article's proposed interpretation, not measured outcomes.

What a strong answer sounds like and what a warning sounds like

Weight names and handoff details above general satisfaction. A client who experienced a strategist's departure can explain who filled the gap, how long it lasted and whether the replacement arrived briefed. Such firsthand detail is more useful than an endorsement when judging continuity.

One caution on interpretation. Turnover in an agency is normal, and a reference who reports one planned handoff with a proper overlap is describing competent management. The warning is churn the client found out about on their own. Where staffing changes caused scope or reporting problems, note them and carry them into the next block.

Compare what a reference tells you about staffing against the publicly documented information on a profile such as DEPT on Growth Partner Index, and note where the two accounts diverge.

Comparison table of staffing reference questions, mapping strong answer patterns against warning signals.
Interpreting staffing answers during a reference call requires distinguishing between normal planned transitions and unmanaged churn.Sources: www.opm.gov, www.crosschq.com · opm.gov

Questions that expose scope creep, budget variance and missed SLAs

For each agency, record at least one specific scope incident, one budget incident and one SLA incident described by a reference, including how the client found out and how the agency responded. This block tests commercial behaviour, and no HR reference template asks about it, because an employee does not carry a statement of work.

Scope: what was promised, what was delivered, what you ended up doing yourself

Ask what the original statement of work promised but never delivered, and what the client had to produce internally. Probe whether extra scope repaired an omission in the original plan. A concrete list, such as landing pages or pixel implementation, identifies workload your team could inherit.

Budget: retainer drift, media fee changes and unbilled overages

Ask who initiated fee changes, which charges surprised the client and how spend variances were explained. Variance alone proves little. A client learning about it only from an invoice reveals a communication problem worth investigating.

SLAs and timing: launches, reporting deadlines and turnaround on changes

'How often were launch dates or reporting deadlines missed, and how were you told?' 'How long did a creative or targeting change typically take to go live?' 'When a deadline slipped, did the agency propose a new date or did you have to chase?' These probe whether the agency treats its commitments as commitments or as targets.

Boundaries: what the agency owned versus what your team had to carry

'Who was accountable when a dependency slipped, such as creative assets, tracking implementation or approvals?' 'Was the division of responsibilities written down, and did it hold?' Question lists written for individual hiring, such as The Hartford's, are built around one person's conduct. The agency version has to ask about the seam between two organisations, because that seam is where most delivery failures live.

Questions that expose scope creep, budget variance and missed SLAs
QuestionOperational risk it testsAnswer that reassuresAnswer that requires follow-up
What in the original statement of work never shipped?Scope shortfallNothing material, or a named item dropped by mutual agreementItems the client still expected and had to raise
What did you produce internally that you expected the agency to handle?Hidden internal workloadLittle, or items agreed as client-owned at kickoffA list of deliverables the client absorbed without a fee adjustment
Did the fee change, and who initiated it?Retainer driftChange proposed in writing with a rationale before it took effectClient noticed the change on an invoice
Were there charges you did not anticipate?Unbilled or surprise overagesNone, or one incident explained and creditedRecurring surprises or disputes over what was included
How did actual spend compare with plan month to month?Pacing disciplineSmall variances flagged by the agency in reportingLarge variances the client found in platform data
How were missed deadlines communicated?SLA behaviourAdvance warning with a revised dateSilence until the client asked
How long did a change take to go live?Turnaround responsivenessConsistent turnaround the client could plan aroundUnpredictable, or dependent on chasing a specific person
Who was accountable when a dependency slipped?Boundary clarityWritten responsibilities that held under pressureEach side assumed the other owned it

Asking these questions in the same order of every reference follows the structure principle in the OPM reference checking guidance and the shared-template rule in GoodHire's guidance. No benchmark for typical scope creep or budget overrun is offered here, because none exists in the evidence gathered for this article.

Structure matters most in this block. If one agency's references describe two unplanned fee changes and another's describe none, that comparison only means something because both sets were asked the identical question.

Matrix detailing questions about scope, budget, and SLAs, contrasting reassuring answers with those needing follow-up.
Differentiating between normal engagement evolution and unmanaged scope creep by analyzing how a reference describes variance.Sources: www.opm.gov · opm.gov

Questions about how problems got solved and how communication felt

Start this part of the agency reference call with the worst month. Ask for an incident and response. If the reference cannot recall one, establish how closely they worked with the agency before weighting their answers.

Ask for the worst month, not the best campaign

Ask what caused the worst month, who noticed first and what happened in the next two weeks. Dates and names make follow-up possible. Crosschq's challenging-period recommendation helps identify someone able to answer.

Escalation: who you called and what happened next

'When something went wrong, who did you escalate to, and did that person have the authority to fix it?' 'Was there ever a problem you had to raise twice?' 'How long did it take to get a named owner for the fix?' Reconstruct the escalation path from the answer:

  • Who noticed the problem
  • Who was called
  • Who had authority to act
  • How much time elapsed

That structure can be compared across agencies in a way an anecdote cannot.

Ownership: did the agency say 'we got it wrong' unprompted?

Self-reported mistakes are the strongest evidence of accountability a reference can supply, because they cost the agency something and cannot be staged for a reference call. Ask: 'Did the agency ever flag its own mistake before you did?' 'How did they handle a campaign or test that failed?' 'Did the post-mortem name a cause, or did it name market conditions?'

Communication cadence and friction

'What was the meeting cadence, and did it hold?' 'Were reporting calls analysis or slide reading?' 'Did tough news arrive by phone or by email?' 'How quickly did you get a substantive reply, as opposed to an acknowledgement?' Communication complaints are common in agency relationships and easy for a reference to soften, so ask for specifics rather than ratings.

In a hypothetical example, a reference describes same-day disclosure of a campaign error, a correction plan two days later and a credit for wasted spend. That provides behaviour and chronology to assess using OPM's framework. 'They were responsive' needs a follow-up: 'Can you give one example?'

None of this is GPI's first-hand campaign experience. It is a questioning method, and its value depends on the caller holding the line until a specific incident is on the table.

Kara Ronin breaks down structured interviewing techniques to evaluate real-world problem-solving behavior, escalation habits, and team communication dynamics.

Questions that test the agency's performance claims and methodology

Bring the agency's case study for that client to the call, and ask the reference to confirm or correct each headline number and how it was measured. This is the block where a reference call does something no public record can: it puts a published result next to the client's own books.

Reconcile the case study with the client's own numbers

'The agency's case study says X. Does that match your internal numbers?' 'Did the growth show up in revenue, or only in platform-reported metrics?' 'Did you run any incrementality or holdout test, and who proposed it?' A reference who says the platform number was accurate and the revenue picture was more modest is giving you a precise and useful answer, and probably an honest one.

Ask how results were measured, and by whom

Ask whose attribution model produced the numbers, how disagreements with finance were resolved and whether the agency accepted methods that could reduce reported results. Attribution assigns credit; it does not establish causation. Apply that distinction to every case study.

Compare against other agencies the reference has used

'Compared with other agencies you have used for this work, where would you place them and why?' Most senior marketers have worked with several agencies. The comparison gives you a third-party ranking against real alternatives, and the 'why' tells you which dimension the reference weights most, which may differ from yours.

Adaptation: what changed when the platforms changed

Ask what changed after a platform or privacy update, whether tactics evolved and who proposed new ideas. This tests adaptability during the engagement, which an older case study cannot show.

The rehire question

'Would you hire them again for the same scope?' LinkedIn's talent guidance treats this as the decisive question and advises listening for hesitation or a negative response as a prompt for further exploration. For an agency, the follow-up to any hesitant answer is 'What would need to be different?' The conditions a reference attaches are often more informative than the yes.

Questions that test the agency's performance claims and methodology
QuestionClaim it testsEvidence a credible reference can supplyGap that should lower confidence
Does the case study match your internal numbers?Headline resultConfirms the figure and names the internal source it matchedHas not seen the case study, or the figure is unfamiliar
Did growth show up in revenue or only in platform metrics?Business impact versus reported efficiencyDescribes what finance saw over the same periodOnly recalls platform dashboards
Did you run an incrementality or holdout test?Causal support for the claimDescribes the test design and who proposed itNo test, and the agency never suggested one
Whose attribution model produced the numbers?Measurement ownershipNames the model and who controlled its settingsDoes not know, or the agency reported from its own tool only
How were disagreements between agency and finance numbers resolved?Reporting integrityDescribes a reconciliation process and outcomeDisagreements were left unresolved or unmentioned
Where would you rank them against other agencies?Relative performancePlaces them with a reason tied to a specific dimensionHas no comparison point
What changed after a platform or privacy shift?AdaptationNames a concrete change and its timingCannot recall any change in approach
Would you hire them again for the same scope?Overall endorsementUnconditional yes, or a conditional yes with named conditionsDeflection, or a yes with heavy qualification

The rehire row follows LinkedIn's guidance on hesitation. As with GPI's category-relevance test, another client's results cannot validate yours. The captured evidence provides no agency ROI benchmark.

Signal diagram mapping rehire answers from unconditional yes to deflection, with recommended follow-ups.
The million-dollar rehire question often reveals more through hesitation or conditions than a simple yes or no.Sources: www.linkedin.com · linkedin.com

Scoring the calls: a template for comparing agencies

Build the scorecard before the first call, score each reference within an hour of hanging up, and review all scorecards side by side before the final agency meeting. The scorecard below is a framework this article proposes for buying committees. It has not been validated as an instrument, and its value comes from consistency of use rather than from the specific scale.

The scorecard fields

Scoring the calls: a template for comparing agencies
Scorecard fieldWhat to recordSource questionScoring rule
Reference role and tenureTitle, dates on account, parts of the work seen directlyCall openerNote only; used to weight all other fields
Category matchCategory, business model, spend band, engagement typeReference request stageYes, partial or no; partial or no reduces weight of results fields
Staffing continuityNames, tenures, handoffs, disclosure timingStaffing block1 to 5; client-discovered churn caps the score at 2
Scope and budget behaviourSpecific scope, fee and pacing incidents and how foundScope block1 to 5; any invoice-discovered change caps the score at 2
Problem resolutionWorst-month incident, escalation path, elapsed time to fixProblem block1 to 5; no recallable incident scores as missing, not as 5
Claims credibilityCase study figures confirmed, corrected or unknown; measurement ownerPerformance block1 to 5; provisional if only agency-selected references reached
Rehire answerYes, conditional or no, plus the stated conditionsClosing questionRecord verbatim; never averaged across references
Verbatim notesExact phrases with names and datesAll blocksKept separate from scores
Follow-up requiredQuestions to put to the agency, with the reference's answer that prompted themAll blocksOpen or closed before final meeting

The field design applies the shared-template rule from GoodHire's hiring guidance and the structure principle from the OPM guidance to a buying committee. Universities such as West Virginia State University publish standard question sets for their search committees for the same reason: consistent questions and consistent recording reduce the pull of the most charismatic reference.

Weighting the four areas for your decision

Weight the areas to the contract you are about to sign. A full-service retainer depends on staffing and scope behaviour, so weight those fields higher. A performance-fee arrangement depends on whether reported results survive reconciliation, so weight claims credibility higher. Write the weights down before the calls so they are not adjusted to fit a preferred outcome.

Reading disagreements between references

Two positive current clients and one critical former client is a pattern, and a useful one. Do not average it away. Take the former client's specific incidents to the agency as follow-up questions and score the agency's response to being challenged as evidence in its own right.

When to ask for one more reference

Keep rehire and claims-reconciliation scores provisional if based on one source or only agency-selected contacts. Request another contact before deciding. The completed scorecards should explain the ranking to a colleague who missed the calls.

Where reference checks sit in the GPI evaluation approach

GPI publishes the methodology behind its Growth Partner Confidence Score so that buyers can see how public agency evidence is weighed. A structured reference call extends the same standard, assessing claims against evidence, methodology and limitations, to information only past clients hold. On timing, the closest available evidence is old and comes from individual hiring: Aberdeen's 2015 research found organisations that moved reference checking earlier in the process were 19% more likely to retain first-year employees. That is a historical finding about employees, and it cannot establish anything about agency relationships today, but it is a reasonable argument for starting reference calls while proposals are still being refined rather than after the decision is emotionally made. On stakes, negligent-hiring cases such as the $60.65 million verdict in November 2024 concerned security personnel, not agencies, and serve only as an extreme illustration that skipped checks carry cost. For the public-evidence half of the assessment, GPI agency profiles such as Agency Jet on Growth Partner Index show the kind of documentation worth reconciling with your reference notes.

Frequently Asked Questions

How many references should we call per agency, and does a former client count more than a current one?

Start with three: two current clients and one former or friction-period client. Treat this as a practical starting point; the captured research does not establish a numerical threshold. A former client's account of problems can complement current endorsements, but weigh firsthand knowledge and relevance more heavily than current or former status. Add calls when accounts conflict.

What should we do when an agency only offers current, visibly happy clients as references?

Ask once more in writing, specifying that you want a client who went through a difficult period, in line with the challenging-period recommendation from executive reference checking. If the agency still declines, run the calls anyway, mark the rehire and claims fields provisional, and treat the refusal as a data point when comparing agencies.

Should we tell the agency what its references said, and how do we use a critical answer in negotiation?

Ask the agency about specific incidents while protecting the reference's identity and confidentiality. Invite its explanation. How the agency responds to being challenged on a scope or staffing incident is useful evidence, and a specific, documented concern is also a legitimate basis for tightening staffing or scope terms in the contract.

Can we reference-check an agency through people the agency did not nominate, and what are the risks of doing so?

You can, through your own network or through former clients visible in the agency's public case studies. Treat it as an option rather than a norm, keep the questions identical to the ones you ask nominated references, and weight the answers with the same caution you apply to any single source. Industry publications such as Greenbook treat reference checking and follow-up as a standard step in agency selection, and unsolicited contacts sit at the edge of that practice.

How do we handle a reference who is clearly reciting rehearsed talking points?

Redirect to the worst-month question and ask for dates and names. Rehearsed references can sustain praise but rarely sustain specificity. If the answers stay general, score the fields as missing rather than positive, and request an additional reference.

Should the reference call cover fees and pricing, or does that belong in the commercial negotiation?

Ask about fee behaviour, such as unplanned changes and surprise charges, because that is conduct a reference observed. Leave the level of fees to the commercial negotiation, where you can compare proposals directly; a reference's rate card from a different spend band tells you little about yours.