Agency reference check questions rarely get the preparation the pitch does. The U.S. Office of Personnel Management's guidance on reference checking notes that structure improves the validity of reference checks, and the logic holds when the candidate is an agency. Use the same written script for every shortlisted partner. This guide explains how to check agency references: which clients to call, what to ask about staffing, scope, communication and results, and how to compare their answers. Before drafting your question list, it helps to see how documented agency evidence is laid out on a GPI profile, for example ACE Agency on Growth Partner Index.
GPI's position is that an agency claim is only as strong as the evidence, methodology and limitations behind it, and that standard should not stop at the public record. A reference call is where a buyer gets to test the case study against the client's own numbers and to ask how a result was measured before deciding what it proves. It is also where attribution stories meet operational reality: who actually ran the account, what slipped, and how the agency behaved when a test failed. The purpose is not to collect endorsements. It is to answer a concrete decision, which is whether this team, working on your category and budget, is likely to behave the way its pitch promised.
Why curated agency references still leak useful signal
An agency chooses its references; you choose the questions to ask agency references. Even a satisfied client can describe account-lead changes, undelivered scope and discrepancies between reported results and finance records.
References are a behavioral interview, not a satisfaction survey
OPM's hiring guidance invokes behavioral consistency: past behaviour informs future expectations. Apply that cautiously to comparable agency accounts. 'Who ran your weekly call in month six?' yields a checkable fact; 'Were you happy?' yields an opinion.
The analogy with individual hiring has limits, and one figure from that world shows why unstructured vetting carries risk there: in Checkr's 2025 survey of hiring managers, 31% said they had personally interviewed a candidate using a fake identity. That is a statistic about job candidates and says nothing about agencies. It does illustrate what happens when a check rests on impression rather than verification.
Structure is what separates signal from politeness
For a marketing agency reference check, adapt GoodHire's hiring recommendation: use a shared question template. Ask every reference the same questions in the same order to make answers comparable.
This article is written for buyers of advertising and media services. The current search results for this keyword are dominated by templates for hiring individuals, such as Rutgers' sample reference questions, and that guidance does not transfer cleanly to a contract with a team. Employment-law and background-screening compliance questions are out of scope here.
Set up the reference call before you ask a single question
Send the agency a written reference request before any call is scheduled. Specify the category match you need, the spend band, at least one former or friction-period client, and the names and roles you expect to speak with. The quality of the call is mostly decided at this stage.
Ask for references that match your category and budget band
GPI's Amazon ads agency buyer guide distinguishes CPG, durables and apparel playbooks. Apply the same relevance test to references: compare category, business model and spend band before weighing an endorsement.
- Category
- Business model
- Spend band
- Engagement type
When two of the four miss, ask for a substitute before you accept the call.
Insist on at least one reference from a hard period
Crosschq recommends a challenging-period reference for executive hiring. Adapt that request to agencies: seek a former client or one whose account had a rough quarter. Record any refusal as a limit on the evidence available.
Agree the format: 30 minutes, one caller, one note-taker, same script
- Confirm the reference's role, account dates and firsthand involvement.
- Explain how notes will be used and who can see them; clarify confidentiality expectations.
- Explain that all references receive the same questions.
- Have the future relationship owner ask questions and a colleague take notes.
- Reserve five minutes for the rehire question and follow-up.
Decide what you will do with a lukewarm answer before you hear one
Write down, before the first call, which answers are disqualifiers and which trigger a follow-up question to the agency. For example:
- A reference who learned about their lead strategist's departure after the fact might trigger a follow-up.
- A fee increase the client discovered on an invoice might be a disqualifier.
Teams that skip this step tend to rationalise warnings once they have already pictured working with a particular agency.

Questions that reveal who actually works on your account
Ask every reference to name the people they worked with most and how long each of them stayed on the account, then reconcile those names against the roster and named roles in the agency's proposal. This block addresses a familiar complaint in agency relationships: the people who won the pitch were rarely seen again after signing.
Pitch team versus delivery team
Ask who pitched, who remained after 90 days, who ran weekly calls and how often the named strategist appeared. Names and approximate dates can be checked against the proposal. Follow up on 'the whole team was great' with a request for specific people.
Turnover and handoffs during the engagement
'How many account leads did you have over the engagement?' 'How were handoffs handled?' 'Did you learn about a staffing change before or after it happened?' The third question is the sharpest, because it tests whether the agency manages its own churn as a client-facing event or as an internal matter the client discovers later.
Seniority and decision rights of the day-to-day lead
'When you needed a strategic call made, who made it and how long did it take?' 'Could your day-to-day lead reallocate budget or change a test plan without escalating?' These questions test whether the seniority in the proposal matched the seniority on the account.
| Question | What you are testing | Strong answer pattern | Warning signal |
|---|---|---|---|
| Who presented in the pitch, and how many were still on your account 90 days later? | Pitch-to-delivery continuity | Names two or three people from the pitch who stayed in named roles | Cannot recall the pitch team, or says none of them stayed |
| Who ran your weekly call? | Identity and stability of the day-to-day lead | One named person for most of the engagement, or a planned handoff explained in advance | Several names, uncertain order, or the client ran the call themselves |
| How many account leads did you have? | Turnover rate on the account | One or two, with reasons for any change | Three or more in a year, or the reference has to count |
| Did you learn about staffing changes before or after they happened? | Whether churn is managed as a client event | Advance notice with a named successor and overlap period | Found out from an out-of-office reply or a new face on a call |
| How were handoffs handled? | Knowledge transfer discipline | Written account history, joint calls, no repeated briefing | Client had to re-explain the business to the new lead |
| Who made strategic calls, and how long did it take? | Decision rights of the account team | Day-to-day lead had authority, or escalation took days rather than weeks | Every decision went to someone the client never met |
| How often did you see the named strategist? | Whether senior time was real or decorative | Regular cadence, present at planning and quarterly reviews | Appeared at the pitch and the renewal |
| Did the team change when your spend changed? | Whether staffing tracks revenue rather than need | Team stayed stable or changes were discussed openly | Senior people disappeared after a budget cut without comment |
The question structure follows the OPM guidance on structured reference checking and the Crosschq recommendation to include a reference from a challenging period. The answer patterns are the article's proposed interpretation, not measured outcomes.
What a strong answer sounds like and what a warning sounds like
Weight names and handoff details above general satisfaction. A client who experienced a strategist's departure can explain who filled the gap, how long it lasted and whether the replacement arrived briefed. Such firsthand detail is more useful than an endorsement when judging continuity.
One caution on interpretation. Turnover in an agency is normal, and a reference who reports one planned handoff with a proper overlap is describing competent management. The warning is churn the client found out about on their own. Where staffing changes caused scope or reporting problems, note them and carry them into the next block.
Compare what a reference tells you about staffing against the publicly documented information on a profile such as DEPT on Growth Partner Index, and note where the two accounts diverge.

Questions that expose scope creep, budget variance and missed SLAs
For each agency, record at least one specific scope incident, one budget incident and one SLA incident described by a reference, including how the client found out and how the agency responded. This block tests commercial behaviour, and no HR reference template asks about it, because an employee does not carry a statement of work.
Scope: what was promised, what was delivered, what you ended up doing yourself
Ask what the original statement of work promised but never delivered, and what the client had to produce internally. Probe whether extra scope repaired an omission in the original plan. A concrete list, such as landing pages or pixel implementation, identifies workload your team could inherit.
Budget: retainer drift, media fee changes and unbilled overages
Ask who initiated fee changes, which charges surprised the client and how spend variances were explained. Variance alone proves little. A client learning about it only from an invoice reveals a communication problem worth investigating.
SLAs and timing: launches, reporting deadlines and turnaround on changes
'How often were launch dates or reporting deadlines missed, and how were you told?' 'How long did a creative or targeting change typically take to go live?' 'When a deadline slipped, did the agency propose a new date or did you have to chase?' These probe whether the agency treats its commitments as commitments or as targets.
Boundaries: what the agency owned versus what your team had to carry
'Who was accountable when a dependency slipped, such as creative assets, tracking implementation or approvals?' 'Was the division of responsibilities written down, and did it hold?' Question lists written for individual hiring, such as The Hartford's, are built around one person's conduct. The agency version has to ask about the seam between two organisations, because that seam is where most delivery failures live.
| Question | Operational risk it tests | Answer that reassures | Answer that requires follow-up |
|---|---|---|---|
| What in the original statement of work never shipped? | Scope shortfall | Nothing material, or a named item dropped by mutual agreement | Items the client still expected and had to raise |
| What did you produce internally that you expected the agency to handle? | Hidden internal workload | Little, or items agreed as client-owned at kickoff | A list of deliverables the client absorbed without a fee adjustment |
| Did the fee change, and who initiated it? | Retainer drift | Change proposed in writing with a rationale before it took effect | Client noticed the change on an invoice |
| Were there charges you did not anticipate? | Unbilled or surprise overages | None, or one incident explained and credited | Recurring surprises or disputes over what was included |
| How did actual spend compare with plan month to month? | Pacing discipline | Small variances flagged by the agency in reporting | Large variances the client found in platform data |
| How were missed deadlines communicated? | SLA behaviour | Advance warning with a revised date | Silence until the client asked |
| How long did a change take to go live? | Turnaround responsiveness | Consistent turnaround the client could plan around | Unpredictable, or dependent on chasing a specific person |
| Who was accountable when a dependency slipped? | Boundary clarity | Written responsibilities that held under pressure | Each side assumed the other owned it |
Asking these questions in the same order of every reference follows the structure principle in the OPM reference checking guidance and the shared-template rule in GoodHire's guidance. No benchmark for typical scope creep or budget overrun is offered here, because none exists in the evidence gathered for this article.
Structure matters most in this block. If one agency's references describe two unplanned fee changes and another's describe none, that comparison only means something because both sets were asked the identical question.

Questions about how problems got solved and how communication felt
Start this part of the agency reference call with the worst month. Ask for an incident and response. If the reference cannot recall one, establish how closely they worked with the agency before weighting their answers.
Ask for the worst month, not the best campaign
Ask what caused the worst month, who noticed first and what happened in the next two weeks. Dates and names make follow-up possible. Crosschq's challenging-period recommendation helps identify someone able to answer.
Escalation: who you called and what happened next
'When something went wrong, who did you escalate to, and did that person have the authority to fix it?' 'Was there ever a problem you had to raise twice?' 'How long did it take to get a named owner for the fix?' Reconstruct the escalation path from the answer:
- Who noticed the problem
- Who was called
- Who had authority to act
- How much time elapsed
That structure can be compared across agencies in a way an anecdote cannot.
Ownership: did the agency say 'we got it wrong' unprompted?
Self-reported mistakes are the strongest evidence of accountability a reference can supply, because they cost the agency something and cannot be staged for a reference call. Ask: 'Did the agency ever flag its own mistake before you did?' 'How did they handle a campaign or test that failed?' 'Did the post-mortem name a cause, or did it name market conditions?'
Communication cadence and friction
'What was the meeting cadence, and did it hold?' 'Were reporting calls analysis or slide reading?' 'Did tough news arrive by phone or by email?' 'How quickly did you get a substantive reply, as opposed to an acknowledgement?' Communication complaints are common in agency relationships and easy for a reference to soften, so ask for specifics rather than ratings.
In a hypothetical example, a reference describes same-day disclosure of a campaign error, a correction plan two days later and a credit for wasted spend. That provides behaviour and chronology to assess using OPM's framework. 'They were responsive' needs a follow-up: 'Can you give one example?'
None of this is GPI's first-hand campaign experience. It is a questioning method, and its value depends on the caller holding the line until a specific incident is on the table.
Questions that test the agency's performance claims and methodology
Bring the agency's case study for that client to the call, and ask the reference to confirm or correct each headline number and how it was measured. This is the block where a reference call does something no public record can: it puts a published result next to the client's own books.
Reconcile the case study with the client's own numbers
'The agency's case study says X. Does that match your internal numbers?' 'Did the growth show up in revenue, or only in platform-reported metrics?' 'Did you run any incrementality or holdout test, and who proposed it?' A reference who says the platform number was accurate and the revenue picture was more modest is giving you a precise and useful answer, and probably an honest one.
Ask how results were measured, and by whom
Ask whose attribution model produced the numbers, how disagreements with finance were resolved and whether the agency accepted methods that could reduce reported results. Attribution assigns credit; it does not establish causation. Apply that distinction to every case study.
Compare against other agencies the reference has used
'Compared with other agencies you have used for this work, where would you place them and why?' Most senior marketers have worked with several agencies. The comparison gives you a third-party ranking against real alternatives, and the 'why' tells you which dimension the reference weights most, which may differ from yours.
Adaptation: what changed when the platforms changed
Ask what changed after a platform or privacy update, whether tactics evolved and who proposed new ideas. This tests adaptability during the engagement, which an older case study cannot show.
The rehire question
'Would you hire them again for the same scope?' LinkedIn's talent guidance treats this as the decisive question and advises listening for hesitation or a negative response as a prompt for further exploration. For an agency, the follow-up to any hesitant answer is 'What would need to be different?' The conditions a reference attaches are often more informative than the yes.
| Question | Claim it tests | Evidence a credible reference can supply | Gap that should lower confidence |
|---|---|---|---|
| Does the case study match your internal numbers? | Headline result | Confirms the figure and names the internal source it matched | Has not seen the case study, or the figure is unfamiliar |
| Did growth show up in revenue or only in platform metrics? | Business impact versus reported efficiency | Describes what finance saw over the same period | Only recalls platform dashboards |
| Did you run an incrementality or holdout test? | Causal support for the claim | Describes the test design and who proposed it | No test, and the agency never suggested one |
| Whose attribution model produced the numbers? | Measurement ownership | Names the model and who controlled its settings | Does not know, or the agency reported from its own tool only |
| How were disagreements between agency and finance numbers resolved? | Reporting integrity | Describes a reconciliation process and outcome | Disagreements were left unresolved or unmentioned |
| Where would you rank them against other agencies? | Relative performance | Places them with a reason tied to a specific dimension | Has no comparison point |
| What changed after a platform or privacy shift? | Adaptation | Names a concrete change and its timing | Cannot recall any change in approach |
| Would you hire them again for the same scope? | Overall endorsement | Unconditional yes, or a conditional yes with named conditions | Deflection, or a yes with heavy qualification |
The rehire row follows LinkedIn's guidance on hesitation. As with GPI's category-relevance test, another client's results cannot validate yours. The captured evidence provides no agency ROI benchmark.

Scoring the calls: a template for comparing agencies
Build the scorecard before the first call, score each reference within an hour of hanging up, and review all scorecards side by side before the final agency meeting. The scorecard below is a framework this article proposes for buying committees. It has not been validated as an instrument, and its value comes from consistency of use rather than from the specific scale.
The scorecard fields
| Scorecard field | What to record | Source question | Scoring rule |
|---|---|---|---|
| Reference role and tenure | Title, dates on account, parts of the work seen directly | Call opener | Note only; used to weight all other fields |
| Category match | Category, business model, spend band, engagement type | Reference request stage | Yes, partial or no; partial or no reduces weight of results fields |
| Staffing continuity | Names, tenures, handoffs, disclosure timing | Staffing block | 1 to 5; client-discovered churn caps the score at 2 |
| Scope and budget behaviour | Specific scope, fee and pacing incidents and how found | Scope block | 1 to 5; any invoice-discovered change caps the score at 2 |
| Problem resolution | Worst-month incident, escalation path, elapsed time to fix | Problem block | 1 to 5; no recallable incident scores as missing, not as 5 |
| Claims credibility | Case study figures confirmed, corrected or unknown; measurement owner | Performance block | 1 to 5; provisional if only agency-selected references reached |
| Rehire answer | Yes, conditional or no, plus the stated conditions | Closing question | Record verbatim; never averaged across references |
| Verbatim notes | Exact phrases with names and dates | All blocks | Kept separate from scores |
| Follow-up required | Questions to put to the agency, with the reference's answer that prompted them | All blocks | Open or closed before final meeting |
The field design applies the shared-template rule from GoodHire's hiring guidance and the structure principle from the OPM guidance to a buying committee. Universities such as West Virginia State University publish standard question sets for their search committees for the same reason: consistent questions and consistent recording reduce the pull of the most charismatic reference.
Weighting the four areas for your decision
Weight the areas to the contract you are about to sign. A full-service retainer depends on staffing and scope behaviour, so weight those fields higher. A performance-fee arrangement depends on whether reported results survive reconciliation, so weight claims credibility higher. Write the weights down before the calls so they are not adjusted to fit a preferred outcome.
Reading disagreements between references
Two positive current clients and one critical former client is a pattern, and a useful one. Do not average it away. Take the former client's specific incidents to the agency as follow-up questions and score the agency's response to being challenged as evidence in its own right.
When to ask for one more reference
Keep rehire and claims-reconciliation scores provisional if based on one source or only agency-selected contacts. Request another contact before deciding. The completed scorecards should explain the ranking to a colleague who missed the calls.
Where reference checks sit in the GPI evaluation approach
GPI publishes the methodology behind its Growth Partner Confidence Score so that buyers can see how public agency evidence is weighed. A structured reference call extends the same standard, assessing claims against evidence, methodology and limitations, to information only past clients hold. On timing, the closest available evidence is old and comes from individual hiring: Aberdeen's 2015 research found organisations that moved reference checking earlier in the process were 19% more likely to retain first-year employees. That is a historical finding about employees, and it cannot establish anything about agency relationships today, but it is a reasonable argument for starting reference calls while proposals are still being refined rather than after the decision is emotionally made. On stakes, negligent-hiring cases such as the $60.65 million verdict in November 2024 concerned security personnel, not agencies, and serve only as an extreme illustration that skipped checks carry cost. For the public-evidence half of the assessment, GPI agency profiles such as Agency Jet on Growth Partner Index show the kind of documentation worth reconciling with your reference notes.
Frequently Asked Questions
How many references should we call per agency, and does a former client count more than a current one?
Start with three: two current clients and one former or friction-period client. Treat this as a practical starting point; the captured research does not establish a numerical threshold. A former client's account of problems can complement current endorsements, but weigh firsthand knowledge and relevance more heavily than current or former status. Add calls when accounts conflict.
What should we do when an agency only offers current, visibly happy clients as references?
Ask once more in writing, specifying that you want a client who went through a difficult period, in line with the challenging-period recommendation from executive reference checking. If the agency still declines, run the calls anyway, mark the rehire and claims fields provisional, and treat the refusal as a data point when comparing agencies.
Should we tell the agency what its references said, and how do we use a critical answer in negotiation?
Ask the agency about specific incidents while protecting the reference's identity and confidentiality. Invite its explanation. How the agency responds to being challenged on a scope or staffing incident is useful evidence, and a specific, documented concern is also a legitimate basis for tightening staffing or scope terms in the contract.
Can we reference-check an agency through people the agency did not nominate, and what are the risks of doing so?
You can, through your own network or through former clients visible in the agency's public case studies. Treat it as an option rather than a norm, keep the questions identical to the ones you ask nominated references, and weight the answers with the same caution you apply to any single source. Industry publications such as Greenbook treat reference checking and follow-up as a standard step in agency selection, and unsolicited contacts sit at the edge of that practice.
How do we handle a reference who is clearly reciting rehearsed talking points?
Redirect to the worst-month question and ask for dates and names. Rehearsed references can sustain praise but rarely sustain specificity. If the answers stay general, score the fields as missing rather than positive, and request an additional reference.
Should the reference call cover fees and pricing, or does that belong in the commercial negotiation?
Ask about fee behaviour, such as unplanned changes and surprise charges, because that is conduct a reference observed. Leave the level of fees to the commercial negotiation, where you can compare proposals directly; a reference's rate card from a different spend band tells you little about yours.

