Before choosing a performance marketing agency, ask what event makes its invoice grow—and what evidence shows that event represents business you would not otherwise have won. GPI’s recommendation is to evaluate the payment rule alongside the measurement plan. A fee tied to reported conversions is not necessarily a fee tied to additional conversions caused by advertising. [Source: arxiv.org]

Use this guide at the finalist stage to compare commercial terms and substantiate measurement promises. For staffing, channel fit and general shortlisting, start with the broader paid media agency selection guide. The decision here is narrower: whether to approve a measurement pilot, require remediation or reject an unsupported claim of incremental business. [Source: growthpartnerindex.com]

Audit the fee trigger before choosing the agency

Start by asking each finalist to translate its proposal into a payment rule: what remains fixed, what increases with media spending and what increases with outcomes? For any outcome component, request the exact conversion definition and supporting evidence. GPI’s inference is that pricing labels alone do not align incentives; the distinction between attributed and causal outcomes still matters when the contract says performance-based. [Source: arxiv.org]

Attribution assigns credit under a rule; incrementality asks what changed because the advertising ran. The April 1, 2026 Predicted Incrementality by Experimentation manuscript explains that last-click counts can overstate advertising’s effect when customers would have purchased anyway, or understate it when advertising influences conversions outside the attribution window. The contractual question is which of those quantities earns the fee. [Source: arxiv.org]

Disclosure: Growth Partner Index is operated by Fieldtrip, a growth firm that operates and invests in marketing services brands. Fieldtrip-affiliated agencies may appear in the index and are labeled. Our methodology is designed to apply the same public-evidence criteria to all listings. Read the GPI ownership and disclosure policy for details. [Source: growthpartnerindex.com]

Normalize every proposal with the same buyer inputs

Build a comparison worksheet using your assumptions, not each agency’s forecast. Set a common currency, comparison period, campaign scope and media spend. Record each finalist’s fixed fee, spend-linked rate and fee per eligible outcome. Ask for any minimum, cap, tier or separate charge that requires an adjustment to the simple arithmetic below. [Source: growthpartnerindex.com]

GPI worksheet formula for an incremental-outcome proposal: total agency compensation = fixed fee + (media spend × spend-linked rate) + (fee-eligible incremental outcomes × outcome fee). This is proposed bid-normalization arithmetic, not observed market pricing, a forecast or a validated scoring benchmark.

Keep attributed conversions, measured incremental conversions and fee-eligible incremental outcomes in separate fields. The last field is the quantity payable under the contract, after applying its outcome definition, period and any threshold or cap. Enter zero only when a component genuinely does not apply. Leave missing terms unresolved rather than interpreting an omission as free service. [Source: arxiv.org]

If a proposal charges on attributed conversions, calculate its contractual fee exposure using your attributed-conversion assumption and label that basis explicitly. Do not silently substitute incremental outcomes to make the quote look comparable. Request revised terms if you want incremental-outcome compensation. Without causal support, retain the attributed count and leave measured incrementality unavailable. [Source: arxiv.org]

Google defines incremental cost per action as total spend divided by incremental conversions. GPI’s proposed comparison adds agency compensation to the numerator: fully loaded incremental CPA = (media spend + agency compensation) ÷ measured incremental conversions. Match the costs and conversions to the same campaigns, population and measurement period. [Source: support.google.com]

Here, fully loaded means media plus agency compensation, not every business cost or a profitability calculation. Document whether measurement and creative work are included in the compensation, and disclose excluded costs alongside the result. Apply the same cost boundary to every finalist rather than comparing an inclusive quote with a narrower service fee. [Source: growthpartnerindex.com]

Before a study, an incremental-conversion input is a buyer-supplied scenario, not a measured result. Run the same scenarios across proposals. If the input comes from a validated prediction model, label the resulting CPA modeled rather than directly measured. This preserves the difference between comparing contractual exposure and establishing campaign effectiveness. [Source: arxiv.org]

For this worksheet, mark incremental CPA not calculable when lift is zero, negative or inconclusive under the agreed decision rule. The PIE manuscript identifies problems with cost-per-incremental-conversion ratios when lift is nonpositive or very small. Do not display a negative CPA as a bargain or fill the denominator with attributed conversions. Separately specify which fixed, spend-linked and outcome fees remain payable. [Source: arxiv.org]

Classify the evidence behind the promised outcome

Attributed reporting counts conversions under a stated attribution rule and window. Ask the agency to document both. That count describes credited outcomes, but alone it does not establish additional business. A change of label from attributed to incremental needs an explanation of the causal evidence connecting the two. [Source: arxiv.org]

Experiment-calibrated prediction learns from previous controlled experiments and applies that relationship to other campaigns. Predicted Incrementality by Experimentation, or PIE, is an example. It is not a fresh experiment for every campaign. Check input timing too: PIE’s richer specifications use post-campaign information, so their reported accuracy should not be presented as the accuracy of a prelaunch forecast. [Source: arxiv.org]

A direct controlled experiment estimates the effect of specified advertising through a treatment-control comparison. Google describes Conversion Lift as comparing outcomes between groups that see the tested ads and groups that do not. Require the agency to state which campaigns and outcomes the comparison covers. GPI’s inference is that evidence of advertising lift does not establish that switching agencies caused an improvement; that requires a comparison designed to answer the agency-change question. [Source: support.google.com]

Give every finalist the same measurement exercise

Include a sanitized account brief in your request for proposal: proposed campaigns, commercial outcome, historical conversion volume, conversion lag, tracking status and available budget. Request a written measurement plan. The exercise below is GPI’s proposed diligence procedure, informed by Google’s account and study inputs—not a platform certification or an established scoring system. [Source: support.google.com]

Begin with the fee-eligible outcome. In a hypothetical lead-generation brief, completed forms might be a secondary diagnostic while accepted sales opportunities are the desired commercial outcome. Ask whether the proposed method can actually observe the latter. Do not let an easier-to-measure event replace the contractual outcome without explicit approval. [Source: support.google.com]

Require the counterfactual: what would have happened without the tested advertising. Ask how treatment and control groups will be constructed, what other activity remains running and how the design addresses bias. IAB and IAB Europe’s November 2025 introduction to their commerce-media guidelines identifies credible counterfactuals, bias control and separation of signal from noise as foundational principles. Those guidelines concern commerce media; this procurement exercise is GPI’s application of the stated principles. [Source: iab.com]

Request separate answers for the smallest effect the proposed study can reasonably detect and the smallest effect worth paying for. Ask for the assumptions, proposed uncertainty interval and pre-agreed decision threshold. Google advises choosing certainty according to business needs and risk tolerance. For procurement, document how the interval will affect payment and scaling, including a positive estimate that still fails the agreed evidence threshold. [Source: support.google.com]

Finish with operating rules. Who can change budgets, bids, audiences or creative? Who approves an extension? What happens to the invoice if the result remains inconclusive? Require a change log and distinguish routine adjustments from changes that alter the question being tested. Agree on these rules before either side knows the result. [Source: support.google.com]

Verify feasibility in the buyer’s account

Generic product availability is not proof of account eligibility. Google’s setup documentation, retrieved September 8, 2026, says Conversion Lift is unavailable to some Google Ads accounts and directs advertisers to their account representative. Request confirmation for your intended campaigns, including whether participation in another active lift study makes a campaign unavailable. [Source: support.google.com]

Inspect the selected conversion actions and tracking diagnostics. Google requires at least one compatible action and says conversions must fire unconditionally. Its user-based setup documentation lists store visits, store sales and offline conversion imports without personally identifiable information as unsupported. Google recommends resolving diagnostic gaps before launch, even though a study can be saved with warnings. [Source: support.google.com]

Ask the finalist to show the feasibility estimate and its inputs: account-level historical data, estimated lift, selected conversion actions, daily campaign budget, duration and holdback percentage. Holdback is the share withheld from the tested advertising. Google notes that larger holdouts increase the sample but carry greater opportunity cost, while smaller holdouts require longer studies to collect enough data. [Source: support.google.com]

Treat the US$5,000 headline carefully. Google’s setup passage specifies budgets above US$5,000 and 1,000 conversions for directional results below 90% confidence. That condition is not an agency fee, a universal access threshold or a precision guarantee. It does not replace account eligibility, tracking readiness or the feasibility estimate for your chosen outcome. [Source: support.google.com]

The timing matters too. Google’s February 2026 article describes a threshold reduction made in 2025, from a previous high of $100,000 to a $5,000 minimum. It is provider messaging about expanded access, not a 2026 launch announcement. Use the account-specific setup requirements—not that headline—to assess the proposed study. [Source: business.google.com]

Tie duration and test governance to conversion lag

Use your average conversion lag—the average time between an impression and a conversion—to justify study duration. Google allows studies as short as seven days but typically recommends more than 14 days, especially for longer conversion lags. Reject a standard-duration package that does not fit your conversion cycle. Unknown lag and volume are inputs to investigate, not reasons to promise a result on a fixed calendar. [Source: support.google.com]

This does not require an indiscriminate freeze on campaign management. Google permits changes but recommends caution and business-as-usual adjustments such as bids or budgets. It specifically warns that updating creative and audiences during a running study can make the findings difficult to interpret. Document permitted changes, approval ownership and how deviations will be reported. [Source: support.google.com]

If a separate team produces the ads, resolve creative-testing ownership alongside the media contract. Who supplies assets, interprets tests and authorizes replacements? Use the guide to evaluate a performance creative agency when those responsibilities sit outside the performance marketing agency’s scope. [Source: growthpartnerindex.com]

Use the chart to challenge attribution-only fees

The accompanying chart and table compare raw attribution with experiment-trained prediction in the April 1, 2026 PIE manuscript. The full model achieved an out-of-sample R² of 0.88 for incremental conversions per dollar, compared with 0.19 for raw seven-day last-click conversions per dollar treated as incrementality. The comparison was weighted by campaign cost. [Source: arxiv.org]

Out-of-sample R² measures predictive performance on observations withheld from model training, relative to a baseline that predicts the study’s mean effect. Higher is better; zero means no improvement over that baseline, and negative values are possible. These values are not percentages, causal lift, revenue gains, agency-quality scores or the share of conversions that were incremental. [Source: arxiv.org]

The population is historical and selective: 2,226 experiment–conversion-event pairs from 998 treatment-control pairs, selected from large-scale tests run by U.S. Meta advertisers between November 1, 2019 and March 1, 2020. Selected experiments had at least one million test-group users and 5,000 test-group conversions. Multiple outcomes could come from the same treatment-control pair; these are not 2,226 advertisers or agencies. [Source: arxiv.org]

The chart’s PIE(Pre) label follows the manuscript’s own input classification. Its data section includes actual campaign spending among those inputs, so do not interpret that specification as a purely prelaunch forecast. More broadly, the chart compares measurement predictions in this dataset, not the forecast accuracy available to a prospective agency before your campaign runs. [Source: arxiv.org]

The procurement lesson is not to discard attribution data. Last-click information contributed strongly when PIE learned its relationship with experimental effects. Challenge the unsupported substitution of attributed outcomes for incremental outcomes—not the use of attribution as an input to a model whose predictions have been validated. [Source: arxiv.org]

Table reading guide: prediction accuracy in historical Meta tests. Each row identifies a measurement method; the numeric column reports out-of-sample R². Higher values mean better prediction in this study. If the table extends beyond your screen, scroll horizontally to read both columns. [Source: arxiv.org]

Predicting incremental conversions per dollar in historical Meta tests (out-of-sample R²). Raw 7-day last click: 0.19; PIE(Pre), source classification: 0.35; PIE plus test-group outcome: 0.46; PIE plus 7-day last click: 0.85; PIE with all listed features: 0.88. Full context and source appear in the matching data table.
Predicting incremental conversions per dollar in historical Meta tests. The April 1, 2026 manuscript compares cost-weighted prediction of incremental conversions per dollar. The population contains 2,226 experiment–conversion-event pairs from 998 treatment-control pairs selected from large-scale tests run by U.S. Meta advertisers from November 1, 2019 to March 1, 2020; selected experiments had at least one million test-group users and 5,000 test-group conversions. PIE(Pre) follows the manuscript’s classification, but its data section includes actual campaign spending, so do not interpret that row as a purely prelaunch forecast. Higher R² indicates better prediction relative to the study’s mean-effect baseline. Values are not percentages, causal lift, agency failure rates or current market benchmarks, and the authors caution against transfer to other settings without validation.arxiv.org
Chart data
GroupValue
Raw 7-day last click0.19
PIE(Pre), source classification0.35
PIE plus test-group outcome0.46
PIE plus 7-day last click0.85
PIE with all listed features0.88
The April 1, 2026 manuscript compares cost-weighted prediction of incremental conversions per dollar. The population contains 2,226 experiment–conversion-event pairs from 998 treatment-control pairs selected from large-scale tests run by U.S. Meta advertisers from November 1, 2019 to March 1, 2020; selected experiments had at least one million test-group users and 5,000 test-group conversions. PIE(Pre) follows the manuscript’s classification, but its data section includes actual campaign spending, so do not interpret that row as a purely prelaunch forecast. Higher R² indicates better prediction relative to the study’s mean-effect baseline. Values are not percentages, causal lift, agency failure rates or current market benchmarks, and the authors caution against transfer to other settings without validation. arxiv.org

Accept modeled incrementality only with validation evidence

PIE is a counterexample to rejecting all modeled measurement. It also carries a clear limitation: the authors caution against transferring their empirical results to other platforms or prediction settings without validation. The approach depends on representative training experiments and stable relationships between inputs and outcomes. [Source: arxiv.org]

Ask an agency proposing modeled incrementality to identify its training experiments, their relevance to your campaigns, validation on unseen observations, input timing and known failure conditions. Request evidence of prediction errors, not just an accuracy headline. Require the agency to explain when it would stop relying on the model and seek new experimental evidence. [Source: arxiv.org]

PIE Figure 7 compares out-of-sample prediction for full sample, repeat experiments, existing advertisers and unseen advertisers.
Figure 7 compares prediction of incremental conversions per dollar across four validation setups. For procurement, focus on the test using a separate group of unseen advertisers when asking whether a model transfers to a new account. These historical Meta results do not validate an agency or your account. Error bars show one standard deviation of the estimated R².Gordon, Moakler and Zettelmeyer · PIE manuscript, Figure 7 · arxiv.org

Keep the research’s own limitations visible. The authors acknowledge that pre-aggregated data prevented within-campaign sample splitting to investigate shared sampling variation between some inputs and outcomes; they argue that the large samples limit the concern. The manuscript includes a Meta-affiliated coauthor. It is a research proof of concept, not independent validation of a finalist agency. [Source: arxiv.org]

Commercial product claims belong in a separate category. In January 2026, Meta reported that its Q4 2025 incremental-attribution model rollout drove a 24% increase in incremental conversions compared with its standard attribution model. This provider claim concerns a different metric and population from the historical PIE comparison. It is neither independent agency evidence nor a benchmark to insert into your fee worksheet. [Source: about.fb.com]

Approve, remediate or reject the measurement promise

Use GPI’s proposed procurement classifications to make the decision. Pilot-ready means the outcome, fee trigger, counterfactual, account feasibility, duration and decision rule are documented. Approve a bounded pilot with agreed spending and authority—not an assumption of successful lift. [Source: support.google.com]

Remediation-required means the design is credible but tracking, compatible outcomes or access remains unresolved. Request a scoped remediation plan with an owner and acceptance criteria. The ability to save a study, or the agency’s familiarity with the product, does not establish readiness to support the compensation promise. [Source: support.google.com]

Unsupported means an incremental-business compensation claim rests only on attributed conversions, a product label or a case-study narrative with no assessable counterfactual. Reject that claim or request different commercial terms. This need not mean rejecting every execution service the agency offers; it means refusing to treat unsupported attribution as evidence of additional business. [Source: iab.com]

An inconclusive study is a possible result, not automatic agency failure. Agree beforehand whether it leads to another study, a revised scope or no outcome-based payment, and what costs remain due. Hold the agency accountable for the promised design and execution without allowing uncertainty to become an unsupported success claim. [Source: support.google.com]

If the mandate extends into lifecycle, analytics and coordination across the funnel, use the separate guide to choose a cross-channel growth marketing agency. For this contract, finish with a simpler test: can both sides explain what earns the fee, what evidence supports it and what happens when that evidence is inconclusive? [Source: growthpartnerindex.com]