Most agency reports answer a question you did not ask. It tells you how many conversions occurred near an ad. The question a budget owner actually has is different: how many of those conversions would have happened without the spend? Marketing incrementality testing is the only measurement method built to answer that, and it has moved into the mainstream. EMARKETER and TransUnion data from July 2025 put adoption at 52.0% of US brand and agency marketers, so asking an agency to support holdout testing is no longer an unusual request. The harder problem is that a lift number is only as trustworthy as the design behind it, and most contracts leave that design to the agency's discretion. This guide sets out what to require, field by field, before any lift figure reaches your dashboard.

An agency saying it runs incrementality tells you very little on its own. You need to know who assigned the control group, how large the holdout was, what success threshold was written down before launch and how wide the interval around the lift estimate is. Our view is that measurement exists to settle a specific budget decision, so the buyer should name the decision first and specify the test second. Attribution describes what happened near the ad; it cannot show what the ad caused. GPI applies one standard to every agency outcome claim: it is only as strong as the evidence, the method and the stated limitations behind it. Put those into the contract and the reporting template, and agency reviews move from persuasion to verification.

TL;DR

  • Platform attribution reports correlation. Only a controlled comparison between an exposed group and an unexposed baseline separates conversions the agency caused from conversions that would have arrived anyway, a distinction that vendor guides and analytics providers both treat as the definition of incrementality.
  • Holdout testing is a mainstream practice, not an enterprise-only one, and an agency that says your budget is too small should be asked to show its reasoning.
  • Require disclosure of test design, holdout share, runtime, confidence level and every input to the incremental ROAS calculation before a lift number appears in a report.
  • Match the design (geo holdout, audience split, synthetic control) to the channel and the decision using a stated level of causal rigor, following the framing in the IAB commerce media guidelines, rather than accepting the platform's native default.
  • Treat self-reported lift as a claim to be evaluated against its methodology and limitations, and write the decision rules for zero or negative results before launch.
  • Independence matters. In a January 2026 survey commissioned by testing vendor Haus, 60% of US senior decision-makers said they trusted independent incrementality testing most among measurement methods.

Incrementality Testing vs A/B Testing: The Distinction Your Agency Must Be Able to Explain

Ask your agency to classify every metric on last month's report as attribution, A/B or holdout-based. If the team cannot do it quickly, the report is probably mixing three kinds of evidence with different strengths. The distinction below is the one an agency should be able to explain without notes.

What an A/B test can and cannot tell you

An A/B test compares two treatments against each other. Both arms are exposed to advertising; the test simply varies the creative, the bid strategy or the landing page. That makes it a good tool for optimizing within a channel. It says nothing about whether the channel deserves the budget at all, because there is no group that saw nothing. An agency reporting a winning variant has learned which ad worked better, not whether either ad produced sales that would not otherwise have happened.

What a holdout test adds: the unexposed baseline

An incrementality test is a controlled experiment that compares people exposed to marketing with a comparable group who were not, and reads the difference in outcomes as the causal effect of the marketing. The unexposed group is the baseline: it shows what the brand would have earned anyway from existing demand, organic search, word of mouth and habit. Matomo describes incrementality as the growth attributable to marketing beyond the overall momentum of the brand. That baseline is the missing ingredient in every last-click or platform-reported ROAS figure.

Why platform attribution belongs in neither category

Attribution assigns credit to touchpoints. It has no control group and tests nothing; it observes and allocates, and its coverage gaps rarely reach the report. For illustration, AppsFlyer's late 2020 and early 2021 analysis found Apple's SKAdNetwork captured about 68% of non-organic installs, so roughly a third of paid activity was invisible to it. The figure applies to iOS app installs under 2021 conditions, not to web or present-day benchmarks, but it documents coverage gaps as a measured limitation. AppsFlyer's guide to incrementality testing positioned holdout experiments as the response.

Incrementality Testing vs A/B Testing: The Distinction Your Agency Must Be Able to Explain
MethodQuestion it answersControl groupWhat it can proveHow agencies commonly misuse it
A/B testWhich of two exposed treatments performs betterNone unexposed; both arms see adsRelative performance within a channelPresented as evidence the channel itself works
Incrementality holdoutWhat the spend caused above baselineUnexposed or geo-excluded groupCausal lift and incremental returnRun without a pre-set threshold, then reinterpreted
Platform attributionWhich touchpoints were near a conversionNoneCorrelation and coverage of tracked pathsReported as return on spend without stating the baseline

The causal definition in the table follows Cometly's description of the exposed-versus-unexposed comparison; the attribution coverage limitation draws on the 2021 AppsFlyer SKAdNetwork data cited above.

A single screening question covers most agency conversations: which of your reported results come from a test with an unexposed control, and which come from attribution?

A three-column comparison table titled 'A/B Testing, Holdouts and Attribution Answer Different Questions'. Columns are 'Method', 'Question it answers', and 'Control group'. Row 1: A/B test / Which of two exposed treatments performs better / None unexposed; both arms see ads. Row 2: Incrementality holdout / What the spend caused above baseline / Unexposed or geo-excluded group. Row 3: Platform attribution / Which touchpoints were near a conversion / None.
Only a controlled holdout isolates conversions caused by advertising from organic baseline momentum.Sources: www.cometly.com, matomo.org · cometly.com

How to Calculate Incremental ROAS and Which Inputs the Agency Must Show You

Add five disclosure inputs to your agency reporting template so that your own analyst can recompute every incremental ROAS figure from the delivered data. The arithmetic is simple; the auditability depends entirely on what the agency is required to show.

The lift and incremental ROAS formulas

The causal logic is the exposed-versus-unexposed comparison already established. Turning it into numbers takes three steps:

  1. Estimate the baseline. Take the control group's conversion rate (or revenue per user, or per region) and apply it to the size of the exposed group. That gives the conversions the exposed group would have produced with no advertising.
  2. Compute incremental conversions. Subtract the scaled baseline from the exposed group's actual conversions. Multiply by average order value, or use incremental revenue directly, to get incremental revenue.
  3. Compute incremental ROAS. Divide incremental revenue by the spend that generated it during the same window.

Lift is usually expressed as incremental conversions divided by baseline conversions. Incremental ROAS is what a budget decision should rest on.

A worked example with hypothetical numbers

These figures are hypothetical and show only the logic. A geo test's exposed markets generate 12,000 orders; comparable control markets, scaled to the same population, imply a 10,000 baseline. Incremental orders are 2,000. At a 60-dollar average order value, incremental revenue is 120,000 dollars. On 40,000 dollars of spend, incremental ROAS is 3.0. The platform might report 400,000 dollars attributed, a ROAS of 10, because it credits every tracked conversion, including the 10,000 baseline orders.

Cometly's guide uses a similarly illustrative contrast, one channel at 3x incremental ROAS and another at 1.2x, to make a point about allocation: the lower channel is not cut outright; budget shifts toward the higher-lift channel until its incremental returns diminish. Those figures are illustrative in the source as well, not measured results.

Inputs that change the answer: baseline, spend window and conversion definition

Platform ROAS and incremental ROAS diverge because attribution counts every conversion it can see while incrementality counts only growth beyond the brand's baseline. Small changes to the inputs move the incremental figure a long way. A control group that is not truly comparable inflates or deflates the baseline. A measurement window that ends before delayed conversions land understates lift; a window that runs past the campaign picks up unrelated demand. Changing the conversion event from purchase to add-to-cart after the fact can turn a null result into a positive one. A wide confidence interval means the point estimate is too imprecise to act on: a reported lift of 20% with an interval spanning minus 5% to 45% supports no budget decision at all.

How to Calculate Incremental ROAS and Which Inputs the Agency Must Show You
MetricFormulaInputs the agency must discloseCommon reporting failure
Baseline conversionsControl conversion rate times exposed populationControl group definition, holdout share, matching methodControl chosen after results were visible
Incremental conversionsExposed conversions minus baselineMeasurement window, conversion event definitionWindow or event changed post hoc
LiftIncremental conversions divided by baselineConfidence interval, statistical powerPoint estimate shown without interval
Incremental ROASIncremental revenue divided by spend in windowSpend attributed to exposed group only, revenue sourceSpend window misaligned with conversion window

The causal structure behind these formulas follows the exposed-versus-unexposed comparison described by Cometly; the baseline concept follows Matomo. The five inputs to require in every report are the control group definition, the holdout share, the measurement window, the conversion event definition and the confidence interval around the lift estimate.

Test Designs to Require: Geo Holdouts, Audience Splits and Synthetic Controls

For each major channel, record three things: which test design the agency proposes, who owns the randomization and the raw output, and what decision the result will trigger. The three design families below differ in who controls the data and how strong a causal claim they support, and an agency should be able to argue for its choice in those terms.

Geo holdout tests

A geo lift test splits markets into test and control regions. Advertising runs in test regions and is withheld or reduced in control regions; the difference in outcomes is the lift. Because it uses only regional sales and spend, it suits television, audio, out-of-home and digital channels without user-level data. Requirements: enough comparable regions for a fair control, a runtime sufficient for effects to appear, and no promotions in only one set of regions. The main weakness is contamination: control regions see the advertising through travel, shared media markets or online spillover.

Audience-level holdouts and platform lift studies

Audience splits randomize at the user level, typically inside an ad platform. Assignment is cleaner than geography allows, and results arrive faster because the sample is large. The tradeoff is ownership: the test is run, measured and reported by the party selling the media. That does not make the result wrong, but it does make it a claim that needs external corroboration, ideally a geo or independent design on the same channel at some point in the year. Require the agency to disclose when a reported lift comes from a platform's own study.

Synthetic control and pre/post designs

When a randomized holdout is impossible, because the channel cannot be switched off regionally or the brand refuses to withhold spend, agencies fall back on synthetic controls, matched markets or pre/post comparisons. A synthetic control builds an artificial baseline from a weighted combination of untreated markets; a pre/post design compares the period before a change with the period after. These are weaker causal claims. They can be informative when labeled honestly and useless when presented as equivalent to a randomized test. The requirement is the label.

Choosing the level of causal rigor for the decision at hand

The IAB's commerce media guidelines frame incremental measurement as a choice of causal rigor matched to the goal, distinguishing campaign optimization, ROI validation and platform calibration as objectives that warrant different methods. That framing is useful for buyers because it turns an abstract argument about rigor into a budget question. A decision about which creative to scale can rest on a platform lift study. A decision to move 30% of annual spend between channels should rest on a randomized geo design, or at minimum on a platform result corroborated by one.

Access has widened. Digital Applied's 2026 review argues that open-source statistical tooling and falling minimum budgets have moved rigorous geo experiments beyond enterprise-only budgets, and it summarizes Google as reporting a cut in minimum test budgets from roughly 100,000 dollars to 5,000 dollars. Treat the budget figure as a vendor claim relayed by a secondary source rather than an audited fact. The practical implication still holds: an agency that says your spend is too small to test should be asked to show the calculation.

Test Designs to Require: Geo Holdouts, Audience Splits and Synthetic Controls
DesignRandomization unitChannels it suitsMinimum conditionsMain weaknessWho controls the data
Geo holdoutRegion or marketAny channel with regional spend control, including offlineEnough comparable regions, adequate runtime, no regional promotionsSpillover between regionsBrand or independent party can own it
Audience split or platform lift studyIndividual userDigital channels inside one platformPlatform support, sufficient audience sizeRun by the media sellerPlatform
Synthetic control or matched marketConstructed baselineChannels that cannot be withheld regionallyLong pre-period of stable dataWeaker causal claim, sensitive to model choicesAgency or analyst who builds the model
Pre/postTime periodLast resort for any channelStable demand across periodsConfounded by seasonality and other changesWhoever holds the outcome data

The causal-rigor framing follows the IAB guidelines; the access argument follows Digital Applied. Descriptions of each design are conceptual and carry no benchmark lift or runtime figures.

Three tradeoffs belong in the RFP regardless of design: the opportunity cost of conversions forgone in the holdout, the contamination risk between regions or audiences, and the overlap with seasonality or promotions that can swamp the signal.

A decision tree for selecting a test design. Root node: 'Does the channel allow user-level randomization?' If Yes, 'Audience split or platform study'. If No, 'Are there enough comparable regions?' If Yes, 'Geo holdout'. If No, 'Synthetic control or pre/post design'.
Match the test design to the channel and the decision using a stated level of causal rigor, following the IAB commerce media guidelines.Sources: www.iab.com · iab.com

The Agency Measurement Requirements Template, Field by Field

Copy the nine fields below into your next agency RFP or quarterly review and score the current agency against each red flag. Each field states what to require, what counts as acceptable evidence and what fails. The fields refer to the test designs by name; the detail on designs sits in the previous section, and guidance on reading the returned results follows in the next.

1. Test cadence and coverage

Require a stated number of channels tested per quarter and a stated share of media spend that has passed through a holdout in the trailing twelve months. Acceptable evidence is a test calendar and a spend-coverage figure the agency updates each quarter. The red flag is a cadence described as "as needed" or "when the client requests". Given that more than half of US brand and agency marketers reported running incrementality tests in mid-2025, an agency with no regular cadence is behind common practice.

2. Design and randomization ownership

Require disclosure, for every test, of who assigns test and control: the agency, the platform or an independent party. Acceptable evidence is a named owner and, where the platform owns assignment, a corroborating non-platform design on the same channel within the year. The red flag is a lift figure with no stated randomization owner. In the January 2026 Haus-commissioned survey of 500 US marketing and finance executives, 60% trusted independent incrementality testing most, 20 points ahead of media mix modeling at 40%. The survey came from a testing vendor with an interest in the result, so read it as a signal about what executives say they want, not as an audit of method quality.

3. Holdout share and minimum runtime

Require the agency to propose the holdout share and runtime before launch, with the reasoning attached. The right values depend on baseline conversion volume and variance, not on an industry percentage, so the acceptable evidence is a power calculation rather than a number. The red flag is a holdout share fixed by convenience, or a runtime cut short once the number looked favorable.

4. Pre-registered hypothesis and success threshold

Require the hypothesis, the primary KPI and the threshold that counts as success to be written and shared before the test launches. This is the single field that prevents an agency from choosing the favorable metric afterward, and it is the buyer's side of the causal-rigor framing in the IAB guidelines: a decision-grade test needs a decision-grade question. Acceptable evidence is a dated document. The red flag is a report whose headline metric differs from the one agreed.

5. Confidence level and power calculation

Require the confidence level to be stated in advance and the power calculation to be shared, showing the minimum detectable effect the design can find at the chosen holdout share and runtime. Acceptable evidence is the calculation itself. The red flag is a test that could not have detected a plausible effect and is then reported as showing no effect.

6. Raw data access and reproducibility

Require that your analyst can recompute lift from data delivered with the report: exposed and control outcomes, spend by group, window definitions and assignment lists or region lists. The exposed-versus-unexposed logic described by Cometly is only as sound as the data behind it. The red flag is a PDF summary with no underlying tables.

7. Reporting format: point estimate with interval

Require every lift and incremental ROAS figure to be shown alongside its confidence interval. Acceptable evidence is the interval printed next to the number, every time. The red flag is a bare point estimate, or an interval buried in an appendix.

8. Decision rules after the result

Require the budget change that follows a positive, null or negative result to be agreed before launch. Acceptable evidence is a written rule such as: a result below threshold at the agreed confidence level triggers a spend reduction of a stated size. The red flag is a null result reframed as "inconclusive" with a recommendation to keep spending.

9. Independence and conflict-of-interest disclosure

Require the agency to disclose which measurement tools it resells, is paid by or has partnership arrangements with, and whether any platform lift study is corroborated by a non-platform design. The red flag is an agency that recommends its own reseller product as the sole arbiter of its own performance.

The Agency Measurement Requirements Template, Field by Field
FieldWhat to requireAcceptable evidence from the agencyRed flag
Cadence and coverageTests per quarter, share of spend tested in 12 monthsTest calendar, coverage figureTesting only on request
Randomization ownershipNamed owner per testOwner listed, corroboration plan for platform testsNo owner stated
Holdout share and runtimeProposed before launch with reasoningPower calculationFixed by convenience
Pre-registrationHypothesis, KPI, threshold dated before launchSigned documentHeadline metric changed
Confidence and powerStated level and minimum detectable effectCalculation sharedUnderpowered test reported as null
Raw data accessRecomputable data deliveredTables with assignment listsSummary only
Reporting formatEstimate with intervalInterval beside every numberBare point estimate
Decision rulesAgreed action per outcomeWritten ruleNull reframed as inconclusive
IndependenceTool relationships disclosedDisclosure statementReseller product as sole judge

Adoption context in the table draws on EMARKETER and TransUnion data from July 2025; the independence field draws on the January 2026 trust survey with the vendor-commissioned caveat noted above.

For the broader standard GPI uses to assess agency outcome claims, see the Growth Partner Confidence Score methodology.

Bar chart showing stated trust by measurement method. Independent incrementality testing is at 60%, and media mix modeling is at 40%. Data from a January 2026 EMARKETER survey commissioned by Haus.
In a January 2026 survey of 500 US marketing and finance executives commissioned by Haus, independent incrementality testing ranked 20 points higher than media mix modeling.Sources: www.emarketer.com · emarketer.com

High-Stakes Fields: Holdout Size, Independence and How to Evaluate the Results You Get Back

Before approving any lift-based budget change, run the returned report through the diagnostic table below and reject any result missing an interval or a pre-registered threshold. Three of the nine fields, holdout share, runtime and independence, are the ones agencies most often ask to soften. Conceding on any of them makes every downstream number unreliable, which is why they deserve separate treatment.

Why holdout share and runtime are where tests quietly fail

A holdout too small for baseline conversion volume produces a control estimate so noisy that any lift sits inside the noise. A runtime too short catches the launch spike but misses delayed conversions, or mistakes one promotional week for the campaign. Neither looks like a failure; both produce a number. Holdout share depends on baseline volume and outcome variance; runtime depends on conversion lag and how many comparable periods separate signal from seasonal noise. No fixed percentage or week count fits every brand, so the power calculation belongs in the template. When an agency proposes shrinking the holdout to protect revenue, ask what minimum detectable effect the smaller design can find and whether that effect matters to the decision.

Reading a lift report: questions that expose weak designs

High-Stakes Fields: Holdout Size, Independence and How to Evaluate the Results You Get Back
What the agency reportsQuestion to askWhat a defensible answer containsFail condition
A lift percentageWhat is the interval around it, and at what confidence level?Interval printed, level stated before launchNo interval, or level chosen after results
A positive incremental ROASWas this the primary KPI pre-registered before launch?Dated pre-registration matching the reported metricMetric or window changed post hoc
A null or inconclusive resultWhat effect size could this design detect?Power calculation showing a plausible effect was detectableUnderpowered test presented as evidence of no effect
A platform lift study resultWho assigned control, and is there a non-platform corroboration?Owner named, corroborating geo or independent test scheduledPlatform-only evidence for a major budget decision
A recommendation to increase spendCan my analyst recompute the lift from delivered data?Raw exposed and control tables providedSummary only

The exposed-versus-unexposed logic behind each row follows the Cometly definition; the causal-rigor framing follows the IAB guidelines.

The habit to build is evaluating the method and the limitations rather than the headline. A modest lift with a tight interval, a pre-registered threshold and delivered data is worth more than a large lift with none of those. GPI's guidance on evaluating agency performance by profit applies the same principle to reported returns.

Handling the 'the holdout costs us revenue' objection

The objection is accurate: the control group forgoes conversions; if the advertising works, that is revenue not earned. The alternative is not certain revenue; it is not knowing whether the spend produces any of it. The holdout is the price of the decision, sized to the budget at risk. A channel carrying a fifth of annual spend justifies a larger, longer holdout than a pilot. Manage risk structurally: stagger tests so only a defined share of spend sits in holdout at once, avoid major promotions and peak seasons, and stop early only under a pre-agreed rule, never because an interim number looks good or bad.

When the result contradicts platform attribution

Incremental ROAS will often come in well below platform ROAS. That gap is expected, because attribution credits baseline conversions that would have occurred without the ad, while incrementality counts only growth beyond the brand's baseline. The gap is not, on its own, a reason to reanalyze. The decision rule agreed in the template applies. If the rule says a channel below threshold loses a stated share of budget, that happens, and the agency's proper response is to propose the next test, not a new interpretation. The allocation logic Cometly illustrates with hypothetical 3x and 1.2x figures, shifting budget toward the higher-lift channel until returns diminish, only holds when each lift estimate has passed the diagnostic checks above.

Where Incrementality Fits Beside MMM and MTA, and What It Costs to Run

Ask the agency to map every reported metric to one of three methods, experiments, models or attribution, and to name which experiments calibrate its model. Each method answers a different question, and a measurement stack is only coherent when each is labeled by what it can prove.

Boundaries: experiments, models and attribution answer different questions

An incrementality test establishes causation for one specific decision in one window. Marketing mix modeling estimates each channel's contribution across the whole mix from historical spend and outcome data, which makes it useful for planning but dependent on statistical assumptions and on having enough variation in past spend. Multi-touch attribution allocates credit across tracked touchpoints and depends on user-level data. Haus's comparison of incrementality testing and traditional MMM and Measured's summary of the tradeoffs among incrementality, MMM and MTA both treat these as complementary rather than competing tools. The buyer's requirement follows from that: experiments should calibrate the model, so that the model's channel coefficients are anchored to measured lift rather than to assumption, and the choice of method should follow the business decision it must answer, in line with the IAB's objective-based framing, not the agency's tooling defaults.

Where Incrementality Fits Beside MMM and MTA, and What It Costs to Run
MethodQuestion it answersData dependencyPrivacy exposureRole in the agency's stack
Incrementality experimentWhat did this spend cause, in this windowAggregate outcomes by group or regionLow for geo and aggregate designsDecision-grade evidence and model calibration
Marketing mix modelHow do channels contribute across the mix over timeHistorical spend and outcomesLowPlanning and allocation, anchored by experiments
Multi-touch attributionWhich touchpoints were near conversionsUser-level trackingHighOperational optimization, labeled as correlation

Method boundaries draw on the Haus and Measured comparisons cited above.

Feasibility for mid-market budgets

The cost objection has weakened. 52.0% of US brand and agency marketers reported running incrementality tests in July 2025 data, a population that includes many mid-market teams. Digital Applied reports that open-source statistical tooling has lowered analysis cost and relays Google's claim of minimum test budgets falling to roughly 5,000 dollars, which remains a vendor figure rather than an audited one. What remains is the cost that no tool removes: forgone conversions in the control group and analyst time to design and verify. Both are modest against the cost of funding a channel for a year that produces no lift.

Durability under privacy and signal loss

Google describes incrementality testing as the industry's standard for understanding advertising impact in a privacy-first way. That is a platform's characterization and Google has an interest in the framing, but the structural reason behind it is sound. Geo and aggregate holdouts compare outcomes across groups and do not need to follow individuals, so they keep working as user-level tracking degrades. Attribution does need to follow individuals, and its coverage gaps are documented; the 2021 SKAdNetwork figure discussed earlier is one dated example. For agency selection the consequence is direct: an agency whose measurement stack is entirely attribution-based carries a durability risk, and the buyer, not the agency, absorbs it when the signal thins.

In this discussion from Humans of Martech, the hosts break down the complementary roles of MMM, MTA, and incrementality experiments, highlighting the operational costs and decision boundaries of testing.

How GPI Reads Incrementality Claims When Evaluating Agencies

Bring the completed template to your next agency review and score the agency on disclosure, not on reported lift. The requirement this article argues for is not that an agency "does incrementality". It is that the agency discloses the design, the inputs and the limitations well enough for you to judge the claim yourself.

Claims, evidence and limitations as the unit of assessment

GPI evaluates agency outcome claims by asking what documented evidence supports them and what limitations the agency states, rather than by the size of the headline result. A lift claim is the same kind of object. Who assigned control, how large the holdout was, what threshold was pre-registered and how wide the interval is are the evidence; the reported percentage is the claim. The market has moved far enough that asking for these fields is asking for current practice: a majority of US marketers already run tests, and senior decision-makers in a vendor-commissioned January 2026 survey said they trust independent testing above other methods.

Using the template as a shared standard between buyer and agency

A workable sequence looks like this:

  1. Audit the current reports and classify every metric as attribution, A/B or holdout.
  2. Insert the nine fields into the next quarterly review or RFP.
  3. Agree decision rules for positive, null and negative results before the first test launches.
  4. Revisit after two test cycles and adjust cadence, holdout sizing and corroboration requirements based on what the data delivered.

A good agency will welcome the standard, because it converts a recurring argument about whose number is right into a shared method both sides can defend. Directory profiles such as DAR Media Marketing and AB Marketing Group show how GPI presents agency evidence for buyers to evaluate; they are reference points for the format of evidence-led evaluation, not statements about any listed agency's measurement practice.

FAQ

How long should a holdout run before the agency reports lift?

Long enough to cover the conversion lag and enough comparable periods to separate the campaign's effect from weekly and seasonal noise. The number comes from the power calculation agreed before launch, not from a rule of thumb. Write the planned runtime and the conditions for early stopping into the pre-registration, and do not accept a report issued before that date.

Can a platform-native lift study satisfy the independence requirement?

Not on its own for a major budget decision. The platform selling the media assigns control, measures and reports, so the result is a claim from an interested party. It is useful for optimization and as one input, and it should be corroborated by a geo or independent design on the same channel at least once a year. The January 2026 survey suggests executives already draw this distinction, with the caveat that the survey itself came from a testing vendor.

What should happen when a test shows zero or negative lift?

The decision rule agreed before launch applies. If the test was adequately powered and pre-registered, a null result is a finding, not an anomaly, and the pre-agreed budget reduction follows. If the agency believes the design was flawed, the remedy is a better-designed retest with the same threshold, not a reinterpretation of the existing data.

How much spend should be in holdout at any one time?

Enough to reach the minimum detectable effect the decision requires, and no more than the business can tolerate as forgone conversions. Both limits are specific to your baseline volume, variance and risk appetite. Agree a ceiling on total spend in holdout across all concurrent tests, stagger tests to stay under it and size each holdout from its power calculation.

How often should each channel be retested?

When something that could change the answer has changed: a material shift in spend level, creative strategy, audience, seasonality or platform behavior. The IAB's objective-based framing is a useful guide, since a channel used for calibration needs periodic refreshes while a channel under active budget review needs a current result. Set a minimum cadence per major channel in the template and trigger additional tests on those change events.

Who should own the raw test data, the brand or the agency?

The brand. Assignment lists or region lists, exposed and control outcomes, spend by group and window definitions should be delivered with every report and retained by the brand. Ownership is what makes the exposed-versus-unexposed comparison described in the definition reproducible by your own team, and it protects continuity if the agency relationship ends.