How to Design an Incrementality Test for Mobile Game UA: Holdouts, Geo-Lift, Sample Size, and Reading the Result
If you switched this campaign off, how many installs would you actually lose? Everyone who spends a UA budget asks that question eventually, and no dashboard answers it. Attribution records which touchpoint came last; it cannot tell you whether the player would have installed without it. The only method that answers the question directly is incrementality testing.
Dissatisfaction with measurement has not gone away. The IAB's State of Data 2026 report, published in February 2026, surveyed more than 400 senior planning and analytics decision-makers at US brands and agencies, and between 60% and 75% said their current measurement approaches do not perform well on core dimensions such as channel coverage and cross-validation. In the same report, advertisers said they expect to scale from three to five incrementality tests a year to eleven or more. The try-it-once phase is over; the question now is how well tests are designed and how often they run.
Why incrementality matters was covered through the lens of rewarded advertising in Rewarded Ads Are Working — But Can You Prove It? The Case for Incrementality Measurement. This post is the next step: a channel-agnostic guide, as of September 2026, to how game UA teams actually design and read an incrementality test.
What Exactly Does an Incrementality Test Measure?
An incrementality test randomly splits people into a group that can see the ad and a group that cannot, and treats the difference in outcomes between the two as the effect the ad created.
Randomization is the whole point. Comparing people who saw an ad with people who did not, after the fact, almost always inflates the effect, because ad platforms deliver more impressions to people who were already likely to convert. Gordon and colleagues, writing in Marketing Science in 2019, analyzed 15 US advertising experiments at Facebook covering 500 million user-experiment observations and 1.6 billion impressions. Observational methods such as exposed-versus-unexposed comparisons and matching failed to reproduce the experimental results, even after conditioning on thousands of behavioral variables. They generally overestimated the effect, and in some cases significantly underestimated it.
So incrementality answers a different question from attribution. Attribution asks which ad an install came from; an incrementality test asks how many installs would disappear without that ad. Results are usually expressed as lift (the percentage increase over the control group) and iROAS (incremental revenue divided by spend). The attribution side of that distinction is laid out in Why Every Dashboard Shows a Different Install Count.
Which of the Four Test Designs Should You Choose?
Game UA incrementality tests come in four broad designs, user-level holdout, PSA, ghost ads, and geo-split, and the right one depends on which channel you are testing and whether that channel supports user-level randomization.
A user-level holdout randomly removes part of the target audience from eligibility. It is the most intuitive design, but because you cannot tell which control-group users would actually have been served the ad, you compare entire eligible groups (intent-to-treat). When only a small share of the treatment group is actually exposed, the effect gets diluted and the required sample grows.
A PSA design shows the control group a public service announcement or unrelated ad instead. It lets you compare exposed people with exposed people, but you pay for the control impressions too, and because modern platforms optimize delivery per ad, the people who end up seeing the PSA are not the same kind of people who see the real ad.
Ghost ads solve both problems. No ad is served to the control group; the platform simply logs the auction moments where the ad would have been shown and uses those users as the comparison. Johnson, Lewis, and Nubbemeyer, in the Journal of Marketing Research in 2017, showed that ghost ads can measure lift as precisely as PSA or intent-to-treat experiments while spending at least an order of magnitude less. The catch is that the platform has to build it, so for an advertiser it is less a design to choose than a property to check for in a network's own lift measurement product.
A geo-split, or geo-lift, test uses regions rather than users as the unit, turning ads off or on in some regions only. It needs no user identifiers and runs on aggregate data, which makes it the fallback when you want to test several networks at once or when a network offers no lift measurement of its own. Meta's open-source GeoLift is the best-known tool, and Google Ads supports geography-based Conversion Lift, including for App campaigns (single-country campaigns only, accessed through a Google account representative).
Design | Control group | Strength | Limitation | Best fit |
|---|---|---|---|---|
User-level holdout | Excluded from targeting | Simple to implement | Diluted when exposure rate is low; needs large samples | Retargeting, owned CRM, push |
PSA | Served an unrelated ad | Compares exposed with exposed | Paid control impressions; delivery optimization skews groups | Simple buys without delivery optimization |
Ghost ads | Would-be impressions logged only | Low cost, high precision | Must be built by the platform | Network-provided lift studies |
Geo-split | Ads paused by region | No identifiers; multiple networks at once | Few regions means low precision; cross-region movement and halo | Cross-network checks, brand-leaning channels, new channel validation |
How Do You Set Sample Size and Test Duration?
Sample size is driven by the control group's conversion rate and the minimum lift you want to detect (the MDE), and duration is the time needed to collect that sample plus the time your success metric needs to mature.
Put simply, the lower the baseline conversion rate and the smaller the lift you want to detect, the faster the required sample grows. The table below shows users needed per group at the conventional settings of 5% significance (two-sided) and 80% power.
Control conversion rate | 5% relative lift | 10% relative lift | 20% relative lift |
|---|---|---|---|
1% | ~640,000 | ~160,000 | ~43,000 |
3% | ~210,000 | ~53,000 | ~14,000 |
These are per-group figures, so the total sample is double, and designs with low actual exposure, such as user-level holdouts, need more on top of that. The practical lesson is that a test built to prove a 5% lift is rarely feasible for a small or mid-sized campaign. Reframing the question from "does this channel work?" to "does this channel deliver enough lift to keep funding, say 20% or more?" cuts the required sample by more than an order of magnitude.
Network lift products have their own floors. Google Ads user-based Conversion Lift requires a minimum campaign budget of $5,000 and at least 1,000 observed conversions, lets you set the holdback anywhere from 1% to 50%, and recommends aiming for 90% certainty. App campaigns are supported, although App campaigns targeting iOS are not.
Duration has two parts. The first is the exposure window. Google allows studies as short as 7 days but typically recommends more than 14, and the GeoLift documentation advises that the test period contain at least one full purchase cycle. For games with strong day-of-week swings, two weeks is a sensible minimum, ideally aligned to whole weeks. The second is the maturation window. If lift is judged on D7 ROAS or D14 payers rather than installs, you have to wait until the users who installed on the last test day have accumulated their D7 data before you read the result. Which ROAS checkpoints actually predict payback is covered in D7 ROAS Benchmarks: The Targets That Tell You a Campaign Will Pay Back.
How Should Lift Be Read Against Attribution and MMM?
A lift result is worth the most over time not as a standalone number but as a correction factor for attribution and as a prior for MMM.
The most common use is an incrementality factor. If the MMP credited a channel with 1,000 installs during the test and the test found 600 incremental installs, that channel's factor is 0.6. In day-to-day operations you multiply attributed numbers by that factor to estimate true contribution, and update it with the next test. A factor well below 1 means the channel is claiming credit for players who would have come anyway; a factor above 1 means attribution is missing some of its contribution, through view-through or halo effects.
The second use is MMM calibration. Google's open-source MMM, Meridian, can parameterize each channel's ROI directly and officially supports feeding past incrementality results in as ROI priors. It includes tooling to translate experiment results into priors, and a channel calibration recommendation that flags channels where the model shows high uncertainty or potential bias as candidates for the next test. In other words, the MMM tells you what to test, and the test corrects the MMM. The overall structure of MMM is covered in Media Mix Modeling for Mobile Games: Measuring What Attribution Can No Longer See.
Tool | Question it answers | Cadence | Relationship to incrementality tests |
|---|---|---|---|
Attribution | Which ad was the last touchpoint? | Daily | Corrected by the incrementality factor from tests |
Incrementality test | How much would we lose without this ad? | Quarterly or at major changes | The ground truth |
MMM | How should budget be split across channels? | Monthly or quarterly | Test results enter as ROI priors |
Read the confidence interval before the point estimate. A result of 15% lift with a 90% interval of -2% to 32% does not mean "the effect is 15%"; it means "we could not rule out no effect." When that happens, rerunning with a larger sample is usually the right call, rather than deciding to cut or keep the channel.
Where Do Game UA Incrementality Tests Most Often Go Wrong?
When a game UA incrementality test goes wrong, the cause is more often operational than statistical, and three issues come up repeatedly: changes during the test, live events, and the app store ranking halo.
First, changing bids or creatives mid-test. Optimizing the treatment campaign halfway through leaves you unable to say what effect you measured. Freeze budget, bids, and creatives for the test window, and handle creative efficiency in separate A/B tests.
Second, live events and seasonality. A major update or collaboration event that overlaps the test window shakes organic inflow, and in a geo-split test regions react to the event differently, contaminating the result. Check the update calendar first and slot tests into the quiet gaps.
Third, the store ranking halo. More paid installs lift chart position, and organic installs follow the ranking. In a user-level holdout, treatment and control share that halo equally, so incrementality tends to be understated; in a geo-split, a country-level chart effect is hard to separate by region. During launch windows or large pushes where ranking effects are strong, consider a country-level design or state the limitation explicitly when reporting the result.
Beyond these, stopping a test the moment the daily numbers look significant (peeking) remains a common mistake. Write down the duration and decision rule before the test starts, and do not call the result before it ends.
When Should a Studio Run an Incrementality Test?
Incrementality testing is not an always-on tool but a decision tool, and the classic moments to use it are before scaling a new channel, before cutting an existing channel's budget significantly, and when attribution and MMM disagree.
In practice: run one before moving a new network from test budget to full budget. Channels prone to claiming credit for players who would have returned anyway, such as retargeting, are worth checking about twice a year. Conversely, if a channel's monthly spend cannot produce the sample in the table above, do not force a test; group several channels into a single geo-split, or postpone. A test that cannot reach a statistical conclusion spends money and muddies confidence.
The need grows once you run more than two or three channels, because more channels means more networks claiming the same player. Channel diversification is covered in Why Over-Relying on Google and Meta Is a UA Risk — And How to Build a Balanced Channel Mix; once you diversify, the next step is using incrementality tests to confirm each channel's real share.
How Do You Design a Holdout for Engagement-Based Campaigns? The Playio Perspective
For engagement-based campaigns, which use playtime or in-game actions rather than the install as the conversion, incrementality shows up properly only when lift is read on post-install behavior rather than install counts.
Engagement-based campaigns reward players for reaching a playtime threshold or a specific level after installing. Measure lift on installs alone and it is hard to answer the objection that the campaign merely rewarded players who would have liked the game anyway. So the comparison metric should be the behavior the campaign is meant to change, such as D7 retention, cumulative playtime, or first-purchase rate, tracked the same way for the control group.
Playio runs an Android-based community of five million gamers, matching games to players by genre preference and play history and tying rewards to reaching a playtime threshold or completing a specific in-game action. Because the game is surfaced inside a single community, a comparison that withholds exposure from part of the matched audience is structurally easier to set up than on an open web spread across many networks. And because the conversion itself is defined as post-install behavior, the metric you read incrementality on and the campaign goal are aligned from the start. Pricing runs on CPI or CPE, and test design can be discussed around your campaign goals.
You can find more details here. (https://playioadsen.oopy.io/bizdeck)
Key Takeaways
As of September 2026, incrementality testing is shifting from an occasional experiment to a regularly scheduled measurement tool. Choose the design based on whether the channel supports user-level randomization: for network-provided lift studies, check whether they use the ghost ads principle; when you need to see several networks together or work without identifiers, use a geo-split. Sample size is set by the control conversion rate and the minimum lift you want to detect, so narrowing the question from "does it work?" to "does it work enough to keep funding?" is the realistic path. Plan at least two weeks of exposure plus the time your success metric needs to mature. Read results with their confidence intervals, and get the most lasting value by using them to correct attribution with an incrementality factor and to calibrate MMM with ROI priors. In game UA, controlling mid-test changes, live events, and the store ranking halo comes before the statistics.
For inquiries about Playio's advertising solutions, reach out at: [email protected]
Sources
IAB, State of Data 2026: The AI-Powered Measurement Transformation, February 2026 (survey with BWG Global of 400+ decision-makers at US brands and agencies): https://www.iab.com/insights/2026-state-of-data-report/
PPC Land, summary of IAB State of Data 2026 (60-75% say measurement approaches do not perform well; tests expected to scale from three to five per year to 11+): https://ppc.land/ai-poised-to-unlock-32-billion-in-marketing-measurement-value-as-current-systems-falter/
Gordon, Zettelmeyer, Bhargava, Chapsky, A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook, Marketing Science 38(2), 2019 (15 experiments, 500M observations, 1.6B impressions; observational methods over- and underestimate): https://pubsonline.informs.org/doi/10.1287/mksc.2018.1135
Johnson, Lewis, Nubbemeyer, Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness, Journal of Marketing Research 54(6), 2017 (same precision as PSA/ITT at least an order of magnitude cheaper): https://journals.sagepub.com/doi/10.1509/jmr.15.0297
Google Ads Help, Set up Conversion Lift based on users ($5,000 budget and 1,000 conversions minimum, 1-50% holdback, 7-day minimum with 14+ days recommended, 90% certainty recommended, iOS App campaigns excluded): https://support.google.com/google-ads/answer/12005564?hl=en
Google Ads Help, Set up Conversion Lift based on geography (App campaigns eligible, single-country targeting, access via account representative): https://support.google.com/google-ads/answer/14097193?hl=en
Meta GeoLift documentation, Walkthrough (test period should contain at least one full purchase cycle): https://facebookincubator.github.io/GeoLift/docs/GettingStarted/Walkthrough/
Google for Developers, Meridian: Set custom ROI priors using past experiments; Channel calibration recommendation: https://developers.google.com/meridian/docs/advanced-modeling/set-custom-priors-past-experiments , https://developers.google.com/meridian/docs/post-modeling/channel-recommendation
Sample size table: calculated with the standard two-proportion test formula (5% two-sided significance, 80% power)