All posts

Incrementality Testing for Shopify Ad Spend

Incrementality Testing for Shopify Ad Spend

META TITLE: Incrementality Testing for Shopify Ad Spend

META DESCRIPTION: What incrementality testing is, how Meta and Google lift tests work, and which ones a Shopify store can actually run without a rep.

SLUG: incrementality-testing

TAGS: incrementality testing, conversion lift, geo lift test, shopify ads, meta ads, google ads, ad measurement

Incrementality testing measures the sales your ads actually caused. It shows ads to one randomised group and withholds them from another. The gap is the real result.

Meta states the question plainly: "What outcomes did this ad or campaign actually cause that wouldn't have happened otherwise?" Google calls its version Conversion Lift, "an incrementality tool" measuring conversions "directly driven by people seeing your ads."

Your ROAS column cannot answer that. An experiment can.

What Is Incrementality Testing?

Here is the short answer to what is incrementality testing. You run a controlled experiment on your own ad spend.

Google describes two groups. People who see your ads sit in the treatment arm. People who do not sit in the control arm. Google calls the difference "the lift, or the increase in conversions that are caused by the presence of the ad."

Meta describes the same design. Its study compares "a test group that sees your ads, and a control group that does not see your ads." Those groups are randomised "when creating a Facebook campaign," not afterwards.

So incrementality in marketing is not a report filter. It is a randomised trial with your budget as the treatment.

Why Reported ROAS Is Higher Than Incremental ROAS

This is why incrementality testing marketing teams exist at all. Platform-reported numbers are not causal numbers.

The strongest evidence here is peer reviewed, not vendor content. Gordon, Zettelmeyer, Bhargava and Chapsky published in Marketing Science in 2019. They used "15 U.S. advertising experiments at Facebook comprising 500 million user-experiment observations and 1.6 billion ad impressions."

They compared those experiments against ordinary attribution-style models. Observational methods "often fail to produce the same effects as the randomized experiments," even "after conditioning on extensive demographic and behavioral variables."

In half their studies, the estimated increase in purchases was "off by a factor of three." Observational methods generally overestimate.

Checkout conversions fared worst, which is exactly the Shopify case. Estimates in "7 of the 14 studies with a checkout-conversion outcome" were off by more than a factor of three. One note: two co-authors were Facebook employees.

Paid search has the same problem. Blake, Nosko and Tadelis ran experiments at eBay, published as NBER working paper 20171. Their branded finding is stark: "brand-keyword ads have no measurable short-term benefits."

The non-brand result was worse than flat. Frequent buyers "whose purchasing behavior is not influenced by ads account for most of the advertising expenses." Average returns came out negative. The modelled version of the same gap sits in our Shopify ad benchmarks.

Incrementality Testing vs A/B Testing: What Is the Difference?

People search incrementality testing vs a/b testing because the two get mixed up.

An A/B test compares two things that both run. It picks a winning creative, audience or landing page. An incrementality test compares advertising against no advertising.

Google files its lift products together for this reason. Brand Lift, Search Lift and Conversion Lift are grouped as lift studies. They "go beyond metrics like clicks and impressions," Google writes.

QuestionRight Test
Which headline performs better?A/B test
Which audience is cheaper?A/B test
Did this campaign create sales at all?Incrementality test
Would we lose revenue if we paused this?Incrementality test, holdout design

Meta Incrementality Testing: How Conversion Lift Works

Meta incrementality testing runs through Conversion Lift. Older guides call it Facebook incrementality testing. It is the same product under the Meta brand.

Meta says a self-serve route exists. "If you meet the criteria, you can run your own Conversion Lift test using the self-serve option in Meta Ads Manager." Meta does not publish those criteria. Treat eligibility as a question for Ads Manager.

Meta also publishes two performance claims there, both vendor claims with no visible study. Businesses running 15 experiments a year saw "30% higher ad performances" than those running zero. Meta also cites a "15% return-on-ad-spend correction with experimentation-based calibration."

Meta's A/B testing guidance adds operating rules. Keep tests running "for at least 2 weeks or up to 30 days." Results appear once there are "at least 100 events observed." One rule stops most contamination. Meta writes: "Do not use the audience for your A/B test for any other campaign you're running at the same time."

Google Incrementality Testing: Which Lift Tests You Can Run Yourself

Google incrementality testing splits into gated tools and self-serve tools. That distinction decides what you can do this quarter.

Conversion Lift itself is gated. Google states it plainly. "Conversion Lift isn't available for all Google Ads accounts. To use Conversion Lift, contact your Google account representative."

Google Ads custom experiments are self-serve. Google recommends a 50% traffic split "to provide the best comparison." Choose the cookie-based option. It "ensures that a given user only views either the original or the experiment."

Custom experiments carry two hard limits. You "can schedule up to 5 experiments for a campaign," but only one runs at a time. They are also "only available for Search, Display, Video, and Hotel Ads campaigns."

Performance Max and Shopping are not on that list. Those are the campaign types most Shopify stores run.

Performance Max experiments only half fill the gap. Google splits them into subtypes. The Uplift experiment is the incrementality one. Google states it "doesn't support Performance Max with GMC feed." It is "only available to advertisers using the Online Sales (Non-feed), Store goals (Offline), and Lead Gen" marketing objectives.

A Shopify store running catalog-driven Performance Max off Merchant Center is outside that. The subtype that does cover Merchant Center is the Upgrade experiment. That one is a Standard Shopping campaign versus a Performance Max campaign. It runs ads in both arms. So it is an A/B test between campaign types, not an incrementality test.

TestSelf-Serve?Covers PMax or Shopping?
Conversion Lift, user basedNo, rep requiredNot directly. Video, Discovery and Demand Gen are listed as available; Shopping and PMax require reaching out to your account representative
Conversion Lift, geo basedNo, rep requiredYes, both listed
Custom experimentsYesNo
PMax Uplift experiment (incrementality)YesNo, excludes PMax with GMC feed
PMax Upgrade experiment, Shopping vs PMaxYesYes, but both arms advertise, so not incrementality

Geo-based lift covers a broader campaign list than the users-based version. That is the real reason a Shopify store gets pointed at the geo variant.

Our Google Ads for Shopify guide argues Performance Max should be judged on incremental return. A Performance Max experiment is how you get that number.

What Is a Geo Lift Test?

A geo lift test splits regions instead of people. Some cities or postal areas keep the ads. Others go dark. You compare total sales across the two sets.

Google offers this as geography-based Conversion Lift. Supported campaign types include Performance Max and Shopping. It still needs a rep, at least one compatible conversion action, and single-country targeting.

Google does publish a number. Its November 2025 measurement update covers the minimum spend for an incrementality experiment. Google says what "once might have cost upwards of $100,000 for a single experiment can now be done for $5,000." Beyond that it says only that geo budgets "tend to be higher than users-based studies." The geo page itself names no figure.

A manual holdout is possible without a rep. Google Ads location targeting reaches regions, cities and postal codes. One caveat hits small towns. Google "only permits targeting for locations that adhere to minimum privacy thresholds."

Google's public geo lift deck adds two design points. Regions are built from aggregated postal codes with boundaries drawn "through the least populous areas possible" that "do not cross well known commutes." Google says this "minimises 'contamination'", where a shopper crosses from a control area into a test one.

Results use "pre-attributed conversions," which are "not attributed/linked to a specific cookie or ad interaction." That is why a geo test survives tracking loss and a pixel report does not.

How Long Should an Incrementality Test Run?

Google's answer is four to six weeks. Google's Experiments FAQs recommend running an experiment "for at least 4-6 weeks or longer if you have a long conversion delay." They repeat the 4 to 6 week figure for Performance Max experiments whose results come back inconclusive.

Do not read week one. Google warns that new experiments "require a 7-14 day ramp-up period where data can be volatile."

Meta publishes no duration for Conversion Lift itself. Its 2-week-to-30-day window and 100-event floor are A/B testing guidance, not lift-test guidance. Do not transplant them either. Google publishes no minimum duration for Conversion Lift proper, so do not transplant the four to six week figure there.

Common Incrementality Testing Mistakes

How to Calculate iCPA and iROAS

Google defines incremental cost per acquisition as "the total spend divided by Incremental Conversions." So iCPA = total spend / incremental conversions.

Incremental ROAS is the mirror image, "the Incremental Conversion Value divided by the total spend." So iROAS = incremental conversion value / total spend.

Compare iROAS against your break-even number, not against reported ROAS. Our break-even ROAS calculator produces that threshold. Pair it with blended MER and your CAC calculation.

Check significance before acting. Google applies "jackknife resampling" to bucketed data, using "twenty buckets" in each arm. Significance is "two-tailed" against a "95% confidence interval." An undecided result means you learned nothing yet.

Does Agency AI Run Incrementality Tests?

No. Agency AI does not run lift studies. It works one layer down, where reported and blended numbers sit together.

The AI Strategist ships with a starter prompt naming this exact problem: "Why does my ROAS not match Meta Ads Manager?" The dashboard side of this is covered in what MER measures.

The Shopify App Store listing shows $59 per month or $492 per year, with a 30-day free trial.

Final Thoughts

Google's Conversion Lift needs an account representative. Meta's self-serve route depends on criteria it does not publish. The self-serve Performance Max experiment is the one most Shopify stores can start with today. Read what it returns against your break-even number, not against the ROAS column in your dashboard. An undecided result is not a green light.

Frequently Asked Questions

What is incrementality testing in marketing?
It is an experiment measuring the outcomes your ads caused. Meta frames it as asking what "wouldn't have happened otherwise." One group sees the ads, a control group does not. The gap is the lift.
How long should an incrementality test run?
Google recommends at least 4 to 6 weeks for Performance Max experiments. Expect a 7 to 14 day ramp-up where Google says data can be volatile. Meta publishes no duration for Conversion Lift. Its 2-week-to-30-day guidance applies to A/B tests, which are a different design.
Can a small Shopify store run an incrementality test?
Not easily. Google's Conversion Lift needs an account representative. The self-serve Performance Max Uplift experiment excludes Performance Max with a Merchant Center feed, which is what most Shopify stores run. Meta offers a self-serve Conversion Lift option to advertisers meeting unpublished criteria.
What is a geo lift test?
It withholds ads from selected regions instead of selected people. Google supports it as geography-based Conversion Lift, including for Performance Max and Shopping. It reports on pre-attributed conversions, so cookie loss does not break it.
Why does my reported ROAS look better than my incrementality result?
Because attribution models tend to over-credit ads. Gordon and co-authors found observational estimates "off by a factor of three" in half of 15 Facebook experiments. eBay's experiments found branded search had "no measurable short-term benefits."

Sources

The definitions of incrementality, Conversion Lift, lift, iCPA and iROAS come from Google Ads Help and from Meta's Conversion Lift page. Both are vendor documentation about their own products. Google Ads Help is the source for the rep-only requirement and the geo-based supported campaign types. It is also the source for the single-country rule and the note that geo budgets run higher than user-based budgets. The $5,000 minimum-spend figure comes from Google's November 2025 incrementality testing update. Google's custom experiments page is the source for the self-serve experiment rules, the 50% split and the five-experiment cap. The same page carries the Search, Display, Video and Hotel Ads limitation and the 10,000-user audience floor. Google's Performance Max experiments pages are the source for the experiment subtypes. They also carry the Uplift experiment's exclusion of Performance Max with a Merchant Center feed and its marketing objective limits. The 4 to 6 week duration and the 7 to 14 day ramp-up come from Google's Performance Max experiments page. The 50% traffic warning and the 20% spend-gap guidance come from the same page and Google's experiments page. The jackknife method, twenty buckets per arm and 95% confidence interval come from Google's experiment statistics page. Meta's test and control design and its self-serve criteria statement are Meta's own documentation. So are the 2-week to 30-day window for A/B tests, not lift tests, the 100-event floor and the audience-reuse warning. Meta's 30% higher ad performance figure and its 15% return-on-ad-spend correction are labelled here as vendor claims. Meta's footnote sources are not shown on that page, so neither should be read as independent research. The independent evidence is peer reviewed. The factor-of-three findings are Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science 2019, and two co-authors were Facebook employees. The branded search and negative-returns findings are Blake, Nosko and Tadelis, NBER working paper 20171. The geo boundary and contamination point and the pre-attributed conversions point come from Google's publicly hosted geo lift deck on services.google.com. Agency AI's starter prompt and its approval rules come from the client's confirmed fact sheet. Pricing and the free trial come from the Agency AI Shopify App Store listing, checked 25 August 2026.

Free to install. Launch campaigns and pay nothing for your first 30 days.

Download Agency AI free on the Shopify App Store