Your Shopify bundle A/B test sample size should come from the effect you need to detect—not a universal number of days or orders. Run the test until each variant reaches its planned sample and the experiment covers the business cycle that matters to your store. “Two weeks,” “1,000 visitors per variation,” and “10 orders per variant” are useful reference points, but none is a complete stopping rule.
The real question is whether your eligible bundle traffic can detect a change large enough to justify acting on. The calculator below turns your baseline conversion, minimum detectable lift, confidence, power, traffic split, and daily eligible visitors into a visitor target and a realistic runtime.
If that plan fits your traffic, Kaching Bundles & Upsells can run the bundle variants with even or custom allocation. If it does not fit, installing a testing tool will not solve the sample-size problem.
Shopify bundle A/B test sample size calculator
Enter the rate for visitors who were actually eligible to see the bundle—not your sitewide conversion rate unless every visitor had the same opportunity. The default values below are hypothetical and show the calculator’s initial state.
Planner inputs
| Input | Your value | What belongs here |
|---|---|---|
| Baseline bundle conversion rate | Eligible visitors who completed the conversion counted by your test | |
| Relative minimum detectable effect | The smallest relative lift worth detecting and acting on | |
| Confidence level | The preselected two-sided significance threshold | |
| Statistical power | The planned chance of detecting the MDE when it is real, under the model | |
| Eligible visitors per day | People who can see the tested bundle widget each day | |
| Control allocation | Share of eligible traffic assigned to Variant A | |
| Variant allocation | Share assigned to Variant B | |
| Minimum business-cycle runtime | The shortest period that covers relevant weekday, weekend, pay-cycle, or campaign behavior | |
| Maximum acceptable runtime | Longest period you can keep the offer and traffic environment reasonably stable |
Live result: 10,297 visitors required per arm · 20,594 total visitors · 42 statistical days · 42 planned days · RUN AS PLANNED.
Live planning result
| Output | Default result | What it means |
|---|---|---|
| Target conversion rate | 4.80% | A 20% relative lift on the 4.00% baseline |
| Minimum sample per arm | 10,297 | Fixed-horizon equal-arm target from the planning approximation |
| Planned control exposure | 10,297 | Eligible visitors assigned to Variant A |
| Planned variant exposure | 10,297 | Eligible visitors assigned to Variant B |
| Total eligible visitors | 20,594 | Combined exposure required at 50/50 |
| Statistical runtime | 42 days | Total visitors divided by 500 eligible visitors per day, rounded up |
| Final planned runtime | 42 days | The longer of statistical runtime and the 14-day business-cycle floor |
| Expected control orders | About 412 | Planning estimate at the 4.00% baseline |
| Expected variant orders | About 494 | Planning estimate at the 4.80% target |
| Kaching consideration-floor check | Expected above floor | Actual observed orders must still reach at least 10 per variant |
| Feasibility state | RUN AS PLANNED | The calculated plan fits inside the 56-day maximum |
The surprising number is not 42 days. It is 10,297 eligible visitors per variant. At a 4% baseline, detecting a move to 4.8% takes far more exposure than the familiar 1,000-visitors-per-variation rule.
That does not make Shopify’s rule useless. It makes it a broad starting point. A sample plan must reflect the rate you already have and the lift you care enough to detect.
Read the result as a feasibility decision, not a prediction
The calculator is answering this question:
If the true eligible-visitor conversion rate rises from 4.00% to 4.80%, how much exposure should a preplanned two-sided, fixed-horizon comparison target under these assumptions?
It is not predicting that Variant B will reach 4.80%. The MDE is a planning threshold, not a forecast or promise.
The four result states turn the arithmetic into an operating decision:
| Result state | What the planner found | Sensible next move |
|---|---|---|
| RUN AS PLANNED | Required exposure and business-cycle floor fit inside the stable test window | Freeze the plan, launch two variants, and wait for the planned endpoint |
| TEST A BIGGER CHANGE | The current MDE takes too long, but a larger commercially meaningful difference fits | Make the variants more distinct without changing several unrelated things |
| EXTEND OR INCREASE ELIGIBLE TRAFFIC | The plan becomes feasible with a longer stable period or more legitimate eligible exposure | Extend the window, expand the eligible product audience, or route more relevant traffic |
| DO NOT RUN A FORMAL A/B TEST | No worthwhile detectable change fits, or the minimum cycle is longer than the stable window | Use qualitative evidence, a guarded rollout, or a reversible merchandising decision |
“Do not test” is sometimes the most useful output. A low-traffic store can spend two months keeping a weak offer frozen and still finish without a clear answer. The calculator lets you see that before the experiment becomes a sunk cost.
Why two weeks, 1,000 visitors, and 10 orders are different rules
These numbers solve different problems.
Shopify’s current A/B testing guide recommends a 50/50 split, a two-week run, and at least 1,000 visitors per variation as a general rule. The same guide also says a test needs enough traffic and enough time to cover normal behavioral variation. Those conditions are not interchangeable.
Kaching separately says each variant needs at least 10 orders to be considered by its documented conversion-rate winner calculation. That is a product threshold. It does not say that 10 orders supply 80% power for your baseline and MDE.
| Rule | What it can help with | What it does not establish |
|---|---|---|
| Two weeks | Covers two weekday/weekend cycles in many stores | Enough exposure to detect your chosen lift |
| 1,000 visitors per variant | Discourages decisions from tiny audiences | Adequate sample at every baseline rate and MDE |
| 10 orders per Kaching variant | Clears the app’s documented consideration floor | A sufficiently powered or economically correct decision |
| Calculated sample | Targets a particular conversion difference under stated assumptions | Coverage of your store’s business cycle |
| Business-cycle floor | Prevents a one-day or one-event sample from defining the result | Statistical power by itself |
Use the calculated exposure and the time floor together. If the sample arrives in eight days but your comparable buying pattern needs 14 days, run for the precommitted 14. If 14 days pass but the planned sample is still short, the test is not finished.
Start with the right baseline conversion rate
Sample-size calculators are extremely sensitive to the event and denominator you enter. “Our store converts at 4%” is not enough information.
For a Kaching bundle test, the useful baseline is usually:
eligible orders divided by visitors who had the opportunity to see the tested bundle widget
Kaching’s CSV documentation defines visitors as people who saw the widget and reports eligible orders separately from bundle orders. That distinction matters:
- Eligible-order conversion asks whether exposed visitors bought a qualifying product.
- Bundle-order conversion asks whether exposed visitors bought through a bundle deal.
- Sitewide conversion includes visitors who may never have reached the relevant product or collection context.
Choose the event that matches your hypothesis, then keep it unchanged in the power plan and final analysis. Do not plan from bundle-order conversion and later declare a winner from a broader storewide order rate.
Use a comparable historical period long enough to smooth an unusual day. Exclude staff tests, known tracking failures, and traffic that could not see the offer. If the coming test targets a narrower collection or market, use the rate and traffic for that audience rather than an all-store average.
Kaching notes that a visit and a later order can appear on different daily CSV rows. Build the baseline from full-period totals instead of averaging daily conversion percentages or judging one day’s row in isolation.
Set an MDE that is worth acting on
The minimum detectable effect is the smallest difference the test is designed to detect with its stated power. It should come from the business decision, not from the lift you hope to see.
The calculator uses a relative MDE. At a 4.00% baseline:
- 10% relative lift means a 4.40% target;
- 20% relative lift means a 4.80% target; and
- 30% relative lift means a 5.20% target.
A 20% relative lift from 4.00% is an increase of 0.80 percentage points, not 20 percentage points.
MDE sensitivity at the default traffic
The table below keeps the 4.00% baseline, 95% confidence, 80% power, 500 eligible visitors per day, and 50/50 split.
| Relative MDE | Target conversion | Required visitors per arm | Total visitors | Statistical runtime |
|---|---|---|---|---|
| 10% | 4.40% | 39,455 | 78,910 | 158 days |
| 15% | 4.60% | 17,923 | 35,846 | 72 days |
| 20% | 4.80% | 10,297 | 20,594 | 42 days |
| 25% | 5.00% | 6,726 | 13,452 | 27 days |
| 30% | 5.20% | 4,765 | 9,530 | 20 days |
Trying to detect a smaller lift is not automatically more rigorous. It can be an expensive way to ask a question that would not change your decision.
Suppose a 10% relative conversion lift would not compensate for the deeper discount, added gift cost, or operational work. Planning around that 10% effect would demand 78,910 visitors in this example and would still answer the wrong business question. Choose the smallest effect that would make you willing to keep the change after checking economics and guardrails.
If the required runtime is too long, do not simply type a larger MDE to make the calculator turn green. The variants themselves need a credible chance of producing a difference that large. “Test a bigger change” means strengthening the actual contrast—for example, a meaningfully different tier ladder or value presentation—not changing one word while pretending a 30% lift is plausible.
What confidence and power mean here
The default settings are 95% confidence and 80% power because they are common planning conventions, not because every Shopify store must use them.
In plain language:
- 95% confidence corresponds to a preselected 5% two-sided significance threshold in this fixed-horizon plan.
- 80% power means the plan targets an 80% chance of detecting the chosen MDE when that effect is truly present and the model assumptions hold.
More demanding settings need more traffic. With the same 4.00% baseline and 20% relative MDE, increasing planned power from 80% to 90% raises the equal-arm target from 10,297 to 13,785 visitors per variant.
Do not change confidence or power after seeing early results. These are design inputs. If your decision carries unusually high downside—such as a deep discount across a large collection—choose the standard with that risk in mind before launch.
The calculation is a normal approximation for two independent proportions, using the two-proportion effect-size and power-planning method documented by Statsmodels. It is appropriate as a transparent planning model, but it does not repair unstable traffic, bot activity, overlapping experiments, repeated cross-device exposure, or a metric that was defined incorrectly.
Daily eligible traffic turns sample size into test duration
A target of 20,594 visitors says little about the calendar until you know how many relevant people enter the experiment each day.
At the default sample:
| Eligible visitors per day | Statistical runtime |
|---|---|
| 250 | 83 days |
| 500 | 42 days |
| 750 | 28 days |
| 1,000 | 21 days |
Use a conservative daily estimate. A launch-week spike, viral post, or paid campaign that will end tomorrow is not a reliable run rate for a six-week experiment.
“Eligible” also matters more than “available.” Sending generic homepage traffic to the store does not help if those shoppers never reach a product where the tested widget appears. More traffic is useful only when it enters the same defined experiment population.
If you plan to expand the bundle across more products to gain exposure, check that the products and shoppers still represent the decision. Mixing unrelated price points, margins, or purchase motives may produce enough visitors while making the average result harder to apply.
A 50/50 split is usually the fastest simple plan
Kaching divides traffic evenly by default and allows custom allocation. A 50/50 split is efficient for a straightforward two-variant comparison because neither arm becomes starved of exposure.
This calculator treats custom allocation conservatively. It first calculates the equal-arm target, then requires the smaller traffic arm to reach that target. The larger arm receives the extra exposure implied by the split.
Using the default 10,297-per-arm target:
| Allocation A/B | Planned A exposure | Planned B exposure | Total exposure | Days at 500/day |
|---|---|---|---|---|
| 50% / 50% | 10,297 | 10,297 | 20,594 | 42 |
| 60% / 40% | About 15,446 | 10,297 | About 25,743 | 52 |
| 70% / 30% | About 24,027 | 10,297 | About 34,324 | 69 |
This is deliberately not an optimized unequal-allocation formula. It is a conservative calendar plan that makes the smaller arm’s bottleneck visible.
There are valid reasons to use an uneven split, such as reducing exposure to a risky variant. The price is usually a longer test. If speed and statistical efficiency are the priorities and both offers are safe to show, keep the first test at 50/50.
Kaching supports up to four variants, but four variants divide the same traffic into smaller streams and add more comparisons. For a merchant asking a single question—“Should I keep this bundle offer or replace it with that one?”—two variants are easier to plan and interpret.
Your business-cycle floor can be longer than the statistical runtime
The calculator asks for a minimum business-cycle runtime because shopper behavior is not identical every day.
Relevant cycles may include:
- weekdays versus weekends;
- payday timing;
- recurring email or ad schedules;
- a normal shipping cutoff;
- a subscription replenishment rhythm;
- a promotion cadence; or
- enough time for the same traffic sources to appear in their usual mix.
There is no universal 14-day law. Fourteen days is the default because it covers two weekly cycles, not because time itself creates statistical validity.
Set the floor from how your store actually trades. A B2B store with weekday purchasing may care about full workweeks. A payday-sensitive product may need a longer window. A short seasonal event may never provide a stable period for the formal test you want.
The planned runtime is the longer of:
- the days needed to collect the calculated exposure; and
- the minimum business-cycle runtime.
Then compare that result with your maximum acceptable runtime. The maximum is the period in which you can reasonably keep the tested offer, product availability, traffic sources, pricing, tracking, and surrounding merchandising stable.
If the business-cycle minimum is 28 days but the offer can remain stable for only 14, the input constraints contradict each other. The correct output is not “run for 14 days and hope.” It is “do not run this formal test under the current conditions.”
Kaching’s 10-order floor is a floor, not your sample-size target
Kaching says each variant must have at least 10 orders to be considered in its conversion-rate winner calculation, which uses a z-test. That protects the badge from evaluating variants with fewer than 10 orders. It does not establish that every test with 10 orders per arm can detect a commercially useful difference.
In the default plan, expected order counts are far above 10:
- control: about 412 orders at 4.00%; and
- variant: about 494 orders at the 4.80% target.
Those are expectations, not guarantees. The live test still needs actual observed counts. If a tracking issue, stockout, or traffic change reduces valid orders, the initial estimate does not override what happened.
Do not reverse the logic and plan around 10 orders. At a 4% rate, 10 expected orders correspond to only 250 visitors in an arm. That is nowhere near the 10,297-per-arm target for the default 20% relative MDE.
The distinction is simple:
- the sample plan asks how much exposure is needed for the chosen comparison;
- the Kaching floor asks whether a variant can be considered by the app’s documented winner logic; and
- the business decision asks whether the observed change is worth keeping after economics and guardrails.
All three can matter, but they are not the same gate.
Do not peek at a fixed-horizon plan and stop on a good morning
The calculator uses a fixed-horizon method. You choose the inputs before launch, collect the planned sample, and analyze at the planned endpoint.
Frequently checking an ordinary fixed-horizon significance result and stopping the first time it looks favorable can raise the chance of a false positive. Research on anytime-valid confidence sequences exists precisely because continuous monitoring requires a different statistical method.
You can and should monitor test health while it runs:
- confirm both variants render correctly;
- check that the allocation is plausible;
- watch for stock, pricing, checkout, or tracking failures;
- stop for genuine customer harm or a broken implementation; and
- record campaigns or incidents that change the traffic environment.
That is operational monitoring, not shopping for a favorable winner.
Kaching uses a browser-local identifier to keep a returning visitor on the same variant in the same browser and device. A person who changes browser or device, or clears local storage, may be assigned again. Treat the planner’s independence assumption as an approximation, not proof of perfect person-level randomization.
What to do when the planned test is too long
An infeasible result does not mean “give up on optimization.” It means the formal comparison is asking more from your traffic than the current design can supply.
Test a bigger, still coherent change
Increase the contrast only when the larger difference would be valuable and plausible. Examples:
- a modest quantity ladder versus a clearly stronger tier structure;
- percentage savings versus a fixed-price bundle when both are economically safe;
- a bundle focused on stock-up value versus one focused on a free gift;
- a compact selector versus a more explanatory bundle presentation.
Keep one decision at the center. Changing discount depth, copy, layout, images, and products together may produce a large result without telling you which lever caused it.
Increase legitimate eligible exposure
You may be able to include more relevant products, keep the widget visible on more qualified pages, or send additional traffic to the same eligible experience. Confirm that the expanded audience still matches the rollout decision.
Extend the runtime
Extend only if inventory, price, campaigns, acquisition mix, and surrounding merchandising can remain stable enough. More days in a drifting environment do not automatically create a cleaner answer.
Use a guarded rollout instead of a formal test
For a low-traffic store, a reversible rollout may be more honest. Launch the economically safe offer to a defined segment or period, monitor conversion and contribution guardrails, collect customer feedback, and retain a rollback condition. Call the result directional. Do not decorate it with statistical certainty it did not earn.
Move from a feasible plan to a clean Kaching test
Once the calculator returns a plan you can support, keep the implementation disciplined:
- Define one primary conversion event and one meaningful A/B difference.
- Record the baseline rate, relative MDE, confidence, power, allocation, sample target, business-cycle floor, and planned end before launch.
- Create Variant A and Variant B in Kaching.
- Use 50/50 allocation unless a real risk justifies a slower custom split.
- Confirm the correct bundle appears on every eligible product and device context.
- Freeze relevant price, inventory, campaign, and merchandising conditions where practical.
- Monitor implementation health without stopping on a favorable fixed-horizon reading.
- At the planned endpoint, aggregate the complete period and verify actual exposure, orders, and allocation.
- Interpret Kaching’s conversion result separately from contribution profit, AOV, returns, and operational guardrails.
Kaching can vary the bundle deal, discounted price, title, image, layout, and button text. Its documentation says bundle visibility and schedule cannot differ between test variants. Design the experiment around what the app can actually randomize.
If your test is feasible and those controls match the change you want to make, you can claim 20% OFF Kaching for your first 3 months. The form on that page is the primary route to receive the deal.
If you prefer to inspect the listing first, view Kaching Bundles & Upsells on the Shopify App Store.
What this calculator deliberately does not decide
The planner sizes a binary conversion comparison. It does not determine which bundle produces more profit.
A deeper discount can raise conversion while lowering contribution per visitor. A free gift can improve uptake while adding product, pick-and-pack, shipping, and return costs. If the decision is economic, analyze contribution profit per assigned visitor after the experiment using consistent cost inputs and a matured measurement window.
The planner also does not:
- provide an anytime-valid or sequential stopping rule;
- adjust for many simultaneous variants or multiple primary metrics;
- guarantee independent person-level assignment across devices;
- correct unstable traffic or broken tracking;
- size AOV, revenue-per-visitor, or profit variance;
- predict the lift your variant will achieve; or
- turn statistical significance into commercial importance.
Its job is narrower and useful: tell you whether a planned conversion test has a credible path to an answer with the traffic and time available.
Frequently asked questions
How many visitors do I need for a Shopify bundle A/B test?
It depends on the baseline eligible-visitor conversion rate, the smallest effect you want to detect, confidence, power, and traffic allocation. In the default example, detecting a 20% relative lift from 4.00% to 4.80% at 95% confidence and 80% power requires about 10,297 visitors per arm.
How long should I run a Shopify A/B test?
Run until you reach the preplanned sample and cover the relevant business cycle. Use whichever requires more time. Then confirm the full plan fits inside a period when products, pricing, traffic, tracking, and campaigns can remain reasonably stable.
Is 1,000 visitors per variant enough?
Not always. It is a broad rule of thumb in Shopify’s current guide, but adequate sample depends on the baseline and MDE. A low conversion rate or small detectable lift can require many times more than 1,000 visitors per variant.
Are 10 orders per Kaching variant enough?
Kaching documents 10 orders per variant as the minimum for a variant to be considered by its conversion-rate winner calculation. That is not the same as a store-specific power calculation and should not be used as a universal stopping target.
What does a 20% relative MDE mean at a 4% conversion rate?
It means the target rate is 4.8%. The relative increase is 20%; the absolute increase is 0.8 percentage points.
Can I stop when Kaching first shows a winner?
Not if you planned the test with this fixed-horizon calculator. Operationally monitor for defects or harm, but analyze the winner at the precommitted endpoint. Continuous result-based stopping requires a method designed for it.
Should I enter sitewide visitors or bundle-widget visitors?
Use visitors who were eligible to see the tested bundle experience. Sitewide sessions can dilute or distort the denominator when much of that audience never encountered the widget.
What if the calculator says the test needs 90 days?
Check whether a larger commercially meaningful variant, more legitimate eligible traffic, or a longer stable window is realistic. If none is, do not run a formal A/B test. Use a reversible rollout and directional evidence instead.
Plan the learning before you spend the traffic
The best time to discover that a bundle test needs 80,000 eligible visitors is before it launches.
Write down the difference worth detecting, calculate the exposure, translate it into days, and apply the store’s real business-cycle floor. When the plan fits, the test has a credible path to a decision. When it does not, redesigning the question is more useful than collecting a small, noisy result and calling it proof.


