Skip to main content
Disclosure
ShopSideK is reader-supported. When you buy through links on our site, we may earn an affiliate commission at no extra cost to you.

Calculating the right Shopify bundle A/B test sample size is challenging. Most merchants test quantity breaks or tier discounts without enough data. They plug total store visits into basic calculators. This error severely shrinks estimated test duration. Merchants then stop early on random conversion spikes.

The correct method is simple. Size your test strictly on daily visitors who view the bundle widget on the product page. Scale your sample size with an inverse-square formula. Finally, keep tests active for a full 14-day business cycle floor.

Our ShopSideK rule is clear. In-app counters like Kaching’s 10-order threshold are software eligibility gates to begin calculating significance. They do not prove statistical power. Merchants must set a fixed testing horizon before launching.

ShopSideK Verdict
My take: Sizing a Shopify bundle test requires decoupling product page traffic from storewide visitors, selecting a realistic Minimum Detectable Effect between 10% and 20%, and locking in a 14-day to 21-day runtime floor to absorb weekday-weekend shopping distortion.
Best for: Revenue-active Shopify stores ($500k–$5M+ GMV) with at least 300 daily eligible PDP visitors seeking to validate quantity tiers and discount depths without falling victim to peeking penalties or false-positive winners.
Watch for: Stopping early based on interim significance, which can inflate false positives, and treating Kaching’s 10-order software threshold as proof of statistical significance.
Next step: Calculate your required sample size using eligible PDP traffic, configure your test in Kaching Bundles, and let the test run uninterrupted across two complete weekly cycles.

The Eligible Traffic Sizing Rule: Why Storewide Sessions Break Test Duration

Many store owners size tests with total store sessions. An analytics screen might show 90,000 monthly sessions. The owner divides by 30 days. They assume they have 3,000 daily visitors for testing. This calculation is wrong. Bundle offers do not appear sitewide. They load on specific product pages.

Your top product might get only 5% of store visits. That leaves only 150 daily eligible visitors (T_E) who see the bundle block. Sizing against 3,000 store sessions makes merchants expect results in 4 days. In reality, collecting the same sample at 150 eligible visitors per day requires 80 days. Ending a test early means acting on random noise.

To measure test duration correctly, count only eligible shoppers reaching the bundle offer:

Duration (Days) = (2 × n) / Daily Eligible Visitors (T_E)

In this formula, n is the required sample size per variant in a 50/50 test. The total visitor target across Variant A and Variant B is 2 × n. Daily eligible visitors (T_E) counts only shoppers who view the bundle widget.

This guide focuses on statistical sample size calculations, MDE sensitivity, duration models, business cycle floors, stopping rules, and low-traffic feasibility gates. For strategic questions outside test duration:

Sizing Your Shopify Bundle A/B Test: Binary Conversion Rates and the Inverse-Square Law

Testing binary conversion rate (CVR) measures whether a shopper orders. The required sample size per variant (n) uses four core inputs:

  1. Significance Level (alpha): Set at 0.05 (95% confidence). This gives a critical value of Z_alpha/2 = 1.96. It limits false positives to 5%.
  2. Statistical Power (1 − beta): Set at 0.80 (80% power). This gives a critical value of Z_beta = 0.84. It ensures an 80% chance of detecting a true lift at the chosen MDE when model assumptions hold.
  3. Baseline Conversion Rate (p): Your current product page conversion rate as a decimal (such as 0.025 for 2.5%).
  4. Relative Minimum Detectable Effect (relative MDE): The target percentage shift in conversion rate. The absolute difference is Delta = p × relative MDE.

The two-sample proportion formula calculates sample size:

n = (Z_alpha/2 + Z_beta)² × 2 × p × (1 − p) / Delta²

With standard constants (1.96 + 0.84)² = 7.84, the math simplifies:

n = 15.68 × (1 − p) / (p × relative_MDE²)

Sample size scales inversely with the square of the effect size (n ∝ 1/Delta²). When you cut your target detectable change in half, the required sample size quadruples. You need four times the traffic.

For instance, consider a 2.5% baseline conversion rate. Detecting a 20% relative change requires 15,288 visitors per variant (30,576 total). Detecting a 10% relative change requires 61,152 visitors per variant (122,304 total). Detecting a 5% change demands over 244,600 visitors per variant. Most product pages get modest daily traffic. Therefore, an MDE between 10% and 20% is typically the most practical target for moderate-to-high traffic PDPs, while lower-traffic catalogs require bolder 20% to 30% targets.

This lookup table provides a rough planning approximation for two-variant 50/50 split tests. It shows how baseline conversion and MDE determine sample size and required days (rounded up) across traffic tiers:

Baseline CVR (p)Relative MDEPer-Variant Sample (n)Total Sample (2n)Days at 500 Daily PDP VisitsDays at 1,000 Daily PDP VisitsOperational Feasibility Window
1.5%10%102,965205,930412 days206 daysInfeasible (> 60 Days)
1.5%20%25,74151,482103 days52 daysExtended Runtime (High Risk)
2.5%10%61,152122,304245 days123 daysInfeasible (> 60 Days)
2.5%15%27,17954,358109 days55 daysExtended Runtime (55–109 Days)
2.5%20%15,28830,57662 days31 daysFeasible at 1,000/day (31 Days)
2.5%30%6,79513,59028 days14 days (Floor)Feasible (14–28 Days / Floor)
3.5%10%43,23286,464173 days87 daysExtended Runtime (> 60 Days)
3.5%20%10,80821,61644 days22 daysFeasible (22–44 Days)
3.5%30%4,8049,60820 days14 days (Floor)Fast Testing (14–20 Days / Floor)

These calculations assume two variants with an equal 50/50 traffic split and independent visitors. If you test three or four variants in Kaching Bundles, or configure custom traffic allocation under Settings, you must recalculate your sample size accordingly.

Interactive Shopify Bundle A/B Test Sample Size & Duration Calculator

Use this interactive planner to calculate sample size, total required traffic, and minimum runtime days for your specific product page traffic:

Planner InputYour ValueInput Description
Baseline bundle conversion rate (p)
%
Your current product page conversion rate as a percentage
Relative minimum detectable effect (MDE)
%
The smallest relative conversion difference you want to detect
Daily eligible PDP visitors (T_E)
visitors/day
Shoppers who view the bundle widget on this product page each day
Confidence level95%Fixed two-sided statistical significance threshold (alpha = 0.05)
Statistical power80%Probability of detecting a true difference at the chosen MDE (beta = 0.20)
Business cycle floor14 daysMinimum required testing window across two full weekly cycles

Live planning result: 15,288 visitors per variant · 30,576 total visitors · 62 statistical days · 62 planned days · STATE 2: MODERATE FEASIBILITY (EXTENDED RUNTIME).

Planning OutputPlanning ResultOperational Meaning
Target conversion rate3.00%A 20% relative target over the 2.50% baseline
Required sample per variant (n)15,288Per-variant exposure target from the two-sample proportion formula
Total required sample (2n)30,576Combined visitor exposure needed across Variant A and Variant B
Statistical runtime62 daysTotal visitors divided by 500 daily eligible visitors, rounded up
Final planned runtime62 daysThe longer of statistical runtime and the 14-day cycle floor
Expected control ordersAbout 382Projected orders at the 2.50% baseline
Expected variant ordersAbout 459Projected orders at the 3.00% target
Kaching eligibility gate checkExpected above floorProjected orders exceed Kaching’s 10-order software threshold
Feasibility state classificationSTATE 2: MODERATE FEASIBILITY (EXTENDED RUNTIME)Takes 62 days. Testing a bolder MDE (25%–30%) or pooling traffic across top product pages can reduce runtime below 42 days.

The Mandatory 14-Day Business Cycle Floor

High-volume stores often make a dangerous mistake. They stop tests after 3 to 5 days once a dashboard shows nominal significance. A store with 3,000 daily eligible visitors reaches 10,000 visitors in days. Yet ending early ignores buyer habits.

Online shopping habits change throughout the week. Monday desktop visitors behave differently from weekend mobile shoppers. Order values also swing around paydays on the 1st and 15th of the month.

New bundle widgets also trigger novelty bias. Regular returning buyers often click fresh on-page widgets out of curiosity. This temporary surge fades once the widget becomes familiar.

To guard against weekday shifts, payday swings, and novelty bias, ShopSideK recommends that every store apply a cycle floor:

Minimum Test Runtime = max(statistical_runtime_days, 14 days)

Statistical math might suggest stopping in 4 days. Even so, run the test for at least 14 days. That covers two full weekly cycles to help absorb weekday-weekend shopping variations. Running for 21 days is even safer. Halting in under 7 days risks keeping false winners that hurt store revenue.

The Mathematical Cost of Dashboard Peeking

Repeatedly checking significance and stopping early when a variant appears to win can inflate false positives. A 95% confidence level (alpha = 0.05) accepts a 5% false positive risk. That risk level assumes one single evaluation at the predetermined sample size.

Repeatedly testing interim results and stopping when nominal significance appears drastically inflates false positives. Because cumulative observations are correlated over time, random variance causes conversion curves to oscillate. An uncorrected run can cross nominal significance purely by chance. The exact increase in false positive risk depends on your inspection frequency and stopping rules.

You can monitor tracking, traffic allocation, and technical problems during the run. However, you must never declare a winner based on interim results. Frequentist z-tests require fixed horizons. Set your sample size n in advance. Commit to a 14-day calendar window. If your test reaches full sample size and time without a statistically clear winner, record the result as inconclusive rather than extending the test indefinitely.

Continuous Metric Sizing: Revenue and Contribution Profit per Visitor

Evaluating tests only on binary conversion rate (CVR) can hurt margins. A bundle offer with a 30% discount might bring more orders. Yet it might produce lower total profit than the full price. To track true financial gains, merchants must track dollar metrics. These include Revenue per Visitor (RPV) and Contribution Profit per Visitor (CPPV). RPV and CPPV require their own variance estimate; the CVR sample-size table does not establish their statistical power.

Continuous metrics combine non-buyers ($0.00 contribution) with buyers who place orders. Those orders might be $25.00, $75.00, or $140.00. Testing continuous metrics requires modeling standard deviation (sigma):

n = 2 × (Z_alpha/2 + Z_beta)² × sigma² / Delta²

Continuous revenue metrics are zero-inflated because non-buyers ($0.00 contribution) heavily outnumber buyers. This creates a wide spread between non-purchasers and multi-unit buyers, expanding the metric variance (sigma²). Because required sample size scales directly with variance (sigma²), detecting small relative changes in Contribution Profit per Visitor typically demands larger sample sizes or bolder MDE targets than binary conversion rate tests.

In practice, because Shopify apps like Kaching export aggregated daily summaries rather than raw visitor-level event logs, merchants rarely have the individual-shopper variance data needed to calculate continuous sample sizes directly. The practical operator approach is to size your test horizon using binary conversion rate (CVR) as the primary statistical anchor, while evaluating Contribution Profit per Visitor (CPPV) as the commercial decision guardrail before rolling out a winner.

ShopSideK applies this exact per-order cost formula:

Contribution Profit = Net Revenue − COGS − Pick/Pack/Packaging − Shipping Subsidy − Payment Processing Fees
Payment Processing Fees = Net Revenue × 0.029 + 0.30 (Illustrative processing fee; replace with your store's actual fee)

The numbers below illustrate this cost scope across two distinct orders (Hypothetical Example):

  • Baseline Single Unit (MATH-001): 1 unit sold at $40.00 retail base ($40.00 net revenue). Subtracting $10.00 COGS, $3.00 pick/pack/packaging, $4.00 shipping subsidy, and $1.46 payment processing fees ($40.00 × 0.029 + $0.30) leaves a stated Contribution Profit of $21.54.
  • Discounted 2-Pack Bundle (MATH-002): 2 units sold at $72.00 net revenue ($36.00 per unit with a 10% volume discount). Subtracting $20.00 COGS (2 × $10.00), $4.00 pick/pack/packaging (consolidated package), $5.00 shipping subsidy, and $2.39 payment processing fees ($72.00 × 0.029 + $0.30 = $2.388) leaves a stated Contribution Profit of $40.61.

This comparison represents an Illustrative Basket Contrast between different order sizes. It is not a causal lift claim. Sizing tests on continuous Contribution Profit per Visitor involves wide dollar spreads. Non-buyers generate $0.00. Single units yield $21.54. Bundles bring $40.61. This spread expands metric variance (sigma²). The required sample size depends on visitor-level variance and the relative effect size you need to detect.

For adjacent profit decisions outside sample sizing:

Platform Architecture and Split Testing in Kaching Bundles

Running a clean test on Shopify requires respecting platform rules. Native Shopify Rollouts under Markets > Rollouts supports Launch and Experiment rollout types for theme and checkout testing (with Experiment rollouts requiring the Grow plan or higher). However, Rollouts change types do not provide native split testing for dynamic product prices or discount bundle offers.

Duplicating products can complicate inventory management. Under Shopify variant inventory rules, stock is tracked per variant. Cloned products create separate inventory pools that can drift out of sync unless managed with external inventory sync tools.

Modern bundle apps avoid this issue. Shopify’s Cart Transform Function API can merge cart lines into a single parent bundle line item. This updates pricing and display on Shopify backend systems natively without theme Liquid code hacks.

Similarly, theme app extensions let apps display dynamic blocks on product pages without editing theme template files. Kaching Bundles holds the Built for Shopify badge on the Shopify App Store, meeting official performance, design, and integration standards.

Setting Up Split Tests in Kaching Bundles

In Kaching Bundles, setting up an experiment happens directly inside the app editor:

  1. Inside the bundle editor, click the Run A/B test button in the bundle preview section to start an experiment.
  2. The app creates Variant A and Variant B by default, and merchants can click Add variant to test up to 4 different split test variants per bundle block.
  3. Traffic is split evenly across variants by default, or merchants can configure custom traffic percentages under Settings using Custom Allocation.
  4. When your variant tiers and styling are ready, click Publish to push the experiment live to your store.

Visitor tracking runs through client storage. Kaching details this in their guide on how A/B split testing works. Kaching assigns each visitor to one variant using a browser storage key called kaching_session_id. A returning shopper sees the same variant across visits. This holds as long as they use the same device and browser without clearing storage data.

Understanding Kaching Analytics and Winner Declaration

Merchants reviewing in-app metrics should interpret data through statistical rules:

  • Winner Calculation: Kaching calculates winning variants based on conversion rate (CR) using a z-test once each variant reaches at least 10 orders.
  • Software Eligibility Gate vs. Statistical Power: Kaching requires at least 10 orders per variant before considering a winner. This in-app eligibility threshold does not establish that your planned statistical sample size or power target has been reached.
  • Inconclusive Outcomes: As shown in Kaching’s guide on why A/B tests do not show a clear winner, if a test lacks data or results are close, the app does not declare a winner. It advises merchants to let the test run longer or test a more impactful change.
  • Dashboard Metrics: Kaching Bundles Analytics tracks Visitors (shoppers who saw the offer), CR (order conversion rate), Revenue / visitor, and Profit / visitor based on Shopify cost per item. Note that in-app profit does not deduct shipping subsidies or processing fees.
  • Time-Series CSV Export: For deeper analysis, merchants can export daily data per Kaching CSV export documentation. The file provides daily rows with date, deal_name, variant, currency, visitors, add_to_carts, atc_rate_%, eligible_orders, bundle_orders, visitor_conversion_%, bundle_conversion_%, total_revenue, added_revenue, aov, and revenue_per_visitor.
  • Timezone Alignment: Kaching records metrics in Central European Time (Europe/Vilnius). It logs visitors on the visit day. It records orders and revenue on the paid day. Visits and orders can fall on different dates. Therefore, metrics are most reliable when analyzed across full multi-week periods.

Pre-Test Feasibility State Machine for Low-Traffic Stores

Sample math may show that a test needs over 42 to 60 days. When an experiment exceeds two months, ad campaigns shift, seasonal buying patterns change, and browser storage attribution degrades over time. Merchants must treat 42 to 60 days as an operational review point. Low-traffic stores (< 500 daily eligible visits) should follow this feasibility state machine:

Eligible PDP Traffic (T_E) Feasibility State Machine:
├── State 1 (T_E ≥ 1,000/day)  → High Feasibility (Standard MDE; Runtime Calculated from CVR)
├── State 2 (300 ≤ T_E < 1,000) → Moderate Feasibility (Recalculate Days; Review MDE Spread)
├── State 3 (100 ≤ T_E < 300)   → Low Feasibility (Cross-PDP Traffic Aggregation or 30%+ MDE)
└── State 4 (T_E < 100/day)    → Infeasible in Window (Bypass Split Test; Guarded Rollout)
  • State 1 (High Eligible Traffic, 1,000 or more visitors/day): High feasibility for standard 50/50 split testing in Kaching Bundles. Calculate required test days using your baseline CVR and target MDE. Always enforce the mandatory 14-day cycle floor regardless of traffic volume.
  • State 2 (Moderate Eligible Traffic, 300 to 999 visitors/day): Calculate duration directly from your baseline CVR and target MDE. An expanded 20% to 25% MDE reduces the required sample compared to 10%, but lower-traffic stores may still require more than 60 days. For example, at a 2.5% baseline CVR and 300 daily eligible visitors, the planning formula yields approximately 102 days at 20% MDE or 66 days at 25% MDE. Always review operational feasibility before launching.
  • State 3 (Low Eligible Traffic, 100 to 299 visitors/day): Testing a single PDP is constrained. Pool your eligible traffic by applying the identical bundle offer across your top 3 to 5 related product pages, or test a radical value proposition targeting an MDE of 30% or higher.
  • State 4 (Micro Traffic, under 100 visitors/day): Formal split testing is often impractical within planned operating windows. Rather than running a months-long inconclusive test, release the offer as a guarded sequential rollout. Track conversion rate and Contribution Profit per Visitor against a prior 30-day pre-launch baseline as a monitoring guardrail.

For adjacent pricing and coupon rules:

Pre-Launch Experimentation QA Checklist

Before pushing a bundle test live, review this 6-point pre-launch checklist:

  1. Isolate Eligible PDP Traffic (T_E): Calculate test duration using only shoppers who visit the product page and view the bundle widget. Never size tests with storewide traffic.
  2. Verify MDE Feasibility: Ensure your chosen relative MDE finishes in 42 days or fewer at your traffic rate. Treat estimates above 42 days as an operational review point to apply the feasibility state machine.
  3. Pre-Commit Fixed Sample Horizon: Calculate the required sample size n before launching. Record this target in your testing sheet and do not stop early.
  4. Lock the 14-Day Cycle Floor: Commit to running the experiment for at least 14 days regardless of early z-test numbers.
  5. Confirm Variant Persistence: Test the storefront in a private browser window. Refresh the page several times to confirm that the kaching_session_id local storage key consistently retains assigned variants across page reloads.
  6. Schedule Weekly CSV Downloads: Set a weekly calendar reminder to export Kaching analytics CSV files. This helps monitor data trends across Central European Time day boundaries.

Structured testing safeguards profit margins against deceptive conversion spikes. Use the CVR test alongside a contribution-profit guardrail when deciding whether to roll out an offer; CVR significance alone does not establish a statistically significant CPPV lift. Lock in your sample size goals before launch. Respect the 14-day business cycle floor. That discipline gives you trustworthy commercial answers.

For a complete breakdown of app setup, features, and pricing tiers, read our detailed Kaching Bundles review.

Chloe Phung

Chloe Phung is a Shopify Specialist and the founder of ShopSideK. As an official Shopify Media Partner, her expertise is rooted in over two years as a Digital Marketing Executive at MyShopKit, where she was a core part of the team behind the Veda Landing Page Builder.Having directly consulted and supported thousands of global merchants to achieve 5-star success, Chloe possesses a deep, "front-line" understanding of conversion rate optimization (CRO), SEO, and strategic app integrations. Today, she leverages her insider knowledge of the Shopify ecosystem to help entrepreneurs transform their stores into high-converting, global brands.

Is a Shopify Bundle Discount Plus Free Shipping Still Profitable? (Double-Dip Math)

Is a Shopify Bundle Discount Plus Free Shipping Still Profitable? (Double-Dip Math)

Does Your Shopify Bundle Increase True Profit or Only Top-Line AOV?

Does Your Shopify Bundle Increase True Profit or Only Top-Line AOV?

How to Test a Shopify Bundle Against No Bundle (Clean Baseline A/B Testing)

How to Test a Shopify Bundle Against No Bundle (Clean Baseline A/B Testing)

Leave a Reply

TABLE OF CONTENTS