Skip to main content
Disclosure
ShopSideK is reader-supported. When you buy through links on our site, we may earn an affiliate commission at no extra cost to you.

When you A/B test Shopify bundle offers, the experiment should answer one commercial question: which offer leaves more contribution profit from the same amount of eligible traffic? Conversion rate, AOV, and revenue per visitor help explain the result, but any one of them can name the wrong winner when discounts and variable costs differ.

Kaching can split visitors between bundle variants, keep the same browser on its assigned variant, and export results by variant. It does not add your COGS, payment fees, fulfillment, shipping, or return costs. The workflow below joins those two sides before you roll an offer out to everyone.

⚡ ShopSideK Verdict

Use Kaching A/B Split Testing when you have two credible bundle offers and enough traffic to compare them concurrently. Treat its conversion result as conversion evidence; calculate the business winner from contribution profit per variant visitor.

  • Best for: Stores with a live bundle, stable traffic, and two economically distinct offers
  • Primary metric: Contribution profit per variant visitor
  • ShopSideK deal: 20% OFF for the first 3 months

Test roadmap

  1. Write the hypothesis, minimum worthwhile lift, guardrails, and stop conditions before launch.
  2. Run two variants concurrently while keeping products, traffic, and other promotions stable.
  3. Aggregate the full-period export, add merchant costs, and return WINNER, CONTINUE, or INCONCLUSIVE.

Claim 20% OFF Kaching for your first 3 months →

The 10-order threshold in Kaching’s winner logic is not a universal stopping rule, and its conversion z-test does not prove that a contribution-profit difference is significant.

Before you A/B test Shopify bundle offers, choose one decision

Do not open the variant editor with a list of colors, badges, images, tier prices, and button labels you would like to try. Start with one decision that would change what you show customers.

A useful hypothesis has a mechanism and a measurable business consequence:

Reducing the three-unit discount from 20% to 12% will preserve enough bundle uptake to increase contribution profit per variant visitor without pushing visitor conversion below our guardrail.

That test changes one economically meaningful element: discount depth. Product eligibility, bundle quantities, page placement, copy, traffic sources, and campaign timing stay fixed.

Other clean comparisons include:

  • a two-tier ladder versus a three-tier ladder with the same entry offer;
  • a percentage discount versus a fixed bundle price that creates equivalent customer savings;
  • a value-led headline versus a savings-led headline with identical tiers;
  • the same offer with two different default tier selections.

Changing the discount, headline, image, layout, and tier count together may produce a winner, but it will not tell you what caused the movement. That is a redesign test, not a clean diagnosis.

When not to run the test yet

Hold the experiment when either variant fails basic margin math, your store cannot reach a useful sample within a practical period, or a major campaign, price change, theme release, or inventory event will affect the variants differently.

Low traffic is not solved by sending 25% of visitors to four ideas. Kaching supports up to four variants, but two variants give each comparison more exposure and make the result easier to interpret. Screen the weak ideas with unit economics first, then test the two you could realistically keep.

This workflow compares two bundle offers. It does not establish a true bundle-versus-no-bundle control, which is a separate experiment design.

Write the experiment charter before you configure the variants

The charter prevents a common failure: changing the definition of “winner” after the results appear. Keep it short enough that the person launching the test will actually use it.

Charter fieldWhat to recordExample
DecisionThe choice the test will makeKeep a 10% or 20% three-unit discount
HypothesisChange, mechanism, and expected result10% preserves uptake while improving contribution per visitor
Control and variantExact customer-visible differenceA: 10% off; B: 20% off
Audience and scopeProducts, Markets, pages, and traffic includedUS product-page traffic to three named products
AllocationIntended traffic share50% A / 50% B
Primary business metricThe number used for the economic decisionContribution profit per variant visitor
Minimum worthwhile effectSmallest lift worth implementingAt least $0.05 more contribution per visitor
Diagnostic metricsNumbers that explain the primary resultVisitor conversion, bundle take rate, AOV, revenue per visitor
GuardrailsOutcomes that must not deteriorate beyond a stated limitRefund cost, cancellations, stockouts, support contacts, cart errors
Stop planSample, minimum runtime, maximum runtime, and no-peeking ruleStore-specific plan set before launch
Maturation ruleWhen orders and return costs are complete enough to analyzeRe-export after the declared order lag; finalize after return allowance matures
Freeze listChanges that require a restart or an explicit analysis notePrices, promotions, product scope, theme, app settings, traffic mix

The minimum worthwhile effect is a business threshold. It answers, “How much extra contribution justifies keeping this variant?” It is related to, but not the same as, the minimum detectable effect used in a sample-size calculation.

If a new layout requires extra design work, support, or ongoing operations, include that implementation cost in the decision threshold. A statistically detectable gain can still be too small to be worth the maintenance.

Use contribution profit per visitor as the business metric

Kaching’s analytics export gives you the pieces needed to calculate visitor conversion, bundle take rate, AOV, and revenue per visitor. Those metrics answer different questions:

  • Visitor conversion: Did a larger share of widget visitors place an eligible paid order?
  • Bundle take rate: Of eligible orders, did more use the bundle discount?
  • AOV: How much revenue did each eligible order contain on average?
  • Revenue per visitor: How much gross revenue did each widget visitor produce?
  • Contribution profit per visitor: How much remained per widget visitor after the variable costs included in your model?

Use this test-level definition:

Contribution profit = total revenue − COGS − payment fees − merchant-funded fulfillment and shipping − matured refund/return adjustment − other variable costs

Then divide each variant’s contribution profit by that variant’s visitor count:

Contribution profit per variant visitor = contribution profit ÷ variant visitors

The denominator matters. Kaching defines visitors as visitors who saw the deal widget on a product page. It is not a generic store-session count. Use the Kaching visitor field for both variants rather than mixing it with sessions from another analytics tool.

Do not subtract the headline discount a second time if Kaching’s total_revenue already contains the discounted selling price. Common fixed overhead can stay outside an A-versus-B comparison when it is identical for both variants. Include an expense when the offer changes it or when your decision specifically needs it.

If your cost sheet is incomplete, ShopSideK’s Shopify profit margin calculator and growth guide can help organize the inputs. The result is only as credible as the cost treatment you apply consistently to both variants.

Set up the A/B test in Kaching

Kaching’s official split-testing guide documents even traffic allocation by default, optional custom allocation, and a maximum of four variants per bundle block. For a first profit-led comparison, use A and B unless there is a strong reason to divide traffic further.

  1. Open the live bundle or create the bundle block you intend to test.
  2. Select Run A/B test in the bundle preview area.
  3. Keep Variant A as the current control. Create Variant B from the same starting setup.
  4. Change only the element named in the charter. Kaching documents tests for bundle deals, discounted prices, titles, layout, product images, and button text.
  5. Use an even split unless the charter explains why a custom allocation is necessary.
  6. Preview both variants and test product selection, quantity, cart, discount, checkout, mobile layout, and analytics events before publishing.
  7. Save the exact launch time, settings, and screenshots for both variants.

Bundle visibility—including Markets and main-product selection—and the schedule cannot be different between test variants according to the current Kaching guide. If your hypothesis requires one of those changes, this in-block test is not the right design.

Kaching also says the app cannot increase the original product price. A price comparison must work through the original product price and the discount architecture, so confirm the customer-visible price and checkout amount in both paths.

If this is the test layer your store needs, unlock 20% OFF Kaching for your first 3 months. Use the ShopSideK form to receive the code, then finish the charter and QA before sending campaign traffic.

Set the stop rule without borrowing a universal number

Kaching’s winner explanation says each variant needs at least 10 orders to be considered and that the app uses a z-test on conversion rate. Ten orders is therefore an eligibility floor in Kaching’s documented logic. It does not tell you that the test has enough data to detect the lift your store cares about.

The required sample changes with:

  • the control conversion rate;
  • the smallest improvement worth acting on;
  • the false-positive and power settings used in the plan;
  • the allocation between variants;
  • the variability and maturity of the business metric;
  • the amount of eligible traffic available each week.

Set the plan before launch. Run long enough to reach the preplanned exposure and cover a complete business cycle that matters for your store, such as both weekdays and weekends. Also set a maximum runtime. A test should not continue indefinitely because the current favorite has not received a winner badge.

Do not stop the first morning a metric turns green. Repeatedly checking a conventional fixed-sample result and stopping on a favorable fluctuation can inflate false positives. If you need a detailed sample calculation or a sequential-testing method, treat that as a separate statistical design task; the Kaching 10-order rule does not replace it.

Check experiment integrity before comparing outcomes

A clean calculation cannot rescue a contaminated test. Complete these checks first.

Compare actual allocation with the plan

A 50/50 allocation will not necessarily produce identical counts, but a large unexplained imbalance deserves investigation. Microsoft’s sample-ratio-mismatch research treats an unexpected allocation ratio as a data-quality symptom that should be diagnosed before effects are trusted.

Do not invent a universal acceptable percentage difference. Compare the observed split with the planned allocation using an appropriate check, then inspect launch timing, missing tracking, product or Market exposure, bot/internal traffic, and variant-specific failures when the deviation is suspicious.

Account for browser-based assignment

Kaching stores a kaching_session_id in browser local storage. The same browser and device should return to the same variant while that value remains. A shopper who changes browser or device, or clears local storage, may be assigned again.

That is a real limit on person-level consistency. Do not describe the experiment as persistent customer-level assignment across devices. It may still be useful for a product-page offer decision, but the limitation belongs in the analysis record.

Freeze competing changes

Keep product price, bundle visibility, inventory treatment, theme placement, cart logic, other discounts, and campaign mix stable. If a change affects both variants equally, record it. If it can affect them differently or alters the tested mechanism, restart the test or label the result inconclusive.

Run operational guardrails

A profit estimate is not permission to ship a broken experience. Compare refund/return cost, cancellations, stockouts, customer-service contacts, incorrect quantities, discount failures, and cart or checkout errors. A variant that clears the profit threshold but breaches a hard operational guardrail does not pass.

Export and aggregate the complete Kaching test period

In Kaching, open Analytics, select the date range and deal block, and choose Export CSV. The current export includes a daily row for each deal and A/B variant, with fields such as visitors, eligible_orders, bundle_orders, total_revenue, aov, and revenue_per_visitor.

Build one summary row per variant:

  1. Filter to the intended deal and test window.
  2. Confirm one store currency and the expected variant labels.
  3. Sum visitors, eligible_orders, bundle_orders, and total_revenue by variant.
  4. Recalculate conversion, bundle take rate, AOV, and revenue per visitor from the summed numerators and denominators.
  5. Do not average the daily percentage or AOV columns. A low-traffic day should not receive the same weight as a high-traffic day.
  6. Reconcile the analysis dates with the charter and note any missing or partial launch day.

Kaching records visitors on the day they visit and paid orders/revenue on the day the order is paid. A short daily row can therefore show visitors without orders or orders without visitors. Use the complete period and apply the same declared order-lag cutoff to both variants.

The help page identifies the reporting timezone as Europe/Vilnius/CET. Use the app’s dates consistently rather than translating them into an assumed fixed UTC offset. If campaign timestamps live in another timezone, document the conversion once and keep it consistent.

Add the variable costs Kaching does not export

The documented Kaching CSV stops at revenue metrics. It does not contain the following merchant economics:

Cost inputPreferred sourceCommon mistake
COGSProduct/variant cost at sale or accounting systemUsing today’s cost for an older test without checking changes
Payment feesPayout or payment-provider recordsApplying one headline percentage to every payment method
Merchant-funded shippingCarrier, 3PL, or order recordsSubtracting customer-paid shipping twice
Pick, pack, and packagingWarehouse or 3PL scheduleAssuming larger bundles cost the same to fulfill
Refund and return adjustmentMatured refund, return, and reverse-logistics recordsSubtracting both full COGS and recovered inventory value
Other variable costStore-specific recordsMixing fixed overhead into only one variant

Map costs to the same variant definition used in the Kaching export. If you cannot join every order to a variant, create a documented high/low cost range for each offer. Do not use one blended average when the deeper bundle clearly ships more units or has a different return pattern.

For returns, use a matured net economic adjustment: refunded revenue plus reverse logistics and nonrecoverable product cost, less any recovered value, using the store’s established accounting treatment. If returns have not matured, either make a provisional decision with a historical allowance or wait. When a realistic return range can flip the winner, the result is inconclusive.

Worked example: the conversion leader loses on profit per visitor

The following Bundle Experiment Winner model is hypothetical. It compares equal exposure to a 10% offer and a 20% offer. All amounts are USD, and the return adjustment is assumed to be mature.

MetricVariant A: 10%Variant B: 20%
Variant visitors10,00010,000
Eligible orders440520
Bundle orders300390
Visitor conversion4.40%5.20%
Bundle take rate68.18%75.00%
Total revenue$15,840.00$17,680.00
AOV$36.00$34.00
Revenue per visitor$1.5840$1.7680
COGS$5,280.00$7,280.00
Payment fees$591.36$668.72
Fulfillment and shipping$2,860.00$3,640.00
Matured return adjustment$475.20$707.20
Contribution profit$6,633.44$5,384.08
Contribution profit per visitor$0.6633$0.5384

Variant B looks stronger in three places. Its visitor conversion is 18.18% higher, its bundle take rate is higher, and its revenue per visitor is 11.62% higher. Variant A has the higher AOV, but that metric alone is also incomplete.

After variable costs, A produces $0.6633 per visitor against B’s $0.5384. B is down 18.83% on the primary business metric and leaves $1,249.36 less contribution over the same 10,000 visitors.

Calculate the cost-to-tie threshold

B would need $1,249.36 less variable cost to tie A at this exposure. Spread across B’s 520 eligible orders, that is about $2.40 per order.

Ask a concrete sensitivity question: could the cost model plausibly overstate B’s costs by at least $2.40 per order relative to A? If the best documented correction is only $0.60, A remains the economic winner. If a missing fulfillment rebate or recovered-return value is worth $2.75, repair the data before deciding.

The $2.40 is a business sensitivity threshold, not a confidence interval. The aggregate export does not provide visitor-level contribution variance, so this model does not calculate statistical significance for profit.

Return WINNER, CONTINUE, or INCONCLUSIVE

Do not force every test into an A-or-B answer.

ResultUse it whenAction
WINNERData-quality checks pass; the prewritten stop condition is met; the contribution-per-visitor gap clears the minimum worthwhile effect and plausible cost uncertainty; guardrails passRoll out the economic winner and keep monitoring the guardrails
CONTINUEIntegrity checks pass, the planned maximum has not been reached, and more exposure can still answer the prewritten decisionContinue without changing variants or the success rule
INCONCLUSIVEAllocation or measurement is compromised, a material setting changed, costs/returns are too immature, or the gap remains inside uncertainty at the planned stopKeep the control or choose the simpler option, repair the design, and write a stronger next hypothesis

Kaching may call one variant a significant conversion winner while your contribution model favors the other. Record both statements. Do not claim that Kaching’s conversion z-test proves the profit difference.

The reverse can also happen: one variant has higher estimated contribution per visitor while Kaching shows no clear conversion winner. Check whether the contribution gap clears your economic and cost-sensitivity thresholds. If it does not, the honest result is not “almost a winner.” It is inconclusive.

When both offers perform similarly, simplicity has value. Keeping the control avoids rollout risk, and choosing the easier variant can be reasonable when the planned analysis cannot distinguish them. State that as an operational decision, not a statistically proven lift.

When Kaching is the right testing layer

Kaching is a documented fit when both candidates can live in one bundle block and you need to compare deals, discounted prices, titles, layouts, images, or button text. It handles the storefront variants, browser-based assignment, traffic allocation, and variant analytics that make this workflow practical.

It is not the complete experiment finance system. It does not add merchant cost fields to the documented export, guarantee cross-device identity, make visibility or schedule variant-specific, or establish a native no-bundle control in the documentation reviewed for this article.

If those limits match your test, review Kaching and claim 20% OFF for your first 3 months. If you prefer the direct route, view Kaching on the Shopify App Store; the ShopSideK form is the route for the discount code.

If you need a different bundle architecture rather than a variant test, compare the approaches in ShopSideK’s best Shopify bundle apps guide before adding another testing layer.

Frequently asked questions

How long should a Shopify bundle A/B test run?

Long enough to reach a preplanned sample and cover the store’s relevant business cycle, but not indefinitely. The correct duration depends on baseline conversion, the smallest worthwhile lift, allocation, eligible traffic, and the analysis method. Kaching’s 10-order-per-variant consideration threshold is not a complete duration rule.

Should I choose the variant with the highest conversion rate?

Only when conversion is the declared primary objective and the economics and guardrails remain acceptable. For a bundle-offer decision, use contribution profit per variant visitor as the business metric and conversion as diagnostic evidence. A deeper discount can convert more visitors while leaving less contribution.

Can I test four Kaching bundle variants at once?

Kaching documents a maximum of four variants per bundle block. Two are usually a better starting point for a profit-led test because each receives more exposure and the difference is easier to interpret. Use more variants only when the traffic plan and analysis support them.

Can Kaching A/B test different prices?

Yes, when the comparison uses discounts. Kaching says it cannot increase the original product price through the app. Verify the original price, discount method, displayed savings, cart, and checkout amount for every variant.

Why do daily Kaching conversion rates look inconsistent?

Kaching records visitors on the visit day and paid orders and revenue on the paid day. Daily numerators and denominators may therefore describe activity that matured on different dates. Sum the complete test-period fields by variant and recalculate the ratios instead of averaging daily percentages.

What if Kaching shows no clear winner?

Follow the charter. Continue only when the planned test still permits more exposure and the data is clean. At the planned stop, use the contribution model, cost sensitivity, and guardrails. If the gap remains too uncertain, keep the control or simpler option and label the experiment inconclusive.

Chloe Phung

Chloe Phung is a Shopify Specialist and the founder of ShopSideK. As an official Shopify Media Partner, her expertise is rooted in over two years as a Digital Marketing Executive at MyShopKit, where she was a core part of the team behind the Veda Landing Page Builder.Having directly consulted and supported thousands of global merchants to achieve 5-star success, Chloe possesses a deep, "front-line" understanding of conversion rate optimization (CRO), SEO, and strategic app integrations. Today, she leverages her insider knowledge of the Shopify ecosystem to help entrepreneurs transform their stores into high-converting, global brands.

Shopify Bundle A/B Test Sample Size Calculator

Shopify Bundle A/B Test Sample Size Calculator

Can Two Product Discounts Apply to the Same Shopify Bundle Item?

Can Two Product Discounts Apply to the Same Shopify Bundle Item?

How to Exclude Shopify Bundles From Discount Codes Without Breaking Mixed Carts

How to Exclude Shopify Bundles From Discount Codes Without Breaking Mixed Carts

Leave a Reply

TABLE OF CONTENTS