When you A/B test Shopify bundle offers, the experiment should answer one commercial question: which offer leaves more contribution profit from the same amount of eligible traffic? Conversion rate, AOV, and revenue per visitor help explain the result, but any one of them can name the wrong winner when discounts and variable costs differ.
Kaching can split visitors between bundle variants, keep the same browser on its assigned variant, and export results by variant. It does not add your COGS, payment fees, fulfillment, shipping, or return costs. The workflow below joins those two sides before you roll an offer out to everyone.
⚡ ShopSideK Verdict
Use Kaching A/B Split Testing when you have two credible bundle offers and enough traffic to compare them concurrently. Treat its conversion result as conversion evidence; calculate the business winner from contribution profit per variant visitor.
- Best for: Stores with a live bundle, stable traffic, and two economically distinct offers
- Primary metric: Contribution profit per variant visitor
- ShopSideK deal: 20% OFF for the first 3 months
Test roadmap
- Write the hypothesis, minimum worthwhile lift, guardrails, and stop conditions before launch.
- Run two variants concurrently while keeping products, traffic, and other promotions stable.
- Aggregate the full-period export, add merchant costs, and return WINNER, CONTINUE, or INCONCLUSIVE.
Claim 20% OFF Kaching for your first 3 months →
The 10-order threshold in Kaching’s winner logic is not a universal stopping rule, and its conversion z-test does not prove that a contribution-profit difference is significant.
Before you A/B test Shopify bundle offers, choose one decision
Do not open the variant editor with a list of colors, badges, images, tier prices, and button labels you would like to try. Start with one decision that would change what you show customers.
A useful hypothesis has a mechanism and a measurable business consequence:
Reducing the three-unit discount from 20% to 12% will preserve enough bundle uptake to increase contribution profit per variant visitor without pushing visitor conversion below our guardrail.
That test changes one economically meaningful element: discount depth. Product eligibility, bundle quantities, page placement, copy, traffic sources, and campaign timing stay fixed.
Other clean comparisons include:
- a two-tier ladder versus a three-tier ladder with the same entry offer;
- a percentage discount versus a fixed bundle price that creates equivalent customer savings;
- a value-led headline versus a savings-led headline with identical tiers;
- the same offer with two different default tier selections.
Changing the discount, headline, image, layout, and tier count together may produce a winner, but it will not tell you what caused the movement. That is a redesign test, not a clean diagnosis.
When not to run the test yet
Hold the experiment when either variant fails basic margin math, your store cannot reach a useful sample within a practical period, or a major campaign, price change, theme release, or inventory event will affect the variants differently.
Low traffic is not solved by sending 25% of visitors to four ideas. Kaching supports up to four variants, but two variants give each comparison more exposure and make the result easier to interpret. Screen the weak ideas with unit economics first, then test the two you could realistically keep.
This workflow compares two bundle offers. It does not establish a true bundle-versus-no-bundle control, which is a separate experiment design.
Write the experiment charter before you configure the variants
The charter prevents a common failure: changing the definition of “winner” after the results appear. Keep it short enough that the person launching the test will actually use it.
| Charter field | What to record | Example |
|---|---|---|
| Decision | The choice the test will make | Keep a 10% or 20% three-unit discount |
| Hypothesis | Change, mechanism, and expected result | 10% preserves uptake while improving contribution per visitor |
| Control and variant | Exact customer-visible difference | A: 10% off; B: 20% off |
| Audience and scope | Products, Markets, pages, and traffic included | US product-page traffic to three named products |
| Allocation | Intended traffic share | 50% A / 50% B |
| Primary business metric | The number used for the economic decision | Contribution profit per variant visitor |
| Minimum worthwhile effect | Smallest lift worth implementing | At least $0.05 more contribution per visitor |
| Diagnostic metrics | Numbers that explain the primary result | Visitor conversion, bundle take rate, AOV, revenue per visitor |
| Guardrails | Outcomes that must not deteriorate beyond a stated limit | Refund cost, cancellations, stockouts, support contacts, cart errors |
| Stop plan | Sample, minimum runtime, maximum runtime, and no-peeking rule | Store-specific plan set before launch |
| Maturation rule | When orders and return costs are complete enough to analyze | Re-export after the declared order lag; finalize after return allowance matures |
| Freeze list | Changes that require a restart or an explicit analysis note | Prices, promotions, product scope, theme, app settings, traffic mix |
The minimum worthwhile effect is a business threshold. It answers, “How much extra contribution justifies keeping this variant?” It is related to, but not the same as, the minimum detectable effect used in a sample-size calculation.
If a new layout requires extra design work, support, or ongoing operations, include that implementation cost in the decision threshold. A statistically detectable gain can still be too small to be worth the maintenance.
Use contribution profit per visitor as the business metric
Kaching’s analytics export gives you the pieces needed to calculate visitor conversion, bundle take rate, AOV, and revenue per visitor. Those metrics answer different questions:
- Visitor conversion: Did a larger share of widget visitors place an eligible paid order?
- Bundle take rate: Of eligible orders, did more use the bundle discount?
- AOV: How much revenue did each eligible order contain on average?
- Revenue per visitor: How much gross revenue did each widget visitor produce?
- Contribution profit per visitor: How much remained per widget visitor after the variable costs included in your model?
Use this test-level definition:
Contribution profit = total revenue − COGS − payment fees − merchant-funded fulfillment and shipping − matured refund/return adjustment − other variable costs
Then divide each variant’s contribution profit by that variant’s visitor count:
Contribution profit per variant visitor = contribution profit ÷ variant visitors
The denominator matters. Kaching defines visitors as visitors who saw the deal widget on a product page. It is not a generic store-session count. Use the Kaching visitor field for both variants rather than mixing it with sessions from another analytics tool.
Do not subtract the headline discount a second time if Kaching’s total_revenue already contains the discounted selling price. Common fixed overhead can stay outside an A-versus-B comparison when it is identical for both variants. Include an expense when the offer changes it or when your decision specifically needs it.
If your cost sheet is incomplete, ShopSideK’s Shopify profit margin calculator and growth guide can help organize the inputs. The result is only as credible as the cost treatment you apply consistently to both variants.
Set up the A/B test in Kaching
Kaching’s official split-testing guide documents even traffic allocation by default, optional custom allocation, and a maximum of four variants per bundle block. For a first profit-led comparison, use A and B unless there is a strong reason to divide traffic further.
- Open the live bundle or create the bundle block you intend to test.
- Select Run A/B test in the bundle preview area.
- Keep Variant A as the current control. Create Variant B from the same starting setup.
- Change only the element named in the charter. Kaching documents tests for bundle deals, discounted prices, titles, layout, product images, and button text.
- Use an even split unless the charter explains why a custom allocation is necessary.
- Preview both variants and test product selection, quantity, cart, discount, checkout, mobile layout, and analytics events before publishing.
- Save the exact launch time, settings, and screenshots for both variants.
Bundle visibility—including Markets and main-product selection—and the schedule cannot be different between test variants according to the current Kaching guide. If your hypothesis requires one of those changes, this in-block test is not the right design.
Kaching also says the app cannot increase the original product price. A price comparison must work through the original product price and the discount architecture, so confirm the customer-visible price and checkout amount in both paths.
If this is the test layer your store needs, unlock 20% OFF Kaching for your first 3 months. Use the ShopSideK form to receive the code, then finish the charter and QA before sending campaign traffic.
Set the stop rule without borrowing a universal number
Kaching’s winner explanation says each variant needs at least 10 orders to be considered and that the app uses a z-test on conversion rate. Ten orders is therefore an eligibility floor in Kaching’s documented logic. It does not tell you that the test has enough data to detect the lift your store cares about.
The required sample changes with:
- the control conversion rate;
- the smallest improvement worth acting on;
- the false-positive and power settings used in the plan;
- the allocation between variants;
- the variability and maturity of the business metric;
- the amount of eligible traffic available each week.
Set the plan before launch. Run long enough to reach the preplanned exposure and cover a complete business cycle that matters for your store, such as both weekdays and weekends. Also set a maximum runtime. A test should not continue indefinitely because the current favorite has not received a winner badge.
Do not stop the first morning a metric turns green. Repeatedly checking a conventional fixed-sample result and stopping on a favorable fluctuation can inflate false positives. If you need a detailed sample calculation or a sequential-testing method, treat that as a separate statistical design task; the Kaching 10-order rule does not replace it.
Check experiment integrity before comparing outcomes
A clean calculation cannot rescue a contaminated test. Complete these checks first.
Compare actual allocation with the plan
A 50/50 allocation will not necessarily produce identical counts, but a large unexplained imbalance deserves investigation. Microsoft’s sample-ratio-mismatch research treats an unexpected allocation ratio as a data-quality symptom that should be diagnosed before effects are trusted.
Do not invent a universal acceptable percentage difference. Compare the observed split with the planned allocation using an appropriate check, then inspect launch timing, missing tracking, product or Market exposure, bot/internal traffic, and variant-specific failures when the deviation is suspicious.
Account for browser-based assignment
Kaching stores a kaching_session_id in browser local storage. The same browser and device should return to the same variant while that value remains. A shopper who changes browser or device, or clears local storage, may be assigned again.
That is a real limit on person-level consistency. Do not describe the experiment as persistent customer-level assignment across devices. It may still be useful for a product-page offer decision, but the limitation belongs in the analysis record.
Freeze competing changes
Keep product price, bundle visibility, inventory treatment, theme placement, cart logic, other discounts, and campaign mix stable. If a change affects both variants equally, record it. If it can affect them differently or alters the tested mechanism, restart the test or label the result inconclusive.
Run operational guardrails
A profit estimate is not permission to ship a broken experience. Compare refund/return cost, cancellations, stockouts, customer-service contacts, incorrect quantities, discount failures, and cart or checkout errors. A variant that clears the profit threshold but breaches a hard operational guardrail does not pass.
Export and aggregate the complete Kaching test period
In Kaching, open Analytics, select the date range and deal block, and choose Export CSV. The current export includes a daily row for each deal and A/B variant, with fields such as visitors, eligible_orders, bundle_orders, total_revenue, aov, and revenue_per_visitor.
Build one summary row per variant:
- Filter to the intended deal and test window.
- Confirm one store currency and the expected variant labels.
- Sum visitors, eligible_orders, bundle_orders, and total_revenue by variant.
- Recalculate conversion, bundle take rate, AOV, and revenue per visitor from the summed numerators and denominators.
- Do not average the daily percentage or AOV columns. A low-traffic day should not receive the same weight as a high-traffic day.
- Reconcile the analysis dates with the charter and note any missing or partial launch day.
Kaching records visitors on the day they visit and paid orders/revenue on the day the order is paid. A short daily row can therefore show visitors without orders or orders without visitors. Use the complete period and apply the same declared order-lag cutoff to both variants.
The help page identifies the reporting timezone as Europe/Vilnius/CET. Use the app’s dates consistently rather than translating them into an assumed fixed UTC offset. If campaign timestamps live in another timezone, document the conversion once and keep it consistent.
Add the variable costs Kaching does not export
The documented Kaching CSV stops at revenue metrics. It does not contain the following merchant economics:
| Cost input | Preferred source | Common mistake |
|---|---|---|
| COGS | Product/variant cost at sale or accounting system | Using today’s cost for an older test without checking changes |
| Payment fees | Payout or payment-provider records | Applying one headline percentage to every payment method |
| Merchant-funded shipping | Carrier, 3PL, or order records | Subtracting customer-paid shipping twice |
| Pick, pack, and packaging | Warehouse or 3PL schedule | Assuming larger bundles cost the same to fulfill |
| Refund and return adjustment | Matured refund, return, and reverse-logistics records | Subtracting both full COGS and recovered inventory value |
| Other variable cost | Store-specific records | Mixing fixed overhead into only one variant |
Map costs to the same variant definition used in the Kaching export. If you cannot join every order to a variant, create a documented high/low cost range for each offer. Do not use one blended average when the deeper bundle clearly ships more units or has a different return pattern.
For returns, use a matured net economic adjustment: refunded revenue plus reverse logistics and nonrecoverable product cost, less any recovered value, using the store’s established accounting treatment. If returns have not matured, either make a provisional decision with a historical allowance or wait. When a realistic return range can flip the winner, the result is inconclusive.
Worked example: the conversion leader loses on profit per visitor
The following Bundle Experiment Winner model is hypothetical. It compares equal exposure to a 10% offer and a 20% offer. All amounts are USD, and the return adjustment is assumed to be mature.
| Metric | Variant A: 10% | Variant B: 20% |
|---|---|---|
| Variant visitors | 10,000 | 10,000 |
| Eligible orders | 440 | 520 |
| Bundle orders | 300 | 390 |
| Visitor conversion | 4.40% | 5.20% |
| Bundle take rate | 68.18% | 75.00% |
| Total revenue | $15,840.00 | $17,680.00 |
| AOV | $36.00 | $34.00 |
| Revenue per visitor | $1.5840 | $1.7680 |
| COGS | $5,280.00 | $7,280.00 |
| Payment fees | $591.36 | $668.72 |
| Fulfillment and shipping | $2,860.00 | $3,640.00 |
| Matured return adjustment | $475.20 | $707.20 |
| Contribution profit | $6,633.44 | $5,384.08 |
| Contribution profit per visitor | $0.6633 | $0.5384 |
Variant B looks stronger in three places. Its visitor conversion is 18.18% higher, its bundle take rate is higher, and its revenue per visitor is 11.62% higher. Variant A has the higher AOV, but that metric alone is also incomplete.
After variable costs, A produces $0.6633 per visitor against B’s $0.5384. B is down 18.83% on the primary business metric and leaves $1,249.36 less contribution over the same 10,000 visitors.
Calculate the cost-to-tie threshold
B would need $1,249.36 less variable cost to tie A at this exposure. Spread across B’s 520 eligible orders, that is about $2.40 per order.
Ask a concrete sensitivity question: could the cost model plausibly overstate B’s costs by at least $2.40 per order relative to A? If the best documented correction is only $0.60, A remains the economic winner. If a missing fulfillment rebate or recovered-return value is worth $2.75, repair the data before deciding.
The $2.40 is a business sensitivity threshold, not a confidence interval. The aggregate export does not provide visitor-level contribution variance, so this model does not calculate statistical significance for profit.
Return WINNER, CONTINUE, or INCONCLUSIVE
Do not force every test into an A-or-B answer.
| Result | Use it when | Action |
|---|---|---|
| WINNER | Data-quality checks pass; the prewritten stop condition is met; the contribution-per-visitor gap clears the minimum worthwhile effect and plausible cost uncertainty; guardrails pass | Roll out the economic winner and keep monitoring the guardrails |
| CONTINUE | Integrity checks pass, the planned maximum has not been reached, and more exposure can still answer the prewritten decision | Continue without changing variants or the success rule |
| INCONCLUSIVE | Allocation or measurement is compromised, a material setting changed, costs/returns are too immature, or the gap remains inside uncertainty at the planned stop | Keep the control or choose the simpler option, repair the design, and write a stronger next hypothesis |
Kaching may call one variant a significant conversion winner while your contribution model favors the other. Record both statements. Do not claim that Kaching’s conversion z-test proves the profit difference.
The reverse can also happen: one variant has higher estimated contribution per visitor while Kaching shows no clear conversion winner. Check whether the contribution gap clears your economic and cost-sensitivity thresholds. If it does not, the honest result is not “almost a winner.” It is inconclusive.
When both offers perform similarly, simplicity has value. Keeping the control avoids rollout risk, and choosing the easier variant can be reasonable when the planned analysis cannot distinguish them. State that as an operational decision, not a statistically proven lift.
When Kaching is the right testing layer
Kaching is a documented fit when both candidates can live in one bundle block and you need to compare deals, discounted prices, titles, layouts, images, or button text. It handles the storefront variants, browser-based assignment, traffic allocation, and variant analytics that make this workflow practical.
It is not the complete experiment finance system. It does not add merchant cost fields to the documented export, guarantee cross-device identity, make visibility or schedule variant-specific, or establish a native no-bundle control in the documentation reviewed for this article.
If those limits match your test, review Kaching and claim 20% OFF for your first 3 months. If you prefer the direct route, view Kaching on the Shopify App Store; the ShopSideK form is the route for the discount code.
If you need a different bundle architecture rather than a variant test, compare the approaches in ShopSideK’s best Shopify bundle apps guide before adding another testing layer.
Frequently asked questions
How long should a Shopify bundle A/B test run?
Long enough to reach a preplanned sample and cover the store’s relevant business cycle, but not indefinitely. The correct duration depends on baseline conversion, the smallest worthwhile lift, allocation, eligible traffic, and the analysis method. Kaching’s 10-order-per-variant consideration threshold is not a complete duration rule.
Should I choose the variant with the highest conversion rate?
Only when conversion is the declared primary objective and the economics and guardrails remain acceptable. For a bundle-offer decision, use contribution profit per variant visitor as the business metric and conversion as diagnostic evidence. A deeper discount can convert more visitors while leaving less contribution.
Can I test four Kaching bundle variants at once?
Kaching documents a maximum of four variants per bundle block. Two are usually a better starting point for a profit-led test because each receives more exposure and the difference is easier to interpret. Use more variants only when the traffic plan and analysis support them.
Can Kaching A/B test different prices?
Yes, when the comparison uses discounts. Kaching says it cannot increase the original product price through the app. Verify the original price, discount method, displayed savings, cart, and checkout amount for every variant.
Why do daily Kaching conversion rates look inconsistent?
Kaching records visitors on the visit day and paid orders and revenue on the paid day. Daily numerators and denominators may therefore describe activity that matured on different dates. Sum the complete test-period fields by variant and recalculate the ratios instead of averaging daily percentages.
What if Kaching shows no clear winner?
Follow the charter. Continue only when the planned test still permits more exposure and the data is clean. At the planned stop, use the contribution model, cost sensitivity, and guardrails. If the gap remains too uncertain, keep the control or simpler option and label the experiment inconclusive.


