The first Shopify bundle test should resolve the most valuable uncertainty you can measure cleanly. That may be the offer mechanic, product mix, discount, tier structure, message, or layout. It is not automatically the easiest setting to edit.
Before choosing any of them, rule out four problems that should not consume A/B-test traffic: missing measurement, broken cart or checkout behavior, unsafe economics, and an experiment your traffic cannot support. Then use evidence from the current bundle to choose one customer-visible variable.
Kaching Bundles & Upsells can execute many of those comparisons inside the bundle block. Its current documentation covers deal, discounted-price, tier, title, image, layout, and button-text variants. The app gives you the testing layer; it does not decide which uncertainty matters most to your store.
⚡ ShopSideK Verdict
Start with: The unresolved decision backed by the strongest evidence—not a universal “price first” or “layout first” rule.
Best for: Merchants with a functioning bundle, several plausible improvements, and enough attention or traffic for only one experiment at a time.
Kaching fit: Kaching Bundles & Upsells can compare the selected bundle variants after measurement, economics, transaction QA, and test feasibility pass.
Decision path: Measure → fix defects → protect economics → check feasibility → prioritize one hypothesis → run the test.
Do not A/B test these four problems
An experiment is useful when both variants are safe, functional choices and the result will settle a real decision. It is wasteful when one side is broken, unmeasured, financially unacceptable, or impossible to evaluate with the available audience.
Run these gates in order.
| Gate | Question to answer | If the answer is no |
|---|---|---|
| Measurement | Can you observe eligible offer exposure and the funnel or business outcome relevant to the hypothesis? | MEASURE FIRST. Repair the baseline before creating a variant. |
| Transaction | Does the control show the right products, variants, quantities, price, discount, and cart or checkout result on priority devices? | FIX FIRST. A defect is not a customer preference. |
| Economics | Are the control and proposed variant inside an approved contribution and operational-risk boundary? | REDESIGN. Do not expose shoppers to an offer you would reject even if it converted. |
| Feasibility | Can the eligible audience support the planned comparison within an acceptable business window? | CHOOSE A LARGER QUESTION OR DO NOT FORMALLY TEST. A cosmetic test does not become useful because it is easy to launch. |
Shopify’s research-led testing guidance makes a similar distinction between issues that should be fixed and choices that deserve testing. It also recommends using analytics and qualitative research to move from an observed problem to a probable cause and then to a proposed change. That sequence matters because “the number is low” is not yet a hypothesis.
A failing gate can be more valuable than a test idea
Suppose mobile visitors add the wrong variant when they choose a three-pack. The first action is not to compare button copy. Fix the product-form or cart behavior and verify it across the affected variants.
Or suppose a deeper discount falls below your contribution floor. Testing whether shoppers like it does not make the economics acceptable. Redesign the discount or product mix first.
The same rule applies to measurement. If you cannot distinguish people who saw the bundle from people who could not see it, a low order count cannot tell you whether the offer is hidden or unpersuasive.
Use evidence to find the unresolved question
A test backlog often starts with nouns: price, headline, image, tier, badge. A useful hypothesis starts with an unanswered question.
For example:
- Do shoppers reject the three-unit commitment, or do they fail to understand why they would need three?
- Does the deeper discount create enough additional bundle uptake to pay for itself?
- Is a complementary bundle weak because the products do not belong together, or because the relationship is not explained?
- Is low mobile uptake caused by presentation, or is the selector difficult to use?
The first version identifies a setting. The second identifies what you intend to learn.
Treat metrics as locators, not diagnoses
Kaching’s current Bundle Analytics documentation lists visitors, add-to-cart activity, checkout rate, conversion, bundle orders, units per order, AOV, revenue per visitor, and profit per visitor when Shopify product-cost data is present. Its CSV export also separates daily deal and A/B-variant rows.
Those fields can show where behavior changes. They cannot, by themselves, prove why.
If exposure is healthy and add-to-cart is weak, possible explanations include the wrong product relationship, an unrealistic quantity, insufficient value, confusing copy, or interaction friction. Choosing layout merely because add-to-cart is low skips the causal question.
Use an evidence ladder:
- DIRECT: The current bundle data and repeated customer evidence point to the same uncertainty.
- SUPPORTED: One credible signal points to the uncertainty and another source does not contradict it.
- GUESS: The idea comes from preference, a competitor screenshot, or a generic best-practice list without store-specific support.
Customer support logs, post-purchase responses, on-site questions, session review, product-use patterns, previous tests, and offer-level analytics can all contribute. One comment is not a trend. A dashboard pattern is not a motive. The strongest candidates usually combine behavioral and qualitative evidence.
Use the Bundle First-Test Diagnostic
The following diagnostic routes the observed pattern to the question you need to resolve. It does not declare the cause.
| Observed pattern | Question to resolve | First candidate to investigate | What not to assume |
|---|---|---|---|
| Offer exposure is low or unknown | Is the bundle unseen, incorrectly targeted, or simply unmeasured? | Measurement and placement | Do not rewrite the offer before confirming people encounter it. |
| Exposure exists, but selection or add-to-cart is weak | Is the product relationship relevant, the commitment realistic, and the value understandable? | Mechanic, product mix, entry tier, or value framing | Do not assume button styling caused the weak response. |
| Add-to-cart is healthy, but bundle orders lag | Is there cart friction, an eligibility error, price inconsistency, or weak value at checkout? | Transaction QA first; then price, tier, or framing if the flow is correct | Do not test around a broken discount or cart. |
| Bundle uptake is healthy, but contribution per visitor is weak | Is the incentive more expensive than the additional behavior it creates? | Discount depth, price presentation, product mix, or tier | Do not let higher AOV hide weaker economics. |
| One tier dominates and the next tier is rarely chosen | Is the next quantity unrealistic, the incremental saving weak, or its use case unclear? | Tier quantity or tier framing—one at a time | Do not change quantity, discount, badge, and default selection together. |
| Returns, stockouts, cancellations, or support contacts worsen | Is the offer creating the wrong order rather than a better order? | Product mix, choice architecture, or operational redesign | Do not optimize bundle take rate while ignoring downstream cost. |
| Results diverge sharply by device or market | Is the experience functioning and comparable in each context? | Segment diagnosis and QA | Do not average a local defect into a global preference. |
Placement is not always the first presentation test
If shoppers rarely encounter the bundle, placement may be the right problem. If exposure is already healthy, moving the widget is a weaker explanation than an offer or value question supported by the next funnel step.
There is also a tooling boundary. Kaching documents layout changes inside a variant, but bundle visibility—including Markets and main-product selection—and the schedule cannot be split-tested in its built-in A/B feature. A placement or targeting hypothesis may therefore need a different test design.
Prioritize hypotheses without a fake precision score
ICE, PIE, and other numeric frameworks can help teams sort long backlogs, but an uncalibrated total can conceal weak evidence. Giving “impact” a four instead of a three does not make lift predictable.
Use ordered labels and explicit tie-breakers instead.
Fill one row for every candidate
| Field | Allowed labels or entry | Question |
|---|---|---|
| Decision | Plain-language action | What will you do differently if the variant wins? |
| Evidence | DIRECT / SUPPORTED / GUESS | What store-specific observation supports the change? |
| Customer-visible variable | One exact difference | Can a merchant describe A versus B in one sentence? |
| Decision impact | FOUNDATIONAL / MATERIAL / COSMETIC | Does the result change the mechanic, products, economics, tier architecture, or only presentation detail? |
| Economic exposure | Low / guarded / unacceptable | What contribution, return, inventory, or support downside needs protection? |
| Feasibility | READY / REDESIGN / NOT FEASIBLE | Can eligible traffic answer this question under the prewritten test plan? |
| Isolation | CLEAN / MANAGEABLE / CONFOUNDED | Will the result teach you which mechanism changed behavior? |
| Reversibility | EASY / MODERATE / HARD | How costly is it to undo the variant? |
| Implementation effort | LOW / MEDIUM / HIGH | What work is required to build and verify it? |
Apply the selection rules in this order
- Remove candidates that fail measurement, transaction, economics, or feasibility.
- Remove GUESS candidates while a DIRECT or SUPPORTED candidate remains.
- When evidence and feasibility are comparable, prefer FOUNDATIONAL over MATERIAL over COSMETIC.
- Prefer the cleaner single-variable comparison.
- Use lower risk, easier reversal, and lower effort as tie-breakers—not as substitutes for decision value.
- If no candidate survives, gather evidence instead of filling the test calendar.
This method deliberately has no total score. Its output is a defendable ordering, not a forecast.
The highest-impact idea can still lose priority
A new bundle mechanic may be foundational, but it should not automatically beat a supported price hypothesis. If the mechanic idea is a guess, requires a different fulfillment process, and cannot be isolated cleanly, the price test may produce the better next decision.
The reverse can also be true. If customer evidence consistently shows that shoppers do not want three identical units, polishing the current three-pack may be less useful than testing a two-unit entry tier or a complementary offer architecture.
Worked example: five ideas, one test slot
Consider a hypothetical supplement store with a live quantity offer:
- one bottle at full price;
- two bottles at 10% off;
- three bottles at 20% off.
The offer receives consistent exposure and adds the correct products and discounts to cart. All existing and proposed prices remain above the merchant’s approved contribution floor. Two- and three-bottle purchases occur, but the merchant believes the three-bottle discount may be giving away more contribution than necessary.
The team proposes five tests.
| Candidate | Evidence | Impact | Isolation | Decision |
|---|---|---|---|---|
| Reduce the three-bottle discount from 20% to 15% | DIRECT: tier uptake exists; economics show meaningful discount cost | MATERIAL | CLEAN | Determines the price of an existing tier |
| Preselect the three-bottle tier | GUESS | MATERIAL | CLEAN | Changes the default |
| Replace “Best Value” with “Most Popular” | GUESS | COSMETIC | CLEAN | Changes one label |
| Add a lifestyle image | GUESS | COSMETIC | CLEAN | Changes visual framing |
| Move the widget higher | WEAK: exposure is already healthy | MATERIAL | MANAGEABLE | Changes placement |
The discount-depth test wins priority. Not because price always comes first, but because this store has:
- direct evidence that shoppers already accept the tier;
- a material economic uncertainty;
- two safe variants;
- a clean difference;
- a result that changes a durable pricing decision.
The hypothesis could be:
For eligible product-page visitors, reducing the three-bottle discount from 20% to 15% will retain enough tier uptake to increase contribution per assigned visitor. Keep the 15% variant only if it clears the prewritten minimum worthwhile contribution difference while visitor conversion, returns, cancellations, and stockouts remain within their guardrails.
The primary business metric is contribution per assigned visitor. Bundle take rate, tier mix, visitor conversion, revenue per visitor, and AOV explain the result; they do not independently choose the winner.
If the same store instead had repeated customer feedback that three bottles felt excessive and almost no one selected that tier, an entry-quantity test could move ahead of discount depth. The evidence changes the priority.
Turn the selected idea into a testable hypothesis
Before opening the app, write seven fields.
| Field | What to write |
|---|---|
| Decision | The exact action the result will authorize |
| Evidence | The behavior and customer input that created the hypothesis |
| Control and variant | One precise customer-visible difference |
| Mechanism | Why that difference should change behavior |
| Primary business metric | The economic outcome that owns the decision |
| Diagnostic metrics | Funnel and offer-behavior numbers that explain movement |
| Guardrails | Outcomes that must not deteriorate beyond an approved limit |
Use this fill-in structure:
For [eligible audience], changing [one variable] from [control] to [variant] should [mechanism], improving [primary business metric] by at least [minimum worthwhile difference] while [guardrails] remain acceptable. If that decision rule is met under the prewritten sample and runtime plan, we will [action].
“Variant B will convert better” is not enough. It does not identify why, how much difference matters, what downside is protected, or what decision follows.
Keep the audience, product scope, measurement definitions, and concurrent promotions stable where possible. If you change several customer-visible elements at once, call it a package test: you may learn which package performs better, but not which element caused the change.
What Kaching can and cannot do for the first test
According to the current Kaching A/B testing guide, a bundle block can have up to four variants with even or custom traffic allocation. Visitor assignment is stored in the browser. Returning visitors can be reassigned if they switch device or browser or clear local storage.
Kaching documents the following testable changes:
- bundle deals and discounted prices;
- discount tiers;
- titles and button text;
- product images;
- bundle-block layouts and other bundle setup choices.
Important boundaries:
- bundle visibility and schedules cannot be split-tested;
- the app cannot increase the original product price through a variant;
- dividing traffic across more variants increases the exposure each arm needs;
- browser-local assignment is not a permanent customer-level identity;
- the app does not decide which hypothesis deserves the traffic.
The current Kaching analytics guide includes funnel and value metrics, and the CSV export can break out A/B variants. Its documented profit-per-visitor field subtracts Shopify product cost from revenue when product cost is available. That may still omit payment fees, fulfillment, shipping subsidy, expected returns, and other costs in your contribution definition.
Kaching’s documented clear-winner logic is based on conversion rate, with at least 10 orders per variant required for consideration and a z-test used for the comparison. Treat that as an app-specific conversion signal—not a universal stopping rule and not proof that one variant produced more contribution profit.
Once your hypothesis passes the four gates, Kaching is a strong fit when the variable lives inside its bundle setup.
Claim 20% OFF Kaching for Your First 3 Months →
What to test first in common bundle situations
These are routing examples, not universal prescriptions.
You have a new live bundle but little offer-level evidence
Measure exposure, add-to-cart, eligible orders, bundle orders, tier mix, and an economic outcome before choosing a cosmetic test. Collect customer questions or support feedback at the same time.
Shoppers see the offer but rarely choose any bundle
Investigate relevance, commitment, value clarity, and interaction friction. Use customer evidence to choose between mechanic, product mix, entry quantity, message, or UX. Low selection alone does not choose one.
Bundle uptake is strong but profit improvement is weak
Prioritize an economic question: discount depth, product mix, tier, or price representation. Keep both variants inside approved floors and choose the result using contribution per visitor plus guardrails.
One tier sells and the next tier does not
Decide whether the unresolved question is quantity or communication. Test the tier breakpoint if use rate or replenishment evidence suggests the commitment is wrong. Test framing if the quantity is defensible but shoppers may not understand the use case. Do not change both in the same diagnostic test.
Mobile performs much worse than desktop
Verify touch targets, selector behavior, variant changes, cart contents, price consistency, and checkout first. Test presentation only after the mobile path is functionally comparable.
Traffic is too low for the proposed change
Choose a larger, more decision-relevant contrast, gather qualitative evidence, concentrate eligible exposure safely, or do not run a formal test yet. Do not replace a feasibility problem with a smaller change that is even harder to detect.
Frequently asked questions
What should I A/B test first in a Shopify bundle?
Test the highest-value unresolved decision supported by current evidence and capable of a clean comparison. First rule out missing measurement, defects, unsafe economics, and inadequate traffic. There is no universal first element.
Should I test bundle price or layout first?
Test price first when the bundle works, shoppers engage with it, and the unresolved question is whether the incentive pays for the behavior it creates. Test layout when evidence points to comprehension or interaction friction. If exposure is low, investigate placement before either.
Can I test more than two Kaching variants?
Kaching currently documents up to four variants per bundle block. More variants divide the eligible audience and can make the result harder to interpret. Use additional variants only when each belongs to the same variable and the traffic plan supports them.
What if my Shopify store has low traffic?
Do not force a small cosmetic A/B test. Consider a larger hypothesis, customer interviews, support-log analysis, usability review, a longer feasible plan, or waiting. The method should match the evidence your traffic can support.
Which metric should choose the bundle winner?
Use the metric tied to the business decision. For offer economics, contribution per assigned visitor is usually more useful than AOV alone. Use conversion, bundle take rate, tier mix, AOV, and revenue per visitor as diagnostic metrics, with returns and operational outcomes as guardrails.
Can Kaching A/B test where the bundle appears?
Kaching documents layout testing inside the bundle setup, but its A/B guide excludes bundle visibility, including Markets and main-product selection. Theme-level placement or targeting may require another test design.
Choose the question before opening the test builder
Run the four gates. Write the unanswered question. Remove candidates based only on preference. Then choose the evidence-backed, high-decision-value variable you can isolate and measure safely.
Only after that should you build the control and variant. The app can divide the traffic; your hypothesis has to make the traffic worth spending.


