To choose Shopify quantity break tiers, do not start with a standard Buy 2, Buy 3, Buy 5 ladder. Start with the quantity customers already buy, find the smallest meaningful step above it, then reject any candidate that creates too much supply, crosses a shipping or packaging cliff, strains inventory, or fails your contribution floor.
That usually makes Buy 2 the first quantity worth examining—not the automatic winner. Buy 3 can be stronger when two-unit orders already happen at full price. Buy 4 or higher needs evidence that customers can use, store, share, or gift that quantity without creating an operational problem.
Kaching Bundles can display and test the breakpoints that survive this process. It cannot infer the correct quantities from your store economics or product-use pattern.
ShopSideK Verdict
Best starting rule: Choose the nearest meaningful quantity above current purchasing behavior that customers can realistically use and that still passes shipping, inventory, and contribution checks.
Do not confuse reach with demand: Orders exactly one unit below a candidate are the closest upgrade pool. They are not a conversion forecast. Orders already at or above the candidate show where you may discount behavior that already exists.
Kaching fit: Kaching Bundles is useful when you want a visible multi-tier product-page offer and the ability to compare whole-offer variants. Native Shopify can be enough for one simple minimum-quantity discount.
Decision path: Build the quantity baseline → screen Buy 2, Buy 3, Buy 4+ → apply product-use and operational vetoes → test only the surviving candidates.
ShopSideK deal: 20% OFF Kaching for the first 3 months.
The quick answer: when Buy 2, Buy 3, or Buy 4 makes sense
| Candidate | Stronger starting evidence | Main risk to check |
|---|---|---|
| Buy 2 | Most relevant orders contain one unit; a second unit has an obvious near-term use; delivery cost changes little | Existing two-unit orders receive a discount, so reach and cannibalization exposure are both high |
| Buy 3 | Two-unit orders form a meaningful group; three units still fit the normal use horizon; shipping stays on the same cost step | The third unit may pull forward the next purchase rather than create incremental demand |
| Buy 4+ | Customers already show bulk, household, sharing, gifting, event, or replenishment behavior; the product stores well | Large behavior jump, storage or shelf-life burden, package changes, inventory pressure, and a higher cash commitment |
This table is a shortlist, not a template. The same product can support Buy 2 on one store and Buy 4 on another because order distribution, price, customer mix, delivery economics, and replenishment behavior differ.
Why copied quantity tiers fail
A quantity break does two jobs at once. It tries to move some customers into a larger basket, and it reduces the price for some customers who may already have bought that quantity.
Those effects are easy to blur when a guide recommends 2/3/5 or 2/4/6 without looking at the store. If 25% of your clean baseline orders already contain two units, a Buy 2 discount reaches a large audience—but it also exposes a quarter of baseline orders to a price reduction on behavior that already happens. If almost nobody buys two units, Buy 3 may avoid more existing orders but ask for too large a jump.
The product changes the answer too. Four rolls of a frequently used household item are not equivalent to four bulky, slow-use products. Price and storage alter the shopper’s commitment. Packaging and carrier rates alter the merchant’s cost.
Research on purchase acceleration gives another reason not to treat the extra units as automatically incremental. Earlier work distinguished between promotions increasing purchase quantity and shortening the time until a purchase. Later research found that stockpiling can also change consumption for some products, with product convenience and salience affecting the result. Those studies are not Shopify tier benchmarks. They support a narrower rule: measure repeat timing and product-specific use instead of assuming every larger order is either pure new demand or pure pull-forward. See Neslin, Henderson and Quelch’s purchase-acceleration research and Chandon and Wansink’s stockpiling study.
Step 1: Build a product-level quantity baseline from Shopify orders
Average units per order is too broad for this decision. A storewide average can combine one-unit hero-product orders, multi-item accessory orders, subscriptions, wholesale purchases, gifts, and unrelated products. You need the order-quantity distribution for the exact product or eligible group that will receive the offer.
Shopify lets you export orders by date as a CSV. The export includes order identifiers and fields such as financial status, created date, line-item quantity, line-item name and SKU, cancelled date, refunded amount, and line-item discount. Shopify also notes that an order with multiple line items appears across separate rows.
Use a representative period with enough paid orders to reveal a stable pattern. There is no universal 30-, 60-, or 90-day window. A replenishable product may need at least a normal repeat cycle; a seasonal product may need a comparable season. Exclude or label periods dominated by a different promotion, stockout, wholesale event, subscription migration, or unusual campaign.
Build the baseline in this order:
- Define the same eligibility unit the future offer will use. If the offer treats variants separately, analyze them separately. If it legitimately pools selected variants or products, group only those eligible lines.
- Group line-item rows back to the order using the order identifier. Sum eligible quantity within each order.
- Apply one written rule for cancelled, refunded, edited, test, wholesale, subscription, and zero-value orders. Do not change the rule because it makes a tier look better.
- Count relevant orders containing one eligible unit, two, three, four, and so on. Keep a four-or-more bucket only if low volume makes individual higher quantities too sparse.
- Divide each quantity count by total relevant orders to create the baseline share.
Shopify’s Orders over time report can provide average units ordered and AOV as context. Its product order and reversal fields can help you spot returns. Those aggregates do not replace the order-level histogram.
Keep the scope consistent with the future offer
Suppose a product has three sizes. Pooling all three in the analysis is wrong if the deal requires three units of one variant. Keeping them separate is wrong if the actual offer lets a shopper mix those sizes toward one threshold.
Write the qualification rule above your worksheet before counting anything. If you cannot reproduce the proposed offer’s counting logic in the baseline, label the analysis provisional.
Step 2: Read two different signals for every candidate quantity
For a candidate quantity q, calculate two shares.
Direct-upgrade pool at q = baseline orders with exactly q−1 eligible units ÷ relevant baseline orders
This finds the customers who were one unit away. For Buy 2, it is the one-unit order share. For Buy 3, it is the two-unit share. It is the closest pool to move, but it does not predict how many customers will accept the offer.
Existing qualifying exposure at q = baseline orders with q or more eligible units ÷ relevant baseline orders
This finds baseline orders that would already meet the threshold. It shows where the new discount could subsidize existing behavior. It is not the exact cannibalization rate because the offer can also change conversion, product choice, timing, and order composition.
The two signals often pull in opposite directions:
- Buy 2 may have the largest direct-upgrade pool and the largest existing exposure.
- Buy 3 may reach fewer nearby orders but protect more two-unit full-price purchases.
- Buy 4 may have little existing exposure but require a behavior jump too large to be credible.
Do not subtract the two percentages to make a “net demand score.” They measure different things and do not share a common economic value. A one-unit upgrade that earns healthy contribution is not interchangeable with a discounted order that would have happened anyway.
Step 3: Convert each candidate into a stock-up horizon
Quantity becomes easier to judge when you express it as time or use cycles.
Stock-up horizon at q = q × usable days or cycles per unit
If one unit normally covers about 30 days, Buy 2 supplies roughly 60 days and Buy 4 roughly 120. For products shared by several people, use household-level consumption. For products with variable use, calculate a conservative and an optimistic range instead of hiding the uncertainty inside one precise number.
Then compare that horizon with observed repeat behavior.
Reorder displacement ratio = stock-up horizon at q ÷ observed median reorder interval
A ratio of 1 means the candidate supplies roughly one observed reorder interval. A ratio of 2 covers roughly two. These are descriptions, not universal pass/fail thresholds. A merchant may intentionally encourage a two-cycle stock-up before a seasonal deadline. The same ratio may be unacceptable for a bulky product, a short-shelf-life item, or a category where delayed repeat purchase damages retention economics.
Use the best store-specific evidence available: actual repeat timing, product instructions, package contents, customer-service questions, subscription cadence, or a documented usage range. Do not invent a consumption rate to justify a tier. If the range is too wide to make the decision, mark the candidate TEST or HOLD.
Add shelf life, storage, and cash commitment
The stock-up horizon still misses three shopper constraints:
- Shelf life: Will the quantity remain usable through the proposed horizon after accounting for transit and time already held in inventory?
- Storage: Can the intended customer reasonably store the physical volume?
- Cash commitment: Even with a better unit price, is the total outlay plausible for this audience?
These checks explain why a higher tier can look attractive in a spreadsheet but remain weak on the product page. A breakpoint should describe a useful purchase, not merely a mathematically larger basket.
Step 4: Let operational and profit constraints veto a tier
Reach tells you where to investigate. It should not overrule a hard constraint.
Find shipping and packaging cliffs
Price each candidate as an actual parcel. Record the product quantity, packaging type, packed dimensions, weight, label cost, pick/pack time, inserts, and any special handling.
Incremental delivery cost at q = shipping, packaging and fulfillment cost at q − the same cost at q−1
A candidate that needs a larger box or crosses a carrier step may be materially more expensive than the preceding quantity. Shopify’s Shipping labels by order report can include the selected package, dimensions, total weight, shipping price, and label cost. Use your warehouse or 3PL invoice when it is the better source of actual cost.
Do not average away a cliff. If Buy 4 moves to a different carton, cost Buy 4 with that carton.
Stress the inventory consequence
A larger tier consumes more units per successful order. Model a range of adoption rather than presenting one guess as a forecast.
Projected inventory coverage at q = available eligible units ÷ projected daily eligible-unit demand under the candidate
Use scenarios such as low, expected, and high uptake, with the assumptions written beside them. No universal number of inventory days is safe. The decision depends on supplier lead time, restock reliability, seasonality, full-price demand, and the cost of a stockout.
Require the profit floor to pass
Run the planned discount for each surviving quantity through your contribution-floor calculation. Include product cost and the variable expenses that change with the order, including the shipping or fulfillment step you just measured.
This article does not choose the discount percentage. The rule here is simpler: a breakpoint that fails the required contribution floor is not a candidate, regardless of how reachable it appears. If you have not set that floor yet, classify the quantity HOLD rather than approving it with incomplete economics.
Quantity Breakpoint Candidate Worksheet: Buy 2 vs Buy 3 vs Buy 4
The following example shows how the signals work together. It is hypothetical—not ShopSideK store data, a Kaching result, or a prediction for your store.
Assume 100 paid, noncancelled baseline orders for one consistently defined product group:
- one unit: 68 orders;
- two units: 21 orders;
- three units: 8 orders;
- four or more units: 3 orders.
One unit is assumed to cover 30 days of normal use. The observed median reorder interval is 45 days, and shelf life is 180 days. The warehouse cost check finds a $4.50 package-and-shipping step between three and four units. Inventory coverage figures below come from an explicitly modeled adoption scenario, not the baseline shares.
| Candidate | Direct-upgrade pool | Existing qualifying exposure | Stock-up horizon | Reorder displacement ratio | Incremental delivery cost | Modeled inventory coverage | Profit-floor check | Class |
|---|---|---|---|---|---|---|---|---|
| Buy 2 | 68% | 32% | 60 days | 1.33 | $0.75 | 45 days | Pass | CORE |
| Buy 3 | 21% | 11% | 90 days | 2.00 | $0.75 | 32 days | Pass | TEST |
| Buy 4 | 8% | 3% | 120 days | 2.67 | $4.50 | 18 days | Fail | REJECT |
Why Buy 2 is CORE, not “proven best”
Buy 2 is the nearest meaningful behavior change for the 68 one-unit orders, and it passes the hypothetical use, shipping, inventory, and profit checks. That makes it the cleanest control candidate.
Its weakness is visible: 32% of baseline orders already contain at least two units. A deep Buy 2 discount could give away contribution on existing behavior. CORE means “best-supported initial candidate,” not “launch at any discount” and not “expected winner.”
Why Buy 3 is TEST
Twenty-one percent of baseline orders sit one unit below Buy 3, while only 11% already meet it. That is a more attractive exposure balance than Buy 2. But three units cover 90 days—twice the observed median reorder interval—and the modeled inventory cushion is thinner.
Buy 3 remains plausible, but the store needs to learn whether customers value a three-unit stock-up and whether the order gain repays delayed repeat purchases and added unit demand.
Why Buy 4 is REJECT
Buy 4 has the lowest existing exposure, but that does not rescue it. The candidate crosses a $4.50 delivery-cost step, leaves only 18 modeled days of inventory, and fails the contribution floor. It also asks most customers to make a much larger jump.
The worksheet therefore does not produce a 2/3/4 ladder. It produces one control candidate, one test candidate, and one rejection.
How to classify your own candidates
Use four labels so uncertainty remains visible:
- CORE: The closest credible behavior change. Product use, delivery, inventory, and contribution all pass. CORE is the initial control candidate, not a guaranteed winner.
- TEST: The candidate is plausible, but reach, current exposure, stock-up horizon, or reorder timing creates a real uncertainty that only a controlled comparison can resolve.
- HOLD: The quantity may fit, but a missing input or temporary constraint prevents a responsible launch. Common examples are unknown packed cost, unreliable replenishment, or no contribution floor.
- REJECT: The candidate fails a hard product-use, shelf-life, storage, shipping, inventory, or profit condition.
Avoid weighted scoring unless you can justify what one point means economically. A 90-day supply cannot be “offset” by a large q−1 share when customers cannot store the product. Keep hard vetoes separate from softer evidence.
Practical rules for Buy 2, Buy 3, and Buy 4+
Start by examining Buy 2 when:
- one-unit orders dominate the clean baseline;
- a second unit has an immediate use, backup, household, gifting, flavor, color, or replenishment rationale;
- packed delivery cost changes little from one to two;
- the total purchase remains accessible; and
- the planned discount passes despite existing two-unit exposure.
Do not choose Buy 2 merely because it is the easiest threshold to explain. If two-unit orders already form a large profitable segment, consider whether a smaller incentive, a higher breakpoint, or no discount protects more contribution.
Put Buy 3 into consideration when:
- two-unit orders are meaningful enough to form a credible direct-upgrade pool;
- three units remain useful within a defensible stock-up horizon;
- the third unit does not trigger a packaging, shipping, handling, or inventory problem; and
- the economics compensate for possible reorder pull-forward.
Buy 3 is especially worth testing when Buy 2 would heavily subsidize full-price multi-unit demand. It is still a larger commitment, so lower existing exposure alone is not enough.
Require stronger evidence for Buy 4 or more when:
- order data already contains credible bulk behavior;
- customers consume, share, gift, resell, use across locations, or replenish the product at that scale;
- shelf life and storage are comfortable;
- the parcel stays on a workable cost step;
- inventory can absorb the unit acceleration; and
- the contribution floor passes at the planned price.
“Best value” copy cannot make an unusable quantity useful. If the business case depends on an assumed customer behavior that is absent from orders, repeat timing, support conversations, or product context, classify the tier TEST or HOLD.
Do not publish every candidate that passes
The worksheet creates a candidate set. It does not decide how many rows your quantity selector should contain.
Two candidates can both pass the economic and operational screen while serving nearly the same shopper decision. Displaying both may add complexity without adding a meaningful choice. Conversely, one simple threshold may be all the product supports.
Choose the smallest visible set that represents distinct, useful purchase missions. Treat tier-count and product-page choice complexity as a separate UX decision; do not turn this worksheet into a rule that every CORE and TEST candidate must appear at once.
Implement the surviving quantities in Shopify or Kaching
Shopify’s native amount-off discount setup supports a minimum quantity of items and can scope a discount to specific products or collections. That can be enough when you have one straightforward threshold and do not need a visible product-page ladder or whole-offer split test.
Kaching is the better fit when the shortlist becomes a visible quantity-break offer. Its documented Quantity Break type supports a percentage or fixed discount per item or a custom total price for the selected quantity. Kaching also documents A/B tests with up to four variants, including different discount tiers, prices, titles, images, and layouts.
That capability does not remove the work above. Kaching does not know your product’s use rate, storage burden, packed cost, supplier lead time, inventory policy, or required contribution. Use it to execute and compare defensible candidates—not to replace the candidate screen.
Claim 20% OFF Kaching for Your First 3 Months →
If you prefer to bypass the ShopSideK form, you can view Kaching on the Shopify App Store.
Measure the ladder without letting AOV make the decision
After launch, rebuild the order-quantity distribution for the offer period and compare it with the clean baseline. Watch:
- share of relevant orders at each quantity;
- conversion rate for comparable eligible traffic;
- contribution profit per assigned visitor;
- incremental shipping and fulfillment cost;
- inventory consumption and stockouts;
- return, cancellation, or reversal behavior; and
- reorder timing for cohorts exposed to the offer.
Kaching’s documented analytics CSV export includes daily rows for deals and A/B variants, with visitors, eligible orders, bundle orders, revenue, AOV, and revenue per visitor. The documented fields do not include product COGS, merchant contribution floor, or a separate profit result for every selected quantity. Combine the variant data with Shopify order quantities and your cost model.
Kaching also states that variant assignment is stored in the visitor’s browser. A returning shopper can receive a different assignment after switching browser or device or clearing storage. Treat the test as a whole-offer visitor experiment under those limits, not as a permanent customer-level split.
Do not declare a winner because AOV increased. A higher tier can raise AOV while shipping, discount cost, inventory pressure, or delayed reorders reduce contribution. The final comparison belongs at contribution profit per comparable visitor versus the no-offer control.
Shopify quantity tier FAQs
What should the first Shopify quantity break be?
The first candidate should be the nearest meaningful quantity above current purchasing behavior that customers can use and that passes delivery, inventory, and contribution constraints. Often that means examining Buy 2 first, but it does not make Buy 2 universally correct.
Is Buy 2 always better than Buy 3?
No. Buy 2 usually has a smaller behavior gap, but it can expose more existing full-price multi-unit orders to a discount. Buy 3 may protect more baseline contribution, yet fail if the third unit creates too much supply, cost, or commitment.
How much order history do I need?
Use a representative period long enough to capture the product’s normal demand and, when relevant, a repeat cycle or comparable season. There is no universal number of days or orders. Separate periods distorted by stockouts, wholesale events, subscriptions, or another promotion.
Can I use average units per order instead of a quantity distribution?
Use it as context, not as the deciding input. Two products can have the same average units but very different shares at one, two, three, and four units. The distribution reveals the direct-upgrade pool and existing qualifying exposure.
How do I estimate cannibalization before launch?
The baseline q-or-higher share shows how much existing behavior is exposed to the new threshold. It does not estimate exact cannibalization. A controlled test with a no-offer baseline and contribution measurement is needed to observe the net effect.
Can Kaching choose the best quantity-break tiers automatically?
Not from the capabilities documented in its current help center. Kaching can implement quantity tiers and compare whole-offer variants. The merchant still supplies order distribution, use, cost, inventory, and profit constraints.
Methodology
ShopSideK built this framework from current Shopify order-export, order-report, shipping-label, and discount documentation; current Kaching Quantity Break, A/B testing, and analytics documentation; and primary research on purchase acceleration and stockpiling. Product and platform claims were fact-checked on July 22, 2026.
The worksheet and classifications are ShopSideK’s decision method. The 100-order example is entirely hypothetical and demonstrates the calculations rather than a merchant result. Behavioral studies are used to explain why repeat timing is uncertain, not to predict a Shopify conversion outcome.


