Brand operations

Do you retire product rules, or judge SKUs on CM2?

Once an AI recommendation layer is live, relevance-guessing rules can go. Constraints that protect cash should stay. Promote and kill SKUs on CM2, built as a reporting view from Shopify costs, not on conversion rate.

Product grid with a co-purchase graph beneath and constraint bars on top, cost waterfalls draining into cash pools or empty troughs.

Key takeaways

Retire manual recommendation rules that guess affinity (also-bought, hand-ranked relatedness). Repurpose the rest as overlays: exclusions, pins, seasonal overrides and complementary pairings. On Shopify’s native surface, complementary recommendations still have to be set up by hand, and new SKUs need a temporary boost until co-purchase data exists. Do not assume high-margin ranking pays; test it. For storefront promote-versus-kill calls, use CM2 (after COGS, shipping, payment fees and fulfilment), not conversion rate and not CM3. CM3, with allocated ads, is the right cut for channel and budget decisions, which the slot cannot make. Build CM1–CM3 as a reporting layer on Shopify’s Cost per Item plus cost lines the shop does not hold. Shopify does not post COGS journals to Xero, so this is not an accounting rebuild. CM4 is not a settled term; treat it as a house convention if you need it. Audit blank or stale costs, bundles, returns and landed cost sitting in overhead before any SKU is killed. Cutting a weak SKU usually returns less margin than the SKU P&L implies, because overhead redistributes.

The AI recommendation layer is live. The manual rules are still sitting on top of it: exclusions, pins, "customers also bought" lists typed in years ago, complementary products in Search & Discovery, seasonal boosts nobody switched off in January. The decision is not whether the model is better than a merchandiser at guessing affinity. On a catalogue with enough order history, it usually is. The decision is which of those rules were doing a relevance job, and which were doing a cash job.

If a rule encodes "this goes with that", it is a candidate for retirement. If it encodes "we cannot afford to push this", it stays, and it should sit above the model as a constraint.

Qualimero’s read of Shopify’s Ajax product recommendations endpoint is the practical starting point, with the caveat that you should check Shopify’s own developer documentation before you treat an agency post as platform law. The endpoint supports two intents, related and complementary, and returns a maximum of ten products per call. Only related recommendations are auto-generated. Complementary products have to be set up by hand in the free Search & Discovery app. That is a concrete reason some manual work cannot be deleted even if you never write another affinity rule.

Zipchat, in a vendor comparison of Shopify recommendation apps, describes business rule overlays as a category feature: teams layer manual logic above AI to exclude SKUs, prioritise high-margin products, pin seasonal items or match on attribute, without throwing away personalisation underneath. Hybrid is the usual architecture, not a halfway house you should be embarrassed about.

What conversion-led merchandising actually costs

Promotion and kill decisions still run on conversion rate in a lot of brands because it is on the dashboard, it moves in a week, and it does not require Cost per Item to be true. The commercial cost is simple: the storefront keeps feeding demand to SKUs that convert, and starves the ones that leave cash after variable cost.

Saras Analytics argues that contribution margin gets harder to calculate as SKU complexity, 3PL costs, shipping zones, returns and channel fees pile up, and that businesses relying on averages overestimate profitability. Leakage shows at order level. A conversion-only promote list scales the average, not the order.

Shopify’s own profit columns will not save you if the inputs are wrong. Ottit, writing on Shopify COGS practice, notes that Shopify tracks unit cost via the Cost per Item field on each variant and uses it for profit columns and inventory valuation snapshots. The Sales by Product and Inventory Valuation reports both depend on that field. A blank or stale value makes every downstream margin number wrong. Landed cost components often sit in operating expense accounts, which distorts the margin you think you have. That is one reason the shop can look healthy while cash is tight: gross margin never saw shipping, payment fees, fulfilment, returns, or the freight you booked in overhead.

StoreHero’s CM ladder is blunt about the next trap. CM1 (revenue minus COGS) and CM2 (CM1 minus shipping, payment fees and fulfilment) can look fine while CM3 (CM2 minus allocated ad spend) turns negative on losing campaigns. A merchandising slot that "wins" on conversion can still be a cash problem once you look below gross margin. The slot did not create the ad spend. It did choose which SKU absorbed the demand you already paid for.

Read that beside how you already look at media. How to read blended media efficiency without losing the plot is the same discipline in a different layer: one rate in a platform is not the business. Should paid media optimise for ROAS or profit? is the acquisition version of the SKU question. Conversion rate is the merchandising version of ROAS. Useful inside the machine. A bad kill switch.

Why deleting every rule, or keeping every rule, both fail

The usual responses do not fix this.

The first is a full delete. MGroup’s Shopify recommendations guide asserts that a recommendation engine will beat any manual merchandising approach. It offers no test data for that claim. Treat it as confidence, not a result. Even if the model is better at co-purchase, complementary intent on the native Shopify surface still needs hand curation. Cold start is also real. MGroup states that every AI recommendation system struggles with new stores, new products and new visitors, and recommends manually assigning new products to recommendation groups, or boosting them with rules, until co-purchase data accumulates. It also suggests content-based filtering for new stores with thin order history. That is a scoped reprieve, not a reason to keep the whole archive.

The second is to keep everything and let rules fight the model. You then pay twice: once in stale merchandising taste, once in a model that is never allowed to use the behaviour you bought it for. Overlays exist so you can constrain without reconstructing relevance by hand.

The third is to switch the sort from conversion rate to gross margin, or to CM1, and call it profit-first. CM1 is close to gross margin. It still omits the variable costs the storefront does control: shipping, payment fees, fulfilment. Polar Analytics’ point stands: gross margin alone omits ad spend, 3PL pick-and-pack, payment fees, returns and discount leakage. You do not need all of those in a slot ranking. You do need the ones the slot can change.

The fourth is to put CM3, with allocated ad spend, on the merchandising decision. StoreHero defaults to CM3 in reporting and warns that stopping at CM2 lets brands scale unprofitable campaigns. That warning is fair for channel and budget decisions. It is the wrong cut for a recommendation slot, because the slot cannot allocate Meta or Google spend. Forcing CM3 into the grid mixes an acquisition problem into a merchandising one. Fix acquisition with the checks in What to check before increasing Meta ad spend, not by hiding paid inefficiency inside a pin.

The fifth is to kill from a SKU P&L before the cost fields are trustworthy. Ottit is clear: if Cost per Item is empty or out of date, the report is fiction. Returns, bundles, write-offs and supplier cost changes belong in the audit, not in the post-mortem.

Are the leftover rules guessing relevance, or protecting cash?

The reframe is diagnosis, not tooling.

Manual recommendation rules were never one system. They were two jobs sharing a UI. Relevance rules try to answer what else people buy. Commercial-constraint rules answer what you refuse to sell into a slot: suppressions, floors, pins, seasonal overrides, complementary pairings you will stand behind.

Once an AI layer owns co-purchase, session behaviour and the related intent, the relevance rules are the ones you can retire. The constraint rules are the only place contribution-margin logic can enter the storefront. The overlay is margin control, not a second guess at taste.

That does not mean you should assume margin-weighted ranking pays. It is an open question. Zipchat sells high-margin prioritisation as a feature. A 2020 arXiv preprint on a search-based recommender in a vehicle marketplace ran a counterfactual where the system ignored differing product margins: expected profit changed very little, and the difference was not statistically significant. The authors conclude the important job is keeping the consumer engaged and away from churn; margin-awareness affected profit only slightly in that setting. Scope limits matter. It is a single study, a different category, simulation estimates from an estimated model, and a preprint. It does not show that margin-aware merchandising fails on a Shopify catalogue. It is a reason to test on your store rather than buy the overlay and declare victory.

RetentionX adds another split: CM2 and CM3 are not directly tied to individual customer behaviour, and CM1 correlates most closely with it. StoreHero and others push attention down to CM3 for paid businesses. Both can be true for different questions. Behavioural relatedness lives nearer CM1. Cash after fulfilment lives at CM2. Cash after ads lives at CM3. Do not ask one number to do all three jobs.

AI can widen retrieval and organise evidence. It does not own the constraint. That is the same split as How AI changes D2C research without replacing judgement: the operator still owns the question, the sources and the decision. An agent that ranks products is not a licence to let it edit the rules that protect cash. Nothing in the public evidence here documents agents autonomously changing merchandising rules, permissions or approval workflow. Keep "AI agent" descriptive of the recommendation layer.

How do you build the CM view and the rule set without redoing the books?

Classify rules first. Build a reporting-layer contribution view from data you already have. Only then run promote and kill.

  1. Tag every live rule as relevance guess, commercial constraint, editorial pin, or cold-start bridge.
  2. Retire relevance guesses once the model has order history for those SKUs.
  3. Keep constraints, pins with an expiry, and complementary pairings you will stand behind.
  4. Give new SKUs a dated boost or group assignment, then retire it when co-purchase exists.
  5. Populate and date-stamp Cost per Item on every live variant before you trust CM1.
  6. Add shipping, payment, fulfilment and returns from Xero or the 3PL in a reporting layer, not in the ledger.
  7. Use CM2 for storefront slots and ranging; use CM3 for channel and budget.
  8. Audit bundles, stale supplier cost and landed cost sitting in overhead before any kill.

Shopify’s complementary intent, if you use the native surface, stays in the editorial or constraint bucket because it is not auto-generated.

Then build CM1 to CM3 as a shadow view. Do not rebuild Xero.

Definitions, stated as the vendor sources state them, not as a universal accounting standard:

  • Contribution margin, in the Polar Analytics formulation, is revenue minus variable costs: the cash that covers fixed costs.
  • CM1 is revenue minus COGS, close to gross margin (StoreHero).
  • CM2 is CM1 minus the variable selling costs the order actually incurs. StoreHero names shipping, payment fees and fulfilment. flinder names logistics, warehousing and payment gateway fees. RetentionX describes CM2 as CM1 minus fulfilment, all costs except marketing.
  • CM3 is CM2 minus allocated marketing. StoreHero uses allocated ad spend. flinder includes digital and offline marketing.

CM4 is not a settled term in this evidence. Polar Analytics names a four-layer breakdown and does not supply an auditable CM4 definition in the material we have. If your finance lead wants a fourth layer, treat it as a house convention (some brands park allocated retention or platform costs there) and label it as such. Do the merchandising work at CM1 to CM3.

Why CM2 for promote versus kill on the storefront: the slot controls which SKU takes demand that already arrived. It does not allocate ad spend. CM2 is the order economics the storefront actually touches. StoreHero is right that CM2 is the better lens for organically acquired orders (email, SEO, returning, referral), and right that a common error is computing CM2 and calling it "contribution margin" as if the ladder stopped there. As practice, not a measured house result: CM2 for merchandising slots and SKU ranging on the site. CM3 for channel, campaign and budget. flinder’s caution still applies. Do not use contribution margin in isolation. A low CM3 paired with a high repeat order rate can still be a valid strategy, though it strains working capital. That is also why a lifecycle decision is not the same as a SKU kill. Which ecommerce lifecycle flow should you fix first? is about earning the relationship after acquisition, not about deleting a variant from the grid.

Shopify holds unit cost on Cost per Item. That is the CM1 base if, and only if, the field is populated and current. Shopify does not post COGS journal entries to Xero or QuickBooks. Ottit is explicit: that handoff needs a third-party app or a manual close. So a CM ladder you build in a spreadsheet, a BI tool or a semantic layer is a reporting layer. It is not an accounting change. You add the cost lines Shopify does not hold (shipping, payment fees, fulfilment, returns, the landed-cost elements sitting in operating expense) from Xero or from the 3PL, then slice by SKU. Polar Analytics advocates defining the formula once, then slicing by SKU, channel, region, cohort. That advice is about consistency, not about buying their layer.

We do not have Xero Central documentation in this set. Do not assume how Xero will display Shopify objects. Assume an integration or a manual export, and keep the books as they are.

Before any SKU loses a slot or gets killed, clear the data:

  • Cost per Item filled on every live variant
  • Supplier cost current, not last year’s
  • Landed cost not hiding in operating expense while you pretend CM1 is complete
  • Bundles not inheriting a blank or double-counted cost
  • Returns and write-offs visible, not only sell-in

If CM1 is weak, the problem is product cost or price. If CM1 is fine and CM2 collapses, fulfilment and shipping are the issue. If CM2 is fine and CM3 collapses, that is acquisition, which is not a recommendation-rule problem.

The kill decision is a capital reallocation, not a cost cut. Wiss, on SKU rationalisation accounting, states that gross margin improvement from cutting a low-contribution SKU is usually smaller than modelled because overhead does not leave with the product; it redistributes across the remaining portfolio. Discontinuation brings inventory write-downs, return provisions, and in wholesale contexts retailer chargebacks and possible slotting obligations. Some of those items do not apply to a pure D2C brand. Write-downs and return provisions often do. Working capital release is real, and slower to show up than the model.

What the public evidence supports, and where it stops

On the platform, the load-bearing mechanic is the related versus complementary split, and the complementary side remaining manual on Shopify’s free Search & Discovery surface, as Qualimero reports it. Verify that against Shopify’s own docs before you brief the team. The ten-product cap is the same class of claim: useful if it matches the endpoint you actually call.

On architecture, rule overlays (exclude, prioritise, pin, match on attribute) are described as a designed feature of the AI recommendation category, not as operator stubbornness. On cold start, the public advice is to assign or boost new products until co-purchase exists. There is no tested shadow-period length in this evidence, and no rollback threshold. Any calendar you write is judgement.

On the CM ladder, the CM1 / CM2 / CM3 cost lines are broadly consistent across StoreHero, RetentionX and flinder, with small differences in how fulfilment and marketing are worded. Finance teams use layered margins rather than one number as complexity grows; that is Saras Analytics’ framing, from a vendor blog, not a benchmark. StoreHero’s reporting default is CM3, and that disagreement should stay on the table for paid-acquisition businesses. The resolution is scope: merchandising slots do not allocate ads. Channel and budget decisions do.

On Shopify plumbing, Cost per Item is the field everything downstream leans on, and Shopify does not post COGS into Xero. That is why a shadow CM view is possible without an accounting rebuild, and why a pretty profit column can still be wrong.

On margin-aware ranking, the vendor category assumes it helps; the vehicle-marketplace preprint found a small, non-significant profit change when margins were ignored. Neither result is a Shopify catalogue test. There is no public lift figure for a margin-weighted recommendation slot against a conversion-only baseline on Shopify data. Inventing one would be the mistake.

On kills, Wiss is the brake: modelled margin gain overstates what you keep, overhead stays, one-off charges land, working capital arrives late. There is no CM2 floor in this evidence at which a SKU should lose its slot. Anyone quoting a threshold is making it up.

Write the promote-or-kill rule before you touch the catalogue

Domain Methods, writing about incrementality testing for ecommerce growth teams, says to choose the outcome that matters and to write the action rule before the test starts. Good test candidates share three traits: material spend, disagreeing reports, and an answer that will change action. Tests consume audience, time, spend and trust. A measurement programme nobody acts on is indefensible.

That piece is about media testing. Transferring it to recommendation slots is our argument, not theirs. If you are going to retire relevance rules or turn on high-margin prioritisation, do not use conversion rate as the outcome. Use contribution margin at the tier that matches the decision, which for the slot is CM2. Write what you will do if the test is flat, including putting a retired rule back.

Open Search & Discovery and the recommendation app. Tag each rule as relevance, constraint, pin or cold-start. Export Cost per Item completeness for the live catalogue. Do not kill a SKU this afternoon.

Useful answers

Questions operators ask

Should we delete our manual product recommendation rules after switching to an AI recommendation engine?
Split the list by job. Rules that guess affinity (hand-built also-bought, collection matching, merchandiser taste) are retirable once the model has order history. Rules that encode a commercial constraint should stay as overlays: exclusions, pins, seasonal overrides, and any high-margin prioritisation you are prepared to test. Two platform facts get in the way of a full delete. Complementary recommendations on Shopify’s native Ajax surface are not auto-generated; they have to be set up by hand in Search & Discovery. New SKUs and thin order history are a cold-start problem, so a dated boost or group assignment is a bridge, not a permanent archive. Hybrid AI-plus-rules is how the tools are built, not a compromise.
Is CM2 or CM3 the right number for deciding which SKUs to promote?
CM2 for storefront merchandising and ranging. CM2 takes revenue minus COGS, then shipping, payment fees and fulfilment: the order economics the slot actually touches. The slot cannot allocate ad spend, so putting CM3 (CM2 minus marketing) on a recommendation pin mixes an acquisition problem into a merchandising one. StoreHero is right to default reporting to CM3 for paid-acquisition businesses and to warn that healthy CM1 or CM2 can hide losing campaigns. Use CM3 for channel, campaign and budget. Use neither number in isolation. flinder notes that a low CM3 with a high repeat rate can still be a valid strategy, though it strains working capital.
Can we build a CM1–CM3 view without changing our accounting in Xero?
Yes, as a reporting layer. Shopify tracks unit cost on Cost per Item and uses it for profit columns and inventory valuation. That field is the CM1 base only if it is populated and current. Shopify does not post COGS journal entries to Xero; the handoff needs a third-party app or a manual close, which is why the ledger can stay untouched. Add shipping, payment fees, fulfilment, returns and any landed cost sitting in operating expense from Xero or the 3PL, then slice by SKU. Define the formula once and reuse it. Polar Analytics names a CM4 layer but does not define it in an auditable way here, so treat any fourth layer as a house convention. We do not have Xero Central documentation in this set; do not assume how Xero will display Shopify objects.
Why do our Shopify margins look healthy when cash is tight?
Often because the profit column is incomplete, not because the brand is secretly fine. If Cost per Item is blank or stale, Sales by Product and Inventory Valuation are wrong. Landed cost (freight, duty, inbound handling) frequently sits in operating expense, so CM1 looks cleaner than the cash. Gross margin also stops before shipping, payment fees, fulfilment and returns. Those land at CM2. Paid demand that looked efficient at CM2 can still be negative at CM3 once ads are allocated. Check the field, then the cost lines below it, before you trust the dashboard.
Will prioritising high-margin products in recommendations reduce conversion?
Unresolved on Shopify catalogues. Vendor tools sell high-margin prioritisation as an overlay you can run without losing personalisation underneath. A 2020 arXiv preprint in a vehicle marketplace found that ignoring product margins changed expected profit only slightly, and not significantly, in a counterfactual simulation; the authors put more weight on keeping the shopper engaged. That is not a Shopify result and it is not a conversion test. There is no public lift figure for margin-weighted slots against a conversion-only baseline. If you turn the overlay on, write the action rule first and use contribution margin (CM2 for the slot) as the outcome, not conversion rate.
How much margin will we actually gain by cutting an underperforming SKU?
Less than the SKU-level analysis implies. Wiss frames rationalisation as capital reallocation, not a cost cut. Overhead does not leave with the product; it redistributes across the remaining portfolio, so modelled gross-margin improvement is usually overstated. Discontinuation can bring inventory write-downs and return provisions. Wholesale extras such as retailer chargebacks and slotting obligations may not apply to pure D2C, but the write-down logic often does. Working capital release is real and slower than the spreadsheet. Audit Cost per Item, bundles and returns before you kill, and do not treat a low CM3 alone as sufficient if repeat rate is carrying the SKU.

About the author

Eddie Cheng

Eddie Cheng founded Penang Media and co-owns VIBAe. He writes from the agency and brand sides of ecommerce growth, connecting paid acquisition with stock, margins, cash flow and contribution profit.

More from Eddie Cheng

The operating context

Growth from the agency and brand sides.

Eddie Cheng writes about profit-first ecommerce growth from both sides of the work: Penang Media, the performance agency he founded, and VIBAe, the footwear brand he co-owns. His articles connect paid acquisition with stock, margins, cash flow and the decisions that determine profitable growth.

About the publication