Separating measured lift from placement credit in on-site personalization
Researchers ran a randomized field experiment at an online retailer, turning product recommendations off for one group of shoppers and leaving them on for another. Shoppers who saw the recommendations bought more of the recommended products and less of everything else, so total sales did not necessarily rise in the short run. The payoff arrived later and somewhere else: the shoppers whose recommendations had been switched off came back less often, and long-run sales fell.1
Almost no recommendation reporting would show you either half of that. A product recommendation carousel earns credit when a shopper clicks a tile inside it and goes on to buy, so a shopper who had already settled on the purchase counts exactly the same as one the carousel persuaded. The Aje Collective reports that 10 percent of its revenue comes from personalized product recommendations.2 A figure like that means one thing if the recommendations created those baskets and something much weaker if they were standing nearby while the shopper bought anyway, and standard reporting cannot separate the two.
Before you spend another quarter tuning a recommendation model, find out which of your product recommendations changed what a shopper would have done on their own. Cart size grows only from those. The placements where that change is largest are the ones most teams treat as housekeeping.
Click attribution cannot tell you what worked
Click attribution is the default across recommendation reporting, whichever platform produces it. A click-attributed report answers only whether a shopper touched a product recommendation on the path to a purchase. It cannot tell you whether that same shopper would have bought the same items without it, because both shoppers produce an identical row in the report. On a product detail page, where the visitor is already looking at an item they came to buy, a large share of carousel clicks come from people who would have converted regardless.
Answering the real question takes a control group. Hold back a slice of traffic and serve those shoppers the same page with a fixed, non-personalized list in the carousel slot. The page length and layout stay identical, so the only thing that differs between the two groups is the personalization, and the difference in revenue per visit is what the personalization produced. Compare that figure to the attributed number your reporting already shows you, and the gap is the portion of your recommendation revenue that was never incremental. Most teams have never run that subtraction, which is why recommendation revenue is one of the least contested numbers in an ecommerce review.
Which metric you compare matters as much as the comparison. Average order value, units per transaction, and revenue per visit move independently, and average order value can climb while revenue per visit falls, because a smaller number of larger orders lifts the average. We worked through that trade-off across all seven AOV tactics earlier this year. For product recommendations, revenue per visit is the baseline metric, since it accounts for how large the basket gets and how often a visit produces one at all.
Cart size grows at the dead ends
At three moments in a visit, a shopper has nothing left to click. A search returns nothing, the product they opened is out of stock, or a cart holds a single item with nothing suggesting a second. In each case, the most likely next event is the end of the session, which makes these the placements where a product recommendation has the clearest claim to whatever revenue follows.
How much larger the return is at those three points is worth measuring on your own site rather than assuming. Compare what personalized product recommendations return on a zero-results page against what the same logic returns on a product detail page, and the difference will have little to do with the algorithm behind either one. It comes down to what would have happened otherwise. A shopper looking at an empty search result was on their way out.
Zero-results pages are the largest of the three for most catalogs, and they hold two different problems: queries you could have served and failed to, and queries your assortment cannot serve at all. We sorted those into two piles and two owners last week. A product recommendation does nothing for the second kind, but it can keep the session alive long enough to sell something else in the catalog.
Teams leave the out-of-stock product page unhandled more often than any other placement, and handling it depends on something a recommendation model cannot supply on its own. The carousel has to know the item is unavailable, which means reading the same live inventory signal that suppressed the product in search results. When a carousel works from a nightly export, it will cheerfully recommend three more items that also sold out yesterday.
Your product-page carousel promotes what would have sold anyway
Most recommendation effort goes into the carousel on the product detail page and the rows on the home page, and those are the two placements with the weakest case. A cross-category field experiment by Dokyun Lee and Kartik Hosanagar, covering 82,290 products and more than a million shoppers at a large North American retailer, tracked what collaborative filtering does to the mix of products sold. Shoppers exposed to it explored more variety individually, but they were pushed toward the same popular items, so niche products gained in absolute sales while losing market share.3 An earlier experiment by the same pair found the algorithms differ sharply from one another, and that one widely used algorithm moved sales volume not at all.4 When a carousel drifts toward the popular item, it is recommending what was likely to sell without it.
Each of those placements still has a job, and the job should be one it can be measured against. A home page row is doing discovery work, so judge it on whether shoppers reach a wider set of categories and return more often, and leave basket size to the placements that can move it.
Product-page bundles behave differently on that same page, and the reason is worth borrowing. A bundle puts an explicit second item in front of the shopper to add, and an add is a different act from a click. When the pairing is specific enough that accepting it means one more item in the cart, the placement earns its slot.
Product data decides how good your recommendations can get
Every product recommendation makes a claim: these two products belong together. The model can only make that claim from what your product records say, and when the attributes are thin, it falls back to co-purchase patterns and returns the best seller. Falling back on popularity is how a carousel ends up recommending the item that would have sold anyway. That habit starts in the product data, which is why retuning the algorithm never reaches it.
The Aje Collective runs three brands on one domain, which makes the problem visible in a way a single-brand catalog rarely does.
Rhyanna Cardillo
Nothing in a co-purchase pattern prevents that pairing, because many Aje customers buy across brands. A product record that carries brand, occasion, and category as real attributes prevents it, because the model can then tell that a shopper working through activewear is not shopping for eveningwear. Without attribute work like that, retuning the model just rearranges which bestseller it returns.
Attribute depth also decides which kinds of product recommendation you can run at all. Behavior-based recommendations draw on browsing and purchase patterns built up over time, while session-based recommendations adapt to what the shopper is doing in the current visit. Both need a catalog carrying attributes specific enough to describe what the shopper is moving toward. A model with only category and price will produce a session-based recommendation that reacts to nothing in particular. That same product data reaches your offsite channels, so attribute work that makes an on-site pairing defensible also improves how your products appear in Google Shopping, marketplaces, and AI assistants.
What to test before peak
Three tests are worth running while you still have time to act on the answers.
- Calibrate attributed against incremental on a page you can control and test. A merchandising campaign in Athos Commerce can carry up to five variations, and one of them can be a true control that applies no boost rules or banners at all. Run that on a busy category or search results page, set the traffic split, and compare revenue per visit between the control and the treatments. What comes back is your own site’s answer to how far an attributed number sits from an incremental one. That ratio is the one to carry into every placement you currently judge on attribution alone. Watch your product data while it runs: if search suppresses an out-of-stock item and the campaign never learns of it, the treated group is being shown products nobody can buy, and a weak result will tell you more about your product records than about your merchandising.
- Check whether your recommendations can see what search already decided. Suppress a product in search, wait an hour, then look at what the carousel on that page surfaces. A carousel that still promotes the suppressed item is reading a different copy of your product data than search is, and tuning either system alone won’t resolve that.
- Audit the product attribute fields your pairing logic depends on before touching the model. List the attributes the recommendation engine draws on, then check what percentage of your catalog populates each one. A pairing the product record cannot support will not survive a control test, no matter which algorithm produced it.
All three tests get easier when search and merchandising read the same product data model that personalization and feed management do, which is the design behind the Athos Commerce intelligent discovery platform. Customer Success, Support, and Education are included at no charge, so your CSM can help you set up the first test.
Which recommendation number to report
Attributed recommendation revenue counts every basket the carousel stood beside. Incremental recommendation revenue counts the baskets it built, and it is the number to take to a CFO, because it is the only one that says whether more investment in product recommendations will return anything.
A test started now returns a defensible baseline before November, when your placements carry the most traffic and receive the least scrutiny. Teams that enter peak season with an attributed number will finish knowing no more than they do today.
Frequently asked questions
What is incremental recommendation revenue?
Incremental recommendation revenue is the revenue a product recommendation created that would not have arrived without it. Measuring it means holding back a portion of traffic, serving those shoppers the same page with a fixed, non-personalized list in place of the personalized one, and comparing revenue per visit between the two groups. The difference is the incremental figure. Attributed recommendation revenue, which most ecommerce reporting shows by default, counts any purchase where the shopper clicked a recommendation, including purchases the shopper had already settled on before seeing it.
How does attributed recommendation revenue overstate impact?
Click attribution is how recommendation reporting works by default across ecommerce tools. It credits a product recommendation carousel whenever a shopper clicks a tile inside it and later buys. On a product detail page, the visitor is often already looking at the item they came to purchase, so a large share of those clicks come from shoppers who would have converted regardless. The report cannot separate them from shoppers the carousel persuaded, which means the attributed number includes baskets the carousel only stood beside.
Which product recommendation placements grow cart size the most?
Recovery placements grow cart size the most. The three are a zero-results search page, an out-of-stock product page, and a cart holding a single item. At each of those points, the shopper was about to leave, so a purchase that follows is one the recommendation can reasonably take credit for.
Why do product recommendations drift toward best sellers?
A recommendation model asserts that two products belong together, and it can only make that claim from the attributes in your product records. When those attributes are thin, the model falls back on co-purchase patterns and returns whatever is popular. A cross-category field experiment covering 82,290 products found that collaborative filtering pushed shoppers toward the same popular items, so niche products gained absolute sales while losing market share. Enriching product data gives the model better inputs, and retraining the algorithm does not.
Should we switch off home page and product-page recommendation carousels?
No. Give each placement a job it can be measured against instead. A home page recommendation row is doing discovery work, so judge it on whether shoppers reach a wider set of categories and return more often, and leave basket size to the placements that can move it. Product-page bundles are worth separating, because they put an explicit second item in front of the shopper, and accepting one puts a product in the cart.
Where should a team start?
Start on a page you can control and test. Run a merchandising campaign on a busy category or search results page with one variation held as a true control that applies no boost rules or banners, compare revenue per visit against the variations that do, and use the difference to calibrate how far your attributed numbers sit from incremental ones. Then check whether your recommendations can see what search already decided: suppress a product in search, wait an hour, and look at whether the carousel on that page still promotes it. A carousel that does is reading a different copy of your product data than search is.
Sources
- Sun, Luping, Yuxin Chen, Xiaona Zheng, Meng Su, and Xiaoquan (Michael) Zhang. “How Do Recommender Systems Benefit Online Retailers in the Long Run? Evidence from a Field Experiment.” MIS Quarterly, published online April 14, 2026, 1-21. https://doi.org/10.25300/MISQ/2026/18885.
- Athos Commerce. “Aje Serves Up Personalized Ecommerce Shopping Experiences.” Athos Commerce case study, accessed August 31, 2026. https://athoscommerce.com/case-studies/aje/.
- Lee, Dokyun, and Kartik Hosanagar. “How Do Recommender Systems Affect Sales Diversity? A Cross-Category Investigation via Randomized Field Experiment.” Information Systems Research 30, no. 1 (March 2019): 239-259. https://doi.org/10.1287/isre.2018.0800.
- Lee, Dokyun, and Kartik Hosanagar. “Impact of Recommender Systems on Sales Volume and Diversity.” ICIS 2014 Proceedings, International Conference on Information Systems, 2014. https://aisel.aisnet.org/icis2014/proceedings/EBusiness/40/.