The four product-data layers most retail catalogs get wrong
A shopper opens an AI assistant and asks for a “coffee color hoodie.” A retailer has that exact product in stock, but the product listing records the color as brown. The assistant cannot connect “coffee” to “brown,” so it never returns the hoodie, and the shopper buys from a competitor whose product data made the match. No error appears anywhere. The sale never happens, and the retailer never learns why.
AI readiness is about preventing losses like that one. Readiness is rarely one thing. Four layers sit between your product data and a completed sale, and each one can pass its own internal check while still breaking the chain. Your product data foundation has to be complete and up to date. Taxonomy has to let an agent traverse categories, synonyms, and variants. Onsite search has to understand what a shopper means, not just match keywords. And the same clean data has to reach every channel where discovery now happens, including AI engines. A weak link in any one layer costs you sales you cannot see, because your dashboards still look healthy.
Gartner predicts that by 2030, 20% of monetary transactions will be programmable, giving AI agents the economic agency to evaluate, recommend, and complete purchases on a shopper’s behalf.1 The retailers that win the next few years are the ones whose product data was ready when the agents arrived. This checklist helps you find out where you stand.
How to use this checklist
Work through the four layers below. Each has four items. Score every item from 0 to 3, where 0 means you have not started, 1 means early progress, 2 means mostly in place, and 3 means fully in place and maintained. Score against what is live in production today, not what is planned for next quarter. Add up all sixteen items at the end for a total out of 48, then read your readiness tier.
Layer 1: Product data foundation
Fewer than 10% of merchants have machine-readable product data, according to Athos Commerce CTO Suhas Gudihal, and most catalogs carry only three or four attributes per product when AI agents need dozens to make a confident match.2 This is the layer everything else inherits, a case we made in Your Data Is Your Storefront. If the underlying product data is thin, stale, or written for a keyword algorithm that no longer exists, no amount of search tuning or syndication downstream can recover it.
Your product feed meets Google’s specifications, so it appears compliant, but compliance only confirms that the required fields are filled. It says nothing about whether an agent can understand what your product is for. A product highlight that reads “Men’s hoodie, charcoal grey, size L” tells an agent what the product is. A highlight that reads “Heavyweight breathable hoodie for cool-weather hiking, relaxed fit through the shoulders” tells the agent what the product is for, which is what a shopper asks about.
Score yourself on the product data foundation:
- Complete required and recommended fields: Every product includes the recommended attributes agents read, such as product highlights, product details, GTIN, material, and dimensions, not just title, price, and image.
- Fresh price and availability: Price and stock status update in near real time, so an agent never recommends a product you can no longer sell at the price shown.
- Machine-readable, structured values: Attributes use standardized, structured values rather than free-text descriptions, so an agent can parse size, color, and material without guessing.
- A single source of truth: One enriched product data set feeds every channel, rather than separate spreadsheets and exports drifting out of sync across teams.
Layer 2: Taxonomy and attribution
Return to the coffee-color hoodie. The product existed, the price was right, and the image was clear. The sale was lost because the retailer’s taxonomy lacked a path from the word a shopper used to the product’s value in the listing. Taxonomy is what lets an agent move from a natural-language request to a specific product, and it is where many otherwise clean catalogs fail.
When taxonomy is weak, the same product gets classified differently across channels, onsite filters return partial results, and synonyms go unmapped. An agent traversing your catalog hits a dead end and moves to a competitor whose category structure and attribute vocabulary let it find an answer.
Score yourself on taxonomy and attribution:
- Standardized category tree: Products sit in a consistent, logical category structure that an agent can navigate without hitting orphaned or duplicate branches.
- Controlled attribute vocabulary with synonyms: Attribute values map to the words shoppers use, so “coffee” resolves to brown, “trainers” to sneakers, and regional terms map to their equivalents.
- Modeled product variants: Color, size, and material variants are related to their parent product, so an agent understands that ten listings are one product in ten forms.
- Taxonomy mapped to each channel’s schema: Your internal categories map cleanly onto Google product categories, marketplace taxonomies, and each destination’s required structure.
Layer 3: Onsite search logic
Score onsite search first, then read why it matters, because the items describe an experience you can test in five minutes on your own site.
- Zero-results rate tracked and near zero: You measure how often onsite search returns nothing, and that rate stays close to zero through synonym handling and query understanding.
- Natural-language and long-tail queries resolve: A query like “waterproof jacket for winter hiking under $150” returns relevant products rather than a literal keyword match or an empty page.
- Intent handled beyond keywords: Search understands what a shopper means, learns from behavior, and improves over time rather than matching strings.
- Ranking transparency and control: Your team can see why a product ranks where it does and adjust it accordingly, rather than treating the ranking order as a black box.
Onsite search is the clearest preview of how an agent will judge your catalog, because agents apply the same logic a good search engine does. In an Athos Commerce and Pixel survey of more than 800 shoppers, 80% said they had abandoned a site because they could not find what they were looking for.3 A shopper gives you a second chance. An agent does not. It reads the response, decides your catalog cannot answer the question, and recommends someone else, with no bounce-back visit to recover.
Layer 4: Offsite syndication and AI visibility
The first three layers make your catalog ready. Offsite syndication makes it visible. Clean, structured product data has to reach every channel where discovery now happens, and each destination wants something slightly different. Google Shopping, marketplaces, social channels, and AI engines such as ChatGPT, Gemini, Perplexity, Microsoft Copilot, and Amazon Rufus all evaluate product data through their own logic. A single generic feed sent everywhere performs on one channel and disappears on the rest. We break down what each engine prioritizes in One Product Feed Won’t Win Five AI Shopping Engines.
Stephanie Brown
Traditional optimization no longer covers you, because the job of the feed has changed. As Mark Batson, Head of GTM Technical Operations at Athos Commerce, puts it, traditional SEO is about your product being found and discovered, while answer-engine and generative-engine optimization is about your product being selected and recommended. Getting found is the baseline. Getting recommended is where the revenue moves.
Score yourself on syndication and AI visibility:
- Feeds reach every revenue channel: Product feeds syndicate to marketplaces, social commerce, and AI engines, not just Google Shopping.
- Data adapted per engine: Product data is tailored to each engine’s priorities, rather than sending a single, identical feed to all of them.
- Answer-engine optimization in place: Your product data is structured and enriched for generative engines so that agents can extract and cite it.
- Cross-engine visibility you can measure: You can see where and how your products appear across the major AI engines, rather than guessing.
Score yourself: the four readiness tiers
Add up all sixteen items for a total out of 48.
- 0 to 16, Reactive. Agents are skipping your products right now, and you are likely losing sales you cannot trace. Start with the foundation layer.
- 17 to 32, Emerging. Your products are visible in places, but data quality and channel coverage leak revenue at every handoff. Close the weakest layer first.
- 33 to 43, Competitive. You are mostly ready. Refine the edges, especially per-engine adaptation and taxonomy consistency across channels.
- 44 to 48, AI-ready. Your product data is complete, searchable, and syndicated across channels. Hold the standard, because readiness erodes as your catalog grows.
Most retail teams score lower than they expect, and the reason is consistent. Each layer passed its internal checks, so no single team owned the gap between them.
Where to start
Fix the foundation layer first, whatever your total. A low score is already costing you. Agents choose between catalogs that can answer a shopper and catalogs that cannot, so every tier below AI-ready means recommendations you are losing without ever seeing them. Taxonomy, search logic, and syndication all inherit the quality of your underlying product data, so a clean, enriched single source of truth is the one investment that raises your score in all four layers at once. Trying to tune search or expand channels on top of thin product data just moves the problem downstream.
The fastest way to see your real starting point is a product feed audit, which Athos Commerce offers at no cost. It diagnoses your data foundation, taxonomy, and syndication coverage in one pass, so you learn which layer is costing you the most before you spend a quarter fixing the wrong one. An all-in-one platform matters here for a structural reason. When search, feed management, merchandising, and generative engine optimization run on a single shared product data model, a fix to the foundation automatically flows through every layer. One shared model also keeps readiness from becoming another manual task, because your team improves the product data once rather than maintaining four layers by hand.
The teams that win
The retailers pulling ahead in agentic commerce have one thing in common. Their product data was complete and well-structured, searchable on their own site, and cleanly syndicated to every channel by the time agents started shopping on their customers’ behalf. Readiness is the prerequisite for everything that follows. Somewhere right now, a shopper is asking an assistant for a coffee-color hoodie. Whether your product is the one it recommends depends on work you can start scoring today.
Frequently asked questions
What does it mean for product data to be AI-ready?
AI-ready product data is complete, machine-readable, and structured so an AI shopping agent can understand what a product is and what it is for. It carries the recommended attributes agents read, including product highlights, product details, GTIN, material, and dimensions, uses standardized values rather than free text, and stays current on price and availability. Readiness spans four layers: the product data foundation, taxonomy, on-site search logic, and off-site syndication to marketplaces, social channels, and AI engines.
How is AI-readiness different from being ready for traditional SEO?
Traditional SEO helps your product appear in search results. AI-readiness gets your product selected and recommended by an agent that returns one answer, not ten links. SEO rewards keywords and rankings, while AI-readiness rewards structured, intent-rich product data an agent can parse, compare, and cite. A product catalog can rank well in Google and still remain invisible to Perplexity or ChatGPT because those engines evaluate product data using their own logic.
Why do AI shopping agents skip products that are in stock?
Agents skip in-stock products when the product data cannot answer the shopper’s question. A missing synonym, a thin set of attributes, or a taxonomy with no path from “coffee” to “brown” sends the agent to a competitor whose data made the match. The product exists, the price is right, and the image is clear, but the agent never surfaces it because it cannot connect the request to the product listing. The loss stays invisible because nothing errors.
Which layer of AI-readiness should retail teams fix first?
Fix the product data foundation first. Taxonomy, onsite search, and offsite syndication all inherit the quality of your underlying product data, so a clean, enriched single source of truth raises your readiness across all four layers at once. A product feed audit is the fastest way to find which layer is costing you the most, since it diagnoses data quality, taxonomy, and syndication coverage in one pass before you commit resources to the wrong fix.
How can a retailer tell if its product catalog is visible to AI shopping agents?
Start by testing your own onsite search with natural-language queries, since agents apply similar logic. Then check whether your product feeds reach beyond Google Shopping, whether your data is adapted to each engine’s priorities, and whether you can measure where your products appear across ChatGPT, Gemini, Perplexity, Copilot, and Rufus. If you cannot see your products across those engines, agents probably cannot either. A feed audit makes that coverage measurable.
Sources & Further Reading
- Gartner. “Gartner Unveils Top Predictions for IT Organizations and Users in 2026 and Beyond.” Gartner, October 21, 2025. https://www.gartner.com/en/newsroom/press-releases/2025-10-21-gartner-unveils-top-predictions-for-it-organizations-and-users-in-2026-and-beyond.
- Gudihal, Suhas. Interview with Athos Commerce, February 17, 2026.
- Athos Commerce and Pixel. Shopper Discovery Survey (800+ respondents), 2026.