2026-07-20

Machine-Readable Product Content

A woman scans products on a warehouse shelf.
By Tiiu Vaartnou, Analytics & Optimization Strategy Specialist, Orium
6 min read

At Stripe Sessions 2026, one of the most attended tracks was titled "Agentic Commerce Under the Hood." The session description didn't lead with payments or checkout, it led with the production challenge of "making catalogs legible to agents."

Legible is an interesting word. Not a dive deep on optimization or a detailed discussion of enrichment, but rather can the machine actually read and understand what you're selling?

For most commerce organizations, the honest answer is probably: partially. And as we’ve discussed in previous articles in this series, partial legibility means partial visibility. Your products appear in some recommendations and get passed over in others, based not on price competitiveness or brand strength, but on whether the data you published was specific enough for an AI system to act on with confidence. The AI Reader Problem

The AI Reader Problem

AI agents need structured, machine-readable product data to function. Without it, your products don't exist in their comparison workflows. When a customer asks an AI shopping assistant to find the best option for a specific need, the agent isn't browsing your website. It's parsing structured signals: attributes, schema markup, review sentiment, pricing, availability, compatibility. If those signals are incomplete or ambiguous, the agent moves on to a product it can evaluate with confidence.

More than half of searches on ChatGPT are discovery-based, and 70% of those include constraints: "I need a carry-on bag that fits under an airline seat, holds a 15-inch laptop, and doesn't look like a travel bag." An agent resolving that query is cross-referencing dimensions, compatibility notes, product category, and probably review sentiment about appearance. Every field it can't find is a reason to prefer a competitor's product.

Adobe's research on AI content visibility puts a number on how widespread this gap already is across retail catalogs, and it's wider than most teams assume: across U.S. retail sites, individual product pages average a score of just 66% on AI content visibility and the gap between the best- and lowest-performing retailers (already at 28 percentage points) translates directly into which brands get recommended and which get skipped.

Four Content Layers, Most Brands at Half Strength

Machine-readable content isn't a one-time implementation task, it's a content strategy that treats AI systems as a primary audience alongside human shoppers, and it requires different thinking about what your product data is supposed to do. Four layers determine AI readiness, and most brands are operating at full strength in only one or two.

Structured product schema tells AI systems what you sell: name, description, brand, SKU, GTIN, images, materials, price, availability. For AI agents, schema is often the primary mechanism for reading product information—ahead of crawled page content—which makes completeness and accuracy non-negotiable.

Review content in structured form carries attributes the product page never states. A review mentioning "perfect for wide feet, didn't cause blisters after 10 miles" is attribute data an agent will factor in for a customer who mentioned foot width. Products with 11 to 30 reviews see conversion rates more than double those of products with zero, which translates to a 255.4% lift. If your reviews live behind a render agents can't parse, that signal is lost entirely.

Intent-ready attributes is the layer most brands are missing almost entirely: "best for," "works with," "compatible with," "use case.” These are the vocabulary agents use to match products to context-rich prompts, and traditional product data wasn't built to carry them.

Trust and sourcing content—material composition, ingredient sourcing, certifications, allergen policies, and quality standards—function as both trust signals and retrieval assets for sustainability- or safety-conscious queries. This content is usually scattered across marketing pages or absent from structured data entirely.

Brand Control Starts at Discovery

Stripe built explicit brand controls into the Agentic Commerce Suite, so businesses can define their checkout experience within an agent-mediated transaction flow. The reasoning is straightforward: businesses want to retain their brand, their upsells, and their customer relationships. Those checkout controls protect brand at the transaction layer, but the transaction layer is the end of the journey. Brand control at discovery, which determines whether a customer encounters your product at all, is a function of machine-readable content.

If your descriptions are vague, your attributes are thin, and your reviews aren't schema-marked, agents will summarize your products however they can from whatever signals are available. That summary may not reflect your positioning or differentiation. Machine-readable content is how you define the narrative instead of ceding it to the model’s inference.

The Real Problem Is Governance, Not Data

There’s a lot of conversation around how data is the heart of where most commerce organizations struggle, but most commerce organizations don't have a data problem— they have a content governance problem that presents as a data problem.

The content usually exists: descriptions are written, attributes are filled in, and reviews are collected. But it’s inconsistent across channels, outdated in ways nobody has tracked, and owned by nobody when it comes to AI readiness. The competitive advantage isn't publishing more content, it's governing what you have well enough that AI systems trust and represent it accurately.

Solving both the ownership problem and the content challenges is easier when the architecture supports it. Most traditional platforms couple content to presentation, which means enriching attributes for a new use case touches the display layer, and adding fields a new protocol requires goes through a backlog. It's the kind of friction that makes clear ownership feel impossible in practice, and creates barriers to effective content management.

A composable architecture removes that friction: content lives in a PIM, is authored once, and flows to every surface. If you add a field a new protocol requires, it propagates everywhere instead of waiting in a queue.

That's what makes AI readiness a continuous capability rather than a recurring project, and it's what makes the audit below worth running against your current stack.

Where to Audit First

  • PDP completeness against AI requirements — not Google Shopping compliance, but conversational attribute coverage.
  • Schema accuracy and coverage — including offer data, review aggregates, and availability, not just presence.
  • Review accessibility — collected consistently, schema-marked, and recent enough to be weighted.
  • Intent-ready attribute gaps — usually concentrated in older, less-trafficked SKUs with the most to gain.
  • Governance ownership — who's accountable for the gap between what you've published and what agents can use?

Vague product descriptions and inconsistent specs stop being minor content quality issues when AI systems are synthesizing them in real time to generate recommendations. They become the reason a machine can't interpret your product confidently, and a machine that can't interpret your product won't recommend it.

The brands making catalogs legible to agents aren't doing it for compliance. It's a competitive moat built from the quality of what you publish. If you want to understand how machine-readable your product catalog is today, talk to our team about a Retail AI Discovery Audit.

Popular Articles