The 9 signals AI shopping agents read from your product data
“Optimize for AI” is vague advice. So let's make it concrete: when an AI shopping assistant evaluates whether to recommend your product, it reads a small, knowable set of signals from your catalog. Here are the nine that matter, grouped the way we score them in Preferd — content signals (what the model reads about the product) and technical signals (how unambiguously it can identify and classify it).
Content signals
- 1 · Title. The single highest-leverage field. A model matching “insulated bottle that keeps drinks hot all day” needs the type and the key attribute in the title: “750 ml Insulated Steel Bottle — 24 h Heat Retention”. Titles made of internal SKUs, ALL-CAPS promo text or bare brand names give the matcher nothing to hold on to.
- 2 · Description. This is the model's interview with your product. Materials, dimensions, weight, use cases, compatibility, care, what's in the box. A useful test: collect the ten questions buyers actually ask (support inbox, reviews, live chat) — a strong description answers all ten in plain language.
- 3 · Tags. Secondary, but useful as disambiguation: occasion, audience, style, season. A handful of specific tags (“trail-running”, “gluten-free”, “office-chair”) beat thirty generic ones.
- 4 · Images + alt text. Modern assistants are multimodal — they genuinely look at product images. Three or more angles give visual evidence; descriptive alt text (“side view, steel bottle with bamboo cap”) doubles as accessibility and as a caption the model can read directly.
Technical signals
- 5 · Product category. Shopify's standard product taxonomy is how every downstream surface — including AI catalog feeds — knows what the item is without inference. An uncategorized product forces the model to guess from text, and guesses don't get recommended.
- 6 · SEO title & meta description. Search-grounded assistants (and classic engines feeding them) often quote this snippet verbatim. It's your one-sentence pitch to a machine.
- 7 · Brand / vendor. A real brand name lets the model connect your product to reviews, comparisons and reputation elsewhere on the web. Placeholder vendors — “Default”, “My Store”, “Vendor” — actively signal an unmaintained catalog.
- 8 · GTIN / barcode. Global identifiers are the strongest possible entity match: they let an assistant say “this exact product” across retailers and review sites. (Handmade and one-of-a-kind goods legitimately lack them — that's fine, the other signals carry the weight.)
- 9 · Structured metafields. Category-specific attributes — size, material, fit, capacity — ride along in the catalog feed as clean key-value data. For a machine reader this is the purest signal of all: no parsing, no ambiguity.
Why weight them at all?
Not all nine are equal. In our scoring model the biggest weights go to the fields that change what an assistant can say (title, description, images) and the fields that change what it can verify (category, GTIN, metafields) — because those map to the two failure modes we see in real answers: the model either can't describe the product convincingly, or can't confirm it's a real, current, identifiable thing.
The signals interact. A perfect description under a junk title still loses, because the title is what gets matched first. That's why worst-first triage beats polishing your already-good products.
Auditing your own catalog
You can audit by hand: open each product and check the nine signals against the list above. For ten products that's an afternoon; for four hundred it isn't realistic, which is exactly the gap GEO tooling fills — score everything, sort worst-first, fix in bulk, re-score to prove the change. Whichever way you do it, re-audit on a schedule: catalogs drift as new products ship without descriptions and seasonal edits overwrite good data. And when you're done fixing, close the loop with measurement — track the AI-sourced traffic that these signals ultimately earn.