Most brands that start tracking Share of Model don’t fail because they lack a tool. They fail because their measurement process is quietly broken, and nobody notices until a board member asks “why does ChatGPT still recommend our competitor” and the dashboard has no honest answer. Bad AI visibility data is worse than no data at all — it sends teams chasing the wrong fixes while the real gaps sit untouched.
We’ve watched this pattern repeat across enough Share of Model rollouts at Salterra to know the mistakes aren’t exotic. They’re the same seven, over and over, and each one is fixable in an afternoon once you know what to look for.
Symptom: the tracked prompt list reads like a brainstorm — a handful of obvious head terms someone typed into a spreadsheet during a Monday meeting. “Best [category] tools,” “top [category] companies,” and not much else.
Why it hurts: a biased, thin prompt set measures a narrow slice of intent and calls it the whole picture. Real buyers ask comparison questions, problem-framed questions, budget-constrained questions, and questions with qualifiers you’d never think to include. If your prompt set skews toward generic “best of” phrasing, you’ll overweight categories where you already do well and miss the long-tail conversational queries where a competitor is quietly winning every mention.
The fix: build the prompt set the way you’d build a keyword list — from actual research, not intuition. Pull real customer questions from sales call transcripts, support tickets, and search console query data. Include problem-first phrasing (“how do I fix X”), comparison phrasing (“X vs Y”), and buyer-intent phrasing (“is X worth it for a small team”). Refresh the list quarterly as language shifts, and document why each prompt is in there so the set doesn’t quietly drift back toward guesswork.
Symptom: the whole Share of Model program lives inside a single tool’s coverage of a single chatbot, usually whichever one the team happens to have open.
Why it hurts: different AI platforms pull from different underlying signals — some lean heavily on their own search index, some rely more on training data, some cite recent web content aggressively and others barely cite at all. A brand can have strong presence in one system and be nearly invisible in another because the underlying retrieval behaves differently. Reporting single-platform results as “our AI visibility” is like reporting one search engine’s rankings and calling it “our SEO performance.”
The fix: track Share of Model across at least three to four major AI platforms with meaningfully different retrieval approaches, and report them separately before you ever roll them into a blended score. When you do blend, weight by where your actual buyers are spending time, not by which platform is easiest to sample. If one platform is dragging the average down, that’s a finding worth investigating on its own — not a number to smooth over.
Symptom: someone runs the same prompt three or four times, gets inconsistent answers, and reports the one where the brand shows up favorably. Or the reporting window conveniently starts right after a good result and ends right before a bad one.
Why it hurts: AI answers are non-deterministic — the same prompt can return a different set of cited brands run to run. That variability is real information; it tells you how stable or fragile your presence actually is. Cherry-picking erases that signal and replaces it with wishful thinking. Worse, it trains the team to trust a number that won’t hold up the next time someone runs the same prompt in front of a client.
The fix: sample every tracked prompt multiple times per platform per period and report the distribution, not a single favorable snapshot — mention rate across N runs, not “did we get mentioned once.” If variability is high on a given prompt, say so in the report. A wide spread on a high-value prompt is itself an actionable finding: it usually means your content presence there is thin enough that the model is guessing rather than confidently retrieving you.
Symptom: the dashboard shows Share of Model climbing, everyone celebrates, and nobody has actually read what the AI is saying about the brand.
Why it hurts: a mention is not automatically good. Models sometimes cite a brand while getting the pricing wrong, attributing a competitor’s feature to you, or listing you as a cautionary example (“Brand X had complaints about…”). A rising mention count built on inaccurate or lukewarm sentiment isn’t visibility you can be proud of — it’s a growing liability that compounds every time someone acts on bad AI-sourced information about you.
The fix: layer sentiment and factual-accuracy scoring on top of raw mention tracking. For every logged mention, capture whether the framing was positive, neutral, or negative, and whether the factual claims about your brand were correct. Flag inaccuracies for correction at the source — usually by strengthening or clarifying the on-site content the model is likely pulling from. Treat “mentioned but wrong” as a worse outcome than “not mentioned,” because it’s actively harder to unwind.
Symptom: the program launches, results start coming in next month, and three months later leadership asks “is this better than before?” — and there’s no before to point to.
Why it hurts: without a baseline, every number is just a number. You can’t tell a genuine content-driven improvement from ordinary run-to-run noise, and you can’t credit (or blame) specific work for a change in visibility. Teams that skip the baseline step almost always end up re-litigating “is this working” arguments with nothing but vibes to settle them.
The fix: before publishing a single piece of content aimed at improving Share of Model, run the full prompt set across all tracked platforms and record it as period zero — including sentiment and accuracy, not just mention rate. Re-run on a fixed cadence after that. When something changes, you’ll be able to trace it back to a specific content push, a competitor’s launch, or an underlying model update, instead of guessing.
Symptom: the team gets excited about being mentioned on broad, high-visibility prompts (“what is [category]”) while ignoring narrower prompts that map directly to buying decisions.
Why it hurts: a mention on a broad definitional prompt feels good in a screenshot but rarely influences a purchase decision — the user asking “what is [category]” is early-funnel and not choosing between vendors yet. Meanwhile the prompts that actually precede a purchase, like “best [category] for a 20-person team” or “[competitor] alternative,” get less attention because they’re less flattering to report on. Optimizing for the wrong prompt tier means winning applause internally while losing the deals that matter.
The fix: tier your prompt set by commercial intent and weight the reporting accordingly. A gain on a comparison or alternative prompt should carry more weight in the executive summary than a gain on a generic definitional prompt, even if the raw mention count moves the same amount.
Symptom: the report is thorough, the gaps are clearly identified — a competitor dominates a whole category of comparison prompts, or the brand is absent from a cluster of use-case prompts entirely — and then the same gaps show up unchanged in next quarter’s report.
Why it hurts: measurement that doesn’t feed action is theater. Share of Model reporting is only valuable insofar as it changes what gets published, updated, or clarified on the site. A gap identified and left alone doesn’t just fail to improve — it tends to widen, because a competitor who is winning a prompt cluster is usually still actively investing in the content that’s winning it.
The fix: treat every reporting cycle as a work order, not just a scoreboard. For each significant gap, assign a concrete content or structural fix — a new page, a clarified claim, an updated comparison, added schema — with an owner and a date. Close the loop by re-testing the specific prompts tied to that fix on the next cycle, not just the aggregate score. If a gap has appeared in two consecutive reports with no assigned fix, that’s a process failure worth escalating on its own.
There's no fixed magic number, but a set that's too small will bounce around from run-to-run noise and mislead you. Most brands need enough prompts to cover definitional, comparison, and use-case intent across their core categories, sampled multiple times each — a handful of prompts run once is not a measurement program, it's a screenshot.
No. Report platforms separately before blending, since they retrieve and cite information differently. A blended average can hide the fact that you're strong on one platform and nearly absent on another, which is exactly the kind of gap this kind of tracking exists to surface.
A mention is just an appearance in an answer. Real visibility accounts for how often you appear, how favorably, and how accurately — a brand mentioned frequently but described with wrong information or negative framing has a visibility problem dressed up as a visibility win.
Enough to catch real change without drowning in noise — a fixed, regular cadence (commonly monthly or quarterly, depending on how fast the category and content are moving) works better than ad hoc checks, because it gives you a consistent basis for comparison against your baseline.
The gap should turn into an assigned content or site fix with an owner and a re-test date on the next cycle. If gaps are only ever logged and never closed, the measurement work stops paying for itself.
Sometimes a single strong page closes a narrow gap, but most durable improvement comes from consistently clear, accurate, well-structured content across the prompts you care about, reinforced over multiple reporting cycles — not a one-time push.
Terry has 30+ years in software and SEO. He’s the founder of Salterra Digital Services and SEO Spring Training, host of the Roundtable SEO Mastermind, and lead instructor at SEO University — teaching the exact tactics his team uses on client work.
This guide is one lesson from the Measuring AI Visibility Share of Model course. Get every lesson, framework and checklist — plus the full 38-course catalog — inside SEO University.
Practitioner-focused training across the full digital marketing stack — from technical SEO to conversion optimization and the AI search era. By Salterra Digital Services, since 2011.