A wave of GEO tools and consultants are telling brands the same story. Track thousands of prompts across every LLM, every day, forever, or you’re invisible in AI search. The pricing follows the pitch. Dashboards with prompt counters ticking upward. Alerts for every fluctuation. Subscriptions that scale with prompt volume, which quietly means the vendor makes more money the more prompts you buy, whether or not you learn anything new from them.
That’s a business model borrowed wholesale from rank-tracking SEO tools and bolted onto a technology that doesn’t behave the same way. Before signing a contract sized around “thousands of prompts,” it’s worth understanding why that number is mostly arbitrary, and what a defensible, right-sized measurement practice actually looks like.
Why Prompt Tracking Isn’t Rank Tracking
Rank tracking works because the thing being measured barely moves, and when it does move, you can usually point to why. Google decides who ranks for “best running shoes” using factors that, beneath all the algorithmic complexity, are still mostly binary. A page is indexed, or it isn’t. A backlink exists, or it doesn’t. A query matches, or it doesn’t. So when a position shifts, there’s normally a real cause behind it, like a competitor picking up a link or an algorithm update rolling out. The search itself doesn’t change meaning between checks, and whatever’s producing the ranking isn’t guessing at anything.
An LLM is guessing at almost everything. It isn’t checking a fixed list of yes/no criteria and returning a position; it’s weighing probabilities across its training data and whatever it retrieves in the moment, and there’s no equivalent to “does this page have a backlink” that a person could audit after the fact. Five specific consequences follow from that.
1. Personalization. Traditional SEO has dealt with a version of this before. A client’s sales manager once searched for their biggest competitor every day out of habit, and Google’s logged-in results quietly started ranking that competitor higher for him personally. Run the same search incognito and the “true” organic position came back. The SEO industry solved this decades ago with a convention: incognito, logged-out, geo-controlled is the baseline everyone measures against, and a personalized result gets treated as the anomaly, not the ground truth.
GEO doesn’t have that convention, and it can’t just borrow it. Logging out of ChatGPT strips account history the way incognito strips cookies, but it doesn’t strip persona. The model can infer who’s asking from the prompt’s own wording, in a single turn, without any login at all. “I’m a developer evaluating CRMs” and “I’m a CMO evaluating CRMs” will shift the answer even across two anonymous sessions with nothing else different. Part of this problem is session-based and can be controlled in the same way rank trackers already handle it. The other part is baked into the prompt itself, and no logged-out setting fixes that.
2. Context and session history. LLMs carry conversational state. A prompt run cold produces a different answer than the same prompt run after three prior turns about budget or a competitor already mentioned. There’s really no such thing as a clean prompt run unless someone deliberately controls for session state on every single test, and most audits don’t bother.
3. Intent variability. “What’s the best noise-canceling headphone under $300?” and “Which headphones should I buy for under $300 that block noise well?” are the same intent phrased two ways. A keyword tool logs them as two distinct queries, each with its own volume. Prompt count isn’t a good stand-in for coverage because of this. Ten well-built prompts covering ten distinct scenarios will teach you more than a hundred paraphrases of the same three.
There’s actual data behind that, and it rhymes with something from search history. Before Google’s Hummingbird update, keyword tools often lumped synonyms and plurals together, “shoe” and “shoes” as one bucket, because the search engine itself had started treating them as the same question. A 2026 study by Peec AI, examining nearly 38,000 AI responses across five engines, found a similar effect with prompts on their own. When real people phrase the same underlying question in their own words, those phrasings cluster together far more tightly than most people assume. Roughly nine out of ten phrasing variations landed close enough in meaning to count, in practice, as the same question. Which brands showed up barely moved as long as the phrasing stayed inside that cluster. The real drop occurred only once the wording drifted far enough to become a genuinely different question, and at that point, brand visibility fell by roughly half.
In practice, that means one well-chosen prompt can often stand in for several close phrasings, rather than requiring a separate test for each. It’s worth deliberately building broader prompts intended to sit within the same cluster as the narrower variations they represent, and treating a strong result there as reasonably indicative of how the narrower ones would perform as well. There’s one place this doesn’t hold up: unbranded, comparison-style questions in the middle of the buying journey are far more sensitive to exact wording than top-of-funnel or branded questions are. Paraphrase count still matters there.
4. Output non-determinism. Run the same prompt five times, and you can get five different answer sets, purely from sampling randomness. A rank tracker checking the same keyword twice in an hour gets the same SERP. A prompt checked twice within an hour might legitimately surface a different set of competitors each time. The fix for that is to repeat the same prompt, not to add more prompts.
5. Judgment-based scoring. A SERP position is a fact. You’re #3, or you’re not. Scoring an AI answer requires a human being to decide whether your brand was the primary recommendation, a passing mention, or framed badly, and that judgment call doesn’t automate cleanly. Every additional prompt in a tracking set adds real review time, not just another row in a spreadsheet. Most “track thousands of prompts” tools quietly skip this part. They can report whether you showed up, but not how you were framed, which is usually the part that actually matters to a buyer.
Add it up and volume ends up looking like rigor without actually being rigor. A vendor running 3,000 prompts a week, without controlling for session state, without repeating runs to check stability, and without a human scoring the tone of the answers, has produced a lot of data and very little insight.
Where Do You Even Get the Prompts?
A question comes up in nearly every GEO workshop, almost always early. Where can I get my prompts? People ask it the way they’d ask for a report, expecting a button somewhere that exports the list. That expectation exists because paid search and Google Search Console spent twenty years training marketers to treat keyword discovery as a data pull. There’s no equivalent for prompts. No engine hands you a report of what people are asking about your category, and most of the GEO tools rushing to fill that gap are selling roughly the same promise the early keyword tools sold: point us at your site and your competitors, and we’ll generate the list for you.
It’s deja vu from the early days of search, except worse in one respect. Back then, nobody expected the list to exist already. It got built through judgment, customer interviews, sales call transcripts, and a sense of what people were actually confused about. Building that list was the strategy work, not a step leading up to it. A few years ago, I ran a workshop for a multinational client that was trying to build a keyword model from scratch. We opened with a six-slide questionnaire meant to force the room to think before anyone touched a tool. Two slides in, someone asked if we could just run software that pulls “all the related words” instead and skip ahead. We pushed through anyway, and by the end of that one session, done entirely by hand, the room had over 100 content ideas and just as many identified gaps. People finally understood how differently their buyers searched depending on how much they already knew about the category. An auto-generated prompt list will give you volume and feel like progress, but it skips the uncomfortable part where you map prompts to real awareness stages and admit where your existing content falls short. That part doesn’t shortcut well now any more than it did then.
A Word on “Visibility Scores”
Seth Besmertnik, CEO of Conductor, made a version of this same argument recently that’s worth repeating. The single unified “visibility score” that most AEO platforms lead with sounds authoritative, but it paper-overs the exact variables described above: persona, journey stage, topic, and whether the prompt is branded. His read on why different tools report different numbers isn’t that the category can’t be measured. It’s that most companies are measuring it too broadly and then hand leadership a number that looks precise but isn’t tied to anything real. A keyword search in the old world had one intent and one SERP that basically everyone saw. A prompt has neither. The same topic asked by different personas at different points in the journey can yield a genuinely different set of answers each time, and averaging across them produces a number that isn’t tied to any decision a real person is making.
So How Many Prompts Do You Actually Need?
Fewer than you’ve probably been told, and built with more care than a scraped keyword list.
Start with coverage, not count. The goal is representative coverage of the scenarios, brand questions, and product battles that actually matter to your business, not exhaustive coverage of every possible phrasing. A reasonable starting set looks like this:
- 10 to 15 Tier 1 (Intent/Scenario) prompts. The real jobs-to-be-done your buyers bring to AI, built with a consistent structure so results stay comparable over time.
- 8-12 Tier 2 (Brand Understanding) prompts. Identity, expertise, competitors, authority. This tier stays small because there are only so many ways to ask what a brand is and whether it’s credible.
- 10-15 Tier 3 (Product) prompts. Comparisons, use-case fit, and price or value questions, focused on your highest-revenue or highest-competition products rather than your full catalog.
That’s roughly 30-40 prompts to establish a baseline. Not thousands. The rest of the rigor comes from how you run them, not how many you run.
- Repeat each prompt three to five times across separate sessions to build a stability score, instead of adding new prompts to pad the set.
- Run the same set across the three or four engines your buyers actually use, since sourcing behavior differs meaningfully between them.
- Re-test on a fixed cadence. Quarterly works for most categories; monthly only makes sense in a fast-moving, heavily contested one.
- Have a human rate framing and positioning, not just whether the brand showed up at all.
Scaling past this baseline can absolutely make sense, but let the baseline tell you where to expand, rather than defaulting to whatever package size a vendor happens to sell. Adding more prompts without adding more review capacity just leaves you with a bigger pile of unscored data.
“How Much Should This Cost?”
There’s a question that comes right behind “how many prompts do we need,” and it’s really the same question in a suit. What should this cost? At my previous agency, clients would always ask how much of my time they’d actually get. My honest, half-joking answer was always the same: as little as possible, but as much as you’re willing to pay for. That’s not really a dodge. The floor and the price tag get set by two completely different things.
The same split is showing up here. Prompt-tracking quotes in the $60K to $ 100K-a-month range aren’t rare anymore, and the number keeps climbing. That’s not because the underlying measurement problem got any harder. It’s because the inputs feeding the quote keep multiplying: more LLMs to cover, more countries and languages, and increasingly a senior executive who’s decided AI visibility belongs on the board agenda and wants comprehensive coverage rather than a narrow index. Very little of that spend is driven by what the measurement actually requires. Most of it is driven by what the mandate can absorb. The floor is the 30- to 40-prompt mini-index approach above, run with real discipline. The ceiling is whatever a stakeholder’s appetite for certainty can fund, and both can be legitimate on their own terms. What isn’t legitimate is a vendor quoting the ceiling price while implying it’s the minimum requirement. Know which one you’re being sold before you sign.
The Honest Trade-off
More prompts, more engines, and more repetitions do genuinely surface things a small set will miss, like a competitor quietly winning over one persona or a category where you’re invisible on only one engine. Vendors selling volume have a point that scale reveals more. Where they’re wrong is treating scale as the first problem worth solving, as if it could substitute for the slower, less automatable work of actually scoring how you’re framed.
Run a 30- to 40-prompt baseline with real discipline around repetition and human scoring, and it’ll tell a mid-sized brand nearly everything it needs to know about where to invest next. That’s a decision worth making before committing to a budget built around a prompt count somebody else picked for you.

