For most of the history of search marketing, measurement started with a relatively straightforward relationship: a person entered a query, a search engine returned a ranked set of results, and we measured whether our page appeared.
There were always complications. Results varied by location, device, personalization, search history, and eventually Google’s increasingly sophisticated understanding of intent. Hummingbird was an important milestone because it demonstrated that Google was becoming better at interpreting the question behind the words rather than simply matching the words themselves.
I remember watching search results change as pages that succinctly answered the inferred question began displacing pages that were optimized around the query but withheld the actual answer behind a lead-generation form or commercial experience.
Even then, however, the basic unit we monitored remained the keyword.
AI Search changes that relationship considerably.
We are now building prompt libraries, running them repeatedly across models, counting brand mentions and citations, and creating visibility scores. That is a logical evolution of search reporting, particularly given the volatility we are seeing in AI-generated answers.
The friction comes when we treat the prompt as simply a longer keyword.
It isn’t always a query to be matched. Increasingly, it is a decision to be solved.
📡 The Signal: Prompt Monitoring Gives Us a New Way to Measure Visibility
There is considerable value in prompt monitoring. If customers are asking AI systems questions about products, services, brands, and purchasing decisions, organizations need a way to observe whether and how those appear in the answers.
Unlike a traditional ranking, however, a brand’s appearance is often the result of a much more complex evaluation.
Consider a prompt I use in a recent webinar on how to get included in AI results:
“What is the best all-inclusive, family-friendly, beachfront resort in Cancun?”
A traditional keyword approach might view this as a collection of related concepts or perhaps as a very long-tail variation of “best all-inclusive resorts Cancun.”
At its simplest, we could represent the prompt this way:
[(All-Inclusive) + (Family-Friendly) + (Beachfront) + (Cancun)] × (Definition of Best)
That looks relatively straightforward, though a bit geeky. It isn’t.
Friction: The Prompt Has an Order of Operations
Those of you who have successfully forgotten most of your high school algebra may need a reminder that an equation has an order of operations. We cannot simply work from left to right, solve each piece independently, and assume that combining the results produces the correct answer.
A multi-criteria prompt creates a similar problem.
Before the system can determine the best all-inclusive, family-friendly, beachfront resort in Cancun, it has to resolve what each of those criteria means and determine how the available properties perform against them.
The decision, represented more formally as an optimization problem, starts to look something like this:

Where:
- (h) = represents “how well” something meets the criteria
- A(h) = represents how well the hotel satisfies all-inclusive (A)
- F(h) = how well it satisfies family-friendly (F)
- B(h) = how well it satisfies beachfront (B)
- Q(h) = overall quality or whatever contributes to best (Q)
- T = the minimum threshold required to qualify
- w = how important (weight) each criterion is to this person
And that’s where things become interesting, as the equation looks more like something from an economics research paper than a vacation question, which is partly the point.
It is not intended to suggest that an LLM literally executes this formula. It is a conceptual model of the operational problem hidden inside the prompt.
Before any of these individual calculations can occur, however, the system must first establish what those variables mean.
What qualifies as all-inclusive? Are meals and drinks sufficient, or should the evaluation include premium restaurants, top-shelf alcohol, activities, airport transfers, children’s programs, water sports, and resort fees?
What constitutes family-friendly? Is a children’s menu sufficient, or should the evaluation consider a kids club, connecting rooms, children’s pools, water slides, babysitting, teen activities, room occupancy, and family dining?
What does beachfront mean? Does the property open directly onto the sand? Do guests walk across the pool deck? Cross a road? Walk two blocks? Is beach access sufficient to qualify as beachfront?
Even Cancun requires interpretation. Are we limiting the answer to the Hotel Zone or Cancun proper, or does Playa Mujeres or the broader Riviera Cancun satisfy the geographic intent?
Only after those definitions are established can the model determine which hotels are actually eligible.
If each is treated as a true independent variable (eligibility gate), a zero on one criterion knocks the property out.
Then:
Best(h)=Eligibility(h)×WeightedDecisionScore(h)
That distinction is important because “best” isn’t another attribute like beachfront.
It is the conclusion produced after resolving the other criteria.
And the criteria aren’t necessarily equally important.
For a family with three young children (the weight of being Family Friendly might be greater than it being Beachfront):
wF>wB
For the family whose definition of a vacation is literally waking up and walking onto the sand:
wB>wF
For a price-sensitive traveler, price may enter the equation even though the person never put it into the prompt.
So the actual problem starts looking more like:
Answer=f(Prompt,Criteria,Definitions,Thresholds,Weights,Person,Context,Evidence)
That is what I think is missing from much of AI Search reporting: decision criteria.
We are measuring whether a brand appears for a prompt without adequately examining why it qualified or failed to qualify for the decision embedded within that prompt.
⚙️ The Friction: From Best of Breed to Best in Show
If the math gave you traumatic flashbacks to high school, let’s use a dog show to offer a more intuitive way to understand the problem.
Before there can be a Best in Show, the dogs are evaluated against criteria appropriate to their breeds. The objective is not simply to identify the dog possessing the greatest number of generically desirable attributes. Each entrant has to be evaluated within an appropriate framework before the winners can be compared at the next level.
A multi-criteria prompt creates a similar operational problem.
The system has to evaluate how well each property satisfies all-inclusive, family-friendly, beachfront, and the geographic requirement. Some of those criteria may function as eligibility gates. An exceptional adults-only resort does not become slightly less attractive for a family because it has a spectacular beach and outstanding restaurants. If family-friendly is a requirement, it may simply be ineligible.
Other criteria are relative. Two properties may both legitimately qualify as beachfront, while one provides immediate access to the sand and another requires walking across the resort grounds. Both survive the qualification stage, but they may receive different evaluations.
Only after those individual criteria have been resolved can the system attempt its equivalent of Best in Show: which of the qualifying properties represents the best combination for the complete decision?
Importantly, the winner does not have to be individually first in every category. A property with the absolute best beach may have weaker family amenities, while the resort with the strongest children’s program may have a merely good beach. The final recommendation requires reconciling those strengths.
This is where “best” becomes the best of the best, rather than simply another word in the prompt.
The Criteria Can Change Each Other
The problem becomes more complicated because the criteria are not necessarily independent.
Consider all-inclusive.
For a couple, premium dining, nightlife, and unlimited top-shelf spirits may substantially increase the attractiveness of an all-inclusive package.
For a family with young children, those same benefits may have considerably less influence. Included meals throughout the day, snacks, water activities, a strong kids club, family-friendly dining, and appropriate room configurations may carry much more weight.
The resort has not changed, and the definition of all-inclusive has not necessarily changed. What changed is the value assigned to the components of all-inclusive because family-friendly is also part of the decision.
This is one of the fundamental differences between matching a keyword and resolving a decision.
The relationship among the criteria can matter as much as the individual criteria themselves.
💥 The Realization: And Then We Add the Person
The weighting becomes even more complicated when we recognize that the person asking the question may influence the evaluation.
I encountered this while preparing a presentation for the International Search Summit. I asked for running-shoe recommendations for someone training for a first marathon. The recommendations included products from ASICS and New Balance, but some of the links went to retailers rather than the manufacturers.

When I asked why, the system explained that the context it had about me suggested I was price-conscious and that those retailers were more likely to offer discounts. In this refresh, it showed me links to ASICS. When I checked, the products were indeed cheaper there.
Some of the conference attendees performed a similar exercise and received different recommendations based on the context available about them.
The products had not changed. The basic need had not changed. What changed was some of the context used to evaluate the decision.
Our already complicated conceptual model therefore expands again:
Answer = f(Prompt, Criteria, Definitions, Thresholds, Weights, Person, Context, Evidence)
Again, this is not intended as an attempt to reverse-engineer an LLM. It illustrates why the words contained in a prompt may not fully explain the resulting answer.
A synthetic prompt tracker can reproduce the words I entered.
It cannot necessarily reproduce me.
🛠️ Prompt Fan-Out Can Break the Math
This distinction becomes particularly important as AI Search tools use fan-out to expand a prompt into related questions.
Take our original request:
“Best all-inclusive, family-friendly, beachfront resort in Cancun.”
A fan-out process might generate:
- Best resorts in Cancun
- Best all-inclusive resorts in Cancun
- Best family-friendly hotels in Cancun
- Best beachfront hotels in Cancun
- Best Cancun resorts for kids
Each is a legitimate question, and measuring them may reveal useful information.
But they are not equivalent to the original decision.
The customer did not ask five independent questions. The customer asked which property best satisfies the intersection of the criteria.
This brings us back to our forgotten high school algebra. Breaking an expression into pieces and solving each piece independently does not necessarily preserve the original order of operations.
The same can be true of prompts.
A resort catering primarily to couples may legitimately be one of the best all-inclusive resorts in Cancun. It should probably disappear when family-friendly becomes an eligibility requirement. A family resort may perform exceptionally well once children’s needs are included while appearing less prominently in the generic all-inclusive evaluation.
If those prompts are aggregated into a single visibility score, we may produce a very precise measurement of several different decisions.
Prompt fan-out tells us what questions surround a decision. Decision decomposition tells us how the decision is actually made.
That distinction matters enormously for reporting.
🏢 Boardroom Moment: Why Prompt Reporting Is Harder Than Keyword Reporting
Traditional keyword reporting primarily asked whether a document appeared for a query and where it appeared. Search personalization and localization complicated the measurement, but the ranked result remained a reasonably useful abstraction.
AI Search asks us to observe something different.
A brand may have to qualify against multiple criteria. Those criteria may contain thresholds. Their relative importance may change depending on other criteria. The definition of “best” may vary by customer. Available evidence may change, retrieval may surface different sources, and the generated answer itself may vary between runs.
The measurement problem is therefore not simply:
Did we match the prompt?
It is increasingly:
Did we qualify for the decision represented by the prompt, and why?
That distinction should influence what we put into AI Search dashboards.
Knowing that a hotel was mentioned in 63% of synthetic runs is useful.
Knowing why it disappeared from the other 37% may be considerably more valuable.
Was it not considered family-friendly? Did another property have stronger evidence of beachfront access? Was its definition of all-inclusive less complete? Did community reviews contradict the hotel’s claims? Did the system interpret Cancun using a different geographic boundary? Did the relative weighting of the criteria change?
A traditional visibility metric tells us what happened. Decision diagnostics can begin to tell us why.
Note: I recently wrote this post on “How many Prompts Do you Need?
From Prompt Visibility to Decision Diagnostics
Imagine replacing a report that says:
Brand mentioned: Yes
Citation present: Yes
Position: 3
with one that can begin to explain:
All-inclusive: Qualified
Family-friendly: Qualified
Beachfront: Did not qualify — available evidence suggests a two-block walk
Cancun: Qualified
Overall recommendation: Excluded because a required criterion was not satisfied
The second report provides the organization with something to investigate.
Perhaps the AI system is wrong, and the hotel needs clearer evidence demonstrating its actual location. Perhaps the hotel’s own language is ambiguous. Perhaps third-party sources contradict its claims. Or perhaps the property really is two blocks from the beach, and the AI has correctly identified a gap between the marketing description and the customer’s definition of beachfront.
Those scenarios require very different actions.
A hotel cannot do much with “we rank fourth in ChatGPT.”
It can do something with “we are consistently excluded because AI systems do not find sufficient evidence that we satisfy this decision criterion.”
That is the transition from visibility measurement to decision diagnostics.
The Signal and the Friction
The signal is that prompt monitoring gives organizations a valuable new way to observe how AI systems represent, consider, cite, and recommend their brands. We should absolutely measure those patterns, particularly over repeated runs and across the models relevant to our customers.
The friction is assuming that the resulting data behaves like keyword rankings simply because we can put it into a dashboard that looks like a rank tracker.
A prompt can contain multiple decision criteria. Those criteria have to be defined and evaluated. Some may determine eligibility, while others influence relative preference. They can interact with one another; their weightings can change depending on the person and context; and the evidence available to support them can change over time.
We should continue monitoring prompts, but the next stage of AI Search reporting should help us understand more than whether a brand appeared. It should help us understand the criteria that determined whether the brand qualified in the first place.


