Originally posted on my Signal & Friction Substack August 14, 2026
As organizations try to understand why brands appear in AI-generated answers, studies such as AirOps’ 2026 State of AI Search are beginning to provide meaningful insights and actions.
There is a great deal of useful information in the report. There is a lot here that I agree with, and some of the findings reinforce things many of us have observed independently. My concern is less with the findings than with what inevitably happens to findings like these once they enter the LinkedIn ecosystem.
The SEO industry has a long history of taking an observed correlation, removing the conditions around it, and turning it into a best practice. That becomes a checklist item in audits, and before long, teams are spending time and money on it because it has become accepted as an optimization practice rather than because anyone has demonstrated that it materially improves the outcome.
AI Search is especially susceptible to this because everyone wants certainty in an inherently uncertain environment. AirOps acknowledges this right at the beginning of the report:
AI Search has no fixed ranking, predictable position changes, or stable “page one.”
That should influence how we interpret the rest of the report. These are signals worth paying attention to, but before turning them into requirements, I think we need to ask what they actually mean for our business.
Four areas stood out to me.
Freshness Should Reflect the Information, Not the Calendar
AirOps found a compelling relationship between freshness and AI citations. Pages not updated quarterly were more than three times as likely to lose citations, more than 70% of cited pages had been updated within the previous year, and more than half had been refreshed within six months.
It is not difficult to imagine how that gets interpreted: update important content every quarter.
This was actually the first thing I flagged when reading the report. There were three sequential headlines emphasizing freshness and the three-month window. To its credit, AirOps adds important qualifications around industry, prompt context, and commercial queries involving current pricing, features, and claims.
Unfortunately, “it depends on the freshness requirements of the information and intent” is not nearly as catchy as “content less than three months old is 3x more likely to get cited.”
That headline is raw meat for GEO bros and content-generation tools.
Does McDonald’s now need to “refresh” its Big Mac page every quarter if the recipe and calorie count have not changed? Does my local landscaper need to rewrite its About Us page every 90 days? What about a page explaining its spring overseeding service if the process is essentially the same as it was last year?
Now compare that with a bank displaying current CD rates. Those rates can change weekly. If I am asking an AI system to recommend the best current CD, a rate page that appears to be a month old could legitimately be less useful than one demonstrating that its rates were verified yesterday. Even if the rate has not changed, the fact that it was recently validated has value.
But the same bank might have another page explaining how to evaluate CDs: APY, maturity, compounding, term length, and early withdrawal penalties. Those concepts do not become stale simply because nobody rewrote the explanation this quarter.
This suggests a different question from “When was this page last updated?”
Does the information or intent have a freshness eligibility requirement, and if so, how current does it need to be?
That also makes me question whether the URL is even the right unit for measuring freshness. A single page can contain today’s interest rate, an offer that expires next month, a policy that changes once a year, and educational information that may remain accurate for a decade. Those facts have very different rates of change despite sharing a single last-modified date or an artificial statement of freshness in dynamically updated, reviewed, and published dates.
There is little value in manufacturing freshness merely to satisfy a quarterly requirement. We need to keep the information as current as the decision it supports requires.
Otherwise, we are maintaining timestamps rather than maintaining knowledge.
Structure Makes Sense. The Checklist Version May Not.
The findings on structure were among the least surprising to me. AirOps found that sequential heading structures are associated with a 2.8-fold higher likelihood of citation. Among ChatGPT-cited pages, 68.7% followed a logical heading hierarchy, 87% used a single H1, and nearly 80% contained lists.
That makes sense.
Well-organized information is easier for people to understand. There is no reason to believe machines would benefit from unnecessary ambiguity either. Headings establish context, sections group related information, and lists and tables make certain types of information easier to scan and compare.
Where I become cautious is in translating that into an “AI-friendly content format.”
We are already seeing recommendations to convert headings into questions, shorten paragraphs, add more lists, and break information into increasingly smaller chunks. Taken far enough, every page starts to look like an FAQ.
That is not what I take from this research.
A heading should tell me what the section is about. A table is excellent when I need to compare things. A list works when the information is actually a list. Some concepts need explanation, evidence, qualification, or a narrative that connects several ideas. Breaking those into five bullets does not necessarily make the information easier to understand.
For me, the useful test is much simpler: can a person or machine understand what this section is about, how it relates to the larger subject, and where the relevant evidence is located?
That is structured content. It is not necessarily templated content.
Three Schema Types? Which Three?
AirOps reports that approximately 61% of cited pages use three or more schema types, and that pages with three or more types have a 13% higher likelihood of citation. It describes schema as providing relevance cues and recommends using multiple types appropriate to the page’s purpose and content
For someone who has advocated for structured data for years, that is encouraging.
My first reaction was also: Which three?
Article + WebPage + BreadcrumbList? Does that count?
What about Organization?
Are the schema types describing the information that was actually cited, or are they describing the page, publisher, and navigation?
And if 61% of cited pages have three or more schema types, what was happening on the other 39%?.
Reading that statement, my first thought was: What are the three types, and how are they contributing? They never list them other than stating to use FAQ schema when the page answers direct user questions.
Article + WebPage + Breadcrumb? Does that count? Organization? Is the schema describing the information being cited or just the page and its navigation? And if 61% of cited pages have three types, what did the other 39% have?
Those details matter if we want to turn the finding into something actionable.
There is another variable worth considering. Organizations sophisticated enough to implement several appropriate schema types may also have stronger technical SEO, better information architecture, clearer entity definitions, better publishing systems, and more mature content governance.
Schema may be contributing directly to machine understanding. It may also be an indicator of a better-organized information environment. Both could be true.
That is why I would be very reluctant to convert this finding into “three schema types produce 13% more AI visibility.”
I would rather know what those types represented, what information they exposed, whether the cited information was contained in the markup, and whether particular schema types or combinations were associated with specific answer types.
That would tell us much more about the mechanism.
And What About FAQ Schema?
AirOps specifically recommends FAQ and QA schema when a page answers direct user questions. This is another area where I think we need to separate a reasonable implementation practice from an AI visibility claim.
Google dramatically reduced the visibility of FAQ rich results in 2023, primarily limiting their regular appearance to authoritative government and health sites. That led some organizations to reduce or remove FAQ markup because the search presentation benefit that originally justified the effort had largely disappeared.
That did not mean Google suddenly stopped understanding structured data, nor does it mean FAQ schema has no value. Google continues to tell us that structured data can provide explicit clues about the meaning of a page.
The question I have is narrower:
Does the FAQ markup itself materially affect AI retrieval, interpretation, or citation?
I increasingly hear the argument that Google is effectively the gateway through which schema reaches AI. Google understands structured data; Google Search helps ground or retrieve information for its AI experiences; therefore schema improves AI visibility.
The first two statements are reasonable. The third is where we have made an inference.
Google’s current guidance says AI Overviews and AI Mode do not require special schema.org markup or additional machine-readable files. Structured data should accurately represent the visible content, but Google does not identify FAQ schema as a requirement for AI visibility.
So I would like to understand what AirOps is actually detecting.
Is FAQ schema improving citation likelihood? Is clearly written question-and-answer content responsible for the difference? Does structured data improve Google’s understanding upstream in a way that subsequently affects retrieval? Or are pages implementing FAQ schema simply more likely to be well structured in the first place?
Those mechanisms lead to very different recommendations.
Google devalued the FAQ rich result, not necessarily FAQ schema as a semantic mechanism. There may well be value here. I just don’t think the correlation gets us all the way to “add FAQ schema for AI.”
You Can Say You’re Great. Does Anyone Else?
The Community and Offsite sections may be the findings I like most in the report.
AirOps reports that approximately 48% of AI Search citations come from user-generated and community sources such as Reddit, LinkedIn, Wikipedia, and YouTube, with meaningful differences among models. During early commercial discovery, approximately 85% of brand mentions come from third-party domains rather than brand-owned properties.
I have been saying some version of this for many years: you can say you are great, but does anyone else say you are great?
That was part of what made PageRank so powerful. A company could make all the claims it wanted on its own website, but references from elsewhere on the web provided another form of validation. AI has access to a much richer version of that evidence.
A hotel can say it is beachfront. Travelers can explain that getting to the beach actually requires walking two blocks.
A software company can accurately describe its platform as highly configurable. Reviews on G2 or discussions on Reddit can confirm that while adding an important qualification: it may be so configurable that you need significant expertise, perhaps even a dedicated employee, to take full advantage of it.
The third party has not necessarily contradicted the brand. It has added information that matters to the decision.
This is why I agree strongly with the report’s emphasis on community and offsite presence. It is also the area where I expect the optimization industry to create the most trouble.
Tell marketers that Reddit influences AI recommendations and Reddit becomes a tactic.
Tell them that third-party mentions correlate with AI visibility and somebody will create a KPI for the number of mentions acquired this quarter.
Then we are right back where we started: link building and article spinning, but with better tools.
AirOps also found remarkably similar correlations between nofollow and dofollow links and AI visibility: 0.509 for nofollow and 0.504 for dofollow. Those are both moderate positive relationships, but what caught my attention is how little difference there is between them.
That does not tell us that nofollow links suddenly “count” for AI or that AI systems treat both link types as equivalent signals. Spearman correlation tells us that two things tend to vary together. It does not tell us why.
A widely recognized brand will naturally accumulate more links of both types. It will probably also accumulate more reviews, discussions, comparisons, citations, videos, articles, and other evidence that AI systems may encounter.
So the interesting question is not whether we should start building nofollow links.
It is what this correlation is actually measuring?
Search engines spent years learning that all links were not equal. Authority mattered. Relevance mattered. Placement and context mattered. Manipulation certainly mattered.
I expect AI systems to face much the same problem with mentions and citations.
Think about an academic paper. A bibliography containing 100 sources tells us something. A precise inline citation supporting a particular assertion tells us considerably more about the relationship between that source and the claim being made.
The same distinction should influence how we think about offsite AI visibility. I am less interested in how many places mention a company than in who mentions it, what they say, the context in which they say it, and which customer decision that evidence helps resolve.
Reporting May Be the Hardest Part
The section on citation volatility may ultimately be the most consequential part of the AirOps report.
Only 30% of brands remained visible across back-to-back answers. Brands that were both mentioned and cited had a 40% greater likelihood of resurfacing, yet only about 28% of answers contained both signals. More than half of the brands that disappeared subsequently resurfaced within two runs.
AirOps reasonably recommends tracking AI visibility over time rather than relying on a single snapshot. I agree. What I am less comfortable with is how quickly we are rebuilding traditional rank tracking around prompts.
Synthetic monitoring is useful. Run the same prompts repeatedly, under controlled conditions, and we can begin to see patterns in how often a brand appears, which sources are cited, and how representation changes.
The problem comes when we start treating that measurement as though it represents a stable position or what every customer will see.
I ran into this while preparing a presentation for International Search Summit. I asked for running shoe recommendations for someone preparing for their first marathon. The answer included shoes from ASICS and New Balance, but some of the links went to retailers rather than the manufacturers.
I wanted to know why.
The system explained that the context it had about me suggested I was price-conscious and that those retailers were more likely to have discounts.

Some attendees ran similar exercises and received different recommendations based on the context available about them. The important point was not whether my answer or theirs was better. The prompt alone did not explain the result.
A synthetic monitoring system can reproduce the words I entered. It cannot necessarily reproduce the context that influenced the answer.
That does not make synthetic prompt monitoring useless. Far from it. But it does mean we need to be careful about what we claim the resulting number represents.
If AirOps is finding that only 30% of brands persist between consecutive answers under its testing conditions, even before we introduce differences in user context, then an AI “ranking” can imply a level of precision the environment itself may not support.
We should measure patterns. We should repeat tests. We should control what we can control. Most importantly, we should document the testing conditions so that management understands what the resulting visibility score actually represents.
That is very different from saying, “We rank third in ChatGPT.”
Don’t Skip the Thinking Step
I came away from the AirOps report with more questions, which I consider a positive outcome. Research like this gives us evidence in an area still overflowing with opinions, anecdotes, and vendor claims. We need more of it.
But research identifies patterns across a population. The next step is determining whether those patterns apply to our particular business, information, customers, and use cases. Only then should we decide what work is justified.
If freshness correlates with citations, identify which information actually has a freshness requirement. If clear structure correlates with visibility, determine whether your information is understandable and retrievable before redesigning every page as an FAQ. If richer schema correlates with citations, understand what that markup represents and how it might contribute before setting a target of three schema types per page. If community and offsite recognition matter, identify the external evidence that actually helps customers make decisions rather than simply accumulating mentions.
The danger is not the research; the danger is skipping that thinking step.
That is how an interesting correlation becomes a best practice, the best practice becomes a checklist, and six months later somebody is reporting that 97% of the website is “AI optimized” without being able to demonstrate what materially improved.
