Claude Penland

By Claude Penland - marketing and business strategy for companies that are good at what they do and hard to find.

Audit the Machine: How to Know You Have Gone Invisible Before the Traffic Tells You

A question-and-answer guide to measuring whether AI assistants can actually see your company. Published by 1000Startups.com. Every figure below is attributed to a named study in the sentence that uses it, and the full source list is at the end.

Everybody is publishing tactics for AI search. Almost nobody is publishing a way to check whether any of it worked. This is the checking part. It is less fun than the tactics and considerably more valuable.

Q: Everyone is shipping AI search advice. What is missing from all of it?

A measurement layer. The advice market is saturated: add schema, write listicles, get on Reddit, refresh your dates. Some of it is even correct. But almost none of it comes with an instrument that tells you whether the needle moved, which means the entire category currently runs on vibes and screenshots.

That gap is the opportunity. Tactics are commodity. A repeatable scoring method is not.

Q: My website is solid and it ranks. Isn’t AI just reading it?

Mostly, no. You are being talked about far more than you are being read, and three independent datasets landed in roughly the same place.

AirOps, analyzing more than 21,000 brands for its 2026 State of AI Search report, found that about 85 percent of brand mentions in AI answers originate on third-party pages rather than the brand’s own domain, with owned domains accounting for roughly 13 percent. OtterlyAI went bigger and got a starker number: reviewing more than one million citations across ChatGPT, Perplexity, and Google AI Overviews for its February 2026 report, it put third-party dependence at about 95 percent. And DerivateX ran a narrower, more surgical test: one buyer-style question for each of 40 B2B SaaS categories, repeated ten times, producing 233 recommendations across 219 tools. When ChatGPT recommended a tool, it cited that tool’s own website 11.6 percent of the time.

Your website is not the scoreboard. It is a reference the referee occasionally consults.

One detail from the DerivateX study deserves its own sentence, because it will reorganize somebody’s budget: review aggregators including G2, Capterra, and TrustRadius accounted for 0.9 percent of all citations, and G2 and Capterra each received zero. The sources ChatGPT reached for instead were independent and niche blogs plus vendor-published content, which made up 81.9 percent of citations. Everyone optimizing for the badge was optimizing for the wrong shelf.

Figure 1. Three studies, three methodologies, one uncomfortable agreement.

Q: I asked ChatGPT about my category and we showed up. Are we winning?

You have one data point, and one data point is a mood, not a measurement.

AirOps found that only about 30 percent of brands stay visible from one answer to the next on the same question, and just 20 percent are still there across five consecutive runs. You are not ranked. You are sampled.

SparkToro’s Rand Fishkin and Gumshoe’s Patrick O’Donnell put a finer point on it in early 2026: 600 volunteers ran 12 prompts through ChatGPT, Claude, and Google AI a combined 2,961 times. Fewer than 1 in 100 runs produced the same list of brands, and fewer than 1 in 1,000 produced that list in the same order. Celebrating a single appearance is like taking one poll and canceling the election.

Fast Eddie Felson had this exact problem, and it cost him more than a marketing budget. Twenty-five hours into the marathon match in The Hustler (1961), Paul Newman’s Eddie is up eighteen thousand dollars on Minnesota Fats. Up is not the same as done. He keeps playing, and by morning the eighteen thousand is gone along with all but two hundred dollars of the stake he walked in with. One good rack is not a standing. Screenshot the win if it makes you happy. Do not file it as a position.

Figure 2. Left: visibility is sampled, not held. Right: the old proxy stopped working.

Q: All right, I’m convinced. What do I build first?

Thirty prompts. Not a tool, not a dashboard, not a vendor contract. Thirty written questions, split evenly:

  • Ten on how your category gets described by people who don’t work at your company.
  • Ten in the language your buyer uses to describe the problem, before they know your category has a name.
  • Ten head-to-head comparisons, including the competitor you find most irritating.

Then write them down and stop editing them. A measurement you keep improving is not a measurement, it is a mood ring with a spreadsheet attached. The prompts can be imperfect. They cannot be moving.

Bert Gordon’s verdict on Eddie is that he has talent but no character. It is the cruelest line in the movie, and it is also an uncomfortably fair description of most marketing measurement: no shortage of ability, no willingness to run the same test twice. Character, in this context, is deeply unglamorous. It is asking the identical thirty questions in March that you asked in February, including the four that made you look bad.

One design note worth stealing: OtterlyAI’s data shows real user prompts average 15.1 words against 8.8 words for prompts marketers guess at. Write yours long and conversational, the way an actual person types at 11pm.

Q: How many assistants, and how often?

Five assistants, monthly, in the same week every month.

Different models read different corners of the internet, and the differences are not subtle. OtterlyAI found that Google AI Overviews pulls about 59.8 percent of its citations from brand-owned sites, while ChatGPT leans on Reddit, Wikipedia, and news for a combined 39.5 percent. A brand can be dominant in one engine and functionally absent from another, and the average of those two numbers describes nobody who exists.

Same week each month matters more than which week. You are looking for change over time, and change over time is only legible against a fixed cadence.

Q: What exactly am I scoring?

Three things. Most teams score one, which is why most teams learn nothing.

MetricThe question it answersWhy it earns its row
Mention rateOut of 150 runs, how many named us?It is the only number most teams track, and on its own it is the least useful of the three.
Citation sourceWhich exact URL got us there?Roughly 85 to 95 percent of the time it will not be a page you own. This column is your real distribution map.
Position stabilityDo we survive the second identical question?This is the leading indicator. It moves before mention rate does, and mention rate moves before traffic does.

If you want a fourth, AirOps found that brands earning both a mention and a citation in the same answer are about 40 percent more likely to resurface across consecutive runs, yet only around 28 percent of answers contain a brand with both. Dual-signal presence is rare and it is sticky. Track it.

Q: My organic traffic is fine. Why should I care right now?

Because a boat does not learn about the reef from the impact.

Organic traffic is the impact. Retrieval share is the depth sounder, and it moves first, quietly, months ahead of anything your analytics package will show you. By the time the traffic chart bends, the decision that bent it was made two quarters ago by a retrieval system you were not watching.

This is the least glamorous argument in the whole piece and it is the one that pays for the program.

Q: Do I really need to log which page produced each citation?

It is the single most useful column you will keep.

When a citation appears, record the exact third-party URL behind it. That list is your real distribution map, and it will look almost nothing like the one in your marketing plan. AirOps found that nearly 90 percent of third-party citations come from listicles, comparison pages, and review roundups, and that roughly 80 percent of cited brands appear within the first three positions of that page.

Which means the work is not always “write more.” Sometimes the work is one email to one editor about moving you from seventh to third in a roundup that already exists.

Q: Should I keep my scorecard private? It feels like an edge.

Publish it. Politely, but publish it.

The ground is moving fast enough that method itself has become scarce. Ahrefs analyzed 863,000 keywords and roughly 4 million AI Overview URLs and found that only 38 percent of cited pages also ranked in Google’s organic top 10 for the same query, down from 76 percent in its July 2025 study. Whoever publishes a credible, repeatable scoring method during a period like this gets quoted as the person who defined it.

And publish the caveats too, because they are what make you trustworthy. Search Engine Journal’s coverage of that Ahrefs study flags that part of the 76-to-38 drop reflects improved citation detection in Ahrefs’ own tooling rather than a pure change in Google’s behavior, which means the two waves are not perfectly comparable. A separate BrightEdge analysis, using different methods, put the top-10 overlap closer to 17 percent. Three numbers, three methodologies, same direction. Say that out loud in your write-up. Sources that disclose their own error bars are the ones that get cited.

Q: We got cited. Can we stop now?

A retrieval position is not a trophy you won. It is a lawn, and it browns.

Seer Interactive’s log-file analysis found that roughly 65 percent of AI bot crawl activity targets content published within the past year, with about 89 percent hitting content from the last three. Amsive’s citation-freshness work found that half of all cited content is under 13 weeks old. AirOps reports that pages not updated quarterly are about three times more likely to lose their citations.

Midway through that same marathon, Fats sets down his cue, walks to the washroom, combs his hair, straightens his tie, cleans his hands, and has talcum powder poured over them. He comes back looking like he just arrived. Eddie stays at the table with the bourbon. Fats wins everything back. Refreshing is not vanity, and it is not a side quest. It is most of the strategy, and it is the least interesting reason anyone has ever won anything.

Set a quarterly refresh cycle on your top pages and a monthly one on anything in a fast-moving category. And refresh substantively. Crawlers can diff your page against the version they saw last time, so bumping the date without changing the content is a trick that stopped working a while ago.

Q: What is this actually worth after a year of doing it?

A dataset nobody else bothered to collect.

Twelve months of thirty fixed prompts across five assistants is 1,800 observations, each tagged with a mention, a source URL, and a stability flag. At that point you are no longer guessing which third-party publishers move your category, because you have counted. You can tell a prospect what happened to their visibility in the eleven weeks after a refresh, with a number.

Tactics are cheap and everywhere. Measurement is rare, dull, repeatable, and it is the thing that turns a service into a product. Score it monthly for a year and the dataset becomes the moat.

Sources

Every figure cited above, in the order it appears.

AirOps, “The 2026 State of AI Search” (21,000+ brands) — https://www.airops.com/report/the-2026-state-of-ai-search

AirOps, “The Influence of Offsite Signals in AI Search” — https://www.airops.com/report/the-influence-of-offsite-signals-in-ai-search

OtterlyAI, “The AI Citation Economy” (1M+ citations, February 2026) — https://www.globenewswire.com/news-release/2026/02/19/3241387/0/en/otterlyai-unveils-groundbreaking-data-ai-search-engines-depend-95-on-third-party-sources.html

DerivateX, “B2B SaaS AI Citation Study” (40 categories, 233 recommendations) — https://derivatex.agency/report/b2b-saas-ai-citation-study/

SparkToro (Rand Fishkin) and Gumshoe, AI recommendation consistency study (2,961 runs) — https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/

Ahrefs, AI Overview citations and organic rankings (863,000 keywords, ~4M URLs) — https://ahrefs.com/blog/search-rankings-ai-citations

Search Engine Journal, coverage and methodology caveats on the Ahrefs findings — https://www.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637/

Seer Interactive, “AI Brand Visibility and Content Recency” (log-file analysis) — https://www.seerinteractive.com/insights/study-ai-brand-visibility-and-content-recency

Seer Interactive, “Content Recency’s Impact on AI Visibility in 2026” (follow-up) — https://www.seerinteractive.com/insights/study-content-recencys-impact-on-ai-visibility-in-2026

OtterlyAI, prompt-length and engine-mix data — https://otterly.ai/blog/ai-keyword-research/

Figures compiled by 1000Startups.com from the studies listed above. Percentages are reported as published; methodologies differ between studies and are not directly comparable.

Claude Penland

Claude Penland builds the marketing and business strategy for companies that are good at what they do and hard to find. Thirty years operating, one exit, eight of them as a practicing casualty actuary.

The free two-page read is genuinely free. Email claude@1000startups.com and I'll send back what I can see from the outside. Or see the work samples and how to work with me.

Leave a Reply