Two thirds of the sources AI cited were notthe business’s own website.
We asked ChatGPT and Claude fifty med-spa questions about St. Louis, with live web search on every call, and recorded every business named and every source cited. This is a pilot — 76 usable answers, one metro, one vertical, one point in time. It is not a trend and the per-business numbers are not stable enough to publish. What is stable is the mechanism, and the mechanism is the finding.
Why this says “pilot” everywhere. The run behind this page was built to test the measuring apparatus, not to produce publishable per-business rankings. Ten runs per prompt is enough to check that the plumbing works and not enough to say a given business appears in a given share of answers — SparkToro measured the same prompt returning the same brand list under 1% of the time. So no business is named here and no per-business percentage appears. The full run is roughly 8,000 calls; when it happens, those numbers become sayable. These ones do not, and publishing them anyway would be the exact failure this study exists to document.
The short version
Across 76 usable answers there were 535 citations. Of those, 363 — 67.9% — pointed at somebody other than the business being discussed: directories, aggregators, professional associations, and competitors who were not on our list. Only 32.1% pointed at a named business’s own site.
That is the whole finding, and it is not the one most businesses are braced for. The instinct when a business hears “AI is answering questions about you” is to go and fix its own website. On this evidence, its own website is a minority of what the machine is reading. The larger share of the answer is being assembled from sources the business does not own and mostly does not know about.
Two further structural facts, both counts rather than rates: 30 of the 50 businesses were never named once in any answer. And being named and being cited turn out to be different things — 6 businesses were named but never cited, and 3 were cited but never named. A business can be the answer without being the source, or the source without being the answer.
What was actually measured
Fifty med spas in the St. Louis metro. Five prompts — three cost-shaped (“How much does Botox cost in St. Louis?”), two recommendation-shaped. Cost-shaped prompts were weighted deliberately, because Whitespark measured that query shape triggering AI Overviews about 97% of the time against 15% for “near me”. Two models, ChatGPT and Claude, both with web search grounded on every call — all 76 usable responses carried non-empty citations. Ten runs per prompt per model, so 100 calls attempted.
Separately, all 46 distinct websites were fetched deterministically and scored for how legible they are to a machine: schema, price disclosure, question headings, crawler access. 44 were reachable.
What those websites tell a machine
The cost-shaped prompts are the ones that trigger AI answers. Here is how much of this market publishes a price in a form the answering system can actually read:
| Publish no prices anywhere a machine can read | 16 of 44 |
|---|---|
| Publish prices only inside images | 11 of 44 |
| Publish prices as machine-readable text | 14 of 44 |
| Emit LocalBusiness-class schema | 25 of 44 |
| Use any question-shaped heading | 12 of 44 |
| Block AI crawlers in robots.txt | 0 of 44 |
Twenty-seven of 44 sites publish no machine-readable price at all — 16 publish none anywhere, and another 11 publish them only inside images, which is the same thing as far as a crawler is concerned. The demand is concentrated on cost questions and most of the market is illegible exactly there.
The last row is worth sitting with. Not one of the 44 blocks AI crawlers. Nobody in this market has opted out, which also means nobody has opted in on purpose. The overall machine-legibility score averaged 63.1 out of 100 across the 41 sites that could be scored, with a range from 22 to 98 — the spread is the story, not the mean.
What this pilot cannot tell you
The instrument partly failed, and here is how. Of 100 attempted calls, 76 were usable. 22 of the 50 Claude calls returned HTTP 400 and 2 timed out, leaving Claude with 26 usable responses against ChatGPT’s 50. That is a defect in our adapter, not a finding about Claude, and it is being fixed before any scale-up. It also means the model split is unbalanced, so this page makes no per-model comparison — the aggregate figures above are weighted toward ChatGPT and should be read that way.
This pilot also says nothing about:
- Google AI Overviews or AI Mode. No API exposes them. They were not measured and nothing here should be read as being about them.
- Personalised or location-signalled sessions. The prompts carried “St. Louis” as text. A real person with location history may see something different.
- Any trend. One metro, one vertical, one point in time. There is no before, so there is no direction.
- Any individual business’s standing. At 10 runs per prompt those numbers move too much to publish.
One earlier reading has also been corrected by this run. An 8 August check appeared to show that every site in the sample blocked us at the firewall. That was our own restricted network, not the market. Of the five genuinely problematic sites here, two domains do not resolve, two sit behind blanket Cloudflare challenges that block ordinary browsers too, and exactly one discriminates by user agent.
What we take from it
If two thirds of what gets cited is not your website, then the work is not only on your website. It is getting into the sources that already get cited — the directories, the associations, the local aggregators — and making your own site legible enough to be usable when it is reached. Those are different jobs from the one most people mean by “SEO”.
We are publishing this at pilot grade because the alternative was publishing nothing until a funded run existed. The numbers here will be superseded by that run, and when they are, this page will say so.
Method and raw data. Every figure above traces to a committed file: the verbatim model responses, the citation source table, the visibility table and the site audit. The method, the findings and the known defects are recorded in full in the study document. If a number here ever disagrees with that document, the document is correct and this page is a defect.
Aroc measures itself in public too — our own Growth Report, including the zero. For what a diagnostic can and cannot reach, see what an SEO audit cannot tell you.