The Surgeon Authority Engine

Resources

What AI models know about finding a surgeon, before they search

Published 2026-07-31 · Measurement & authority · Authority Engine

A woman in a navy blazer standing at an office window with a mug, looking out at a residential street

We put fifteen questions to two AI models in July 2026, with web search switched off, to see what they already carry about finding surgical care. Two things hold no matter how you count the run: not one of the thirty answers named an individual surgical practice, and the names that did come back were patient directories. What we're not publishing this month is a rate. One of the two models was the wrong instrument, and once you set its answers aside the denominator depends on a judgement call about which ones count. Not yet.

Read that as what these models have absorbed, not as what a patient sees today. With retrieval off, a model answers from what it learned, so this is a measure of which names are established enough to come back unprompted. Directories are. Practices are not.

Below is what we asked, what we found, what we could not measure, and what it means if you run a practice. Our own share of voice is in it too, and it's zero.

How we measured it

We keep a fixed set of fifteen questions, the ones a surgeon, an administrator, or a patient would type when trying to find care or find help getting found, and we re-ask them monthly. The set stays identical, because a trend only means something if the test doesn't move.

This month we put those questions to two models, both with web retrieval disabled: one through a local command line with tools off, one through an API call without grounding. The first is a coding assistant and was the wrong tool for the question, which is the first limit below.

Four limits bound everything above, and they matter more than the numbers.

One of the two models was the wrong instrument. We put it through a coding assistant's interface. It declined six of its fifteen questions outright and prefaced most of the rest with a note that the topic sat outside what it's built for, though several of those went on to answer and to name brands. That leaves no clean line between an answer and a decline, and a rate you can only produce by drawing that line by hand is not a measurement. So the counts are held this month. A general-purpose model goes in next month, and the rate returns when the panel can produce it without a judgement call.

Retrieval was off, so this is not what ChatGPT tells a patient today. A model with search enabled can go and read a practice website mid-answer. These two couldn't. That makes this a floor rather than a forecast: it measures which names are already established in the model, which is a harder bar than being findable. It also means a practice absent here is not necessarily absent from a live answer, and we haven't tested the live case.

Five live-retrieval surfaces are unmeasured. ChatGPT, Perplexity, Google's AI results, Claude on the web, and Copilot all need a logged-in browser session, which an automated monthly run can't reach. Those cells are recorded as unmeasured, never as zero. Counting an unmeasured answer as an absence would manufacture the very finding we're reporting, so we don't. The live surfaces remain an open question for us.

The brand count is a floor, not a ceiling. We count brand mentions by matching against a dictionary of known names, so a brand we failed to include wouldn't be counted. The direction of that error is worth noting: an incomplete dictionary would tend to produce the conclusion that the models name almost nobody, and that's not what we found.

We publish the findings but not the question list or the scoring detail. Those are the working parts of a measurement service, and a panel that gets copied stops being a stable instrument.

Which brands came back?

We're holding the naming rate this month rather than publishing one we can't reproduce. What the run does show, and what a re-run two days later showed again, is the shape: when these models named anything, they named patient directories. Healthgrades, WebMD, Vitals, Zocdoc and Yelp all came back. No individual surgical practice did, in any of the thirty answers, under any way of counting the run.

Those five are all patient-facing directories. They are what the models reach for when the subject is finding surgical care, and they are the names a patient would meet first.

A caution on reading any of this: a brand appears because an engine mentioned it, not because an engine recommended it. Some mentions are recommendations, some are examples, and some are the engine listing where a patient might look. We're reporting mentions, which is the narrower claim. Treat the list as a map of who occupies the answer space, not as a ranking of anyone's quality.

Why this matters if you run a practice

The intuitive expectation is that these models have nothing to say about surgical care. What we found instead was that the models had names ready, and none of them belonged to a practice.

Be careful about what that does and does not establish. We counted which brands appear in an answer. We didn't measure whether the model recommended them, and we didn't test why. What the count supports is narrow and still useful: directory names are established in these models to a degree that practice names are not.

Two things follow.

The first is that your directory records are load-bearing whether or not you think of them as marketing. If a model reaches for Healthgrades when the subject is your specialty, then whatever Healthgrades says about you is doing work you didn't supervise. A stale listing there is not a tidiness problem. The labels engines read are documented publicly. Google publishes its structured data policies and its local business markup, and both are worth reading before anyone sells you a service built on them.

The second is the more interesting one. No practice was named in any of the thirty answers, so in this test the position of "the practice a model names without being told" is unoccupied. Unoccupied is different from contested. It suggests the work required to be named at that level is not yet being done by anyone in this answer space.

What does our own baseline look like?

We ran this panel on ourselves too. Our share of voice across the thirty answers is zero. Not one of them named us.

We're publishing that for the same reason we'd tell you yours: a number you won't say out loud is a number you can't steer by. This is our first baseline, taken deliberately before the work it's meant to measure, and the only thing that'll make it meaningful is what it reads next month and the month after.

If your own agency or marketing vendor can't tell you your current AI-citation baseline, that's worth asking about before you spend anything else. Not because the number is bad. Because not having it means nobody is measuring the surface patients are actually using.

Next month

The same fifteen questions, every month, and we'll add the live consumer assistants as soon as we can reach them reliably. We'll report what moves, including when nothing does, and we'll say when a change in the models rather than a change in the work explains a result.

If you'd like your own number rather than ours, our free visibility audit gives you the same measurement for your own practice: your number, the answers behind it, and where each one came from. We review every request, and qualifying practices get the audit back within 24 hours, at no charge.

For the reader-run version of this check, our guide on why AI doesn't recommend your practice walks through how to ask the questions yourself and how to read what comes back. For where this sits in the way patients actually choose a surgeon, see how patients find a surgeon. The build and the measurement behind it are on our AI search visibility page.

Correction, 1 August 2026. This page first reported that seventeen of thirty answers named a brand. Fifteen of those thirty came from a model that was the wrong tool for the question: it declined six of them outright and prefaced most of the rest with a note that the topic sat outside its scope, though several of those went on to answer. Any rate therefore rested on a line drawn by hand between an answer and a decline, and a different reasonable line gives a different number. We've withdrawn the rate rather than restate it at a new denominator, and the panel will publish one again when it can produce it without that judgement call. Two findings did not change and do not depend on where the line falls: no answer named an individual surgical practice, and our own share of voice is zero.

Common questions

Do AI models name individual surgeons?

In our July 2026 test, no. We asked two models with web search switched off, and not one of the thirty answers named an individual surgical practice. Patient directories did appear: Healthgrades, WebMD, Vitals, Zocdoc and Yelp. We're not publishing a naming rate for the month, because one of the two models was the wrong instrument and the denominator would depend on a judgement call about which of its answers count. It is not a statement about what a live assistant tells a patient today, which we have not measured.

Why couldn't you measure all the engines?

The five live-retrieval surfaces we could not reach (ChatGPT, Perplexity, Google's AI results, Claude on the web, and Copilot) all need a logged-in browser session, which an automated monthly run cannot open. We recorded those cells as unmeasured rather than as an absence. Counting an unmeasured answer as "did not name us" would have made our own visibility look worse and the finding look stronger than the evidence supports, so we report only what we measured and label the rest.

How many questions did you ask?

Fifteen questions, the kind a surgeon, an administrator, or a patient would type when looking for surgical care or for help getting found, put to two models with retrieval switched off.

We publish the findings but not the question list or the scoring detail. A measurement panel that gets copied stops being a stable instrument, and the value to a reader is in the result rather than in the apparatus.

Does appearing in an AI answer mean a brand was recommended?

No. A brand appears in our data because an engine mentioned it somewhere in the answer. Some of those mentions are recommendations, some are illustrative examples, and some are the engine pointing at where a patient might look. We report mentions because that's what we can verify. Read the results as a map of who occupies the answer space, not as a quality ranking.

Could your brand counts be too low?

Yes. We match brand names against a dictionary, so any brand we didn't include wouldn't be counted, which makes our numbers a floor. The direction of that bias is worth noting: an incomplete dictionary would tend to produce a finding that the engines name almost nobody, and we did not find that. Directories came back repeatedly; no practice did.

What's a good share of voice for a practice?

There isn't a published benchmark worth quoting, and we won't invent one. What's useful is your own trend against your own baseline on a fixed set of questions. A practice moving from zero to being named in a handful of answers has changed something real. A percentage compared against an industry average nobody has measured properly tells you nothing you can act on.