Blog/Honest Comparison
HONEST COMPARISON

What a Real AI Visibility Audit Should Include

A broad blank dark docking chassis on a low rail in a deep blue void, three of its four bays each holding a blank dark upright plate standing proud, the fourth bay standing empty with its open recess washed warm amber from within

Key Takeaways

  • Demand the question set before the run, not after. The questions are the measuring instrument. If they can be edited later, the day-90 comparison and any guarantee attached to it mean nothing.
  • Refuse any audit built on a single pass. We ran one identical 25-question set twice, an hour apart, on August 1st 2026. The two lists of rival names had nothing in common. One run cannot separate a real competitor from a fluent guess.
  • Ask to see the sentence each name was judged from. A list of competitors is an assertion. The quoted sentence that produced each name is evidence, and it is the difference between reviewing an audit and agreeing with one.
  • Expect one specified page, not a list of themes. The useful output names your worst question and writes the page that wins it: the URL, the heading, every section in order, and the questions it must answer.
  • Read the method section first. How it was measured, how many engines, how many passes, what was suppressed and why. An audit that will not tell you how it counted is asking for trust it has not earned.

Almost everyone selling an AI visibility audit right now is selling the same shape of thing: a score, a list of competitors, and some advice. The scores look alike. The advice looks alike. The prices do not.

What separates them is not the cover page. It is whether the thing underneath can be checked by the person who bought it.

This is the list we would hand a buyer who was about to pay someone else. It is written so you can hold any vendor to it, including us.

What should a real AI visibility audit include?

Eight parts: a frozen question set, more than one assistant, more than one pass, the sentence behind every named rival, a named competitor set, one specified page to build, a stated method, and a re-runnable baseline.

Every one of those is checkable by a non-technical buyer in about ten minutes. That is deliberate. An audit you cannot check is a report you have to take on trust, and the whole reason a business orders one is that it is tired of taking things on trust.

It starts with a frozen question set

The questions are the instrument, so they have to be fixed before the first run and shown to you. Everything measured afterward is measured against them.

A good set is written the way buyers speak, not the way the industry writes. "Who does small group strength coaching in Denver" is a question. "Fitness coach visibility ranking factors" is not something a human being types. Ask to read the 25 questions before anybody runs anything, and if a single one of them contains your business name, ask why: questions that already name you are testing recall, not discovery.

Freezing them matters because of what comes later. A re-run at day 90 only means something if it asks the same questions. A question quietly reworded between the two runs turns an honest comparison into a decorative one.

Why one run is not a measurement

Because the same question, asked twice, can produce two different worlds. This is the finding that changed how we build audits, and it is the one we would most want a buyer to know.

SAME 25 QUESTIONS, ONE HOUR APARTSOURCE: OUR OWN TEST RUN, 2026-08-01 · CHIP COUNTS ILLUSTRATIVE
First pass, names returned
Names both passes agreed on
Nothing. The two lists had no name in common
Second pass, an hour later

The middle band is what a single run cannot give you. Names are only worth reporting where two independent passes, and two different assistants, produced them both.

answerhalo.com
The empty band is the measured result. The number of bars in the outer two bands is illustrative and stands in for the names each pass returned.

On August 1st 2026 we ran one identical set of 25 questions twice, an hour apart, with nothing changed in between. The two lists of rival businesses that came back had no overlap at all. Not a small overlap. None.

Carry the limits of that with the number: it is one question set, in one niche, on one day, and it is our own test rather than a published study. It is enough to establish the shape of the problem, which is that a language model asked to name businesses will produce a fluent, confident, different answer on request. A vendor who runs your questions once and prints the result has handed you one roll of a die and called it a measurement.

The same logic applies across tools. Assistants disagree with each other as readily as they disagree with themselves, which is why a single-tool audit describes that tool rather than your business. They are not even built the same way: Perplexity's help center describes a product that searches and cites sources for every answer, while Google's guidance on its AI features describes a system that issues several related searches behind one question. Two different machines, asked the same thing, reading different pages.

Evidence, not assertion

Ask for the sentence each competitor name was judged from. A list of names is something you can only agree with. A quoted sentence is something you can check.

This is the difference between reviewing an audit and rubber-stamping one. When a report says a rival was named in your place, the useful version shows you the line where the assistant said it, so you can see whether the assistant actually recommended them or merely mentioned them in passing. Those are very different findings and they look identical in a bar chart.

There is research pointing the same direction from the content side. The Princeton-led study GEO: Generative Engine Optimization tested content changes against generative search answers and found that adding citations, quotations and statistics raised how visible a source became. What engines reward in a page is exactly what a buyer should demand in a report: attribution rather than assertion.

The eight parts, checked against any vendor

Here is the whole list with the tell for each one. Print it, and use it on whoever you are considering.

The eight parts of an AI visibility audit, what each looks like when it is real, and what the thin version looks like instead
Part Real version Thin version
The question set 25 buyer-worded questions, shown to you and frozen before the run Questions never shown, or written in industry language
Coverage Several assistants, each reported separately One tool, presented as your visibility
Passes Two independent runs, with agreement required One run, printed as fact
Named rivals Each name carries the sentence it was judged from A bare list of competitor names
The score A count you can recompute: named in N of the questions asked An index out of 100 with no stated method
The next action One page specified section by section, with the questions it answers Themes, pillars, and a content calendar
Method How it was counted, what was suppressed, and why No method section at all
The baseline Stored so the same questions can be re-run and compared Nothing to compare against later

Two things belong on the other side of the ledger, because a good audit refuses to promise them. No honest vendor can guarantee a specific citation by a specific date, because they do not control the assistant. And a promise measured in pages per week is measuring their output, not your result.

Also worth saying plainly: plenty of businesses do not need to buy one of these at all. If a free readout shows one obvious gap, fix the gap. That case is set out in when a free check is enough, and the ongoing version of the question is covered in monitoring tools versus done-for-you work. Whether the paid version earns its price is argued through in is a paid AI visibility audit worth it.

If you want the free version first, that is the sensible order. The free check asks ChatGPT 25 real buyer questions from your niche at $0 and emails a plain-English readout: your score, who got named in your place, and the one fix to start with. It runs as soon as you ask for it, and the report lands within the hour between 7am and 5pm Central, or by 10am the next morning outside those hours.

Questions we hear the most

What should a real AI visibility audit include?

A frozen question set in buyer language, more than one AI tool, more than one pass, the quoted sentence behind every named competitor, one specified page to build, a method section, and a baseline you can re-run against later.

Why does the question set have to be frozen?

Because every later comparison is made against it. If questions can be edited between the first run and the re-run, the before and after numbers are measuring two different things and any promise attached to them is unfalsifiable.

Why is one pass not enough?

Assistants are not deterministic. We ran the same 25 questions twice, an hour apart, on August 1st 2026, and the two lists of rival names had no overlap at all. Agreement across passes is what separates signal from noise.

How many AI tools should an audit cover?

More than one, because they disagree routinely. A business can be named consistently by one assistant and absent from another, so a single-tool result describes that tool rather than your visibility.

What should an audit never promise?

A guaranteed citation by a fixed date, a score with no stated method, or a page count per week. None of those are within a vendor control, and the last one measures output rather than any result.

Do I need a paid audit to get started?

No. Start with a free check and only pay when you want the full picture. Plenty of businesses read a free readout, see the one obvious gap, and fix it themselves without buying anything.

What does the AnswerHalo free check include?

It asks ChatGPT 25 real buyer questions from your niche and emails a plain-English report: your score, who got named in your place, and the one fix to start with. It costs $0 and it runs as soon as you ask for it.

What does the $500 audit add on top of that?

It runs 125 checks across five AI tools as two independent passes, names rivals only where two engines and both passes agree, and ends with the page that wins your worst question, section by section. It arrives by this time tomorrow.

SEE WHERE YOU STAND

Your buyers are already asking. Find out what AI tells them.

The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.

Get my free check →