Key Takeaways
- Stop reading one run as a measurement of one day. A single run is a sample of a system that generates a fresh answer every time. Read it as evidence, weighted by how extreme the result is, not as a reading off a gauge.
- Trust the extremes and distrust the middle. Named in zero of 25 answers is a finding on the first run. Named in 6 of 25 versus 9 of 25 is the same situation described twice, and no amount of staring settles it.
- Count questions, not repeats. Asking one question 25 times measures the machine. Asking 25 different buyer questions once each measures your business, which is the thing you wanted to know.
- Size the run to the decision. To find out whether you exist in these answers, 25 questions is plenty. To prove something you changed made a difference, you need the same questions run the same way twice.
- Write down the method before you look at the number. Same wording, fresh session, logged out, same tools, same day of the week. Without that, the second run measures your setup rather than your progress.
Somebody asks ChatGPT to recommend a coach in their city, sees a competitor's name, and concludes something about their business. Then they ask again an hour later and get a different list.
Both answers are real. Neither is a measurement yet. The question underneath is how many looks it takes before a number is telling you about your business rather than about the machine.
The answer is smaller than people expect for the first reading, and larger than they expect for the second.
How many checks before the number means anything?
About 25 different buyer questions, asked once each, is enough for a first reading. Below ten, a single different answer swings the result far enough to mislead you. Above 25, you are mostly buying the ability to detect smaller changes later.
That is a sample-size answer, and it depends on one thing being true: the 25 have to be 25 different questions, not one question asked 25 times. Those measure different things, and only one of them is about you.
It also depends on what you are asking the number to do. "Do I appear at all" needs far fewer observations than "did last quarter's work move anything", because the first is a question about a level and the second is a question about a difference. Differences are always harder to see.
What one run can already settle
More than people assume, when the result is lopsided. If your name does not appear in any of 25 answers, you have learned something on the first run, and a second run is unlikely to rescue it.
The reason is that the failure modes are not symmetric. There are many ways to be absent from an answer and only a few ways to be present in one. A business that assistants can find easily tends to show up somewhere across 25 different phrasings, even on a bad day. Scoring zero across all of them is a pattern, not a coincidence.
The middle is where single runs stop being useful. Named in 6 of 25 on Tuesday and 9 of 25 on Thursday is, in most cases, the same situation described twice. Treating the gap as progress or decline is the most common misreading of these reports, and it is covered along with the others in how to read an AI visibility report.
Why the variation exists at all is worth knowing before you size anything. OpenAI's own explanation of what ChatGPT is describes a system that generates a response rather than retrieving a stored one. Perplexity's help center describes answers built from sources fetched at the moment you ask, and Google says much the same about its AI features. Two runs an hour apart can differ without anything about your business having changed. The longer version of that mechanism is in why the same question gives different answers.
The arithmetic of one different answer
The size of the run sets how much a single answer can move the score. That is arithmetic, not statistics, and it is the most useful thing to have in your head when you read one of these numbers.
- 5 questions 20 points
One different answer and the score has moved a fifth of its range.
- 25 questions 4 points
The size of the free check. One answer is a nudge, not a swing.
- 125 checks 0.8 points
The size of the paid audit, spread across five tools.
Each track is the full range of the score. The lit block is the share of that range that moves when exactly one answer comes back differently. A small run is a coarse instrument, and no amount of care in reading it makes it finer.
Five questions is a coarse instrument. One answer flipping takes the score from 40 to 60, which reads as a transformation and is nothing of the kind. This is why a quick test on a handful of prompts produces such confident and such wrong conclusions.
Twenty-five questions puts one answer at 4 points. That is fine enough that the number carries information and coarse enough that you will not chase single-mention wobbles. It is the size the free check runs for exactly this reason.
A hundred and twenty-five checks puts one answer at 0.8 points, and buys something else besides precision: spread across five tools. Two businesses with the same total can have completely different shapes underneath, and only the wider run shows you whether you are losing everywhere or losing in one place.
| Run size | One answer moves the score by | What it can settle | What it cannot |
|---|---|---|---|
| 5 questions, one tool | 20 points | Whether you are obviously present or obviously absent | Anything about degree, direction or change |
| 25 questions, one tool | 4 points | Whether you appear at all, and who appears instead of you | Whether one tool differs from another |
| 125 checks, five tools | 0.8 points | Where you are weak by tool, and which questions fail everywhere | Whether a change you made caused a difference, without a second run |
| The same run, twice, months apart | Unchanged per run | Whether the picture moved, and in which direction | Why it moved, which needs the rival list either side |
Spread the checks across questions, not repeats
Twenty-five different buyer questions beats one question asked 25 times, every time. The repeats measure the tool's variation. The spread measures your visibility, which is the thing you were trying to find out.
Build the list from how buyers actually talk. A mix of broad questions ("best strength coach for runners"), local ones ("strength coach in Denver"), situation ones ("coach for someone coming back from a knee injury"), and vetting ones ("what should I ask a coach before signing up"). Spread across those four shapes matters more than the exact count, because your business can be invisible in one shape and fine in another.
Then hold the method still. Same wording every time, a fresh session for each question, logged out where you can be, and the same tools in the same order. You are the easiest variable to remove and the one most likely to flatter you, because your own account history makes you look more findable than a stranger would find you.
Record the businesses named, not just whether you were. The rival list is stable in a way the score is not: the same three or four names tend to recur across phrasings and across runs, and that recurrence is a stronger signal than any single total.
What it takes to prove something changed
Two runs of the same 25 questions, run the same way, far enough apart for the work to have landed. That is the real reason to care about sample size, because detecting a difference is much harder than detecting a level.
A working rule for a 25-question run, offered as a rule of thumb rather than a statistical test: a move of one or two mentions is noise, a move of four or more is worth investigating, and anything between the two calls for a repeat before you act. The same rule scaled to 125 checks moves those boundaries proportionally.
Leave enough time between runs. Pages have to be found and read before they can be used, and that takes longer than publishing feels like it should. Re-running weekly mostly samples the tools' own variation and produces a chart that goes up and down for no reason, which is a good way to spend a quarter reacting to nothing. The cadence question has its own post: how often you should check.
Here is the admission this post owes you. For most small coaching and consulting businesses, the first run comes back at or near zero, and every argument above about sampling is beside the point. Sample size matters when there is a number to be precise about. When the answer is "you do not appear anywhere", one run has already told you what you needed, and the useful next step is the rival list rather than a second measurement.
If you have not taken the first reading yet, the free check is 25 questions put to ChatGPT for your niche at $0, and it emails a plain-English readout: your score, who got named in your place, and the one fix to start with. It runs as soon as you ask for it, and the report lands within the hour between 7am and 5pm Central, or by 10am the next morning outside those hours.
Questions we hear the most
How many times should you check your AI visibility before the number means anything?
Around 25 different buyer questions, asked once each, is enough for a first reading. Fewer than ten and a single different answer swings the result. More than that mostly buys you the ability to detect small changes later.
Is one check ever enough?
Yes, when the result is extreme. Being named in none of 25 answers tells you something real on the first run, because a business that is easy to find rarely scores zero by accident. Middling results need a second run.
Should I ask the same question several times or ask more questions?
Ask more questions. Repeating one question measures how much the tool varies. Asking 25 different questions your buyers actually ask measures how visible your business is, which is the thing you are trying to find out.
Why do I get a different answer every time I ask?
Because the answer is generated when you ask it, from sources retrieved at that moment, and because your account history can change what you see. The text differs run to run even when the underlying picture has not moved at all.
How big a change counts as a real change?
Rough rule: on a 25-question run, treat a move of one or two mentions as noise and a move of four or more as worth investigating. Anything in between needs a repeat run before you act on it.
How many questions does the AnswerHalo free check ask?
Twenty-five real buyer questions from your niche, put to ChatGPT, at $0. The readout names who got mentioned in your place and the one fix to start with, and it lands in your inbox within the hour.
What does the $500 audit run that the free check does not?
It runs 125 checks across five AI tools instead of 25 on one, which separates a question you lose everywhere from one you lose in a single place. It arrives in your inbox by this time tomorrow.
How often should I re-run the whole thing?
Quarterly for most businesses, or after you publish something substantial. Weekly re-runs mostly measure the variation in the tools, and they tempt you into reacting to movement that was never real.
Your buyers are already asking. Find out what AI tells them.
The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.
Get my free check →