Key Takeaways
- Compare runs, never single answers. One answer is a coin toss. A count across the same questions asked the same way is the only thing that can move in a meaningful direction.
- Learn your own noise before you judge anything. Run the same questions twice in one week without changing anything. Whatever moves in that week is the size of a meaningless change.
- Track four things, not one score. How often you are named, in how many distinct questions, against which competitors, and from which sources. They improve at different speeds.
- Keep the questions frozen. Rewriting a question mid-series destroys the comparison. Add new questions as a separate list rather than editing the original.
- Give it ninety days before you conclude anything. Published pages take weeks to be found and read, so a month of flat results is normal rather than evidence that the work failed.
You published four pages, claimed two directory listings, and asked ChatGPT about your niche this morning. It named you. Last month it did not.
That feels like progress and it is not evidence of anything. The same question asked twice in one afternoon can return different names, so a single good answer proves roughly what a single coin toss proves.
This is how to tell a real improvement from ordinary variation, using a spreadsheet and about an hour a month.
How do you tell if your AI visibility is improving?
By comparing counts from a frozen list of questions asked on a fixed schedule, not by comparing answers. If you are named in 4 of 25 questions in March and 11 of 25 in June, that is a signal. If one answer changed, it is not.
Three rules make the comparison hold. The questions stay identical word for word. Each question gets asked more than once per run. And you record the count rather than the feeling you got from reading the answers.
That is the whole method. The rest of this post is about the traps, and the biggest one has nothing to do with your work.
The part everybody skips: measuring your own noise
Run your questions twice in the same week without changing anything. Every difference you see between those two runs is meaningless movement, and any later change smaller than that is meaningless too.
This takes one extra hour and it is the single most useful thing in the whole method, because without it every result looks like a result. Perplexity's help center describes answers built from sources retrieved at the moment you ask, which is the mechanical reason the number wanders on its own.
- Up by one mention Inside the noise
- Down by two mentions Inside the noise
- Up by six mentions A real move
- Down by five mentions A real move
The shaded middle is the range your own two same-week runs wandered across. Anything landing inside it is noise, however satisfying it looks. The four moves plotted here are illustrative, chosen to show where the line sits rather than taken from a real measurement, and your band will be a different width from anybody else's.
A rough guide while you build your own band: on a 25-question list asked three times each, moves of one or two mentions are usually noise and moves of five or more usually are not. Those numbers are illustrative and yours will differ, which is exactly why you measure rather than borrow them.
The four numbers worth tracking
Four, and they move at different speeds. Collapsing them into one score hides which part of the work is paying off.
| What you record | How to get it | What an increase means | Typical speed |
|---|---|---|---|
| Mention rate | Times you were named, out of total asks | You are entering more answers than before | Slowest to move |
| Question coverage | Distinct questions where you appeared at least once | You match a wider set of buyer situations | Moves before mention rate |
| Competitor gap | Your count minus the most-named competitor's | You are gaining on the businesses being chosen instead | Moves last |
| Source list | Which sites the answers cite when they name anybody | New places are carrying your name | Moves first, often weeks earlier |
| Near misses | Answers naming your category but not you | Falling is good here: you were on the edge and got in | The early warning line |
The last two rows are the ones to watch when you are impatient. A new site appearing in the source list, or a near miss turning into a mention, is real progress weeks before the headline count reflects it. The method for reading source lists is in how to find out which websites AI cites in your niche, and the spreadsheet version of all of this is in how to track your AI visibility without buying software.
How long before a change should show up
Weeks, not days, and ninety days before you conclude anything. A published page has to be found, read, and then preferred over whatever the tool was citing before, and those are three separate delays stacked on each other.
Google Search Central's guidance on its AI features is explicit that pages have to be crawlable and indexable to be used at all, which is the first delay. OpenAI publishes the same requirement from its side, documenting the crawlers it uses and how a site can allow or block them. The second delay is whether the page is good enough to be chosen, and the third is that a lot of what these tools know is not refreshed the moment you publish. That last one is unpacked in how often ChatGPT updates what it knows.
A flat month is therefore normal and not a verdict. What would worry us is a flat quarter with no movement in the source list either, which usually means the pages are not being found rather than not being liked.
How to tell if it is genuinely getting worse
A drop bigger than your noise band, holding across two consecutive runs, with a competitor rising into the space you left. One of those on its own is a bad week. All three together is a trend.
The usual causes are boring. A page that used to be cited went offline or got rewritten. A directory listing lapsed. A competitor published the comparison page you never got around to. Check the source list first, because it will usually name the problem before you have to guess.
How often to do all of this is its own question, and the short answer is quarterly for most businesses and monthly while you are actively changing pages. That is argued properly in how often you should check whether AI recommends your business.
If you would rather start from a measurement somebody else took, the free check runs 25 real buyer questions from your niche through ChatGPT at $0 and emails a plain-English readout within the hour. It runs as soon as you ask for it, and the readout gives you the first row of your own record.
Questions we hear the most
How do you tell if your AI visibility is improving?
Ask the same frozen list of questions on a fixed schedule, count how often you are named, and compare counts rather than answers. A change only counts as real if it is bigger than the movement you get from running the same questions twice.
Why do my results change when nothing has changed?
Because these tools assemble answers from sources fetched at the moment you ask, so the same question can return different names an hour apart. That variation is normal and is exactly why single readings prove nothing.
How many times should I ask each question?
Three times is the practical minimum, five is better. You are counting how often your name appears out of the total, so a single ask gives you a number that cannot be compared to anything.
How long does it take for changes to show up in AI answers?
Usually weeks rather than days, and ninety days is a fair window for a full judgment. A new page has to be found, read, and then chosen over whatever was being cited before.
What does the free 25-question check include?
It asks ChatGPT 25 real buyer questions from your niche at $0, then emails a plain-English readout within the hour: how often you were named, who was named instead, and the first thing to fix. It runs as soon as you ask for it.
Does the $500 audit give me something I can compare later?
Yes. It is a single dated measurement of 125 checks across five AI tools, in your inbox by this time tomorrow, and every answer is stored with its date. That fixed baseline is what a later run gets compared against.
Can I track this without buying software?
Yes. A spreadsheet with one row per question and one column per run does everything the paid tools do for a business of this size, as long as you keep the questions and the schedule fixed.
Should I track every AI tool or just ChatGPT?
Start with one and add others once the habit holds. Tracking five tools badly is worse than tracking one properly, and the tools tend to move in the same direction over a quarter even when they disagree on any given day.
Your buyers are already asking. Find out what AI tells them.
The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.
Get my free check →