Blog/Field Notes
FIELD NOTES

Settled or Improvised? What Testing 125 Questions Across Five Assistants Taught Us (2026)

A broad dark board in a deep blue void holding two rows of five square glass tiles, with the second and fourth tiles lit warm amber in both rows so the lit tiles line up in two columns

Key Takeaways

  • Treat one answer as an anecdote, not a finding. When we ran the same 25 questions twice, an hour apart, the two lists of competitors named had nothing in common. Every name was plausible, and the set meant nothing.
  • Look for the questions that come back the same. Some questions are settled. Asked twice across five assistants in August 2026, the question of which AI SEO companies to trust returned twenty of the same companies both times.
  • Count a name only when it repeats. In our audits a competitor reaches the report only if both passes named it and at least two of the five assistants named it. A name that appears once is not counted.
  • Score yourself on the worse of two runs. We report the lower of the two passes for your own score, so a lucky mention on one run never inflates it. Do the same when you test yourself.
  • Accept that strict filters drop some real names. A rule that throws out one-off mentions will sometimes throw out a real competitor. We chose that on purpose: a missing name can be added later, and an invented one misleads.

Ask an AI assistant who the best consultant for your problem is, and it will give you names. Ask again an hour later, and it may give you different ones. Both answers read as confident.

We learned how much that matters by testing. The audit we sell asks 25 buyer questions on five assistants, which is 125 checks, and it asks every one of them twice. This post is what testing 125 questions across five assistants taught us, and why the design ended up that way.

Here is the short version, the run that shared nothing, the question that came back the same, the two filters a name has to pass before we report it, and what this means for any test you run yourself.

What testing 125 questions across five assistants taught us

That a single answer is not a finding. Some questions are settled, and keep returning the same names when asked again and on other assistants. Others are improvised, and return a different plausible list every time. The only way to tell them apart is to ask twice and ask in more than one place.

The two observations below come from our own test runs: one on a frozen set of 25 questions, and one on a question about our own industry that we published in full. Neither publishes any client's data.

The run that shared nothing

On August 1, 2026, we ran the identical 25 questions twice, one hour apart, on the same frozen question set. The two lists of competitors named had 0% overlap. Not one name appeared in both runs.

Every name was individually plausible. Any one of them, read alone, would have looked like a finding. Together they said nothing. Our reading is that the answers were being improvised: nothing on the web settled those questions, so each run filled the slots with whatever fit.

That run is why a paid audit that asks each question only once would have been a coin flip dressed up as a method. The same caution applies when you ask an assistant about yourself once. Why two assistants can say opposite things about you is covered in why two AI assistants say opposite things about your business.

The question that came back the same

Some questions are settled, and they look completely different. When we asked five assistants twice in August 2026 which AI SEO companies to trust, twenty companies came back in both runs, and one firm, iPullRank, was named by all five assistants.

We published that run, with the firms named and how many assistants named each, on our page about how to tell a legitimate AI SEO company. It is the most settled result we have published.

The difference between the two runs is not the assistants. It is the question. Our reading is that where the web holds a lot of consistent material about who the answer is, assistants converge. Where it holds little, they improvise, and your job as a business is to be the consistent material.

Assistants differ in how they build answers, too. Google says of its own AI features that AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary. That is two products from one company. Five assistants from five different companies have even less reason to agree.

The two filters a name has to pass

In our audits, a competitor reaches the report only if it clears two filters: both passes named it, and at least two of the five assistants named it. Everything else is dropped. For your own score, we report the lower of the two passes.

ILLUSTRATIVE: EVERY NAME IN ONE TWO-PASS RUNWHICH NAMES GET REPORTED
Named by 1 assistant2 assistants3 assistants4 assistants5 assistants
One pass only 145210
Both passes 64311

Reported: named in both passes by at least two assistants

9 of 37 names reach the report. 28 are dropped.

answerhalo.com
An illustrative two-pass run, with every competitor name counted by how many of the five assistants named it and whether it appeared in one pass or both. Named in one pass only: 14 names by one assistant, 5 by two, 2 by three, 1 by four and none by five. Named in both passes: 6 by one assistant, 4 by two, 3 by three, 1 by four and 1 by all five. Only the both-pass names from at least two assistants are reported, 9 of 37. The counts are illustrative, not a measurement.

In the illustrative run above, 37 different names turn up somewhere. Only 9 clear both filters. The 22 names that appeared in one pass only are dropped, and so are the 6 that appeared in both passes but on a single assistant. Those numbers are illustrative, not a measurement, but the shape is the point: most names in a raw set of answers do not repeat.

The rules our audit applies to a two-pass, five-assistant run, and what each one protects against
Rule What it removes What it protects against What it costs
Ask the identical questions twice Nothing on its own; it creates the second copy Mistaking an improvised answer for a settled one Twice the answers: 250 stored instead of 125
A competitor must appear in both passes Names that turned up once A report built on a one-off answer Occasionally drops a real competitor
A competitor must be named by at least two assistants Names only one assistant believes in One assistant's quirk becoming your rival list Can hide a rival that matters on the one assistant your buyers use
Your own score is the lower of the two passes A lucky mention A score that flatters you Your score may look slightly worse than a single good run
Every answer is stored and dated Nothing; it keeps the evidence A number nobody can check None to you

What this means for any test you run yourself

Ask every question at least twice, on more than one assistant, with the wording held exactly the same, and count only the names that repeat. That is the whole lesson in one sentence. Everything in our audit design is that sentence applied at scale.

Freeze the wording. Write your questions down and paste them in. A change of one word makes the second run a different experiment. How to set up a run across several assistants in one afternoon is in how to test your business across five AI assistants.

Mark every name twice. Once for which run it appeared in, and once for which assistant. The names with two marks on both counts are your real competitors. The rest are noise until they repeat.

Treat a missing name gently. If you appear once and not the second time, you are not on the list yet. That is a finding, and it usually means the web does not hold enough consistent material about you for the question to settle your way.

Research on this subject measures visibility the same careful way. The GEO paper from Princeton and collaborators reports that content changes can boost visibility by up to 40% in generative engine responses, on its own benchmark. That is a result measured on a benchmark, not a promise about any one answer.

The damaging admission: our filters are a choice, and they have a price. Requiring two passes and two assistants will sometimes leave out a competitor who really does show up for your buyers, especially on the one assistant they happen to use. We accepted that trade because a missing name can be added later, and an invented one in a paid report cannot be taken back. How to read the list that results is in how to read an AI visibility report.

To see a first reading of where you stand, start with the free check. It asks ChatGPT 25 real buyer questions from your niche at $0 and emails a plain-English readout: your score, who got named in your place, and the one fix to start with. It runs as soon as you ask for it, and the report lands within the hour between 7am and 5pm Central, or by 10am the next morning outside those hours.

Questions we hear the most

Do AI assistants give the same answer to the same question twice?

Often not. When we asked the same 25 questions twice, an hour apart, the competitors named in the two runs had no overlap at all. Some questions are more settled than that, but you cannot tell which from a single answer.

What is a settled question in AI answers?

A question where the answer keeps naming the same businesses when you ask again and when you ask a different assistant. The names that repeat on both counts are the ones worth taking seriously.

How many times should you ask AI a question to test it?

At least twice, on more than one assistant, with the wording held exactly the same. Two runs are the minimum that can separate a settled answer from an improvised one. One run cannot.

Why would an AI assistant name a business that does not repeat?

Because an answer is written fresh each time, and when nothing clearly settles a question, several plausible names can fill the same slot. Each name can be reasonable on its own and still not be the stable answer.

Which five AI assistants does AnswerHalo test?

ChatGPT, Claude, Gemini, Perplexity and Google's AI Overviews. The $500 audit asks your 25 buyer questions on all five, which is 125 checks, and asks every one twice, so the report stores 250 answers.

How does AnswerHalo decide which competitors to show in an audit?

A competitor is shown only if both passes of the run named it and at least two of the five assistants named it. Names that appear once are left out, so the list reflects answers that repeat rather than answers that happened once.

Does the free check run twice too?

No. The free check asks ChatGPT 25 real buyer questions from your niche once, at $0, and emails a plain-English readout within the hour. It is a first look. The two-pass, five-assistant run is the $500 audit, which lands by this time tomorrow.

SEE WHERE YOU STAND

Your buyers are already asking. Find out what AI tells them.

The free check asks ChatGPT 25 real buyer questions about your niche and sends you the report within the hour, at $0. If it shows you are already getting named everywhere, we will say so in plain words.

Get my free check →