AI search
How we test AI search visibility
We test AI search visibility by asking the systems the questions our customers actually ask, recording the full answers and which sources were cited, and repeating it quarterly with a fixed question set. It is manual and unglamorous. It is also the most reliable signal currently available, because reporting for AI citation is genuinely immature; and running it consistently turns it into original data.
Why this exists
A lot of what is sold as “AI search optimization” is unmeasured. Schema gets added, an FAQ block appears, a file called llms.txt gets published, and the invoice says “AI optimization”; with nothing checking whether any of it changed anything.
We did not want to sell that. So before offering it as a service, we built a test we could run on ourselves and publish, including the parts where we do not show up.
The method
1. The question set
Thirty questions across four categories, fixed so results are comparable over time:
| Category | Example |
|---|---|
| Direct brand | “What is Novarius?” |
| Service + qualifier | “Who builds websites for contractors?” |
| Problem-led | “Why isn’t my contractor website getting calls?” |
| Local | “Web design companies in Miami” |
The question set does not change between runs. Changing the questions makes results incomparable; that is the mistake that turns this from data into anecdote.
2. The systems
Google AI Overviews, Google AI Mode, ChatGPT Search, Perplexity, and Bing Copilot.
3. What gets recorded
For each question and each system: the full answer text, every source cited in order, whether novarius.us appeared and where, whether we were cited, mentioned without citation, or absent; and the date.
4. Controls
Logged out, fresh session, no personalisation. Same day and same session across all five systems. Location set explicitly where the system allows it. Answers screenshotted, because these systems change and are not reproducible after the fact.
What we expect to find
Recorded before the first run, so it can be checked against reality rather than rationalised afterwards.
We expect to be largely absent at first. A new domain with no authority and no third-party presence is not a likely citation source. Anyone who launches a site and immediately reports AI citations is describing a brand query, not a discovery query.
We expect brand queries to work before anything else. “What is Novarius?” should resolve once the site is indexed; that is retrieval, not endorsement.
We expect problem-led questions to be the first real opportunity. They are where genuinely specific content can outrank authority.
We expect the local questions to be the hardest. Those answers are likely assembled from directories, which is consistent with what we found researching the Miami market; those searches are held by aggregators, not agencies.
The part that makes this hard
Research into how these systems select sources shows a consistent bias toward earned media; independent third-party sources; over content a business publishes about itself.
That is an uncomfortable finding for anyone selling on-page AI optimization, because it means on-page work alone is not sufficient. No amount of schema on your own site makes you the kind of source these systems prefer to quote.
What appears to move it: content genuinely worth citing, a consistent business identity everywhere it appears, and third-party presence you did not write.
The third is the slowest and hardest. We would rather say that than sell a package that quietly omits it.
Results
The first run happens once this site is live and indexed. Results are published here with the date, and every subsequent quarterly run is appended so the trend is visible; including runs where nothing improved.
How to run this yourself
It takes about two hours and needs no tools:
- Write 20-30 questions a customer might actually ask. Include brand, service, problem and local.
- Open a logged-out browser session.
- Ask each system each question.
- Record the answer, the sources cited, and whether you appear.
- Screenshot everything.
- Repeat quarterly with the same questions.
The value is entirely in the repetition. One run tells you almost nothing. Four runs tell you whether anything you did mattered.
