Core Concept

AI Visibility Evaluation

The repeatable measurement of whether AI systems can find a business, understand it correctly, and recommend it: producing a result you can act on, argue with, and check again later.

What AI Visibility Evaluation Means

An AI visibility evaluation is a structured measurement of how AI systems perceive a business. It asks three questions in order: can these systems find the business at all, do they understand what it is and who it serves, and when someone asks a question this business should be the answer to, is it named?

The word that carries the weight is evaluation. An evaluation is not a look at something. It is a measurement made under a stated method, producing a result that means the same thing next month as it did this month, and that someone else could reproduce. Everything that separates a useful AI visibility evaluation from an impressive-looking one comes back to that definition.

This matters commercially because AI systems have become an allocation layer. When a system decides which three businesses to name in an answer, the businesses it does not name do not get a lower ranking. They are absent. There is no page two. So the question of what those systems currently believe about a business has moved from interesting to operational, and answering it casually is worse than not answering it.

Evaluation, Audit, Score, Monitoring: Four Different Things

These four words get used as though they are interchangeable, and the confusion is expensive because each answers a different question and none substitutes for another.

Evaluation
A measurement under a stated method, at a point in time, producing a comparable result. Answers: where does this business actually stand?
Audit
An inspection against a checklist, usually manual and usually diagnostic. Answers: what is technically wrong here?
Score
The compressed output of an evaluation. A score without its method is a number with an opinion attached, not a measurement.
Monitoring
Repeating an evaluation on a schedule so change becomes visible. Monitoring without a stable evaluation underneath measures its own noise.

The practical consequence: an audit can tell you your schema is malformed and still leave you with no idea whether any AI system names you. An evaluation can tell you that you are invisible and not tell you why. Both are needed and they are not the same purchase.

What Makes an Evaluation Trustworthy

Five properties. An evaluation missing any of them will still produce a number, which is precisely the problem.

1. It is deterministic

The same inputs produce the same result. This sounds obvious and is the single most commonly violated rule in this category, because the easiest way to build an AI visibility checker is to ask a language model what it thinks of a business and report the reply. Language models are sampled, not queried. Ask the same question five times and you can get five different answers, so a tool built that way is reporting sampling variance as though it were a finding about your business. Run it twice and watch the number move for no reason.

2. It has no owned-domain floor

The evaluator must be capable of returning a bad result about itself, its own properties, and its own customers. Any evaluation whose scoring quietly guarantees a flattering outcome for anything is marketing wearing a lab coat. We hold this against our own domains: every property in our ecosystem grades 0.5 on the corroboration measure, because references between sites the same party controls are navigation, not independent evidence, and pretending otherwise would corrupt every other number on the page.

3. The method is published

You cannot argue with a verdict whose reasoning is hidden, and an evaluation you cannot argue with is not a measurement, it is an assertion. Every finding should arrive with what was checked, what was found, and what rule turned that into a result.

4. It is dated, and comparisons respect the method

AI systems change underneath the measurement. A result carries the date it was taken and the version of the method that took it. If the method changed, the comparison to last quarter is withheld rather than shown, because a number that improved because we changed the ruler is worse than no number.

5. Unknown is a real output

Not measured and measured zero are different findings, and collapsing them is how an evaluation lies without anyone deciding to. If a signal could not be read, the result says so. A tool with only two states will always render its uncertainty as the tidier one.

The test to apply to any tool in this category, including ours: run it twice on the same business, an hour apart, changing nothing. If the number moves, it was never measuring the business.

What an Evaluation Has to Answer

A result that says "your AI visibility is 62" has told a business nothing it can act on. A complete evaluation resolves four separate questions, and they fail independently:

  • Retrieval. Can these systems reach and read the business at all? A site that blocks AI crawlers, or renders its substance only in JavaScript, is invisible for reasons that have nothing to do with its quality.
  • Comprehension. Having read it, do they describe the business correctly? Being confidently misdescribed is a distinct failure from being unknown, and it is more damaging, because it is being actively recommended to the wrong people.
  • Corroboration. Does anything the business does not control agree? This is the hardest measure to move and the most valuable, and it is why our own ecosystem scores badly on it and says so.
  • Selection. When a real buying question is asked, is the business named, and who is named instead? This is the only measure that touches revenue directly, and a business can pass the first three and still fail this one.

The Objection Worth Taking Seriously

Is compressing a business's standing into an evaluation reductive? Yes, and anyone selling one should say so. A number cannot capture whether a firm is good at its work, and there is a real risk in this whole category that businesses start optimising for the measurement instead of for being worth recommending. That failure already happened once to search, and the industry it produced is why "SEO" is a word many business owners now flinch at.

Two things keep an evaluation honest against that. First, it measures whether AI systems can accurately perceive what is already true, not whether a business has performed some ritual. Everything it rewards, being findable, being described correctly, being independently vouched for, is something a business would want regardless of who is reading. Second, the largest single block of the measurement is corroboration, which cannot be manufactured on your own website by definition. A measurement whose hardest component is other people's opinion of you is difficult to game without becoming better.

Where the objection lands and we accept it: an evaluation is a snapshot of a moving system, and it should be read as evidence rather than as a verdict. Anyone presenting it as a final judgement is overselling it.

How AIOInsights Runs the Evaluation

AIOInsights exists to run AI visibility evaluations and nothing else. It reads what AI systems can actually reach, tests how the business is described, checks whether independent sources corroborate the claims, and asks the buying questions that matter in that category to see who gets named.

The evaluation engine is AIOTruth, which is deterministic by construction, and the comparison layer is AIOStory. Every finding shows its rule, every result carries its date, and anything that could not be read is reported as unread rather than as a zero. Where the evaluation finds a gap that needs ongoing work rather than a fix, the honest recommendation is ongoing responsibility for it, which is what Digilu does.

Get Evaluated

One evaluation, the full method shown, every finding dated.

Free Trust Check