AI Visibility Evaluation
The repeatable measurement of whether AI systems can find a business, understand it correctly, and recommend it: producing a result you can act on, argue with, and check again later.
What AI Visibility Evaluation Means
An AI visibility evaluation is a structured measurement of how AI systems perceive a business. It asks three questions in order: can these systems find the business at all, do they understand what it is and who it serves, and when someone asks a question this business should be the answer to, is it named?
The word that carries the weight is evaluation. An evaluation is not a look at something. It is a measurement made under a stated method, producing a result that means the same thing next month as it did this month, and that someone else could reproduce. Everything that separates a useful AI visibility evaluation from an impressive-looking one comes back to that definition.
This matters commercially because AI systems have become an allocation layer. When a system decides which three businesses to name in an answer, the businesses it does not name do not get a lower ranking. They are absent. There is no page two. So the question of what those systems currently believe about a business has moved from interesting to operational, and answering it casually is worse than not answering it.
Evaluation, Audit, Score, Monitoring: Four Different Things
These four words get used as though they are interchangeable, and the confusion is expensive because each answers a different question and none substitutes for another.
The practical consequence: an audit can tell you your schema is malformed and still leave you with no idea whether any AI system names you. An evaluation can tell you that you are invisible and not tell you why. Both are needed and they are not the same purchase.
What Makes an Evaluation Trustworthy
Five properties. An evaluation missing any of them will still produce a number, which is precisely the problem.
1. It is deterministic
The same inputs produce the same result. This sounds obvious and is the single most commonly violated rule in this category, because the easiest way to build an AI visibility checker is to ask a language model what it thinks of a business and report the reply. Language models are sampled, not queried. Ask the same question five times and you can get five different answers, so a tool built that way is reporting sampling variance as though it were a finding about your business. Run it twice and watch the number move for no reason.
2. It has no owned-domain floor
The evaluator must be capable of returning a bad result about itself, its own properties, and its own customers. Any evaluation whose scoring quietly guarantees a flattering outcome for anything is marketing wearing a lab coat. We hold this against our own domains: every property in our ecosystem grades 0.5 on the corroboration measure, because references between sites the same party controls are navigation, not independent evidence, and pretending otherwise would corrupt every other number on the page.
3. The method is published
You cannot argue with a verdict whose reasoning is hidden, and an evaluation you cannot argue with is not a measurement, it is an assertion. Every finding should arrive with what was checked, what was found, and what rule turned that into a result.
4. It is dated, and comparisons respect the method
AI systems change underneath the measurement. A result carries the date it was taken and the version of the method that took it. If the method changed, the comparison to last quarter is withheld rather than shown, because a number that improved because we changed the ruler is worse than no number.
5. Unknown is a real output
Not measured and measured zero are different findings, and collapsing them is how an evaluation lies without anyone deciding to. If a signal could not be read, the result says so. A tool with only two states will always render its uncertainty as the tidier one.
The test to apply to any tool in this category, including ours: run it twice on the same business, an hour apart, changing nothing. If the number moves, it was never measuring the business.
What an Evaluation Has to Answer
A result that says "your AI visibility is 62" has told a business nothing it can act on. A complete evaluation resolves four separate questions, and they fail independently:
- Retrieval. Can these systems reach and read the business at all? A site that blocks AI crawlers, or renders its substance only in JavaScript, is invisible for reasons that have nothing to do with its quality.
- Comprehension. Having read it, do they describe the business correctly? Being confidently misdescribed is a distinct failure from being unknown, and it is more damaging, because it is being actively recommended to the wrong people.
- Corroboration. Does anything the business does not control agree? This is the hardest measure to move and the most valuable, and it is why our own ecosystem scores badly on it and says so.
- Selection. When a real buying question is asked, is the business named, and who is named instead? This is the only measure that touches revenue directly, and a business can pass the first three and still fail this one.
The Objection Worth Taking Seriously
Is compressing a business's standing into an evaluation reductive? Yes, and anyone selling one should say so. A number cannot capture whether a firm is good at its work, and there is a real risk in this whole category that businesses start optimising for the measurement instead of for being worth recommending. That failure already happened once to search, and the industry it produced is why "SEO" is a word many business owners now flinch at.
Two things keep an evaluation honest against that. First, it measures whether AI systems can accurately perceive what is already true, not whether a business has performed some ritual. Everything it rewards, being findable, being described correctly, being independently vouched for, is something a business would want regardless of who is reading. Second, the largest single block of the measurement is corroboration, which cannot be manufactured on your own website by definition. A measurement whose hardest component is other people's opinion of you is difficult to game without becoming better.
Where the objection lands and we accept it: an evaluation is a snapshot of a moving system, and it should be read as evidence rather than as a verdict. Anyone presenting it as a final judgement is overselling it.
How AIOInsights Runs the Evaluation
AIOInsights exists to run AI visibility evaluations and nothing else. It reads what AI systems can actually reach, tests how the business is described, checks whether independent sources corroborate the claims, and asks the buying questions that matter in that category to see who gets named.
The evaluation engine is AIOTruth, which is deterministic by construction, and the comparison layer is AIOStory. Every finding shows its rule, every result carries its date, and anything that could not be read is reported as unread rather than as a zero. Where the evaluation finds a gap that needs ongoing work rather than a fix, the honest recommendation is ongoing responsibility for it, which is what Digilu does.
See the method in full: how the evaluation is scored