AI Visibility Evaluation

What an AI visibility evaluation cannot tell you

The limits, stated plainly, because a tool that answers everything confidently has decided not to say which answers it made up.

It cannot tell you whether you are any good

An AI visibility evaluation measures whether AI systems can find, correctly understand, and recommend a business. It measures nothing about whether the business does good work. A well-run firm with poor signals scores badly. A mediocre one with excellent signals scores well. Anyone presenting this as a quality judgement is overselling it, and that includes us.

It cannot predict revenue

Visibility is upstream of enquiries, which are upstream of sales, and both of those steps depend on things no evaluation can see: the offer, the price, the sales process, the market. Improving visibility puts a business in front of more of the right people. What happens next is not measured here, and a tool claiming a revenue projection from a visibility score has invented the causal chain in between.

It cannot tell you what a specific person will be shown

Answers vary by phrasing, by context, by conversation history, and by chance. What can be measured is a rate across many samples in a defined field. What cannot be measured is what one buyer saw yesterday, and any tool reporting that as a fact is reporting one draw from a distribution.

The pattern in all three: the limits are all about causation and specificity. The evaluation is good at describing a current state and poor at predicting a particular future, which is the honest shape of nearly all measurement.

It cannot see private surfaces

Some of what happens is not observable from outside: internal ranking signals, personalisation, and whatever an engine does that it has not disclosed. An evaluation measures the inputs and the outputs, not the mechanism, and inferring the mechanism from the outputs is speculation that should be labelled as such.

It cannot tell you it is complete

New surfaces appear. A measurement built for the current shape of AI answers will eventually be measuring a shrinking share of what matters, and it will not announce that it has. This is why the method is versioned, why comparisons across a method change are withheld, and why an evaluation is evidence rather than a verdict.

Why saying this makes the rest usable

If every finding arrives with the same confidence, a reader has to treat all of it as equally uncertain, which in practice means treating all of it as equally ignorable. Marking which findings are firm and which are estimates is what makes the firm ones worth acting on.

Back to AI Visibility Evaluation