Why the four measures have to be read together
Each one fails independently and each has a different fix. A single composite score hides which of the four is actually broken.
The problem with one number
An AI visibility evaluation produces a composite, and a composite is useful for tracking movement over time and useless for deciding what to do. Four businesses can score identically for four completely different reasons requiring four unrelated responses.
So the score is the headline and the four measures underneath it are the finding.
The four failure shapes
Retrieval failure. Nothing is being read. Everything downstream is meaningless, and no amount of content or corroboration work will register while it persists. Cheapest to fix, and it must be fixed first.
Comprehension failure. Being read and described wrongly. The fix is clarity and consistency across sources, not more publishing. Publishing more of an unclear description makes it worse.
Corroboration failure. Read and understood correctly, with nothing independent supporting it. The slowest to fix and usually the largest single gap. It cannot be solved on your own site by definition.
Selection failure. Everything is in order and competitors are still named. This one is comparative rather than absolute, and the response depends entirely on who is being named instead.
The order matters. Retrieval, then comprehension, then corroboration, then selection. Working on a later one while an earlier one is broken produces effort that cannot show up in the result, which is how businesses conclude that none of this works.
The most expensive misdiagnosis
Treating a corroboration gap as a content problem. It presents as low visibility, and the intuitive response is to publish more. Publishing does not create corroboration, because corroboration is by definition what other people say. Months of content production later, the number has not moved, and the conclusion drawn is usually that AI visibility is unmeasurable rather than that the wrong lever was pulled.
Why they are weighted unequally
Corroboration carries the most weight because it is the hardest to fake, and retrieval acts closer to a gate than a component, since failing it caps everything else. Any weighting is a judgement and should be published so it can be argued with. A model that hides its weights is asking to be trusted rather than checked.