How to compare AI visibility tools honestly
Six questions that separate a measurement from a plausible number, and they work on any tool including this one.
Why comparison is hard right now
Scores between tools are not comparable, because each one defines its own field, asks its own questions, and weights its own signals. A business scoring 72 on one and 41 on another has learned nothing about itself. What can be compared is method, and method is usually the thing least prominently displayed.
The six questions
1. Does it survive the two-run test? Same inputs, an hour apart, changing nothing. If the number moves, it is reporting its own variance. This one question eliminates most of the category.
2. Does it ever return "not measured"? A value in every cell every time means it either checks nothing that can fail or fills gaps with zeros.
3. Does it show the rule behind a finding? A verdict you cannot argue with is an assertion, not a measurement. Every finding should say what was checked, what was found, and what turned that into a result.
4. Can it produce a bad result about its own maker? Any scoring that quietly guarantees a flattering outcome for the vendor, its clients, or its own properties is marketing in a lab coat.
5. Is the question it asks a question a buyer would ask? Testing a business against its own brand name proves nothing. A favourable score is very often a favourable question.
6. Does it date its results and withhold comparisons across a method change? A number that improved because the ruler changed is worse than no number.
Apply all six to this evaluation. We publish that our own ecosystem domains grade 0.5 on corroboration, that unmeasured signals are excluded rather than zeroed, and that comparisons are withheld when the method changes. Those are checkable claims, and checking them is the correct response to any of this.
Two things that look like quality and are not
The number of signals checked. Two hundred checks sounds thorough. Most of the weight in any honest model sits in a handful of things, and a long list is frequently padding that makes a score feel precise.
Speed. An instant score is usually a single model call. The parts that take time, sampling many runs and verifying third-party sources, are the parts that make the result mean anything.
What a fair comparison looks like
Run two tools on the same business, then compare their findings rather than their scores. Where they disagree, ask each one why. The tool that can answer, in terms you can check, is the one worth using, regardless of which number was higher.