AI Trust Signals: What 466 Websites Show
Almost every website we evaluate has already fixed the machine-readable layer. Schema is present, the entity is consistent, the structure is clean. The pillar that decides whether an answer engine has a reason to name the business is the one almost nobody has touched.
The short version
Across 466 websites we evaluated independently, reputation signals scored 2.15 out of 10 on average. Entity consistency, over the same 466 sites, scored 8.43. That is not a small gap and it is not a rounding artifact: 410 of the 466 sites, 88 percent, scored below 5 on reputation signals, while only 32 sites, 7 percent, scored below 5 on entity consistency.
The pattern is the same one in every direction we cut the data. The technical work is largely done. The trust work is largely not.
What was measured, and on what
These are sites we chose to evaluate. Nobody asked us to, nobody paid for it, and none of them are Digilu clients or Digilu properties. They were selected as part of ongoing competitor research across a wide spread of fields, from electricians and estate planning attorneys to CRM platforms and browser games. Each site was scored under the published rubric on six pillars, using only what a public visitor, a crawler or an answer engine can access from the public web.
Every published evaluation is readable. The full set is at public signal research, and it is grouped by field: electricians, estate planning, insurance agencies, HVAC, CRM software and thirty more. Nothing below is a claim you have to take on trust: the underlying page for every site in the sample is published.
The six pillars, across 466 sites
| Pillar | Average | Scored below 5 | Scored 8 or above |
|---|---|---|---|
| Reputation signals | 2.15 | 410 of 466 (88%) | 49 (11%) |
| Local presence | 2.93 | 346 of 466 (74%) | 54 (12%) |
| AI discoverability | 6.81 | 129 of 466 (28%) | 194 (42%) |
| Semantic clarity | 7.22 | 51 of 466 (11%) | 208 (45%) |
| Authority structure | 7.95 | 61 of 466 (13%) | 320 (69%) |
| Entity consistency | 8.43 | 32 of 466 (7%) | 332 (71%) |
Read the last column. Seven out of ten sites are already at 8 or better on entity consistency and on authority structure. One in ten gets there on reputation.
Why this is not just an artifact of our scoring
A single blended average across a corpus collected over months is worth very little on its own, because the engine that produced the scores changed during that period. A score is a measurement made by a specific version at a specific moment, and versions are not interchangeable. So the honest test is whether the finding survives when the corpus is cut by engine version rather than pooled.
It does, in every version present:
| Engine version | Sites | Reputation signals | Entity consistency | Authority structure |
|---|---|---|---|---|
| 2.9.3 | 69 | 2.78 | 8.69 | 7.91 |
| 2.7.0 | 40 | 3.15 | 8.53 | 8.20 |
| 2.5.0 | 20 | 0.90 | 8.04 | 7.46 |
| 2.6.0 | 10 | 1.80 | 8.44 | 8.20 |
| 2.6.1 | 10 | 0.20 | 8.06 | 6.05 |
| 2.9.1 | 10 | 0.60 | 7.28 | 6.54 |
| Not recorded | 307 | 2.08 | 8.44 | 8.06 |
Reputation signals are the weakest pillar in every row. The highest any version reaches is 3.15. Entity consistency never drops below 7.28.
One thing in that table is a fault of ours and not of the sites: 307 of the 466 evaluations were stored without an engine version. That is 66 percent of the corpus carrying a score whose provenance we cannot fully reconstruct, on a property whose entire argument is that a measurement without its version is not a measurement. It is recorded here because leaving it out would make the research look tidier than it is. The finding does not rest on those rows: it holds in each of the six versioned subsets on its own.
Local presence only means something for some businesses
Local presence averaging 2.93 across the whole corpus is close to meaningless as stated, and reporting it that way would be dishonest. A physics laboratory, a browser game and an author have no service area, so scoring them on local presence measures nothing about how well they are doing. Splitting the corpus by whether the business has a service area at all:
| Kind of business | Sites | Reputation signals | Local presence |
|---|---|---|---|
| Has a service area (electricians, HVAC, insurance agencies, estate planning, mediation, doulas, gyms, cleaning) | 168 | 2.79 | 4.53 |
| No service area (software, platforms, games, publishers, laboratories) | 298 | 1.79 | 2.02 |
For the 168 businesses where local presence is a real question, 4.53 is the honest figure, and it is still below half marks. For the other 298 the local presence number should be ignored, which is why it is separated here rather than folded into a single headline.
Reputation is weak in both groups. It is the one pillar with no excuse available to it: every business in this corpus, local or not, can be written about, reviewed, cited and referenced by somebody other than itself.
The finding that should be uncomfortable
Ten of the sites in the corpus sell reputation software. Their own reputation signals average 2.20. Their entity consistency averages 9.17.
They are not unusual. Sixteen sites in the AI optimization field score 2.75 on reputation against 9.13 on entity consistency. Nine CRM platforms score 2.22 against 9.21. The companies closest to this problem professionally show the same shape as everyone else, and in some cases a sharper version of it: near perfect on the parts a developer can ship, near zero on the part that requires other people to say something.
That is the honest reading of why this gap exists, and it is not that anyone is lazy. Entity consistency, schema and site structure are buildable. One competent person can finish them in a sprint and they stay finished. Reputation signals cannot be built by the business that wants them. They accumulate, from other people, over time, and they decay if nothing keeps producing them. Work that can be completed gets completed. Work that has to be maintained gets postponed.
Where the field is strongest, and why it fits
The highest scoring field in the corpus is divorce mediation: 11 sites, 4.91 on reputation, 7.05 on local presence, 8.14 overall. Commercial cleaning reaches 4.25 on local presence. Business insurance reaches 6.33.
These are all fields where a customer will not proceed without some external reason to believe the business, and where that reason has historically been produced in public. The pattern is not that some industries are better at marketing. It is that some industries were already forced to accumulate third-party evidence, and that evidence is exactly what an answer engine can read.
What this means if you run one of these businesses
The practical consequence is a reordering, not a new task list. If a site is already at 8 on entity consistency and 2 on reputation, more schema returns very little. The constraint has moved.
- Check where you actually are before acting on any of this. The distribution above is a corpus, not a diagnosis of your site, and the whole argument of this piece is that averages do not tell you your own number. The free Trust Visibility Check returns your six pillar scores.
- Expect reputation to be the low one, and treat that as normal rather than as a failure. It is the low one for 88 percent of the sites here, including for the companies who sell reputation tooling.
- Understand that it is ongoing work. That is the uncomfortable part and the reason it stays undone. There is no version of this that is finished in a sprint, because the signals are produced by other people and they age.
Limits of this research
Stated plainly, because a data piece without them is an advertisement:
- This is not a survey. The 466 sites were chosen by us during competitor research. They are not a random or representative sample of any industry, and no figure here should be read as "the average electrician" or "the average CRM platform". It is 29 electricians we looked at, and their pages are published.
- Six engine versions plus 307 unversioned rows. Covered above. The pooled averages are indicative; the per-version table is the load-bearing evidence.
- A score is a measurement of published signals on a date, not a judgement of the business. A firm with a superb local reputation and no public trace of it will score low here, correctly, because the score reports what is legible to a machine, which is the entire question being asked.
- Local presence is not scored meaningfully for a business without a service area, and has been separated rather than blended for that reason.
- We have not shown that reputation signals cause AI recommendation. This piece reports what the sites carry. Our separate share of voice research measures what engines actually answer. Connecting the two properly needs both, over time, and we are not claiming it here.
Method
466 evaluations of public websites, stored between June and August 2026, restricted to independent research and excluding every client site, every Digilu property and every report a visitor requested about their own business. Six pillars per site, scored deterministically from publicly observable signals under the rubric described in the methodology. Pillar values are stored per evaluation with the engine version that produced them; where the version was not recorded it is reported as not recorded and never assumed. Nothing is imputed and no missing value is treated as a zero. Figures above are rounded to two decimals.
Every individual evaluation behind these numbers is published and linked from the research directory.