The short version

Across 466 websites we evaluated independently, reputation signals scored 2.15 out of 10 on average. Entity consistency, over the same 466 sites, scored 8.43. That is not a small gap and it is not a rounding artifact: 410 of the 466 sites, 88 percent, scored below 5 on reputation signals, while only 32 sites, 7 percent, scored below 5 on entity consistency.

The pattern is the same one in every direction we cut the data. The technical work is largely done. The trust work is largely not.

What was measured, and on what

These are sites we chose to evaluate. Nobody asked us to, nobody paid for it, and none of them are Digilu clients or Digilu properties. They were selected as part of ongoing competitor research across a wide spread of fields, from electricians and estate planning attorneys to CRM platforms and browser games. Each site was scored under the published rubric on six pillars, using only what a public visitor, a crawler or an answer engine can access from the public web.

Every published evaluation is readable. The full set is at public signal research, and it is grouped by field: electricians, estate planning, insurance agencies, HVAC, CRM software and thirty more. Nothing below is a claim you have to take on trust: the underlying page for every site in the sample is published.

The six pillars, across 466 sites

PillarAverageScored below 5Scored 8 or above
Reputation signals2.15410 of 466 (88%)49 (11%)
Local presence2.93346 of 466 (74%)54 (12%)
AI discoverability6.81129 of 466 (28%)194 (42%)
Semantic clarity7.2251 of 466 (11%)208 (45%)
Authority structure7.9561 of 466 (13%)320 (69%)
Entity consistency8.4332 of 466 (7%)332 (71%)

Read the last column. Seven out of ten sites are already at 8 or better on entity consistency and on authority structure. One in ten gets there on reputation.

Why this is not just an artifact of our scoring

A single blended average across a corpus collected over months is worth very little on its own, because the engine that produced the scores changed during that period. A score is a measurement made by a specific version at a specific moment, and versions are not interchangeable. So the honest test is whether the finding survives when the corpus is cut by engine version rather than pooled.

It does, in every version present:

Engine versionSitesReputation signalsEntity consistencyAuthority structure
2.9.3692.788.697.91
2.7.0403.158.538.20
2.5.0200.908.047.46
2.6.0101.808.448.20
2.6.1100.208.066.05
2.9.1100.607.286.54
Not recorded3072.088.448.06

Reputation signals are the weakest pillar in every row. The highest any version reaches is 3.15. Entity consistency never drops below 7.28.

One thing in that table is a fault of ours and not of the sites: 307 of the 466 evaluations were stored without an engine version. That is 66 percent of the corpus carrying a score whose provenance we cannot fully reconstruct, on a property whose entire argument is that a measurement without its version is not a measurement. It is recorded here because leaving it out would make the research look tidier than it is. The finding does not rest on those rows: it holds in each of the six versioned subsets on its own.

Local presence only means something for some businesses

Local presence averaging 2.93 across the whole corpus is close to meaningless as stated, and reporting it that way would be dishonest. A physics laboratory, a browser game and an author have no service area, so scoring them on local presence measures nothing about how well they are doing. Splitting the corpus by whether the business has a service area at all:

Kind of businessSitesReputation signalsLocal presence
Has a service area (electricians, HVAC, insurance agencies, estate planning, mediation, doulas, gyms, cleaning)1682.794.53
No service area (software, platforms, games, publishers, laboratories)2981.792.02

For the 168 businesses where local presence is a real question, 4.53 is the honest figure, and it is still below half marks. For the other 298 the local presence number should be ignored, which is why it is separated here rather than folded into a single headline.

Reputation is weak in both groups. It is the one pillar with no excuse available to it: every business in this corpus, local or not, can be written about, reviewed, cited and referenced by somebody other than itself.

The finding that should be uncomfortable

Ten of the sites in the corpus sell reputation software. Their own reputation signals average 2.20. Their entity consistency averages 9.17.

They are not unusual. Sixteen sites in the AI optimization field score 2.75 on reputation against 9.13 on entity consistency. Nine CRM platforms score 2.22 against 9.21. The companies closest to this problem professionally show the same shape as everyone else, and in some cases a sharper version of it: near perfect on the parts a developer can ship, near zero on the part that requires other people to say something.

That is the honest reading of why this gap exists, and it is not that anyone is lazy. Entity consistency, schema and site structure are buildable. One competent person can finish them in a sprint and they stay finished. Reputation signals cannot be built by the business that wants them. They accumulate, from other people, over time, and they decay if nothing keeps producing them. Work that can be completed gets completed. Work that has to be maintained gets postponed.

Where the field is strongest, and why it fits

The highest scoring field in the corpus is divorce mediation: 11 sites, 4.91 on reputation, 7.05 on local presence, 8.14 overall. Commercial cleaning reaches 4.25 on local presence. Business insurance reaches 6.33.

These are all fields where a customer will not proceed without some external reason to believe the business, and where that reason has historically been produced in public. The pattern is not that some industries are better at marketing. It is that some industries were already forced to accumulate third-party evidence, and that evidence is exactly what an answer engine can read.

What this means if you run one of these businesses

The practical consequence is a reordering, not a new task list. If a site is already at 8 on entity consistency and 2 on reputation, more schema returns very little. The constraint has moved.

  • Check where you actually are before acting on any of this. The distribution above is a corpus, not a diagnosis of your site, and the whole argument of this piece is that averages do not tell you your own number. The free Trust Visibility Check returns your six pillar scores.
  • Expect reputation to be the low one, and treat that as normal rather than as a failure. It is the low one for 88 percent of the sites here, including for the companies who sell reputation tooling.
  • Understand that it is ongoing work. That is the uncomfortable part and the reason it stays undone. There is no version of this that is finished in a sprint, because the signals are produced by other people and they age.

Limits of this research

Stated plainly, because a data piece without them is an advertisement:

  • This is not a survey. The 466 sites were chosen by us during competitor research. They are not a random or representative sample of any industry, and no figure here should be read as "the average electrician" or "the average CRM platform". It is 29 electricians we looked at, and their pages are published.
  • Six engine versions plus 307 unversioned rows. Covered above. The pooled averages are indicative; the per-version table is the load-bearing evidence.
  • A score is a measurement of published signals on a date, not a judgement of the business. A firm with a superb local reputation and no public trace of it will score low here, correctly, because the score reports what is legible to a machine, which is the entire question being asked.
  • Local presence is not scored meaningfully for a business without a service area, and has been separated rather than blended for that reason.
  • We have not shown that reputation signals cause AI recommendation. This piece reports what the sites carry. Our separate share of voice research measures what engines actually answer. Connecting the two properly needs both, over time, and we are not claiming it here.

Method

466 evaluations of public websites, stored between June and August 2026, restricted to independent research and excluding every client site, every Digilu property and every report a visitor requested about their own business. Six pillars per site, scored deterministically from publicly observable signals under the rubric described in the methodology. Pillar values are stored per evaluation with the engine version that produced them; where the version was not recorded it is reported as not recorded and never assumed. Nothing is imputed and no missing value is treated as a zero. Figures above are rounded to two decimals.

Every individual evaluation behind these numbers is published and linked from the research directory.