The evidence register
every claim, labelled and dated
By Jānis Plūme, Founder, Outbound Pros · 8 min read · 2026-08-06
Quick answer
This is the table every other page on this site defers to. Each load bearing claim we publish appears once, with one of four confidence labels, the source that produced it, the date we last verified it, and the date it stops being trusted. Rows are ordered by recheck date, soonest first, so the claim closest to expiry is the first thing you read. Contested rows never carry their figures, because a table row travels without its caveat once something extracts it.
The register exists because the evidence base under this whole category is small, ageing, and mostly re-cited rather than re-run. Publishing the labels is the only way to make that visible, including when it is inconvenient for us. Nobody has to catch us. The table catches us.
The register
Sorted by recheck date ascending. Rows with no scheduled recheck sit at the bottom, because there is nothing to re-run: either the source never published a method, or nobody outside the platform has the answer at all.
| Claim | Label | Source | Verified | Recheck |
|---|---|---|---|---|
| Major AI crawlers fetch JavaScript files and execute none of them, so a client rendered page returns an empty shell | Strong | Vercel and MERJ, 17 December 2024, server logs on nextjs.org cross validated on two unrelated stacks | 6 August 2026 | September 2026 |
| The byte counts we publish for each route type on the parent domain, read against a shell measured on the same domain the same day | Confirmed | First party fetch with the documented GPTBot user agent string, raw response bodies, 6 August 2026 | 6 August 2026 | October 2026 |
| Structured data is not required for Google generative AI features, and there is no special markup to add for them | Confirmed | Google guide to optimizing for generative AI features, 15 May 2026, mythbusting section | 6 August 2026 | November 2026 |
| Google Search does not use llms.txt | Confirmed | Google documentation, restated by John Mueller comparing it to the keywords meta tag | 6 August 2026 | November 2026 |
| The citations carrying most influence over an answer are longer, more structured, and richer in extractable evidence such as definitions, numerical facts, comparisons and procedural steps | Confirmed | arXiv 2604.25707, 602 controlled prompts, 21,143 search layer citations across ChatGPT, Google and Perplexity | 6 August 2026 | December 2026 |
| The five Quotable Unit properties separate a section that gets lifted from one that does not | Strong | Our own heuristic, assembled from the 2026 trait list above. No controlled study tests these five as a set | 6 August 2026 | December 2026 |
| Bingbot renders JavaScript, so a single page app can sit indexed in Bing and still serve an empty shell to the crawler that extracts the quote | Confirmed | Microsoft evergreen Bingbot documentation, plus the separate fact that Bing supplies candidate URLs to ChatGPT search | 6 August 2026 | December 2026 |
| On broad B2B category queries, the majority of citations go to third party sources rather than to vendor sites | Strong | Several 2026 measurements agree on the direction. The percentages circulating are vendor published and stay off every page here | 6 August 2026 | January 2027 |
| Schema markup helps Microsoft language models understand content | Confirmed | Fabrice Canel, principal product manager at Bing, on the record at SMX Munich, March 2025 | 6 August 2026 | January 2027 |
| llms.txt adoption reached 10.13% of roughly 300,000 domains, zero among the top 1,000 by traffic, with no citation effect once authority, schema density and recency were controlled for | Confirmed | SE Ranking study, May 2026, sample and controls disclosed | 6 August 2026 | February 2027 |
| Ranked list formats take a disproportionate share of citations relative to their share of the URL set | Contested | Evertune, May 2026, roughly 25,000 unique URLs and around 400 million citations across six surfaces. A platform reporting its own telemetry | 6 August 2026 | February 2027 |
| The comparison table, FAQ schema, content recency and expert quotation citation multipliers | Contested | Single vendor in each case, no published method, no replication file, no denominator | 6 August 2026 | None scheduled. There is no method to re-run |
| The overlap between ChatGPT citations and Bing rankings | Contested | Two vendor studies disagree by more than an order of magnitude and neither published a replication file | 6 August 2026 | None scheduled. There is no method to re-run |
| A Wikipedia page causes correct entity resolution, rather than correlating with the notability that produced both | Contested | No causal study exists. Wikipedia notability guideline for organizations rules most B2B companies out in any case | 6 August 2026 | None scheduled. There is no method to re-run |
| Whether any non Google engine weights structured data at all | Unknown | Perplexity and Anthropic have published nothing either way. Anyone stating a position here is inferring | 6 August 2026 | When a platform publishes |
| Whether entity markup changes retrieval, or only changes how a correct retrieval gets labelled | Unknown | No published evidence separates the two. The distinction decides whether schema is a lever or a label | 6 August 2026 | When a platform publishes |
On the build date shown at the bottom of this page, no row above has passed its recheck date. When one does, it moves to the top of the table, its label drops to contested, and it carries a visible past recheck marker until somebody re-verifies it or removes the claim from the site. That demotion is not a judgement on the source. It is a statement that we have stopped standing behind the freshness of our own citation of it.
What the four labels mean
The label decides what may be printed, not how interesting the claim is. A contested claim can be the most important thing on a page and still appear without a single digit.
| Label | What has to be true | What may appear on a live page |
|---|---|---|
| Confirmed | The platform stated it in its own documentation, or a study measured it with disclosed methodology and a sample worth trusting | The figure, with the source named in the same sentence |
| Strong | Several independent sources agree, at least one of them primary, and no credible contradiction was found | The figure, with the source named and a recheck date attached |
| Contested | Sources disagree, or the only sources are vendors selling the answer | The claim by name, never its digits. Describe the disagreement instead |
| Unknown | Nobody outside the platform knows, and saying otherwise is inference dressed as fact | A sentence saying so. This is the most valuable label in the set |
The contested rule is the one that costs us something. It means the numbers most competitors lead with cannot appear here at all, including on the page whose whole purpose is to take them apart. A refusal list printed with the digits intact is still a distribution channel for the digits, because an answer engine lifting the left column of a table carries whatever is in it and leaves the refutation behind.
What happens when a row expires
A recheck date is a commitment, not a decoration, so the failure mode has to be visible rather than quiet. Four things happen, in this order, and none of them requires a reader to notice first.
- The row sorts to the top of this table, above every claim still inside its window.
- Its label drops to contested, whatever it was before, and it carries a past recheck marker.
- Any figure it carries comes off the pages that cite it, because contested claims travel by name and not by number.
- Either somebody re-verifies it against the original source and sets a new date, or the claim comes off the site. There is no third option and no quiet extension.
One row is scheduled to test this on us fairly soon. The measurement underneath our entire first gate, that major AI crawlers do not execute JavaScript, is roughly twenty months old and most 2026 writing re-cites it rather than re-running it. If a rigorous public re-test showed the crawlers now render, the first gate of our model collapses and we would say so here within a week. We intend to re-run it ourselves across the group and publish whatever comes back, including a result that contradicts the argument this site is built on.
Why publish the register instead of keeping it internal
Two reasons, and the second one is the less flattering of the two. First, it is the only honest way to sell measurement. A practice that grades other people on evidence quality and keeps its own evidence private is asking to be taken on trust, which is the exact thing it tells clients not to do.
Second, it is a citation strategy and we would rather say so than pretend otherwise. A page stating plainly what is unknown is doing something a language model cannot synthesise from the rest of the web, because the rest of the web is confident. That makes it the kind of source a careful retrieval system reaches for. The same logic runs the other way on the page where we grade our own parent agency, which spends most of its length on who it suits badly.
Common questions about the register
Why is there no figure in some of the claims?
Because those rows are contested, and a contested claim goes on the site by name and never by number. The reason is mechanical rather than moral: a retrieval system lifting a table row carries the digits and drops the caveat sitting next to them, so printing a debunked figure inside its own debunking still ships the figure.
What makes a source primary here?
The platform describing its own system, or a study publishing its method and sample. A vendor describing a market is not primary, however large its dataset, and a secondary write up of a paper is not the paper. Where the only available source is vendor telemetry, the direction can go on a page and the effect size cannot.
Who sets the recheck dates?
We do, and they are editorial judgements rather than measurements. The rule of thumb is that platform documentation gets a long window because it changes slowly and announces itself, our own measurements get a short one because we control the re-run, and anything built on a single study gets a window shorter than the study is already old.
What is not in this register?
Campaign performance figures from the agency side of the group, which live on the properties that own them, and anything we have never published. The register indexes claims that appear on this site. It is not a list of everything we believe.
What should I do if a row looks wrong?
Send the source. A claim here is only as good as its citation, and a correction with a link is the cheapest thing anyone can give us. Rows have been demoted before they expired for exactly that reason.
The checker publishes its full rubric and every threshold on the page, for the same reason this table exists.
Last updated: 2026-08-06
See what a crawler sees
on your own site
Paste a URL and get the extractability read: what an assistant can actually retrieve, and what it cannot.
Free. No signup, no email capture.