Retrieval Readiness Index · 2026-09-03
Of 101 homepages measured for AI retrieval, 1 is ready.
Every score here was taken by running the same scorer against a live page on one day. The index measures whether an answer engine can fetch a page, lift a passage out of it, and attribute it. It does not measure whether the page ranks. How it is measured.
Where the cohort sits
Counts of the 101 scored sites per band. The 22 sites we could not read are not in this distribution, because a page that could not be measured is not a page that scored zero.
Access is solved. Comprehension is not.
Median score per dimension across the scored sites. Crawlers are mostly let in and pages mostly load fast. What fails is everything an engine has to do after that: chunk the page, find a dated answer, and read the structure.
What each dimension asks
The eight dimensions this sweep measured, with the weight each carries in the total. A dimension that could not be judged goes unscored rather than to zero, so the number of sites behind each median differs. A ninth, Query alignment, needs a specific question and is explained below.
- Accessweight 18
- Can AI crawlers fetch the page at all. Measured on 96 of 101.
- Timeweight 8
- How long the page takes to become readable. Measured on 101 of 101.
- Entity Resolutionweight 14
- Does the brand resolve to a known entity. Measured on 101 of 101.
- Answerabilityweight 16
- Are answers stated directly and early, or buried. Measured on 83 of 101.
- Extractionweight 16
- Does the page chunk cleanly into liftable passages. Measured on 101 of 101.
- Freshnessweight 8
- Is the page dated, and recently. Measured on 43 of 101.
- Structureweight 12
- Is the markup shaped so a machine can navigate it. Measured on 100 of 101.
- Token costweight 8
- How much budget the page burns to yield its answer. Measured on 101 of 101.
The index
All 123 domains, sortable. Measured is how many of the eight dimensions returned a number for that site: a score resting on three is weaker evidence than one resting on eight.
| Domain | Score | Meas. | Flags | Access | Extract | Answer | Entity res. | Tokens | Time | Struct | Fresh |
|---|
Method
Each domain's homepage was fetched and scored by the same instrument on 2026-09-03, scoring version 4. 123 domains attempted, 101 scored, 22 unmeasurable, 0 errors. The cohort is homepages across crypto, fintech, SaaS, media, retail, travel, health, education and enterprise.
Homepages only
This is a homepage sweep, and homepages are built to route people rather than to answer questions. Structure and Answerability read low on almost every brand for that reason. A product or pricing page from the same site will usually score higher. Read these numbers as a floor, not as a verdict on the whole domain.
Bands
80–100 Retrieval-ready. 60–79 Retrievable, with gaps. 40–59 At risk. 0–39 Largely invisible. Nobody has scored 100 and nobody will: Entity Resolution alone requires a Wikidata entry and a Knowledge Graph presence, so the practical ceiling sits well below the nominal one. The top of this cohort is 80.
What unscored means
Partial scans, not zeros. A site whose page we could not read has not been measured and is absent from the distribution. 22 of 123 sites returned a partial scan: one or more sub-measurements timed out or was blocked, so no overall score was formed. Those rows are in the table with their partial dimensions intact and are excluded from every statistic on this page.
Dimensions not measured here
Query alignment is scored only when a specific question is supplied, and this sweep supplied none, so it is null on all 123 rows and is left out of the table. Freshness carried a number on only 43 of 101 scored sites, because an undated homepage goes unscored on it rather than taking a zero.
Change from the previous run
Homepages only, all measured with the live scorer on one day. Same cohort as the 2026-08-16 run, re-measured because undated homepages stopped taking a flat zero on Freshness and now go unscored on that dimension, which lifts most of the cohort. Per-dimension scores are kept this time so the next rubric change can be modelled before it is deployed.
Reproducing it
The scorer runs two scans at a time with a retry pass, and refuses to write a file if more than a tenth of its rows came back from cache. A run that finishes too fast has not measured anything. Scores move with the rubric as well as with the sites, so a score is only comparable to another score from the same version.