Open dataset — AI-readiness scores
Every AgentFit audit that produced a usable measurement, as one file. 6001 entry URLs across 5961 hosts, scored 0–100 on 28 criteria in six categories. CC BY 4.0. Regenerated daily.
Download
dataset.csv 1.1 MB, 23 columns, RFC 4180 · dataset.json 3.0 MB, with a population envelope
Generated 2026-07-28. Both files are the same rows from the same snapshot.
What is in it
One row per audited entry URL: the host, the entry URL, the date of the most recent valid audit, the total score, the six category scores with their maxima, the rubric version, how many valid audits that URL has, and the run id — which resolves to the full public report at /r/{run_id}, evidence and all. Every host in the file has a live page at /browse/site/{host}.
What is not in it: nothing fetched from the audited sites — no page text, no evidence snippets, no HTML, no titles. No IP addresses and no submitter information; the file holds only numbers we computed and identifiers we assigned. Per-criterion detail stays on the report pages.
Columns
| Column | Meaning |
|---|---|
host | Hostname, lower-case, punycode. Not unique: a few hosts were audited under more than one entry URL. |
entry_url | The normalised entry URL. Unique — this is the file's primary key. |
is_canonical_for_host | true on the row that /browse/site/{host} shows as the host's entry URL. |
audited_on | Date of the most recent valid audit, UTC. |
first_audited_on | Date of the first audit of this entry URL, of any validity, UTC. |
ok_run_count | How many valid audits this entry URL has — how thick the evidence is. |
run_id | The run every score in the row comes from. Opens at /r/{run_id}. |
validity | Always ok. The column keeps the filter visible inside the file itself. |
rubric_version | Which rubric measured the row. Rows with different versions are not comparable. |
total_score | Total score of that run. |
total_max | The scale the total sits on, in the file rather than in the documentation. |
A_discovery … F_agent_surface | Score of one category (A–F) in that run. |
A_discovery_max … F_agent_surface_max | What that category was worth under this row's rubric version. |
license | Constant: the licence and the page carrying the full notice. It is repeated on every row so that a CSV separated from this page still says what it is. |
Rubric epochs
Category budgets changed between rubric versions: under v2 the six categories were worth 13/20/24/19/14/10; under the current rubric they are 14/21/17/23/21/4, over 28 criteria instead of 30. Rows with a different rubric_version were produced by different instruments.
Group by rubric_version before you compute anything. This is the same rule /diff enforces when it refuses to present a cross-version delta as a change in a site.
Rows per rubric version: v4: 5920 · v3: 66 · v2: 15
Valid only
A row exists only where the audit actually obtained enough responses to measure (validity = ok). Runs blocked by anti-scraper rules, or that hit an unreachable host, are excluded: their totals measure our access, not the site.
Of the 6928 entry URLs in the public index, 362 have no valid run and are absent. A further 565 have one, but from before scoring versions were stamped — a score with no ruler is not reproducible, so they are absent too, and will return as the corpus is re-audited.
Population
These are URLs people submitted themselves, not a curated list of API documentation. The distribution describes what was submitted, not the market. Do not read it as a ranking of the API industry.
It also skews low, and in one direction: many submitted URLs are not API documentation at all and score near zero — about a fifth of the hosts sit below 10 points and roughly half below 20. Any percentile or ranking you compute from this file therefore flatters its subjects compared with a set of real documentation sites. Say so if you publish one.
Living export
The file is generated from live data and changes as sites are re-audited. That is deliberate: a quarterly frozen release would keep distributing rows we have been asked to remove.
The cost is that there is no version number to cite — so cite the date you downloaded it, and keep your copy. The download is named agentfit-dataset-YYYY-MM-DD.csv for exactly that reason.
Licence
AgentFit AI-readiness dataset © 2026 Gumeniuk Stanislav is licensed under CC BY 4.0. To view a copy of this licence, visit https://creativecommons.org/licenses/by/4.0/
Scores are facts we computed. In some jurisdictions facts and thin compilations attract no copyright at all; where they do, or where a sui generis database right applies, CC BY 4.0 is the licence. Where no such right subsists, no permission is needed and the attribution below is a request rather than a condition.
CC BY 4.0 covers the data in /dataset.csv and /dataset.json only — not the AgentFit service, the site content, the rubric text, or the source code.
How to cite
Copy one of the three forms below. The retrieval date in them is the date of the snapshot on this page; if you downloaded the file earlier, cite that date instead.
Plain text
AgentFit AI-readiness dataset by Gumeniuk Stanislav (agentfit.dev), licensed under CC BY 4.0. Retrieved 2026-07-28 from https://agentfit.dev/dataset
HTML
<a href="https://agentfit.dev/dataset">AgentFit AI-readiness dataset</a> by Gumeniuk Stanislav, <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>, retrieved 2026-07-28.
BibTeX
@misc{agentfit_dataset,
title = {AgentFit AI-readiness dataset},
author = {Gumeniuk, Stanislav},
year = {2026},
howpublished = {\url{https://agentfit.dev/dataset}},
note = {Rolling dataset, CC BY 4.0. Retrieved 2026-07-28}
}
Independence
Independent, automated assessment. AgentFit is not affiliated with the audited sites and did not ask permission to publish. Scores are heuristic, produced by one automated fetch on one date, and describe what a machine could read — not the quality of the product behind the docs. The rows AgentFit owns are not in the export: agentfit.dev and gumeniuk.com are excluded from every collection and every published figure, because the one site we cannot assess at arm's length is our own.
Removal
If a row refers to a URL you control and you want it gone, email [email protected] with the report link. Within 30 calendar days the report is removed from /browse and the direct link returns 404; the export follows within a further 24 hours, because it is regenerated daily. Copies other people already downloaded cannot be recalled — that is true of any open licence, and it is why this is a live export rather than a set of permanent quarterly releases.
How the criteria are scored · Browse the corpus · Terms · Privacy