🏆 Finalist — NIH Data Sharing Index (“S-Index”) Challenge

FAIR Agent

Coaching you to share your data better

Enter a DOI, upload a paper, or paste its text. The FAIR Agent reads the full text, scores how Findable, Accessible, Interoperable, and Reusable the data is, and hands you concrete steps to improve it. A parallel quality metric — independent of the DataRank citation score. See the FAIR showcase →

Or try:

How the FAIR score works

FAIR measures how well a paper and its underlying data follow the FAIR principles — Findable, Accessible, Interoperable, Reusable. The headline 0–100 score comes from the fact-shaped criteria that independent models agree on; the four F/A/I/R sub-scores are shown as the advisory full picture.

Findable

Rich, machine-readable metadata, a persistent DOI, and presence in repositories and indexes so the work and its data can be discovered.

Accessible

The paper and data are openly retrievable — Open Access, deposited files, and a clear protocol for how to obtain them.

Interoperable

Data uses standard formats and vocabularies and is linked through standard identifiers — dataset DOIs, accessions, and registered relations.

Reusable

A clear open license, provenance, versioning, and enough methodological detail for others to reproduce and reuse the work.

How it's scored. An AI agent reads the paper's freely available full text and answers a fixed checklist of criteria drawn from published standards — the RDA FAIR Data Maturity indicators, the F-UJI/FAIRsFAIR metrics, and the NIH Data Management & Sharing Policy. For each criterion it must answer yes, partially or no and quote the sentence in the paper that justifies it. We check the quote really appears in the text; a verdict the model can't point to is downgraded, not trusted. The score itself is then computed from those verdicts in code — the model never picks a number. We only score data papers we can read in full, never from an abstract alone, and if the agent can't complete its assessment we publish no score rather than a guess.

Headline vs. advisory. We measured that two capable models agree on the fact-shaped criteria — is there a repository, an accession, a licence, a code link — but diverge on the judgment ones (is the documentation adequate?), because that answer isn't in any single sentence. So the published headline is computed only from the fact-shaped criteria, the ones a score can stand behind; the judgment criteria are still assessed and shown, but as advisory guidance that doesn't move the number. The score is a declared instrument: what a specific, pinned model reports, reproducibly — not a claim of a single universal truth.

Calibrated across papers. The headline score is standardized into a percentile over every paper evaluated by the current version of the agent, so you can see where a paper stands relative to the rest of the corpus.

Independent of DataRank. FAIR is a parallel quality metric. It is never folded into the citation-based DataRank score — it measures data stewardship, not citation impact.

How we check it

FAIR is validated against things it should be able to tell apart, and we publish the results whether or not they flatter the score. Two checks matter most: how far two independent models agree on the same paper, and whether the score is merely restating citation counts. On the second, the honest answer is that it isn't — across 2,850 recent full-text data papers the rank correlation with citations is ρ = 0.07, close to unrelated. That is the argument for measuring data sharing separately rather than assuming citation impact already captures it.

Read the full measurement and its limits — reproducibility across models, the citation check, and what the score does and doesn't separate.