🏆 Finalist — NIH Data Sharing Index (“S-Index”) Challenge
Press & Mentions

Blog post · Methodology

Does DataRank do more than re-count citations?

DataRank is built from citation counts, so of course it correlates with them — a correlation on its own proves nothing. The question worth answering is whether the 1-hop citer network changes any decision a citation sort would have made.

20 July 2026 · 12,931 ranked data papers · algorithm datarank_citation_only_1hop_v6

Verdict

The claim holds. DataRank agrees with citations globally (ρ = 0.99), but that figure is carried almost entirely by the long tail. Inside any citation band the agreement falls to ρ = 0.52–0.74, it separates 82% of papers that citations leave tied, and it reverses the verdict on 29.6% of pairs in the top 500.

Two caveats belong with every use of these numbers: the citer fetch is capped at 100 per paper, so head-of-corpus comparisons are made on top-100-citer samples; and 26.9% of ranked papers have no citer network at all and are scored citation-only.

A ranking can be arithmetically perfect and still measure nothing. There are two ways DataRank could be worthless: it could be a monotone function of citations, reordering nothing; or its deviations from citation order could track how completely we happened to fetch each paper's citers, in which case it would be measuring our own data collection rather than the literature. What follows tests both, on the production corpus, read-only.

A

It agrees with citations, as it must

Spearman against raw citation counts is 0.9875; Pearson against log1p(citations), the form the score is actually built from, is 0.8010. DataRank is not an orthogonal measure — a reader who trusts citations is not being asked to abandon them.

Taken alone, both numbers are uninformative. A ranking that merely relabelled citations would produce exactly this. Everything below exists to distinguish the two cases.

B

The 0.99 is a tail artifact

Recompute the same correlation inside citation bands and it falls apart — which is precisely where a reader actually compares two papers.

1–5
0.927
6–20
0.742
21–100
0.717
101–500
0.519
501+
0.570
Citation bandPapersSpearman within band
1–53,4150.9271
6–203,2780.7415
21–1002,3400.7171
101–5006570.5188
501+3740.5699

The global figure is dominated by the fact that a 3-citation paper ranks below a 3,000-citation paper under both measures. Strip that trivial ordering out and the network term is doing roughly half the work among well-cited papers.

What would have falsified this: within-band ρ staying near 0.99. That would mean the network only re-derives citation order at every scale, and the model is decoration.

C

It separates papers that citations cannot

A citation count is a coarse integer. Across the 10,064 cited papers in scope there are only 722 distinct citation values but 8,329 distinct DataRank values. 95.5% of those papers share their count with at least one other paper, and a citation sort orders them arbitrarily — DataRank gives 82% of that tied mass a distinct score.

The largest tie is 472 papers all with exactly 5 citations. DataRank spreads them across 408 distinct values, from 0.269 to 1.349 — a 5× spread among papers a citation sort declares identical. Two from inside that group, with the same citation count and the same number of citers:

Higher

Global diversity of policy, coverage, and demand of COVID-19 vaccines

5 cites · 4 citers · DataRank 1.349

Lower

Nationwide Temporal Trends in Adverse Pregnancy Outcomes

5 cites · 4 citers · DataRank 0.269

Same citation count, same number of citers — the entire difference is who cited them. This is the cleanest sense in which DataRank carries strictly more information than the count it is built from.

Excluded from this section

The 2,867 papers with zero citations. They form one degenerate tie that DataRank also scores 0.000, so counting them would overstate both the problem and the resolution rate.

D

It reverses the verdict on 3 in 10 head-to-head pairs

Among the top 500 papers by DataRank there are 36,889 inverted pairs out of 124,750 — cases where the paper with fewer citations wins. These prove the network term changes decisions rather than decorating them, and the pattern in the winners is consistent: canonical data infrastructure, the resources whose citers are themselves heavily cited.

The Protein Data Bank: a computer-based archival file for macromolecular structures

8,643 cites · DataRank 20.48

ImageNet: A large-scale hierarchical image database

62,145 cites · DataRank 18.66

The NCBI Taxonomy database

1,556 cites · DataRank 16.87

KEGG: Kyoto Encyclopedia of Genes and Genomes

39,714 cites · DataRank 16.11

The SWISS-PROT protein sequence database and its supplement TrEMBL

3,226 cites · DataRank 17.81

Systematic and integrative analysis of large gene lists using DAVID

37,684 cites · DataRank 16.22

Ensembl 2005

348 cites · DataRank 13.49

The SILVA ribosomal RNA gene database project

34,334 cites · DataRank 12.87

The obvious objection: maybe the loser simply had fewer citers fetched, and DataRank is measuring our data collection. Requiring both papers to have retrieved their full citer budget barely moves the count.

Min. completeness, both papersSurviving pairsRetained
≥ 0%36,889100.0%
≥ 50%36,889100.0%
≥ 75%36,81299.8%
≥ 90%36,34198.5%
≥ 100%35,16595.3%

95.3% of the inversions survive the strictest possible cut. The disagreement is not produced by one side having been sampled less thoroughly.

Read "100% sampled" precisely

Completeness means the fraction of the 100-citer budget retrieved, not the fraction of a paper's real citers. For ImageNet at 62,145 citations, 100% still means we read its top 100. That sample is principled and deterministic — the fetch is sorted cited_by_count:desc, so the highest-signal citers are always the ones kept — but it is a sample, and head-of-corpus comparisons should be described that way.

E

Reordering is concentrated where readers look

Seven of the top ten papers are not the top ten by citations, and the median top-100 paper sits 65 places from where a citation sort would put it.

Top 10
70%
Top 25
68%
Top 50
68%
Top 100
53%
Top 500
22%

Share of each top-N list that a pure citation ranking would not have included. Reordering is strongest at the very head and decays outward — the opposite of what a noise process would produce, and the region where a ranking's choices carry the most weight.

F

The movers are recognisable to a domain reader

No statistic can make this judgement. Restricted to well-cited papers (≥100 citations), the biggest risers are database and challenge-benchmark papers: CIBEX moves from citation rank #722 to #93, the PhysioNet/CinC ECG challenge from #880 to #142, the NIH Roadmap for Medical Research from #1016 to #220. That is the infrastructure signature the 1-hop model was designed to capture.

The fallers are the more useful list, because they are where the method's failure mode shows up rather than its success. Chromatin Architecture of the Human Genome falls from #374 to #2799 — with 505 citations and one citer fetched.

Do not read the biggest fallers as findings

A paper with 505 citations and 1 citer fetched did not have a weak network — we failed to collect it. A starved fetch and a genuinely weak network are indistinguishable to the engine. The interesting faller is the UniProt website API (#765 → #2162), which had a full 100-citer fetch and still fell: that is a real result about its citing literature.

G

The check designed to kill the claim

If rank movement tracked how completely we fetched each paper's citers, DataRank would be measuring our own data collection, and everything above would be an artifact. Measured across 9,459 papers with a citer network, that correlation is +0.32.

A correlation near zero would mean the reordering is independent of fetch depth; near 1.0 would mean DataRank measures how hard we looked. +0.32 is present but not dominant — collection depth explains some of the movement and cannot be dismissed, which is why the starved-fetch rows above are flagged rather than reported as results. Only 4 of 1,038 well-cited papers (0.4%) got a single citer or fewer.

The larger disclosure is this: 3,472 of 12,931 ranked papers have no citer network at all. They score citation-only by design — the offline fallback sets N(p) = 0 — so they can fall relative to their citation rank but can never rise. Any published claim should quote this alongside the corpus size.

What would have falsified the whole story: within-band ρ near 0.99, inversions collapsing under the completeness filter, or this correlation running above ~0.6. None of the three occurred.

All figures are computed read-only against the production database over the canonical data_only scope — is_dataset papers only. Reproduce with python -m scripts.showcase_datarank_robustness and python -m scripts.validate_datarank.

A note on the numbers: the citer cap is inferred from the stored data (100 here), not read from DATARANK_MAX_CITERS, which is the cap the next recompute would use. Assuming the setting rather than measuring the data silently rescales every completeness figure above.

More on how the score is computed in the methodology.