Pilot corpus only
This score is computed over theSindex pilot corpus and does not cover the full scientific literature. Scores are relative to papers we have ingested — papers, authors, and institutions outside the pilot are not represented. Methodology.
An author's DataRank is the sum of the DataRanks of all 1 indexed paper attributed to them. A prolific author with many moderate-impact papers can outrank one with a single high-impact paper.
Author scores recompute whenever paper DataRanks are refreshed, so this number lags the underlying paper scores by at most one batch run.
Read the full methodology →The highest-impact dataset this researcher has shared, ranked by DataRank — the single contribution doing the most to lift their data-sharing standing.
Daniel Fu, Quentin Anthony, Yonatan Oren, Adams, Shane, Anton Alexandrov +14 more
Driven by 5 papers — median percentile 78. Top paper: “RedPajama: an Open Dataset for Training Large Language Models”.
Daniel Fu, Quentin Anthony, Yonatan Oren, Adams, Shane, Anton Alexandrov +14 more
Stefano Massaroli, Éric Nguyen, Daniel Y. Fu, Tri Dao, Stephen A. Baccus +4 more
Khaled K. Saab, Michael Poli, Tri Dao, Karan Goel, Christopher Ré +1 more
Elliot L. Epstein, Éric Nguyen, Armin W. Thomas, Michael Zhang, Tri Dao +3 more
Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, Volodymyr Kuleshov +1 more