🏆 Finalist — NIH Data Sharing Index (“S-Index”) Challenge

Percy Liang

Stanford University

ORCID: 0000-0002-0458-6139
Computer Science

Pilot corpus only

This score is computed over theSindex pilot corpus and does not cover the full scientific literature. Scores are relative to papers we have ingested — papers, authors, and institutions outside the pilot are not represented. Methodology.

Top 4%percentile
0.919Author DataRank

Indexed papers

1in pilot corpus
datarank_citation_only_1hop_v6· scope data_onlyMethodology
Why this DataRank?

An author's DataRank is the sum of the DataRanks of all 1 indexed paper attributed to them. A prolific author with many moderate-impact papers can outrank one with a single high-impact paper.

Author scores recompute whenever paper DataRanks are refreshed, so this number lags the underlying paper scores by at most one batch run.

Read the full methodology →

Top data-sharing exemplar

The highest-impact dataset this researcher has shared, ranked by DataRank — the single contribution doing the most to lift their data-sharing standing.

Top 22%18 citations

Daniel Fu, Quentin Anthony, Yonatan Oren, Adams, Shane, Anton Alexandrov +14 more

Papers

Driven by 2 papers — median percentile 78. Top paper: RedPajama: an Open Dataset for Training Large Language Models.

Top 22%18 citations

Daniel Fu, Quentin Anthony, Yonatan Oren, Adams, Shane, Anton Alexandrov +14 more

LinkBERT: Pretraining Language Models with Document Links

Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)(2022)10.18653/v1/2022.acl-long.551
N/A
0DataRank · unranked
303 citations

Jure Leskovec, Percy Liang, Michihiro Yasunaga