Pilot corpus only
This score is computed over theSindex pilot corpus and does not cover the full scientific literature. Scores are relative to papers we have ingested — papers, authors, and institutions outside the pilot are not represented. Methodology.
Driven by 2 papers. Top paper: “Automating expert-level medical reasoning evaluation of large language models”.
Wenya Xie, Jiaxi Li, Zaifu Zhan, Meijia Song, Han Yang +14 more
Shuang Zhou, Yu Hou, Yiran Song, Min Zeng, Fang Tian +10 more