Pilot corpus only
This score is computed over theSindex pilot corpus and does not cover the full scientific literature. Scores are relative to papers we have ingested β papers, authors, and institutions outside the pilot are not represented. Methodology.
Driven by 1 paper. Top paper: βMedCalc-Bench: Evaluating Large Language Models for Medical Calculationsβ.
Qiao Jin, Guangzhi Xiong, S.M. Dunn, Serina Applebaum, Zain Anwar +12 more