A benchmark study on current GWAS models in admixed populations is a dataset published in Briefings in Bioinformatics (2023). On theSindex it has a DataRank of 0.538, placing it in the top 36.1% of the data-sharing corpus. It has been cited 15 times, with 12 citing works in its 1-hop citation network. Its calibrated FAIR score is 67/100.
Ranks in the top 36% for downstream scientific impact
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 .”
The synthetic dataset is assigned a BioStudies accession (S-BSST936), which is a persistent identifier scheme.
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 .”
The synthetic data are held in BioStudies, a curated repository registered in re3data. [majority verdict 'yes' (4/5 passes agreed)]
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 .”
The dataset identifier appears only in the body text of the data availability statement, not as a reference-list entry.
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“DATA AVAILABILITY The analysis code to produce the major results presented in the paper is available at https://github.com/ZikunY/Benchmark_GWAS . The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 . Request for genetic and phenotype data of Peruvian cohort can be submitted to [email protected] .”
The data availability statement points to a repository record for the synthetic data (BioStudies with accession).
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“We generated a cohort of synthetic individuals ( N = 19 234) that simulates (i) a large sample size; (ii) two-way admixture (Native American and European ancestry) and (iii) a binary phenotype.”
The dataset content is described in running prose rather than in an itemised section, table, or list.
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 .”
The synthetic data are openly accessible at a public repository with no stated precondition.
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“The synthetic data used in the simulation study are available at https://www.ebi.ac.uk/biostudies/studies/S-BSST936 .”
The paper describes where the data can be accessed but does not label the access level with a standard term. [majority verdict 'partial' (4/5 passes agreed)]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
“Request for genetic and phenotype data of Peruvian cohort can be submitted to [email protected] .”
The gatekeeper for the sensitive human Peruvian cohort data is a natural person (the corresponding author).
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
The paper does not state when the data become available or how long they persist. [majority verdict 'no' (4/5 passes agreed)]
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
No file format for the released data is mentioned in the paper.
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
No community-standard data or metadata checklist, ontology, or schema is named for the released data.
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
“Genomes Project C, et al. A global reference for human genetic variation. Nature 2015;526:68–74. 26432245 10.1038/nature15393”— not found in the paper; verdict downgraded
The paper cites the 1000 Genomes Project reference dataset with a DOI, which is a qualified reference to an external resource. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
No reuse license is stated for the data; the Creative Commons license applies only to the article.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
No version token or date is provided for the released datasets.
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
“The analysis code to produce the major results presented in the paper is available at https://github.com/ZikunY/Benchmark_GWAS .”
The paper provides a machine-resolvable GitHub URL for the study's own code.
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“The National Institutes of Health (NIH), National Institute on Aging (NIH-NIA) supported this work [R56AG069118, R56AG066889, R56AG059756, U19AG074865, R01AG082009].”
Award numbers are provided for the named funder.
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“We then used RFMix2 (v2.0.3), a discriminative approach that estimates both global and local ancestry using random forests, to inference global ancestry”— not found in the paper; verdict downgraded
The paper names specific instruments and software (HAPNEST, Shapeit, RFMix2, Infinium Global Screening Array-24 BeadChip) used to produce the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
No documentation object (README, codebook, data dictionary) is mentioned as accompanying the deposited data.
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
Base Score Contribution
0.416
From this paper's citation signal
Citation Network Contribution
0.122
From 8 citing papers with measurable signal
Ranked by each citer's contribution to N(p) — log1p(Cq) divided by its reference count — out of 12 citers.
National Institute on Aging
Grant: R56AG069118
National Institute on Aging
Grant: R56AG066889
National Institute on Aging
Grant: R56AG059756
National Institute on Aging
Grant: U19AG074865
National Institute on Aging
Grant: R01AG082009
NIA NIH HHS
Grant: RF1 AG082009
National Institutes of Health
Grant: 3U19AG074865-04S9
Recruitment and Retention for Alzheimer's Disease Diversity Genetic Cohorts in the ADSP (READD-ADSP)
National Institutes of Health
Grant: 1R56AG069118-01
Genetic and environmental risk factors in mestizos and indigenous populations of Peru: the role of Native component in Alzheimer's disease
National Institutes of Health
Grant: 1RF1AG082009-01A1
Polygenic Risk Scores for Alzheimer's Disease in Hispanic/Latinx Populations
National Institutes of Health
Grant: 1R56AG066889-01
Whole genome sequencing of the Mexican Health Aging Study (MHAS) cohort
National Institutes of Health
Grant: 5R56AG059756-02
Genetics of Alzheimer's Disease in Mexico
National Institutes of Health
NIH HHS
NIH HHS
FWCI
3.28
Citation Percentile
0.9%
Citation Trend
Fields of Study
MeSH Terms
Keywords