A high-quality genome assembly for a desert-adapted rodent, Merriam’s kangaroo rat (Dipodomys merriami) is a dataset published in Journal of Heredity (2025). On theSindex it has a DataRank of 0, placing it in the top 100% of the data-sharing corpus. Its calibrated FAIR score is 67/100.
Ranks in the top 100% for downstream scientific impact
Linked data & code
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).”
The paper provides BioProject accessions (PRJNA851460, PRJNA851459) which are persistent identifiers registered in identifiers.org. [majority verdict 'yes' (4/5 passes agreed)]
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140.”
The paper names NCBI Short Read Archive and Dryad Data repository as holders of the data. [majority verdict 'yes' (4/5 passes agreed)]
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).”
The dataset identifiers appear only in the body text (Data Availability Statement), not in a reference-list entry. [majority verdict 'partial' (4/5 passes agreed)]
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate). Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140. Assembly scripts and other data for the analyses presented can be found at the following GitHub repository: www.github.com/ccgproject/ccgp_assembly . Preliminary annotation and mitochondrial genome sequence are available on the Dryad Data repository at https://doi.org/10.5061/dryad.x0k6djhtc .”
The data-availability statement points to repositories (NCBI, Dryad) with accessions and DOIs, meeting Colavizza category 3. [majority verdict 'yes' (3/5 passes agreed)]
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“Table 1. Metrics for the primary and alternate assemblies of Merriam’s kangaroo rat (Dipodomys merriami) genome.”— not found in the paper; verdict downgraded
The paper includes an itemised inventory of assembly metrics in a table, providing structured description of the dataset. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).”
The data are stated to be available in public repositories without any precondition such as embargo, registration, or request. [majority verdict 'yes' (4/5 passes agreed)]
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).”
The paper does not explicitly label the access level, but describes the action of accessing the data via repository identifiers, allowing inference of open access. [majority verdict 'partial' (4/5 passes agreed)]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
“Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140.”
The data are non-sensitive rodent genome sequences, and no gatekeeper is named; the data are openly deposited.
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
“Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).”
The paper states the data are available in the present tense but gives no persistence commitment. [majority verdict 'partial' (3/5 passes agreed)]
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
The paper does not name the file format of the released data (e.g., FASTA, FASTQ).
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
No data or metadata community standard from the specified list (e.g., MIAME, BIDS, ontologies) is named and applied to the data; BUSCO is a tool, not a standard format or schema.
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
“We used liftoff ( Shumate and Salzberg 2021 ) to lift over gene coding sequences, exons, and mRNAs from D. spectabilis (NCBI: GCF_019054845.1) to D. merriami”
The paper provides an identifier (GCF_019054845.1) for a third-party resource (D. spectabilis genome) used in the analysis. [majority verdict 'yes' (4/5 passes agreed)]
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
No reuse licence is stated for the data; the CC BY-NC 4.0 licence applies only to the article text.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
Neither a version token nor a date is provided for the deposited data; the BioProject IDs and DOI are not versioned. [majority verdict 'no' (3/5 passes agreed)]
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
“Assembly scripts and other data for the analyses presented can be found at the following GitHub repository: www.github.com/ccgproject/ccgp_assembly .”
The paper provides a machine-resolvable code-forge URL for the study's own code.
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“This work was supported by the California Conservation Genomics Project, with funding provided to the University of California by the State of California, State Budget Act of 2019 [UC Award ID RSI-19-690224].”
The paper includes an award/grant number (RSI-19-690224) attached to a named funder.
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“High molecular weight (HMW) genomic DNA (gDNA) was extracted from 78 mg of liver tissue (male, MVZ:Mamm:240054, JLP29074) using the Nanobind Tissue Big DNA kit as per the manufacturer’s instructions (Pacific BioSciences—PacBio, Menlo Park, CA).”
The paper names specific instruments, kits, and software used to produce the data (e.g., Nanobind Tissue Big DNA kit, PacBio Sequel II, Illumina NovaSeq 6000). [majority verdict 'yes' (4/5 passes agreed)]
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
The paper does not mention a README, data dictionary, or codebook that accompanies the data. [majority verdict 'no' (4/5 passes agreed)]
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
University of California
Grant: RSI-19-690224
NIH HHS
Grant: S10 OD018174
NIH HHS
Grant: S10 OD010786
MeSH Terms
Keywords
Data from: A high-quality genome assembly for a desert-adapted rodent, Merriam’s kangaroo rat (Dipodomys merriami)