Complete genome sequence of a virulent barcoded Mycobacterium tuberculosis str. Erdman commonly used for non-human primate infection studies is a dataset published in Microbiology Resource Announcements (2025). On theSindex it has a DataRank of 0.104, placing it in the top 77.8% of the data-sharing corpus. It has been cited 1 time, with 1 citing works in its 1-hop citation network. Its calibrated FAIR score is 58/100.
Ranks in the top 78% for downstream scientific impact
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“The complete genome sequence and annotation are available on GenBank under accession CP172229.”— not found in the paper; verdict downgraded
The paper provides a GenBank accession for the genome sequence, which is a persistent identifier. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“Sequencing reads are on SRA under accession numbers SRX26089191 and SRX26089190 , both of which are associated with BioProject PRJNA1161419 . The complete genome sequence and annotation are available on GenBank under accession CP172229 .”
The paper names SRA, BioProject, and GenBank as repositories holding the data. [majority verdict 'yes' (3/5 passes agreed)]
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“Sequencing reads are on SRA under accession numbers SRX26089191 and SRX26089190 , both of which are associated with BioProject PRJNA1161419 . The complete genome sequence and annotation are available on GenBank under accession CP172229 .”
The dataset identifiers appear only in the body text (Data Availability section), not in the reference list. [majority verdict 'partial' (3/5 passes agreed)]
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“All relevant code and methods are available on GitHub: https://github.com/maxgmarin/erdman-asm-explore . Sequencing reads are on SRA under accession numbers SRX26089191 and SRX26089190 , both of which are associated with BioProject PRJNA1161419 . The complete genome sequence and annotation are available on GenBank under accession CP172229 .”
The data availability statement points to archived data in public repositories with accessions, corresponding to Colavizza category 3.
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“The final genome assembly, 4,416,075 bp with 65.61% GC content, was annotated by lifting over high-confidence annotations from the H37Rv genome (GenBank AL123456.3 ) using RATT (v1.5), with unannotated regions filled using de novo annotations from Bakta (v1.9). The resulting annotated genome contains 4,011 coding sequences and harbors a barcoding plasmid at the L5 phage integration site at genomic positions: 2,764,911–2,770,126.”
The dataset's content and extent are described in running prose (size, GC content, coding sequences), but not in an itemised inventory (table or section heading). [majority verdict 'partial' (4/5 passes agreed)]
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“Sequencing reads are on SRA under accession numbers SRX26089191 and SRX26089190 , both of which are associated with BioProject PRJNA1161419 . The complete genome sequence and annotation are available on GenBank under accession CP172229 .”
The data are stated to be available in public repositories (SRA, GenBank) with no stated precondition. [majority verdict 'yes' (3/5 passes agreed)]
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“Sequencing reads are on SRA under accession numbers SRX26089191 and SRX26089190 , both of which are associated with BioProject PRJNA1161419 . The complete genome sequence and annotation are available on GenBank under accession CP172229 .”
The paper describes the locations of the data (SRA and GenBank) but does not use an explicit access-level label such as 'open access' or 'publicly available'; the access level must be inferred from the action of depositing in public repositories. [majority verdict 'partial' (3/5 passes agreed)]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
The data are a bacterial genome sequence, not human or sensitive data, and no gatekeeper is mentioned.
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
No temporal commitment or retention statement is made for the data. [majority verdict 'no' (3/5 passes agreed)]
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
The paper does not name a file format for the released data (e.g., FASTA, FASTQ).
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
The paper does not name a community standard (e.g., MIAME, MIxS, an ontology) for the data. It uses standard tools but not a named standard.
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
“The final genome assembly, 4,416,075 bp with 65.61% GC content, was annotated by lifting over high-confidence annotations from the H37Rv genome (GenBank AL123456.3 ) using RATT (v1.5)”
The paper gives an identifier (GenBank AL123456.3) for the H37Rv reference genome, a resource other than its own dataset. [majority verdict 'yes' (4/5 passes agreed)]
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
No reuse license is attached to the data; the article's CC BY license does not apply to the data itself.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
No version token or date is provided for the data snapshot; the GenBank accession is given without a version suffix, and no release date is stated. [majority verdict 'no' (4/5 passes agreed)]
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
“All relevant code and methods are available on GitHub: https://github.com/maxgmarin/erdman-asm-explore”
The paper provides a machine-resolvable code repository URL (GitHub) for the study's own code.
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“This project was funded in whole or in part with Federal funds from the National Institute of Allergy and Infectious Diseases, National Institutes of Health, Department of Health and Human Services, under Contract No. 75N93019C00071. M.G.M. is currently supported by the National Library of Medicine/NIH grant (T15LM007092).”
The paper includes award/grant numbers (75N93019C00071, T15LM007092) attached to named funders.
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“DNA was commercially sequenced using long-read (Oxford Nanopore) and short-read (Illumina) platforms (SeqCenter, Pittsburgh). Illumina sequencing libraries were prepared using the Illumina DNA Prep kit following the standard protocol and sequenced on a NextSeq 2000, generating a total of 29,432,744 reads (2 × 151bp paired-end). Adapter trimming was performed using bcl-convert (v4.0.3). Nanopore sequencing libraries were prepared using the ONT Native Barcoding Kit 24V14 (SQK-NBD114.24) and sequenced on a MinION device with R10.4.1 flow cells. Basecalling with Guppy (v6.4.6) and trimming with porechop (0.2.3_seqan2.1.1) yielded 104,882 reads with median length of 1,351bp and an N50 of 8,928. The nanopore reads were de novo assembled into a closed, circular chromosome using the Flye assembler (v2.9) and medaka polisher (v1.9.1). The resulting long-read assembly was further polished with short reads using the PolyPolish software (v0.5) and rotated to start at the dnaA locus using Circlator (v1.5).”
The paper names specific instruments, kits, and software versions used to produce the data (e.g., Illumina NextSeq 2000, ONT MinION, Flye v2.9).
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
No documentation object (README, data dictionary, codebook) is mentioned as accompanying the deposited data.
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
Base Score Contribution
0.104
From this paper's citation signal
Citation Network Contribution
0
From 0 citing papers with measurable signal
This paper's DataRank is currently driven only by its base citation score. None of the citing papers had measurable citation signal.
Learn more about DataRank methodology →HHS | National Institutes of Health
Grant: 75N93019C00071
HHS | National Institutes of Health
Grant: T15LM007092
National Institutes of Health
Grant: 5T15LM007092-07
HARVARD-MIT-NEMC RESEARCH TRAINING IN HEALTH INFORMATICS
FWCI
1.59
Citation Percentile
0.8%
Citation Trend
Fields of Study
Keywords
Sustainable Development Goals