Draft Genomes Sequences of 11 Geodermatophilaceae Strains Isolated from Building Stones from New England and Indian Stone Ruins found at historic sites in Tamil Nadu, India is a dataset published in Journal of Genomics (2022). On theSindex it has a DataRank of 0.104, placing it in the top 77.8% of the data-sharing corpus. It has been cited 1 time. Its calibrated FAIR score is 38/100.
Ranks in the top 78% for downstream scientific impact
Linked data & code
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The paper provides BioProject accession numbers (PRJNA...) which are persistent identifiers in a recognized scheme. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“The draft genome sequences of these bacterial strains have been deposited in GenBank under the accession numbers listed in Table 2.”
The paper names GenBank (and DDBJ/ENA) as the repository holding the data.
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The dataset identifiers appear only in the body text (data-availability statement and Table 2), not as a reference-list entry. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“The draft genome sequences of these bacterial strains have been deposited in GenBank under the accession numbers listed in Table 2. Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The data-availability statement points to a repository record with accession numbers and BioProject identifiers (Colavizza category 3). [downgraded to 'partial' — no verifiable quote from the paper]
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“Table 2. Genome Statistics.”
The paper includes an itemized table (Table 2) that lists each genome's accession, number of reads, contigs, coverage, assembly size, N50, CDSs, G+C content, rRNAs, and tRNAs, serving as a structured inventory of the dataset.
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The paper states the data are available at public repositories with no stated precondition such as an embargo, registration, or application. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The paper describes the action of accessing the data via the repositories but does not apply a standard access-level label such as 'open access' or 'restricted access'. [downgraded to 'no' — no verifiable quote from the paper]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
“Both the assembly and raw reads are available at DDBJ/ENA/GenBank under BioProject numbers: PRJNA478225, PRJNA478231 PRJNA478233, PRJNA478236, PRJNA478237, PRJNA478240, and PRJNA480027.”— not found in the paper; verdict downgraded
The data are bacterial genomes and are not sensitive human-subject data; the paper names no gatekeeper of any kind, and the data are openly available.
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
The paper does not state how long the data will persist or when they become available beyond the deposit. [majority verdict 'no' (4/5 passes agreed)]
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
The paper does not name any file format for the deposited data (e.g., FASTQ, FASTA) despite mentioning raw reads and assemblies.
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
“Gene Ontology (GO) domains were assigned to each aligned genome protein sequence”
The paper explicitly names the Gene Ontology (GO) and COG as community standards applied to the data. [majority verdict 'yes' (3/5 passes agreed)]
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
The paper does not provide a persistent identifier (DOI, accession, or version) for any external resource that the data depend on; it only names tools and databases in prose.
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
The paper does not state any license or reuse terms for the data; the CC BY 4.0 license applies only to the article itself, not the deposited data.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
The paper does not provide a version token, release date, or access/download date for the dataset; the accession numbers are not versioned.
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
The paper does not mention any code written for the study and offers no locator (URL, DOI, or repository) for software.
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“Innovative Programs to Enhance Research Training (IPERT) from the National Institute of General Medical Sciences R25GM125674”
The paper lists specific grant numbers (R25GM125674, P20GM103506, DBI-1229361) attached to named funders. [majority verdict 'yes' (4/5 passes agreed)]
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“Sequencing was completed on an Illumina HISeq 2500 HiSeq2500 platform (Illumina Inc., San Diego, CA) to produce 250 bp paired -end reads at the Hubbard Center for Genome Studies (UNH, Durham, NH).”
The paper names specific instruments (Illumina HiSeq 2500) and software versions (SPAdes v3.13, Trimmomatic v0.36) used to produce the data. [majority verdict 'yes' (4/5 passes agreed)]
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
“Table 2. Genome Statistics.”
The variable definitions (genome statistics) are provided inside the article in Table 2, but no documentation object (README, codebook) is said to accompany the data.
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
Base Score Contribution
0.104
From this paper's citation signal
Citation Network Contribution
0
Citation network not refreshed for this result
This paper's DataRank is currently driven only by its base citation score. Citation network data was not refreshed for this result.
Learn more about DataRank methodology →NIGMS NIH HHS
Grant: R25 GM125674
NIGMS NIH HHS
Grant: P20 GM103506
National Institutes of Health
Grant: 5R25GM125674-02
Bioinformatics "Train the Trainer" (T3): The Integration of Bioinformatics into an Undergraduate Biology Curriculum
National Institutes of Health
Grant: 2P20GM103506-06
Administrative Core
National Science Foundation
Grant: 1229361
MRI: Acquisition of an Illumina HiSeq 2000
FWCI
0.18
Citation Percentile
0.4%
Citation Trend
Fields of Study
Keywords
Sustainable Development Goals