Towards complete and error-free genome assemblies of all vertebrate species is a dataset published in Nature (2021). On theSindex it has a DataRank of 6.0, placing it in the top 3.3% of the data-sharing corpus. It has been cited 3,155 times, with 100 citing works in its 1-hop citation network. Its calibrated FAIR score is 67/100.
Ranks in the top 3% for downstream scientific impact
Linked data & code
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“PRJNA489243”
The accession PRJNA489243 is a BioProject identifier, which is a persistent identifier scheme accepted at 'yes'.
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“All raw data, intermediate and final assemblies are publicly available via GenomeArk (https://vgp.github.io/genomeark), archived on NCBI/EBI BioProject under accession PRJNA489243”
GenomeArk and NCBI/EBI BioProject are named repositories, both of which are proper nouns listed in the class-1 repository list.
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“archived on NCBI/EBI BioProject under accession PRJNA489243”
The dataset identifier (PRJNA489243) appears only in the body text (data availability section), not in the reference list.
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“All raw data, intermediate and final assemblies are publicly available via GenomeArk (https://vgp.github.io/genomeark), archived on NCBI/EBI BioProject under accession PRJNA489243 with annotations, and browsable on the UCSC Genome Browser (https://hgdownload.soe.ucsc.edu/hubs/VGP/).”— not found in the paper; verdict downgraded
The DAS points to a repository record (GenomeArk and NCBI BioProject) with an accession number, satisfying Colavizza category 3. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“All raw data, intermediate and final assemblies are publicly available”
The dataset content is described in a single sentence in running prose, not an itemised inventory. [majority verdict 'partial' (4/5 passes agreed)]
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“All raw data, intermediate and final assemblies are publicly available via GenomeArk (https://vgp.github.io/genomeark), archived on NCBI/EBI BioProject under accession PRJNA489243 with annotations”
The text gives a route to the data at GenomeArk and NCBI BioProject with no stated precondition; the data are 'publicly available'. [majority verdict 'yes' (4/5 passes agreed)]
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“publicly available”
The data availability statement explicitly labels the data as 'publicly available', which is a natural-language synonym for 'open access'. [majority verdict 'yes' (4/5 passes agreed)]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
The data are genome assemblies from non-human species and are openly available; no gatekeeper is mentioned because no access control is needed.
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
The paper states the data are publicly available but does not specify a retention period or persistence commitment. [majority verdict 'no' (3/5 passes agreed)]
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
No file-format token (e.g., FASTA, FASTQ, VCF) is named for the released data; the text only refers to 'assemblies' generically.
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
“BUSCO vertebrate gene set”
The paper names a community standard (BUSCO) applied to the data. [majority verdict 'yes' (3/5 passes agreed)]
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
“GRCh38”— not found in the paper; verdict downgraded
The paper gives the human reference genome assembly identifier GRCh38, which is an identifier for a resource other than the study's own data. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (3/5 passes agreed)]
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
The paper states a Creative Commons Attribution 4.0 license for the article, but no license is explicitly applied to the data themselves.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
No version token or date is given for the data; the assemblies are referred to by accession (PRJNA489243) but without a version or release date. [majority verdict 'no' (4/5 passes agreed)]
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
“https://github.com/VGP/vgp-assembly”
The study's code is given a machine-resolvable locator (GitHub URL), which is an authoritative versioned locator.
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“WT207492”
The paper gives a Wellcome Trust grant number (WT207492) in the acknowledgements, which is an alphanumeric award identifier.
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“PacBio continuous long reads (CLR) or Oxford Nanopore long reads”— not found in the paper; verdict downgraded
The text names specific sequencing technologies (PacBio CLR, Oxford Nanopore) used to produce the data, which are proper-noun instruments. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
“Extended Data Table 1”
Variable definitions are provided in an article table, not in a separate documentation object shipped with the data. [majority verdict 'partial' (3/5 passes agreed)]
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
Base Score Contribution
1.2
From this paper's citation signal
Citation Network Contribution
4.8
From 100 citing papers with measurable signal
Ranked by each citer's contribution to N(p) — log1p(Cq) divided by its reference count — out of 100 citers.
Villum Fonden
Grant: 00025900
NIGMS NIH HHS
Grant: R01 GM130691
Wellcome Trust
Grant: WT207492
Biotechnology and Biological Sciences Research Council
Grant: BBS/E/T/000PR9818
NIDCD NIH HHS
Grant: R21 DC014432
Wellcome Trust
Grant: 108749/Z/15/Z
NHGRI NIH HHS
Grant: R01 HG010485
Medical Research Council
Grant: MR/T021985/1
Intramural NIH HHS
Grant: ZIA HG200398
Biotechnology and Biological Sciences Research Council
Grant: BBS/E/T/000PR9817
NHGRI NIH HHS
Grant: U41 HG007234
Wellcome Trust
Grant: 207492/Z/17/Z
European Research Council
Grant: 681396
NHGRI NIH HHS
Grant: U24 HG007234
NHGRI NIH HHS
Grant: U41 HG002371
NHGRI NIH HHS
Grant: R44 HG008118
FWCI
147.53
Citation Percentile
1.0%
Citation Trend
Fields of Study
MeSH Terms
Keywords
Sustainable Development Goals
Additional file 4 of Genomic insights into body size evolution in Carnivora support Peto’s paradox
Additional file 4 of Genomic insights into body size evolution in Carnivora support Peto’s paradox
Additional file 5 of Genomic insights into body size evolution in Carnivora support Peto’s paradox
Additional file 5 of Genomic insights into body size evolution in Carnivora support Peto’s paradox
Additional file 10 of Telomere-to-telomere assembly of a fish Y chromosome reveals the origin of a young sex chromosome pair
Additional file 10 of Telomere-to-telomere assembly of a fish Y chromosome reveals the origin of a young sex chromosome pair
Additional file 1 of Telomere-to-telomere assembly of a fish Y chromosome reveals the origin of a young sex chromosome pair
Additional file 1 of Telomere-to-telomere assembly of a fish Y chromosome reveals the origin of a young sex chromosome pair
Additional file 1 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 1 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 2 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 2 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 3 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 3 of Accurate long-read de novo assembly evaluation with Inspector
Additional file 10 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel
Additional file 10 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel
Additional file 2 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel
Additional file 2 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel
Additional file 8 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel
Additional file 8 of Genomic variations and epigenomic landscape of the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel