PanCNV-Explorer: Deciphering copy number alterations across human cancers is a dataset published in bioRxiv (Cold Spring Harbor Laboratory) (2026). On theSindex it has a DataRank of 0, placing it in the top 100% of the data-sharing corpus. Its calibrated FAIR score is 25/100.
Ranks in the top 100% for downstream scientific impact
DataRank reads this dataset's downstream impact straight off the citation graph — no black box, no proprietary weighting. How is this computed?
FAIR checklist signals are shown for context only and do not affect DataRank scoring.
Full FAIR picture · advisory
The headline score is computed from the scored criteria — the fact-shaped checks (a repository, an accession, a licence) that two independent models agree on. The advisory criteria below are real FAIR guidance but rest on judgment calls that models read differently, so they inform without moving the number.
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.”— not found in the paper; verdict downgraded
The paper gives a web address for the data, not a persistent identifier from a PID scheme. [downgraded to 'no' — no verifiable quote from the paper]
RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit · RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier' · FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.”— not found in the paper; verdict downgraded
The data are hosted on an institutional web server, not a named repository from the curated list. [downgraded to 'no' — no verifiable quote from the paper]
RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed ( · NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived · NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.”— not found in the paper; verdict downgraded
The dataset's identifier (URL) appears only in the body text, not as a reference-list entry. [downgraded to 'no' — no verifiable quote from the paper]
FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first- · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes · FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'
Advisory · not in the published score
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/. The source code of PanCNV-Explorer, including the database generation, the web server, the web front end and the CNA annotation modules, is available at https://gitlab.gwdg.de/MedBioinf/mtb/cnv-database. The annotation modules for CNVs and structural variants were added to the Onkopus framework and are publicly available through APIs at https://mtb.bioinf.med.uni-goettingen.de/onkopus/api. All data sources of the copy number alteration annotation are publicly available.”— not found in the paper; verdict downgraded
The data availability statement points to a web instance and code repositories, not to a repository record with an accession, so it is a partial category. [downgraded to 'no' — no verifiable quote from the paper]
Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li · Springer Nature research data policy — Data Availability Statements: standard statement templat · RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes
“we created three versions of the database: First, we generated the database according to consensus segments, where we calculated the occurring CNV segments from the existing start and end positions and then applied clustering to reduce the number of CNV segments to 200.000. Second, we computed fixed-sized bins of 10.000 base pairs across the whole genome. And third, we generated gene-level CNV segments, were CNV start and end positions corresponded to the positions of protein-coding genes.”
The dataset's content is described in running prose, not in a section, table, or enumerated list, so it is a partial description. [majority verdict 'partial' (4/5 passes agreed)]
RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential) · FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability' · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.”— not found in the paper; verdict downgraded
The data are stated to be publicly available with no stated precondition. [downgraded to 'partial' — no verifiable quote from the paper]
RDA-A1.1-01D — 'Data is accessible through a free access protocol' · FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'
Advisory · not in the published score
“A public instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.”— not found in the paper; verdict downgraded
The paper labels the data as 'public' in the data availability statement, an explicit access-level label. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]
FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data' · RDA-A1-01M — metadata contains information to enable the user to get access to the data · COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl
The data are aggregated from public sources and are not sensitive human-subject data; no gatekeeper is named.
NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee · RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and · NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse
The paper does not mention any retention period or availability timing beyond the present availability.
NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines · NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy' · RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'
“The completed database is downloadable via the web interface in BED and VCF format.”
BED and VCF are open, community-standard file formats.
FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co · RDA-R1.3-02D — data is expressed in a machine-understandable community standard · RDA-I1-01D — data uses a knowledge representation expressed in a standardised format
Advisory · not in the published score
“we converted all CNV files into the standard Variant Call Format (VCF) [35] terminology for copy number events”
VCF is a community-standard data format, and the paper states it was used. [majority verdict 'yes' (3/5 passes agreed)]
RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential) · RDA-R1.3-01D — 'Data complies with a community standard' · RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'
“All downloaded databases were stored in the genome assembly GRCh38.”
The paper references the GRCh38 genome assembly and GENCODE version, which are identifiers for external resources. [majority verdict 'yes' (4/5 passes agreed)]
RDA-I3-01M — '(meta)data include references to other (meta)data' · RDA-I3-03M — 'metadata includes qualified references to other metadata' · FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'
No license for the data is stated anywhere in the text.
RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu · RDA-R1.1-02M — 'Metadata refers to a standard reuse licence' · RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'
No version token or date is given for the dataset itself.
DataCite Metadata Schema 4.6 — the 'Version' property · RDA-R1.2-01M — provenance information (which version was used is provenance) · NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'
“The source code of PanCNV-Explorer, including the database generation, the web server, the web front end and the CNA annotation modules, is available at https://gitlab.gwdg.de/MedBioinf/mtb/cnv-database.”— not found in the paper; verdict downgraded
The paper provides a GitLab URL for the source code, a machine-resolvable locator. [downgraded to 'partial' — no verifiable quote from the paper]
NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code' · FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear · FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)
“This work was supported by the Gemeinsamer Bundesauschuss (01NVF20006), the Volkswagen Foundation (11-76251-12-1/19), the Deutsche Krebshilfe (70114018), the Deutsche Forschungsgemeinschaft (KFO5002) and the Bundesministerium für Bildung und Forschung (BMBF) (01KD2437, 01KD2401B, 01KD2208A, 01KD2414A).”
The paper lists specific award numbers for each funder.
DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award · Crossref Funder Registry — canonical funder identifiers for funding metadata · RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco
Advisory · not in the published score
“To retrieve cancer-specific CNAs, we downloaded copy number segmentation files from The Cancer Genome Atlas (TCGA) from 33 cancer types using the TCGAbiolinks package”— not found in the paper; verdict downgraded
The paper names specific tools and databases (TCGAbiolinks, TCGA, etc.) used to produce the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]
RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa · FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati · W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance
No documentation object (README, codebook, or schema) is named as accompanying the data, and no variable-definition table is inside the article.
RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu · FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data' · NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t
Calibrated FAIR score — a parallel quality metric, independent of the DataRank citation score. See the full evaluation →
Gemeinsamer Bundesausschuss
Grant: 01NVF20006
Volkswagen Foundation
Grant: 11-76251-12-1/19
Deutsche Krebshilfe
Grant: 70114018
Deutsche Forschungsgemeinschaft
Grant: KFO5002
Bundesministerium für Bildung und Forschung
Grant: 01KD2437
Bundesministerium für Bildung und Forschung
Grant: 01KD2401B
Bundesministerium für Bildung und Forschung
Grant: 01KD2208A
Bundesministerium für Bildung und Forschung
Grant: 01KD2414A
Deutsche Forschungsgemeinschaft
Grant: unidentified
unidentified
Fields of Study
Keywords