Export ClinVarbitration decisions to Parquet
Source:R/export.R
rclinvarbitration_export_clinvarbitration_parquet.Rdschema = "compatibility" writes the seven columns in Centre for Population
Genomics ClinVarbitration's clinvar_decisions.tsv: contig, position,
reference, alternate, clinical_significance, gold_stars, and
allele_id.
Arguments
- con
A DuckDB DBI connection initialized with
rclinvarbitration_init().- path
New output
.parquetfile path.- release_id
Imported ClinVar release label to export.
- assembly
Genome assembly:
"GRCh38"or"GRCh37".- profile_id
Policy profile identifier, normally
"default".- submitter_exclusions
Additional submitter names to exclude from this export. Matching is case-insensitive and ignores surrounding whitespace. These exclusions are combined with any exclusions already stored for
profile_id; imported source submissions are not deleted.- schema
Output schema: the upstream-compatible seven-column relation or the canonical scalar evidence relation.
Details
schema = "tidy" writes the canonical scalar clinvar relation.
record_kind distinguishes variations, alleles, assembly locations, source
assertions, conditions, genes, observations, citations, text, attributes,
allele policy decisions, and disease-level disease_decision policy rows.
Every row has its own stable record_key; repeated source elements are rows
rather than lists or structs. release_id is a
required Parquet column, so a reopened export retains its source identity.
The compatibility source is the allele-level policy view joined through
clinvar_vcf. Both GRCh37 and GRCh38 are supported, including distinct X/Y
locations for one AlleleID. Primary NC_ placements take precedence when
present. An allele available only on alternate placements retains its exact
sequence accession and is not mislabeled as a primary-chromosome VCF record.
The file is schema-compatible with the upstream TSV/Hail decision resource, but is not claimed to be byte-for-byte equivalent: this package derives submissions and locations from VCV XML, whereas upstream uses ClinVar's tab-delimited submission and variant summaries. PM5 is deliberately not exported; Rduckhts/DuckHTS own downstream consequence and PM5 processing.