This document records where RClinVarbitration intentionally or unavoidably differs from:
658b9f241eb2d43aa11214b153b19c1e18a16337; andIt also reports observed differential results. A difference is not called a bug unless source release, code version, policy, grouping, coordinate rules, and output key are aligned. RClinVarbitration outputs are alternative derived observations and are not official ClinVar classifications.
| Area | Upstream 2.2.11 | RClinVarbitration | Status and consequence |
|---|---|---|---|
| Normal input | NCBI submission_summary plus variant_summary |
the same flat reports; optional complete VCV XML/XML.GZ | Flat import is the compact default. XML feeds the same table with richer attributable evidence but different entity semantics. |
| Runtime | Python, Pandas, Hail, Spark, Nextflow, bcftools | R, DuckDB SQL, package-owned C extension | Intentional implementation change; semantic parity is tested separately from performance. |
| Decision scope | allele/VariationID | disease plus allele views | Intentional extension. Disease grouping can produce multiple decisions for one allele. |
| SCV versions | flat rows; no XML assertion identity step | highest SCV version per assertion identity and decision group | Intentional XML adaptation; can change counts where versioned rows are exposed. |
| Qualified Illumina exclusion | declared, but the pinned Python inner-loop continue does not remove the row |
removes benign evidence from normalized submitter illumina laboratory services; illumina |
Intentional correction to documented policy, not bug-for-bug compatibility. |
| Submitter exclusion matching | lower-cased flat submitter names and CLI/config values | trimmed, case-insensitive names; profile-wide or classification-qualified rules | Intentional extension. Imported evidence remains present. |
| Disease identity | absent from the decision key | canonical identifier, trait-set, name, then package entity fallback | Intentional extension with known grouping heuristics; see open limitations. |
| Outputs | TSV, Hail Table, VCF, PM5 relation | one scalar DuckDB table, SQL views, and Parquet exports | Intentional. Hail, VCF file rendering, VEP, and PM5 generation are out of scope. |
| Alternate-accession coordinates | builds CHROM from the flat report’s Chromosome field |
uses ChromosomeAccession for non-primary placements, matching the official ClinVar VCF |
Intentional correction. Upstream can label an alternate-locus position as if it were on a primary chromosome. |
| Star and 60/20 rules | pinned Python implementation | equivalent SQL rules | Expected parity when input rows, order, grouping, and exclusions are identical. |
| Strong-review choice | first strong row encountered | min_by classification in retained source order |
Expected parity. It does not impose practice-guideline priority over an earlier expert-panel row. |
rclinvarbitration_import_flat() is the ordinary compact release
import. rclinvarbitration_reproduce_clinvarbitration_parquet() remains
a direct seven-column oracle for the upstream output contract. VCV XML
is optional enrichment, not a second public storage model.
NCBI ClinVar applies its own aggregate germline classification and review-status rules. RClinVarbitration instead applies the pinned 2016 evidence window, classification bins, exclusions, strong-review source-order rule, 60/20 rule, and reduced star calculation. Therefore:
clinvar_variants.aggregate_classification and
clinvar_rcv_assertions.classification are source values;clinvar_policy_decisions.policy_classification is a package-derived
value;The importer retains VCV records, alleles, assembly locations, genes, RCVs, SCVs, conditions and names, cross-references, observations, citations, attributes, and selected attributable text. It deliberately does not persist:
clinvar_text additionally projects condition names and string-valued
attributes for discovery. Those rows are attributable normalized
projections, not original XML serialization.
Each location row keeps the declared assembly, assembly accession,
source chromosome label, exact sequence accession, one-based VCF
position, reference, and alternate allele. GRCh37 and GRCh38 rows
coexist; the importer does not infer one assembly by lifting the other.
clinvar_vcf uses conventional primary-contig names (1/chr1,
M/chrM) only for primary NC_ accessions or flat-report rows that
do not provide a sequence accession. NT_, NW_, and other alternate
placements retain their accession as contig. X and Y placements for
the same AlleleID remain separate rows, including pseudoautosomal
placements. Consumers must choose the desired assembly and placement
rather than collapsing on AlleleID alone.
The seven-column decision export prefers primary NC_ placements when
they exist. An allele available only on an alternate placement keeps
that accession. The canonical location rows and clinvar_vcf retain
every source placement regardless of this export selection.
The March 2026 flat report exposed 40 upstream decision rows for which
this distinction changes CHROM. The pinned upstream implementation
emitted chr17, chr19, chr22, or chr9; NCBI’s 9 March 2026 GRCh38
VCF emitted the same 40 AlleleIDs and identical POS/REF/ALT values on
NT_187661.1, NT_187693.1, NT_187633.1, or NW_009646201.1.
clinvar_vcf and the seven-column compatibility export follow the
official VCF accessions. The direct upstream reproducer remains
available when deliberate bug-for-bug comparison is required. The
coordinate
audit
records the official VCF URL, digest, release dates, and comparison
counts.
No public upstream rationale for discarding ChromosomeAccession was
found. The choice entered in the starter
commit.
PR 3
later restricted the derived chromosome names after a ChrUn row broke
the parser, and PR
4
prevented X/Y records from overwriting each other. Neither change reads
or discusses ChromosomeAccession; no public issue, pull request,
commit message, or review comment documents an alternate-locus
coordinate policy.
Where an XML element does not provide a stable public identifier, the parser constructs a scoped entity ID from its VCV/assertion context and source ordinal. These keys support joins inside one imported release. They must not be presented as NCBI accessions or assumed stable across a changed parser policy. Public VCV, RCV, SCV, Variation, Allele, Gene, and external condition IDs remain separate typed columns.
Disease grouping prefers an explicit condition database/identifier. Canonical cross-reference preference is MedGen, MONDO, OMIM, Orphanet, MeSH, UMLS, then OMIM phenotypic series. Fallbacks use ClinVar trait-set ID, normalized name, and finally the package condition entity.
This can split synonymous conditions whose cross-references differ or collapse same-named conditions that lack identifiers. Disease decisions therefore need identifier-level review before clinical use.
clinvar_hpo_terms, clinvar_semantic_documents, and
clinvar_literature_links normalize discovery surfaces. Their presence
does not assert that an HPO term describes the proband under review,
that a text row supports the submitted classification, or that a cited
publication contains admissible evidence. Context IDs and source
spans/rows must be retained through any downstream claim and review
process.
The strongest current differential uses the NCBI March 2026 monthly VCV
XML and the March 2026 archived submission_summary/variant_summary
files. Both RClinVarbitration paths used policy
cpg-clinvarbitration-2.2.11, GRCh38, and the seven-column key
(contig, position, reference, alternate, allele_id).
Source SHA-256 digests:
| Source | SHA-256 |
|---|---|
ClinVarVCVRelease_2026-03.xml.gz |
8c369922c38958bdba0c99225d2db794cd02995930b98cfce7a4754faf65f7c8 |
submission_summary_2026-03.txt.gz |
dfc875bc831292b857d8d0a85eb57157452e12f04fbc3591addbf59208de727f |
variant_summary_2026-03.txt.gz |
0d6c0c8760529befdfc1fbfcfa90cfb0aa11bfc5fe72176ac9bc9820884f710e |
Observed results:
| Metric | XML path | Flat reproduction |
|---|---|---|
| Rows | 4,125,736 | 4,125,389 |
| Shared keys | 4,125,382 | 4,125,382 |
| Path-only keys | 354 | 7 |
| Shared-key classification or star disagreements | 16 | 16 |
| Exact agreement among shared keys | 99.9996% | 99.9996% |
All 377 differences are now source-row classified:
| Source-backed class | Shared disagreements | XML-only keys | Flat-only keys |
|---|---|---|---|
Flat ClinicalSignificance is -; XML carries the current germline classification |
15 | 354 | 0 |
| Flat repeats one SCV/version with divergent rows; XML has one current assertion identity | 1 | 0 | 0 |
XML uses unbinned legacy Affects; flat uses Pathogenic |
0 | 0 | 4 |
XML uses unbinned no known pathogenicity; flat uses Benign |
0 | 0 | 1 |
| Flat location is a nested compound/haplotype allele; the XML compatibility export is top-level-allele only | 0 | 0 | 2 |
| Total | 16 | 354 | 7 |
The duplicate case is VariationID 548128: the flat source repeats
SCV000783104.86 as Pathogenic, Uncertain significance, and
not provided, whereas the XML has one current SCV assertion classified
Uncertain significance. The flat algorithm therefore reaches P/LP
while XML assertion identity produces VUS. The two nested-allele cases
are CYP2C19 compound records 633845 and 633881. Their child locations
and AlleleIDs are retained in the XML relations, but are deliberately
absent from the top-level allele compatibility export.
These are not unexplained arithmetic failures. They demonstrate that NCBI’s same-month XML and flat products are not row-equivalent policy inputs. The 377-key receipt contains the variant-summary physical line, flat-submission physical lines, XML record/entity ordinals, source values, outputs, and assigned class for every key. Detailed flat submission, XML SCV, and XML entity receipts are pinned by the differential manifest.
The upstream TSV stage at commit
658b9f241eb2d43aa11214b153b19c1e18a16337 was rerun against the exact
March archive files identified above. The upstream decision module was
loaded unmodified; inert Hail/loguru stand-ins allowed execution through
write_dicts_as_tsv() and stopped before Hail, VCF, and PM5
post-processing. The package’s direct flat reproducer was then compared
on the complete seven-column key and values.
| Metric | RClinVarbitration flat path | Pinned upstream execution |
|---|---|---|
| Rows | 4,125,389 | 4,125,389 |
| Shared keys | 4,125,389 | 4,125,389 |
| Path-only keys | 0 | 0 |
| Classification or star disagreements | 0 | 0 |
| Exact agreement | 100% | 100% |
The upstream TSV SHA-256 is
35ffe949d488b6d0a79bba3a3295e810f4b79e193ffdb5b4da15faedfe1970ee. The
oracle manifest
records both input digests, producing commit, algorithm digest,
configuration, runtime, output byte counts and digests, and comparison
result. The executable
tools/run_pinned_upstream_flat_oracle.py
repeats the TSV-only run. This is the algorithm-conformance result; the
XML comparison above is a source-projection differential.
The published upstream reference was Zenodo
19196770, file
clinvarbitration_26-03.release.tar.gz (SHA-256
f7a2c7695d73b5c2d88350a38faa5122f1499a2dae27b5ed28bcf4a6a8b7c69b). The
direct flat reproduction produced:
| Metric | RClinVarbitration flat path | Upstream TSV |
|---|---|---|
| Rows | 4,125,389 | 4,135,355 |
| Shared keys | 4,118,490 | 4,118,490 |
| Candidate-only keys | 6,899 | — |
| Reference-only keys | — | 16,865 |
| Shared-key disagreements | 1,863 | 1,863 |
| Classification mismatches | 1,295 | 1,295 |
| Star mismatches | 1,149 | 1,149 |
| Both classification and star mismatch | 581 | 581 |
| Exact agreement among shared keys | 99.9548% | 99.9548% |
None of the candidate-only keys had an allele ID or locus represented in the reference-only set, and vice versa. This indicates a source-snapshot difference rather than coordinate formatting. The Zenodo artifact does not bundle the exact NCBI inputs or their content digests, and its March 24 publication could have used mutable current flat files rather than the March 5 archive. It also does not pin the producing commit in its artifact metadata.
The exact-input rerun above now resolves the concern raised by the 16,865 reference-only keys: with identical archived inputs, the package and pinned upstream code have zero key or value differences. The Zenodo result remains a release-snapshot comparison and must not be used as an algorithm-conformance failure count.
A second complete pass counted source elements in the same March XML and
compared them with import counts. The counter is
tools/audit_xml_structure_coverage.c;
the element
receipt and
coverage
manifest pin the
source and counter digests.
| XML structure | Source elements | Typed representation |
|---|---|---|
VariationArchive |
4,478,872 | 4,478,872 clinvar_variants rows (100%) |
ClinicalAssertion |
6,813,140 | 6,813,140 clinvar_scv_assertions rows (100%) |
ObservedIn |
6,869,122 | 6,869,122 clinvar_observations rows (100%) |
Sample |
6,869,122 | no sample relation; selected fields flattened into observations |
Method / MethodType |
6,877,237 / 6,877,237 | no method relation; one scalar method_type per observation row |
ObservedData / child Attribute |
6,494,488 / 6,494,488 | attributes retained in generic clinvar_attributes; no observed-data relation |
MolecularConsequence |
24,908,913 | no dedicated typed relation |
FunctionalConsequence |
564,026 | no dedicated typed relation |
The flattened sample fields cover all 6,869,122 Origin, Species, and
AffectedStatus children and 123,764 NumberTested children. The
source also contains sample details with no typed column: 73,583 Age,
135,617 Sex, 48,536 Ethnicity, 26,335 GeographicOrigin, 547,099
Tissue, and 47,233 FamilyData elements. Methods additionally contain
182,688 TypePlatform and 52,989 MethodAttribute elements without
typed columns. Since there are 8,115 more methods than observations, the
scalar method projection also cannot represent method multiplicity. Of
the molecular-consequence structures, 433 nested XRefs survive as allele
XRefs, but their consequence-node association is not retained.
This quantifies the selected-projection boundary; it does not rebrand generic attribute retention as complete typed coverage.
Tinytests exercise classification bins, unknown exclusion, qualified and
profile exclusions, SCV deduplication, old-versus-modern evidence,
strong-review source order, 60/20 decisions, VUS handling, stars,
disease grouping, allele-level output, and XML fixture import.
Release-differential fixtures now cover stable public disease keys
despite changed labels/entities, highest SCV version replacement,
disappearance of a withdrawn assertion, compound child alleles, and
GRCh38 MT to chrM export.
Projection checks use real source contexts. VCV000158424/SCV000192942
pins HPO HP:0002282 to assertion condition 2 from the complete
2026-07-02 release. VCV000091629/SCV000827729 pins PMID 21735045 to
assertion 1599586 and its PubMed URL. The fixture
manifest records those
expectations. These tests validate linkage only; neither a term nor a
citation is treated as clinically relevant merely because it exists.
Update this document whenever parser coverage, arbitration semantics, grouping, or a differential result changes. Render with:
Rscript -e 'rmarkdown::render("docs/ERRATA.Rmd", output_format = "github_document", quiet = TRUE)'