The duckvep-coding haplotype contract

Status: signed; issue #2 is closed for this coding-only correctness contract. bcftools csq is the executable authority for whole-haplotype consequences, and NMD uses ejc50.

The separately measured 2× throughput criterion is met by the fused native reader on the pinned one-core HG002 workload; details and receipts are in the qualification report and #34. Bounded current profiles under #11 and #50 are described in section 3; their broader issue scope remains open. #28 tracks phased structural composition.

Scope: coding-only whole-haplotype prediction, not general haplotype prediction. Maintainer approval is required before classifiers change. Keep independent-event and haplotype classifiers separate, sharing projection/edit/translation facts only where their semantics agree. bcftools csq is the pinned test oracle, not a runtime dependency; DuckVEP extends its existing native path.

1. Authority per output

bcftools csq (pinned) is the authority for whole-haplotype consequences. Pin the bcftools and htslib commits, the Ensembl 116 GRCh38 GFF3 (sha256 08e881d96ab6385a2c31f063a018be4b2c36860b323f2724be07022deeef21ce), the reference FASTA and the -p phasing mode, and record them in an oracle receipt. csq classifies the combined haplotype directly: same-codon compounds, frame opening and restoration, stop gain and loss, and start loss, with phase awareness. That is exactly the layer VEP 116 and Haplosaurus do not define.

2. Supported domain

Support strict, complete phased calls of any ploidy (a haploid call has one lane; an unphased slot makes the call incomplete); literal A/C/G/T SNPs/MNVs/indels of any length, bounded by the workspace limits; nonoverlapping edits contained within coding exons; CDSs in any supported genetic code with no curated RNA/peptide edits or recoding. A complete CDS has a start codon (ATG in the standard code, any start codon of another code), a terminal stop and no internal stop. A CDS whose start is not annotated (cds_start_NF) is read without a start test, including one that begins inside a codon and is stored padded with N as Ensembl translates it; one whose end is not annotated (cds_end_NF) has no stop and may end in a partial codon. Support both strands, multiple coding exons and their internal phases, reference lanes, multiple transcripts and small multisample fixtures. Splicing is fixed to the selected model. Each transcript/sample must have one unambiguous heterozygous phase domain. An unphased heterozygous call that is the sample’s only heterozygous or missing call on the transcript needs no phase (the haplotypes are the one with the allele and the one without) and is read in slot order. Phased calls without PS use the documented implicit domain, not inferred cross-block phase.

The 36-field result retains cds and protein as raw compatibility replay and exposes conditional protein contrasts separately in prediction_reference_protein and prediction_protein; nominal_length_diff is the final field. Typed recoding preserves an annotation only at an unchanged reference codon; a changed codon uses ordinary translation within the conditional pair. This does not establish biological recoding competence: raw replay remains unchanged, NMD is unknown, and no unconditional consequence is emitted. Legacy untyped recoding remains unsupported. Every result also retains carriers, contributors, statuses, haplotype_consequences, haplotype_impact, nmd_prediction, prediction_status and reason. Keys include model/transcript identity and sample/phase/lane. Preserve source-record identity, ALT ordinal, original operands, genotype evidence and normalized edit identity through every block; sequence deduplication must not erase contributors.

The reduced SO set reports reached HIGH start/stop/frame effects; otherwise it reports one whole-protein category below. This deliberately omits lower-severity subeffects. Start loss suppresses other biological predictions; otherwise frame and stop terms may coexist. Among LOW classes, terminal-retained precedes synonymous. IMPACT is the maximum severity of emitted terms. Frame SO is normalized-edit-path-sensitive: pin/test that convention rather than claiming representation-independent labels.

Situation Exact policy output Oracle
Same-codon cis substitutions; corresponding trans control Separate trans lanes. With no start/stop/frame effect: identical peptide → synonymous_variant/LOW; changed peptide → missense_variant/MODERATE. Classify the combined sequence, not either SNP. Pinned csq via the csq→SO map; independent R goldens for edge cases.
Open/restored frames Premature first-stop codon intersects a displaced-frame interval, or frame remains displaced at CDS exhaustion → frameshift_variant/HIGH. Restored before termination → no frameshift SO; identical peptide → synonymous/LOW, otherwise protein_altering_variant/MODERATE. Pure frame-preserving insertion/deletion uses inframe_insertion/inframe_deletion/MODERATE. Pinned csq via the csq→SO map; independent R goldens; track cumulative frame offsets and first stop.
Stop created/removed New first stop before the homologous reference terminator → stop_gained/HIGH. Reference termination abolished without an earlier gained stop → stop_lost/HIGH. Downstream edits remain contributors, not expressed effects. Pinned csq via the csq→SO map; independent R goldens; include compensating edits both before and after the first stop.
Start/terminal codons Canonical start abolished → start_lost/HIGH; initiation and NMD unknown. Changed terminal codon remaining a stop, without a higher effect → stop_retained_variant/LOW. When the stop is lost the protein is read through into the transcript’s stored 3′ flank, to the next stop or the end of the flank; nothing is invented beyond the stored sequence. Pinned csq via the csq→SO map; independent R goldens; terminal synonym, loss and restoration controls.

Reference-only lanes have an empty SO set and NULL IMPACT. Other mixed frame-preserving replacements use protein_altering_variant; unchanged peptide uses synonymous.

NMD policy (ejc50): for a newly premature stop, return trigger when J−S>50, otherwise escape. S is the stop’s final nucleotide; J the penultimate exon’s final nucleotide, both 1-based in edited spliced-transcript coordinates. Intronless transcripts escape. Known termination without a new premature stop → not_applicable; incomplete phase/exon topology, lost initiation or unavailable termination → unknown. Return stop/junction coordinates and joint contributor attribution. Test 49/50/51 bases, intronless/last-exon cases and indel-shifted junctions independently. The prediction itself is an EJC-distance heuristic, not clinical NMD truth, VEP’s allele-position plugin, or the NMD_transcript_variant biotype term. Two exceptions to the junction rule are reported beside it in nmd_exceptions without changing it: start_proximal, the stop codon lies within the first 100 coding bases (reinitiation; the distance Ensembl’s NMD plugin uses), and long_exon, the stop lies in an exon of the edited transcript longer than 407 bases (Lindeboom, Supek and Lehner, 2016). Whether an exception is an escape is the consumer’s decision.

Failure contract: missing alleles, unphased heterozygosity or unresolved cross-PS phase → incomplete_input; contradictory edits → edit_conflict; other overlaps/ambiguous same-gap insertions → unsupported_overlap. Preserve reference-mismatch/projection reasons; excluded contexts → unsupported_context. Whole SO/IMPACT are NULL and NMD unknown on these paths; any compatibility replay stays explicitly conditional. Retain every contributor, including omitted, shadowed and post-stop sources. Malformed identities/budget overflow fail explicitly rather than truncate.

2a. Parts of the domain outside the csq comparison

csq is the authority on its comparable domain (diploid calls, the standard genetic code). Four parts of the supported domain lie outside it. The classifier is the same there; these are the reasons it applies and the checks that stand in for csq:

Part of the domain Why the classifier applies Check
Calls that are not diploid A lane is one chromosome copy; its prediction depends on its edits only. Each lane of haploid, triploid and tetraploid calls equals the diploid lane carrying the same edits (duckvep_haplotype_eligibility.test; the property suite for ploidies 1 to 4).
Alleles over 50 bases The classifier reads the edited CDS at any edit length within the workspace limits; 50 bases is not a semantic boundary. A 51-base insertion with a hand-derived protein. The independent-event path is compared with executable VEP on random alleles up to 100 bases. There is no csq comparison at these lengths.
Genetic codes other than the standard one Translation, start and stop tests follow the transcript’s code. The first residue is the initiator, so a change between two start codons is not a peptide change. Four edits read under NCBI table 2 and under the standard code on the same CDS, with outcomes derived by hand from the published tables (TGA Trp, AGA stop, ATA Met and start).
Transcripts with an unannotated CDS start or end With no annotated end nothing can be lost or retained: a stop is new, a frame still displaced where the annotation ends is a frameshift, an edit confined to the trailing partial codon is an incomplete_terminal_codon_variant, and NMD is unknown for a new stop. With no annotated start there is no start test, and an unchanged peptide with an edit in the unknown first codon is a coding_sequence_variant. Hand-derived cases (duckvep_haplotype_incomplete.test). On HG002, single-edit haplotypes are compared with the per-variant annotation, which is exact against VEP: 91,212 of 91,708 are identical and the rest fall in named policy differences (for example the whole-haplotype view adds stop_gained to a frameshift).

Stop-loss extension and NMD-exception evidence. A lost terminal stop extends translation through the transcript’s stored 3′ flank to the next stop or the end of that sequence. Stop-loss protein lengths match per-variant protein HGVS (itself exact against VEP) for all 376 HG002 single-edit paths with a numbered new stop (extTer N or fsTer N). A separate 20 paths reach the stored-flank boundary and match VEP’s open-ended Ter? notation. The HG002 domain oracle receipt records the input, reference, GFF3, bcftools/htslib pins and oracle-output digest. Hand-derived cases check the NMD exceptions, including the 100-base boundary (duckvep_haplotype_nmd_exceptions.test).

Single unphased heterozygous site. On HG002 the rule resolves 782 of 1,110 unphased-heterozygous paths. The SQL fixture checks the phased twin and unresolved neighbours with two unphased sites or an unphased site beside phased evidence (duckvep_haplotype_eligibility.test).

The open domain boundary is tracked in #50. Two or more unphased heterozygous sites on one transcript, an unphased site beside phased or missing evidence, unresolved cross-PS phase, overlapping edits or ambiguous same-gap insertions, general curated RNA/peptide edits, legacy untyped recoding, a complete CDS with an internal stop or a missing start/stop without its matching annotation flag, and non-strict phase policies remain outside this strict decoded-call prediction contract. The separate duckvep_haplotype_arrangements relation enumerates bounded diploid alternatives; it preserves observed GT/PS separately from hypothetical arrangement identity and includes an explicit reference lane, but does not make those alternatives observed cis calls for this contract. Raw source-record mode is a separate conditional replay interface, not an input to this contract. Unsupported paths retain explicit statuses, reasons, and contributor identities. The single-site phase rule, supported ploidies and genetic codes, correctly flagged cds_start_NF/cds_end_NF transcripts, and translation past a variant-lost stop through stored 3′ sequence are supported above.

3. Bounded extensions and exclusions

Regression gates cover supported-domain validation, original-operand ownership, sequence replay, and the measured resource limits.

4. Scale contract

The scale profile is a 5M-physical-variant, single-sample GRCh38/MANE-selected-model job, not 5M expanded call rows. Account for every input as emitted, explicitly outside coding scope, or unavailable. The 2,621,440-row fixture contains only 2,560 physical events; its favorable sharing is not genome-scale evidence.

Enforce 16 GiB per-job process memory, including R, and 4 GiB aggregate DuckVEP-native memory, including resident models, indexes and all workers—not merely workspace_limit. Use native admission/accounting, bounded windows, spillable DuckDB staging, a temporary-disk quota and an external job limit. Table-backed/chunked output is the scale interface; eager R collection is not the throughput benchmark. Capacity failure, cleanup and connection reuse are acceptance checks.

Acceptance gate: at least 2× faster than pinned bcftools csq on identical input, transcript model and allocated cores, under the same caps. Compare medians of three fresh processes on one host. The recorded 2026-09-28 baseline was 22.2 s for the phased HG002 GRCh38 v4.2.1 genome (4,023,088 records; malformed MHC records dropped), setting an 11.1 s target for that measurement round. Acceptance uses the same-round csq median because timings vary by host and load. csq skips records outside transcripts; DuckVEP’s fused reader scans the same full input and decodes calls only for coding records.

Measure separately: (A) preordered input through prediction and complete output materialization; (B) identical unsorted input including decoding, discovery, staging and sorting. Include unavoidable internal sorts in A. Report cold model-load and warm execution separately; all phases obey memory caps. Record physical sources/ALTs, projections, calls, carrier states, unique paths, translated bases, output rows/bytes, peak active window, model/native/DuckDB/RSS peaks, spill bytes and full-output checksums. Use real phased input plus dense/long-transcript and low-sharing stress controls. Run three fresh processes per mode: median-time gates, caps on every run, complete failures reported.

Throughput qualification (2026-09-30)

On the pinned full-HG002 workload with one core, the Ensembl 116 model and the qualification caps, fresh-process medians were 17.25 s for csq and 8.21 s for the fused-reader CLI path (2.10×, DuckVEP model load included); the R path took 8.60 s (2.01×). Both meet the 2× gate. The qualification report and receipts record the model, caps, pins and run conditions.

Mode-B measurement (2026-09-29)

The DuckDB CSV-reader path measured 12.58 s against csq’s 19.66 s (1.56×), below the 2× gate. The qualification report records its method and conditions.

5. Regression and sign-off gates

The contract is checked by pinned csq comparisons and independent base-R goldens for same-codon, frame/restoration, start/stop, and NMD outcomes; status and provenance cases cover missing or unresolved phase, overlaps/conflicts, reference mismatch, excluded contexts, and contributor conservation. Native properties, SQL tests, installed-R tests, and independent-event regressions protect the public implementation. The throughput and scale gates use the measured profiles and conditions in section 4 and the linked receipts; complete outputs, per-run resource caps, explicit capacity failures, cleanup, and connection reuse are required.

Issue #2 is closed for the signed duckvep-coding correctness contract: supported outputs have zero unexplained missing, extra, or discordant results, while unsupported cases retain their statuses and contributor identities. This sign-off covers the pinned csq authority and mapping, EJC50 heuristic, stated domain, and resource limits. It does not claim general haplotype, compound-HGVS, or structural-composition compatibility.