gnomad_qc.v5.annotations.generate_variant_qc_annotations

Script to generate annotations for variant QC on gnomAD v5.

usage: gnomad_qc.v5.annotations.generate_variant_qc_annotations.py
       [-h] [--environment {rwb,batch}] [--app-name APP_NAME]
       [--driver-cores DRIVER_CORES] [--driver-memory DRIVER_MEMORY]
       [--worker-cores WORKER_CORES] [--worker-memory WORKER_MEMORY]
       [--overwrite] [--test] [--test-n-partitions [TEST_N_PARTITIONS]]
       [--generate-trio-stats] [--union-trio-stats] [--validate-trio-stats]
       [--chrom CHROM] [--generate-sibling-stats]
       [--vcf-ht-checkpoint-path-override VCF_HT_CHECKPOINT_PATH_OVERRIDE]
       [--repartition-for-join]
       [--info-ht-path-override INFO_HT_PATH_OVERRIDE] [--generate-ac-info-ht]
       [--union-ac-info-hts]
       [--union-input-ac-info-ht-paths UNION_INPUT_AC_INFO_HT_PATHS [UNION_INPUT_AC_INFO_HT_PATHS ...]]
       [--union-contigs UNION_CONTIGS [UNION_CONTIGS ...]]
       [--union-allele-strata UNION_ALLELE_STRATA [UNION_ALLELE_STRATA ...]]
       [--union-n-partitions UNION_N_PARTITIONS] [--create-sites-vcf-ht]
       [--create-final-info-ht] [--use-tmp-info-paths]
       [--ac-info-ht-checkpoint-path-override AC_INFO_HT_CHECKPOINT_PATH_OVERRIDE]
       [--max-alleles MAX_ALLELES] [--min-alleles MIN_ALLELES]
       [--lowqual-indel-phred-het-prior LOWQUAL_INDEL_PHRED_HET_PRIOR]
       [--export-info-vcf] [--use-local-allele-agg]
       [--validate-local-allele-agg] [--explode-partitions]
       [--chunk-start CHUNK_START] [--chunk-stop CHUNK_STOP]
       [--read-subintervals-per-chunk READ_SUBINTERVALS_PER_CHUNK]
       [--read-subintervals-scale READ_SUBINTERVALS_SCALE] [--scout-alleles]
       [--scout-n-partitions SCOUT_N_PARTITIONS]
       [--scout-rows-per-partition SCOUT_ROWS_PER_PARTITION]
       [--scout-byte-weight] [--scout-byte-weight-cap SCOUT_BYTE_WEIGHT_CAP]
       [--scout-limit-intervals SCOUT_LIMIT_INTERVALS]
       [--create-variant-qc-annotation-ht] [--impute-features]
       [--n-partitions N_PARTITIONS] [--export-true-positive-vcfs]
       [--transmitted-singletons] [--sibling-singletons]

Named Arguments

--environment

Possible choices: rwb, batch

Environment where script will run.

Default: “rwb”

--app-name

Job name for batch/QoB backend.

--driver-cores

Number of cores. Applies to Batch environment only. Hail default is 1 if unspecified.

--driver-memory

Memory for driver node. Applies to Batch environment only. Hail default is ‘standard’ if unspecified.

--worker-cores

Number of cores. Applies to Batch environment only. Hail default is 1 if unspecified.

--worker-memory

Memory for worker nodes. Applies to Batch environment only. Hail default is ‘standard’ if unspecified.

--overwrite

Overwrite output files.

Default: False

--test

Write to test path.

Default: False

--test-n-partitions

Use only n partitions of the VDS as input for testing purposes (default: 2).

--generate-trio-stats

Calculate trio stats for a single chromosome (–chrom). Run once per chromosome, then combine with –union-trio-stats.

Default: False

--union-trio-stats

Union the per-chromosome trio stats HTs into the combined trio stats HT.

Default: False

--validate-trio-stats

Validate the combined trio stats HT (union integrity + value sanity).

Default: False

--chrom

Single chromosome/contig (e.g. ‘chr20’). Required for –generate-trio-stats (trio stats are computed one chromosome at a time). With –generate-ac-info-ht, restricts the VDS read to the contig: without –scout-alleles the contig span is subdivided into read intervals via –read-subintervals-scale or –read-subintervals-per-chunk; with –scout-alleles the scout pass is restricted to the contig. Mutually exclusive with chunk processing and –test-n-partitions for the AC-info step.

--generate-sibling-stats

Calculates sibling stats.

Default: False

--vcf-ht-checkpoint-path-override

Optional override path for the reformatted sites VCF HT. By default the path is derived (sites_vcf[_test].ht in the durable annotations bucket, or the 30-day temp bucket with –use-tmp-info-paths or on test runs); written by –create-sites-vcf-ht and read by –create-final-info-ht.

--repartition-for-join

With –create-final-info-ht, co-partition the VCF and AC info HTs onto identical interval boundaries before the join so it requires no shuffle (removes the join shuffle stage). Best for full-genome runs.

Default: False

--info-ht-path-override

Optional override path for the info HT output. If set, this path is used instead of the default resource path.

--ac-info-ht-checkpoint-path-override

Optional override path for the AC info HT checkpoint. By default the path is derived automatically from the run’s parameters (test/–test-n-partitions, –chunk-start/–chunk-stop, –min-alleles/–max-alleles, union), landing in the durable annotations bucket, or in the 30-day temp bucket with –use-tmp-info-paths or on test runs; the –union-ac-info-hts output and the –create-final-info-ht input share the derived ‘union’ path.

--max-alleles

Optional maximum number of alleles (including the reference) allowed at a locus when creating the info HT. Variants at loci with more than this many alleles are filtered out. Default is None (no filtering).

--min-alleles

Optional minimum number of alleles (including the reference) allowed at a locus when creating the info HT. Variants at loci with fewer than this many alleles are filtered out. Default is None (no filtering).

--lowqual-indel-phred-het-prior

Phred-scaled prior for a het genotype at a site with a low quality indel. Default is 40. We use 1/10k bases (phred=40) to be more consistent with the filtering used by Broad’s Data Sciences Platform for VQSR.

Default: 40

--export-info-vcf

Export info ht as VCF.

Default: False

--use-local-allele-agg

Compute the AC aggregation and AS_pab_max in local-allele space, so per-genotype cost scales with each genotype’s local alleles instead of every global alt. Output is identical to the original aggregation. Recommended for high-allele strata (e.g. –min-alleles 10 and above); at low-allele loci the original default is slightly cheaper. (Formerly –run-opt2.)

Default: False

--validate-local-allele-agg

Before writing, diff the –use-local-allele-agg aggregation (AC arrays and AS_pab_max) against the original on the loaded VDS and abort if they do not match. Use with a small input (e.g. –scout-alleles or –test-n-partitions).

Default: False

Split info HT workflow

Create the info HT in independent steps: per-stratum AC info HTs (–generate-ac-info-ht, one run per –min-alleles/–max-alleles range), union them (–union-ac-info-hts), reformat the sites-only VCF once (–create-sites-vcf-ht), and join to write the final info HT (–create-final-info-ht). The VCF is never interval-filtered, so split (min-repped) VCF loci can never be dropped at interval boundaries.

--generate-ac-info-ht

Run only the AC info aggregation on the VDS (no VCF import or join). Combine with –min-alleles/–max-alleles, –use-local-allele-agg, and scout/chunk processing. The checkpoint path is derived automatically from the run’s partition/allele parameters (see –ac-info-ht-checkpoint-path-override to override).

Default: False

--union-ac-info-hts

Union per-allele-count-stratum AC info HTs (from separate –generate-ac-info-ht runs) and write the result to –ac-info-ht-checkpoint-path-override. Fails if the input strata overlap (both allele bounds are inclusive, so e.g. max 9 pairs with min 10). Mutually exclusive with –generate-ac-info-ht.

Default: False

--union-input-ac-info-ht-paths

Space-separated paths to the per-stratum AC info HTs to union with –union-ac-info-hts. At least two paths are required. Alternatively derive the list with –union-contigs/–union-allele-strata.

--union-contigs

With –union-ac-info-hts, derive the input paths instead of listing them: one per contig x allele stratum, using the same path derivation as –generate-ac-info-ht (so run flags like –test must match the generate runs). Pass contig names (e.g. chr1 chr2 chr3) or ‘autosomes’ for chr1-chr22. Requires –union-allele-strata; mutually exclusive with –union-input-ac-info-ht-paths. Derived inputs are existence-checked up front, so a forgotten stratum run fails before any compute.

--union-allele-strata

Allele strata as MIN:MAX specs (inclusive, either side optional) matching the –min-alleles/–max-alleles of the generate runs, e.g. ‘:9’ ‘10:100’ ‘101:’. Used with –union-contigs.

--union-n-partitions

Optional number of partitions for the unioned AC info HT. Default is None (keep the union’s partitioning).

--create-sites-vcf-ht

Import and reformat the full AoU annotated sites-only VCF into a HT (adds AS_lowqual) and write it to the derived sites VCF HT path (see –vcf-ht-checkpoint-path-override to override). No VDS is loaded and no interval filtering is applied.

Default: False

--create-final-info-ht

Join the AC info HT at –ac-info-ht-checkpoint-path-override (e.g. the –union-ac-info-hts output) with the reformatted VCF HT to write the final info HT to –info-ht-path-override (or the default resource path). Reads the reformatted VCF HT at the derived path (or –vcf-ht-checkpoint-path-override override), so run –create-sites-vcf-ht first. Combine with –repartition-for-join for a shuffle-free full-genome join.

Default: False

--use-tmp-info-paths

Derive AC info HT and sites VCF HT checkpoint paths under the 30-day temp bucket instead of the default durable annotations bucket. Use for full-VDS experiments that should self-expire; –test/–test-n-partitions runs always use the temp bucket, without needing this flag. This flag exists (rather than reusing –test) because –test also switches the input to the 10-sample test VDS. Every step of a workflow (–generate-ac-info-ht, –union-ac-info-hts, –create-sites-vcf-ht, –create-final-info-ht) must agree on this flag for the derived paths to line up.

Default: False

Chunk processing parameters

--explode-partitions

Derive locus sub-intervals for chunk processing.

Default: False

--chunk-start

Start partition index for chunk processing.

--chunk-stop

Stop partition index for chunk processing (exclusive).

--read-subintervals-per-chunk

Number of locus sub-intervals to subdivide each chunk into, per contig. Default is None (1 sub-interval per contig, legacy behavior). In contig mode (–chrom without –explode-partitions) this count is spread across the contig’s VDS partition spans with a floor of one interval per span, so the realized total is at least the number of spans regardless of this value. Mutually exclusive with –read-subintervals-scale.

--read-subintervals-scale

Derive the sub-interval count automatically from the VDS partition count instead of stating a total. In chunk mode (–chunk-start/–chunk-stop) the total is this multiplier times the chunk’s partition count (e.g. 15 over a 3-partition chunk gives ~45), allocated across contigs proportionally to their position span. In contig mode (–chrom without –explode-partitions) each of the contig’s VDS partition spans is sliced into ceil(scale) intervals, so fractional values round up and the realized total is ceil(scale) times the number of spans. Use instead of –read-subintervals-per-chunk when the partition count is not known up front. Mutually exclusive with –read-subintervals-per-chunk.

Scout target loci by allele count

Two-pass mode: a cheap rows-only pass finds loci within the –min-alleles/–max-alleles range, then the VDS is re-read targeting only those loci with a scaled (not per-row) number of partitions.

--scout-alleles

Enable the two-pass scout mode when creating the info HT. Requires at least one of –min-alleles or –max-alleles.

Default: False

--scout-n-partitions

Approximate number of intervals/partitions to spread the scouted target loci across on re-read. Takes precedence over –scout-rows-per-partition.

--scout-rows-per-partition

Target number of scouted loci per interval/partition on re-read. Used only when –scout-n-partitions is not set. Default is 20, calibrated for heavy high-allele strata; raise it when scouting cheap low-allele loci.

--scout-byte-weight

Shrink scout chunks inside VDS partitions whose entries part-file is disproportionately large (bytes per target locus vs the median partition), so entry-dense regions get proportionally more, smaller intervals. Uses only storage metadata (one JSON read + one parts listing). Requires –scout-alleles; incompatible with –test.

Default: False

--scout-byte-weight-cap

Maximum byte-weight up-weighting (minimum chunk shrink factor) per partition. Default is 16. A partition denser than this many times the median is clamped here, so its chunks stay larger than its density warrants; if a run still has slow stragglers in a known entry-dense region, raise the cap to let that partition shrink further, or lower –scout-rows-per-partition to shrink every chunk. Only meaningful with –scout-byte-weight.

Default: 16.0

--scout-limit-intervals

TEST ONLY: truncate the scout-derived interval list to the first N intervals (genomic order), so the re-read covers only the start of the contig’s target loci. The scout pass and byte-weighting still run over the full contig. Requires –scout-alleles and –ac-info-ht-checkpoint-path-override (a partial-contig output must not land on a derived production path).

Variant QC annotation HT parameters.

--create-variant-qc-annotation-ht

Creates an annotated HT with features for variant QC.

Default: False

--impute-features

If set, imputation is performed for variant QC features.

Default: False

--n-partitions

Desired number of partitions for variant QC annotation HT.

Default: 5000

Export true positive VCFs

Arguments used to define true positive variant set.

--export-true-positive-vcfs

Exports true positive variants (–transmitted-singletons and/or –sibling-singletons) to VCF files.

Default: False

--transmitted-singletons

Include transmitted singletons in the exports of true positive variants to VCF files.

Default: False

--sibling-singletons

Include sibling singletons in the exports of true positive variants to VCF files.

Default: False

Module Functions

gnomad_qc.v5.annotations.generate_variant_qc_annotations.generate_ac_info_ht(vds)

Compute AC and AC_raw annotations for each allele count filter group.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.validate_local_allele_agg(vds)

Diff the local-allele-space aggregation against the original on a subset.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.create_sites_vcf_ht(...)

Import the AoU annotated sites-only VCF and reformat it into a HT.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.join_vcf_and_ac_info_hts(...)

Annotate the AC info HT with the reformatted sites VCF HT's annotations.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.union_ac_info_hts(...)

Union per-allele-count-stratum AC info HTs into a single AC info HT.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.run_generate_trio_stats(mt, ...)

Generate trio transmission stats from a VariantDataset and pedigree info.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.validate_trio_stats(ht, ...)

Validate the combined trio stats HT for union integrity and value sanity.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.run_generate_sib_stats(mt, ...)

Generate sibling stats from a VariantDataset and relatedness info.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.create_variant_qc_annotation_ht(...)

Create a Table with all necessary annotations for variant QC.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.get_tp_ht_for_vcf_export(ht)

Get Tables with raw and adj true positive variants to export as a VCF for use in VQSR.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.scout_target_loci(...)

Find loci whose allele count falls within a target range (cheap rows-only pass).

gnomad_qc.v5.annotations.generate_variant_qc_annotations.group_scout_loci_into_intervals(...)

Group scouted target loci into a bounded set of locus intervals.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_scout_intervals(args)

Run the scout pass and derive partition-scaled re-read intervals.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_chunks(args)

Derive per-contig locus sub-intervals for a chunk of the VDS.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_contig_intervals(args)

Derive ~equal-data read intervals covering a single contig.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.main(args)

Generate all variant annotations needed for variant QC.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.get_script_argument_parser()

Get script argument parser.

Script to generate annotations for variant QC on gnomAD v5.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.generate_ac_info_ht(vds, max_alleles=None, min_alleles=None, use_local_allele_agg=False)[source]

Compute AC and AC_raw annotations for each allele count filter group.

Function also adds AS_pab_max and allele_info annotations.

With use_local_allele_agg, both the AC aggregation and AS_pab_max are computed in local-allele space: per-genotype cost scales with the number of local alleles instead of global alts, which is much faster at high-allele loci and produces output identical to the original. The local-allele path adds per-genotype hash-map overhead that can be a net loss at low-allele loci, so the original dense aggregation remains the default.

Parameters:
  • vds (VariantDataset) – VariantDataset to use for computing AC and AC_raw annotations.

  • max_alleles (int) – Optional maximum number of alleles (including the reference) allowed at a locus. Variants at loci with more than this many alleles are filtered out. Default is None (no filtering).

  • min_alleles (int) – Optional minimum number of alleles (including the reference) allowed at a locus. Variants at loci with fewer than this many alleles are filtered out. Default is None (no filtering).

  • use_local_allele_agg (bool) – If True, compute the AC aggregation AND AS_pab_max in local-allele space. This removes the ~n_global_alts-per-genotype costs in the AC aggregation and pab_max_expr, the dominant costs at high-allele loci. Recommended for high-allele strata (e.g. min_alleles >= 10). Default is False (original aggregation).

Return type:

Table

Returns:

Table with AC and AC_raw annotations split by high quality, release, and unrelated.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.validate_local_allele_agg(vds, n_partitions=5, max_alleles=None, min_alleles=None, use_local_allele_agg=False, pab_atol=1e-09)[source]

Diff the local-allele-space aggregation against the original on a subset.

Builds the pre-split AC / AC_raw arrays and AS_pab_max with both the original and the local-allele strategy and asserts they match at every locus.

Comparisons handle missingness and NaN explicitly (rather than letting hl.agg.all silently skip them), verify the key sets are identical in both directions (so row-set drift is caught, not swallowed), and fail on empty input (so a zero-locus run is never a vacuous PASS). AC arrays (integers) are compared exactly; AS_pab_max (float) is compared to within pab_atol.

Parameters:
  • vds (VariantDataset) – VariantDataset to validate on.

  • n_partitions (int) – Number of leading variant_data partitions to sample. Pass None to validate every loaded partition.

  • max_alleles (int) – Optional max alleles filter applied before aggregation.

  • min_alleles (int) – Optional min alleles filter applied before aggregation.

  • use_local_allele_agg (bool) – If True, validate the local-allele AC aggregation and AS_pab_max against the original aggregation.

  • pab_atol (float) – Absolute tolerance for the AS_pab_max float comparison. Default is 1e-9.

Return type:

bool

Returns:

True iff the key sets match and every compared field matches at every sampled locus (and the input was non-empty).

gnomad_qc.v5.annotations.generate_variant_qc_annotations.create_sites_vcf_ht(vcf_path, header_path, lowqual_indel_phred_het_prior=40, test=False)[source]

Import the AoU annotated sites-only VCF and reformat it into a HT.

Converts the array-typed AS_* annotations to scalars, drops fields redundant with their AS_* versions, and adds AS_lowqual.

Parameters:
  • vcf_path (str) – Path to the annotated sites-only VCF.

  • header_path (str) – Path to the header file for the VCF.

  • lowqual_indel_phred_het_prior (int) – Phred-scaled prior for a het genotype at a site with a low quality indel. Default is 40. We use 1/10k bases (phred=40) to be more consistent with the filtering used by Broad’s Data Sciences Platform for VQSR.

  • test (bool) – Whether to use just the first two partitions of the loaded VCF.

Return type:

Table

Returns:

Reformatted sites VCF Table.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.join_vcf_and_ac_info_hts(ac_info_ht, vcf_ht, ac_info_ht_path=None, vcf_ht_path=None, use_repartition_for_join=False)[source]

Annotate the AC info HT with the reformatted sites VCF HT’s annotations.

The join is driven by the AC info side’s keys (the AoU VCF has variants not present in the samples we consider high quality, and those rows are dropped); AS_pab_max is moved from AC_info into info to match the final schema.

Parameters:
  • ac_info_ht (Table) – AC info HT (split, e.g. a generate_ac_info_ht checkpoint or a union_ac_info_hts output).

  • vcf_ht (Table) – Reformatted sites VCF HT (from create_sites_vcf_ht).

  • ac_info_ht_path (str) – Readable path for ac_info_ht. Required when use_repartition_for_join is True (both sides are re-read with shared _intervals).

  • vcf_ht_path (str) – Readable path for vcf_ht. If None and use_repartition_for_join is True, the VCF HT is first written to a temp path.

  • use_repartition_for_join (bool) – If True, co-partition the VCF and AC info HTs onto identical interval boundaries (via repartition_for_join) so the join is partition-aligned and requires no shuffle. Best for full-genome runs. Default is False.

Return type:

Table

Returns:

Info HT.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.union_ac_info_hts(ac_info_ht_paths, n_partitions=None, expected_strata=None)[source]

Union per-allele-count-stratum AC info HTs into a single AC info HT.

The per-stratum AC info HTs come from separate –generate-ac-info-ht runs over disjoint allele-count ranges (e.g. –max-alleles 9, –min-alleles 10 –max-alleles 100, –min-alleles 101), so the result is a row union, not a key join. All inputs must share identical row and key schemas. When every input carries the ac_info_ht_parameters global (written by –generate-ac-info-ht), the union replaces it with an ac_info_ht_strata array holding each input’s parameters in input order; otherwise globals are taken from the first input.

After the union, asserts that the total row count equals the sum of the input row counts and that all keys are distinct, so accidental locus overlap between strata (e.g. overlapping min/max allele ranges – both bounds are inclusive) fails loudly instead of silently duplicating rows.

Parameters:
  • ac_info_ht_paths (List[str]) – Paths to the per-stratum AC info HTs to union. Must contain at least two paths.

  • n_partitions (int) – Optional number of partitions for the unioned Table. Default is None (keep the union’s partitioning).

  • expected_strata (List[Dict]) – Optional list (aligned with ac_info_ht_paths) of dicts of expected ac_info_ht_parameters global fields (e.g. contig, min_alleles, max_alleles, test). Each input’s recorded parameters must match its expectation – a path name is only a convention, but the globals are written by the run itself, so this catches a table sitting at the right path with the wrong contents. Inputs without the provenance global fail verification. Default is None (no verification).

Return type:

Table

Returns:

Unioned AC info HT (checkpointed to a temp path).

gnomad_qc.v5.annotations.generate_variant_qc_annotations.run_generate_trio_stats(mt, fam_ped)[source]

Generate trio transmission stats from a VariantDataset and pedigree info.

Parameters:
  • mt (MatrixTable) – Dense trio MatrixTable.

  • fam_ped (Pedigree) – Pedigree containing trio info.

Return type:

Table

Returns:

Table containing trio stats.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.validate_trio_stats(ht, per_chrom_hts, expected_chroms)[source]

Validate the combined trio stats HT for union integrity and value sanity.

Raises on hard failures (row/contig loss, schema mismatch, invariant violations); logs biological plausibility metrics for eyeballing.

Parameters:
  • ht (Table) – Combined (unioned) trio stats HT.

  • per_chrom_hts (Dict[str, Table]) – Per-chromosome trio stats HTs keyed by contig.

  • expected_chroms (List[str]) – Contigs the union should cover (autosomes).

Return type:

None

gnomad_qc.v5.annotations.generate_variant_qc_annotations.run_generate_sib_stats(mt, relatedness_ht)[source]

Generate sibling stats from a VariantDataset and relatedness info.

Parameters:
  • mt (MatrixTable) – Input MatrixTable.

  • relatedness_ht (Table) – Table containing relatedness info.

Return type:

Table

Returns:

Table containing sibling stats.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.create_variant_qc_annotation_ht(info_ht, trio_stats_ht, sib_stats_ht, impute_features=True, n_partitions=5000)[source]

Create a Table with all necessary annotations for variant QC.

Annotations that are included:

Features for RF:
  • variant_type

  • allele_type

  • n_alt_alleles

  • has_star

  • AS_QD

  • AS_pab_max

  • AS_MQRankSum

  • AS_SOR

  • AS_ReadPosRankSum

Training sites (bool):
  • transmitted_singleton

  • sibling_singleton

  • fail_hard_filters - (ht.AS_QD < 0.5) | (ht.AS_FS > 60) | (ht.AS_MQ < 30)

Parameters:
  • info_ht (Table) – Info Table with split multi-allelics.

  • trio_stats_ht (Table) – Table with trio statistics.

  • sib_stats_ht (Table) – Table with sibling statistics.

  • impute_features (bool) – Whether to impute features using feature medians (this is done by variant type).

  • n_partitions (int) – Number of partitions to use for final annotated Table.

Return type:

Table

Returns:

Hail Table with all annotations needed for variant QC.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.get_tp_ht_for_vcf_export(ht, transmitted_singletons=False, sibling_singletons=False)[source]

Get Tables with raw and adj true positive variants to export as a VCF for use in VQSR.

Parameters:
  • ht (Table) – Input Table with transmitted singleton and sibling singleton information.

  • transmitted_singletons (bool) – Whether to include transmitted singletons in the true positive variants.

  • sibling_singletons (bool) – Whether to include sibling singletons in the true positive variants.

Return type:

Dict[str, Table]

Returns:

Dictionary of ‘raw’ and ‘adj’ true positive variant sites Tables.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.scout_target_loci(vds, checkpoint_path, min_alleles=None, max_alleles=None, contig=None)[source]

Find loci whose allele count falls within a target range (cheap rows-only pass).

This is the first pass of a two-pass “scout then re-read” strategy: it scans only the row metadata of the VDS variant data (entries are pruned, so no genotype aggregation runs) to identify the loci of interest, and writes them to a checkpointed Table that stays distributed (no driver collect of every locus).

Parameters:
  • vds (VariantDataset) – VariantDataset to scout. Only variant_data row keys are read.

  • checkpoint_path (str) – Path to checkpoint the target loci Table.

  • min_alleles (int) – Optional minimum number of alleles (including the reference) at a locus. Loci with fewer are excluded. Default is None (no minimum).

  • max_alleles (int) – Optional maximum number of alleles (including the reference) at a locus. Loci with more are excluded. Default is None (no maximum).

  • contig (str) – Optional contig (e.g. “chr1”) to restrict the scout to. Default is None (whole input).

Return type:

Table

Returns:

Table keyed by locus, alleles containing only the target loci.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.group_scout_loci_into_intervals(target_ht, n_partitions=None, rows_per_partition=None, parent_sizes=None, byte_weight_cap=16.0)[source]

Group scouted target loci into a bounded set of locus intervals.

Rather than one interval (and therefore one partition) per target locus, the sorted target loci are grouped into chunks so that each resulting interval – and thus each partition produced when the VDS is re-read with these intervals – holds many target loci. This keeps the partition count in a healthy range regardless of how many target loci there are.

The derivation stays distributed: only the chunk-boundary loci reach the driver, ~2 per interval, so the collect scales as 2 * n_loci / rows_per_partition rather than n_loci. The 20-row default is calibrated for small high-allele strata; a large low-allele scout should raise rows_per_partition (or set n_partitions) to keep the boundary collect and the interval count bounded.

With parent_sizes (from _partition_entry_sizes), chunking is byte-weighted: each locus is assigned to its parent VDS partition, each parent’s entries-bytes-per-target-locus is compared to the median parent, and parents above the median get proportionally smaller chunks (down to chunk / byte_weight_cap). Equal-locus chunks alone leave a residual skew – a locus in an entry-dense (e.g. pericentromeric) parent carries up to ~17x the genotype volume of a median locus – and the part-file size is a free storage-metadata proxy for exactly that density. Weighted intervals additionally never span parent partitions, so reads align to storage boundaries. A locus at a parent boundary may be assigned the neighboring parent’s weight; the signal is regional, so this is immaterial.

Intervals never span contigs.

Parameters:
  • target_ht (Table) – Table of target loci (e.g. from scout_target_loci).

  • n_partitions (int) – Desired (approximate) number of intervals/partitions. If set, takes precedence over rows_per_partition.

  • rows_per_partition (int) – Target number of loci per interval/partition. Used only when n_partitions is None. Defaults to 20 if neither is set, calibrated for the heavy high-allele strata scout mode targets.

  • parent_sizes (List[Dict]) – Optional per-parent-partition spans and entries part-file sizes from _partition_entry_sizes, enabling byte-weighted chunking. Default is None (uniform equal-locus chunks).

  • byte_weight_cap (float) – Maximum up-weighting (minimum chunk shrink factor) for a dense parent. Default is 16.

Return type:

Tuple[List[Interval], Dict]

Returns:

Tuple of the hl.Interval list (genomic order, each interval a contiguous chunk of target loci within one contig) and, when parent_sizes was given, a dict with the weighting provenance (median_bytes_per_locus, cap, and weights: one entry per up-weighted parent with its global parent_index, n_target_loci, weight, and chunk_size); None otherwise.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_scout_intervals(args)[source]

Run the scout pass and derive partition-scaled re-read intervals.

Parameters:

args – Parsed CLI args.

Return type:

Tuple[List[Interval], Dict]

Returns:

Tuple of the locus intervals covering the scouted target loci and the byte-weighting provenance dict (None unless –scout-byte-weight).

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_chunks(args)[source]

Derive per-contig locus sub-intervals for a chunk of the VDS.

Parameters:

args – Parsed CLI args.

Returns:

List of locus intervals covering the chunk.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.compute_contig_intervals(args)[source]

Derive ~equal-data read intervals covering a single contig.

Each VDS partition’s position span on the contig (from the partition bounds in the table metadata, a metadata-only lookup) is subdivided into equal-width slices, so interval width adapts to data density: every interval holds about 1/scale of one ~equal-data partition. Equal-width slicing of the full contig span instead gives severely skewed jobs on contigs with large variant-free regions (empty sub-second reads over the p-arm/centromere next to multi-partition slices in dense pockets). Regions between partition spans hold no variant rows and get no interval, matching chunk mode’s variant-bounds semantics.

With –read-subintervals-scale, each partition span gets scale slices; otherwise –read-subintervals-per-chunk gives a total spread evenly across the contig’s partition spans.

Parameters:

args – Parsed CLI args.

Return type:

List[Interval]

Returns:

List of locus intervals covering the contig’s data, in genomic order.

gnomad_qc.v5.annotations.generate_variant_qc_annotations.main(args)[source]

Generate all variant annotations needed for variant QC.