Search Shortcut cmd + k | ctrl + k
duckhts

DuckDB extension for reading HTS file formats via htslib

Maintainer(s): sounkou-bioinfo

Installing and Loading

INSTALL duckhts FROM community;
LOAD duckhts;

Example

-- Load the extension
LOAD duckhts;

-- Read a VCF/BCF file (tidy FORMAT columns)
SELECT CHROM, POS, REF, ALT, SAMPLE_ID
FROM read_bcf('test/data/formatcols.vcf.gz', tidy_format := true)
LIMIT 5;

-- Read a BAM/SAM file
SELECT QNAME, RNAME, POS, READ_GROUP_ID, SAMPLE_ID
FROM read_bam('test/data/rg.sam.gz')
LIMIT 5;

About duckhts

DuckHTS provides DuckDB table functions for common high-throughput sequencing (HTS) formats using htslib. Query VCF, BCF, BAM, CRAM, FASTA, FASTQ, GTF, GFF, and tabix-indexed files directly in SQL.

The extension also includes sequence utility UDFs, SAM flag predicate helpers, and HTS metadata helpers.

Full function reference includes signatures, return schemas and examples.

Functions included in this extension:

Diagnostics

  • duckhts_htslib_version: Return the runtime version reported by the htslib library loaded with DuckHTS. Rduckhts uses this value to reject a downstream linking receipt whose source/header version does not match the loaded library.
  • duckhts_htslib_features: Return the htslib runtime feature bitfield reported by hts_features(). Use duckhts_htslib_feature_string() for the corresponding build description.
  • duckhts_htslib_feature_string: Return htslib's runtime build-feature description, including configured transports, compression libraries, compiler, and build flags. DuckHTS snapshots it once while loading the extension so parallel SQL calls read immutable text.
  • duckhts_simd_backend: Return the current DuckHTS SIMD dispatch label. For explicit scalar or concrete backend requests this is the requested policy; for auto it is the single selected backend when all logical kernels resolve to the same backend, or mixed when per-kernel auto-dispatch resolves to multiple backends. Use duckhts_simd_kernel_info() for per-kernel details.
  • duckhts_simd_requested_backend: Return the current explicit SIMD backend request, usually auto unless SELECT backend FROM duckhts_simd_set_backend('auto'|'scalar'|backend) was called. The selected per-kernel backend may differ under auto-dispatch across x86, ARM, wasm, and scalar-only builds.
  • duckhts_simd_backend_compiled: Return whether a concrete DuckHTS SIMD backend was compiled into this build. This is independent of whether the current CPU/runtime supports executing that backend; for example avx512 can be compiled but not CPU-supported on the running host.
  • duckhts_simd_backend_cpu_supported: Return whether the current CPU/runtime supports a concrete DuckHTS SIMD backend, independent of whether DuckHTS compiled an implementation for it. Availability is the intersection of compiled and CPU-supported.
  • duckhts_simd_backend_available: Return whether a concrete SIMD backend is usable in the current process. Availability means the backend is compiled into DuckHTS and supported by the current CPU/runtime. auto is a selection request rather than a concrete backend and is not reported as available here.
  • duckhts_simd_info: Report compiled, runtime-supported and selected status for each concrete DuckHTS SIMD backend.
  • duckhts_simd_kernel_info: Return one row per logical DuckHTS SIMD kernel showing the concrete backend selected by the current immutable dispatch table, the selected capability, the requested backend policy, whether scalar was used as a per-kernel fallback, and the dispatch mode. This is the authoritative diagnostic for mixed auto-dispatch when different kernels resolve to different backends.
  • duckhts_simd_set_backend: Explicitly select the DuckHTS SIMD dispatch policy for this process using a one-row table-function call and return the current dispatch label in a backend column. Use auto for per-kernel runtime dispatch or scalar for a portable baseline; unavailable platform-specific requests such as avx512 on non-AVX-512 CPUs raise an error instead of silently falling back.
  • duckhts_duckdb_type_supported: Return whether the currently open DuckDB runtime advertises a logical type with the given name through duckdb_types(). This is a catalog-level runtime probe for feature gating SQL/macros across DuckDB versions.
  • duckhts_duckdb_supports_variant: Return whether the currently open DuckDB runtime advertises the VARIANT logical type. Use this to gate optional SQL that depends on DuckDB VARIANT support.
  • duckhts_duckdb_supports_geometry: Return whether the currently open DuckDB runtime advertises the GEOMETRY logical type. Use this to gate optional SQL that depends on DuckDB GEOMETRY support.

Variant Annotation

  • duckvep_ensembl_regions: Match tiled FASTA sequence to one Ensembl core assembly and assign dense model-local sequence-region ordinals.
  • duckvep_ensembl_transcripts: Build validated VEP-116 Ensembl core transcript models from core tables and matching tiled FASTA sequence.
  • duckvep_ensembl_regulation_features: Prepare VEP-116 RegulatoryFeature and MotifFeature intervals for a DuckVEP model.
  • duckvep_model_receipt: Create a deterministic provenance receipt and semantic hash for prepared DuckVEP model relations.
  • duckvep_model_load: Load a validated immutable consequence model under a name in the current DuckDB database; return one TRUE row.
  • duckvep_model_drop: Remove a named resident DuckVEP consequence model and release its transcript and regulation-feature interval indexes, sequences, and cached worker state. Returns FALSE when the name is absent or the model is in use by an annotation vector.
  • duckvep_allele_geometry: Separate uploaded, VEP-116 feature and minimized-edit geometry for one literal biallelic allele.
  • duckvep_transcript_projection: Project independent literal alleles and existing DuckVEP annotations into typed, unshifted VEP-116 transcript display fields.
  • duckvep_repeat_alleles: Prepare bounded literal reference and alternate alleles from exact ordered repeat descriptions.
  • duckvep_breakend_geometry: Parse one raw VCF 4.5 breakend ALT into mate coordinates, orientation and retained replacement sequence.
  • duckvep_haplotypes: Replay literal phased CDS/protein paths with carriers, source contributors, coding blocks, aligned differences and optional protein HGVS.
  • duckvep_phase_call: Assign decoded GT/PS allele slots to haplotype lanes under strict or pinned VEP-116 phase policy.
  • duckvep_annotate: Annotate independent literal alleles, exact typed structural events and paired breakends against a resident VEP-116-compatible model.
  • duckvep_so_terms: Return VEP-116 Sequence Ontology terms, consequence-mask bits, impact, severity rank and evaluator tier.

Readers

  • read_bcf: Read VCF/BCF with header-typed INFO/FORMAT, typed CSQ/ANN/BCSQ annotations, sample selection and optional tidy sample rows.
  • read_geno: Read one row per VCF/BCF record with typed arbitrary-ploidy GT/PS calls, selected FORMAT fields and optional original VCF genotype text.
  • read_bcf_samples: Read the typed VCF/BCF sample catalog as sample_index UINTEGER and sample_name VARCHAR without reading records. Indices are zero-based positions in the original header, remain stable under selection and join read_geno calls from the same unchanged file. NULL or '-' selects all; an empty string selects none; comma-separated names include samples; '^' excludes them. Names are validated by HTSlib, selected rows retain header order, and unknown names error.
  • read_bam: Read SAM/BAM/CRAM alignments with optional typed SAM tags, auxiliary maps and packed sequence, quality or CIGAR output.
  • read_fasta: Read full FASTA records or indexed regions with text or packed sequence output.
  • read_bed: Read BED3-BED12 interval files with canonical typed columns and optional tabix-backed region filtering. scan_mode := 'sequential' forces full-file streaming/counting instead of index-backed count paths and is incompatible with region.
  • fasta_nuc: Compute bedtools nuc-style nucleotide composition for supplied BED intervals or generated fixed-width bins over a FASTA reference. A failed reference fetch fails the query with the file and zero-based half-open interval; requested intervals are not silently omitted. For bgzipped FASTA, gzi_path may point to an explicit .gzi sidecar when it is not colocated with the FASTA.
  • read_fastq: Read single-end, paired-end, or interleaved FASTQ files with optional legacy quality decoding. By default, FASTQ qualities are interpreted as modern Phred+33 input. Use sequence_encoding := 'nt16' to return SEQUENCE as UTINYINT[] and quality_representation := 'phred' to return QUALITY as UTINYINT[] instead of VARCHAR. input_quality_encoding accepts 'phred33', 'auto', 'phred64', or 'solexa64'. scan_mode := 'sequential' forces raw streaming/counting instead of index-backed count paths.
  • read_bigwig: Read stored BigWig signal intervals as CHROM, START0, END0 and VALUE.
  • read_gff: Read GFF annotations with optional raw scalar and parsed list/pair attributes, strict GFF3 validation and indexed region selection.
  • read_gtf: Read GTF annotations with optional raw scalar and parsed list/pair attributes and indexed region selection.
  • read_tabix: Read tabix-indexed text with optional header handling, inferred types and region selection.
  • fasta_index: Build a FASTA index (.fai) and return a single row with columns success (BOOLEAN) and index_path (VARCHAR).
  • hts_union_query: Generate a UNION ALL BY NAME query string that reads every file matching a glob pattern through the named reader function. The result includes a 'filename' column identifying the source file for each row. Assign to a variable with SET VARIABLE and execute via query(getvariable(…)). Optional params string is appended to each reader call. In R, use the typed rduckhts_*_multi() helpers instead, which accept file vectors with optional per-file parameters and create DuckDB tables directly.
  • hts_region_union_query: Generate UNION ALL BY NAME SQL over separate per-region scans of one HTS file.

Converters

  • duckhts_bcf_convert_parquet_sql: Build COPY SQL for read_bcf() output with Parquet metadata, VCF header text and selected columns, filters or partitions.
  • duckhts_bam_convert_parquet_sql: Build COPY SQL for read_bam() output with Parquet metadata, SAM header text and selected columns, filters or partitions.
  • duckhts_gff_convert_parquet_sql: Build COPY SQL for read_gff() output with Parquet metadata, GFF/tabix header text and selected columns, filters or partitions.
  • duckhts_tabix_convert_parquet_sql: Build COPY SQL for read_tabix() output with Parquet metadata, header text and selected columns, filters or partitions.

Coverage

  • read_pileup: Construct a region-scoped BAM pileup with one row per covered position, emitting chrom, 1-based position, depth, observed bases, and Phred+33 qualities after SAM flag and MAPQ filtering. This is a compact htslib pileup view, not samtools mpileup text parity.
  • bam_bin_counts: Count BAM or CRAM read starts into fixed-width bins. Returns one row per bin across the selected contig span, including zero-count bins, with total, forward, and reverse counts; rmdup := 'streaming' applies the WisecondorX-style larp/larp2 consecutive-position deduplication, rmdup := 'flag' drops SAM duplicate-flagged reads, and stats := 'gc', 'mq', or 'gc,mq' adds per-bin pre/post-filter GC and MAPQ sufficient statistics, including reference GC when reference is provided.
  • duckhts_bam_bed_coverage: Compute samtools coverage-like regional summaries for BAM or CRAM input over a BED target set, returning one row per BED interval with DuckHTS-specific pre/post-filter read counts, covered bases, percentage covered, mean depth, mean baseQ, mean mapQ, and strand-specific post-filter summaries in read mode. Indexed BAM/CRAM input is required in the current implementation. decompression_threads controls htslib worker threads for BAM/CRAM decoding; use 0 to disable them.
  • duckhts_mosdepth: Write mosdepth-compatible coverage files from indexed BAM/CRAM.

Intervals

  • duckhts_cgranges_create: Create an empty session-scoped cgranges registry entry that can be populated with intervals and finalized for overlap queries.
  • duckhts_cgranges_add: Append an interval to a session-scoped cgranges registry entry before finalization. Labels may be BIGINT-like, DOUBLE, VARCHAR, or BOOLEAN.
  • duckhts_cgranges_index: Finalize a populated cgranges registry entry and build its immutable overlap index for subsequent queries.
  • duckhts_cgranges_destroy: Destroy a session-scoped cgranges registry entry and release its indexed interval storage when it is not in active use.
  • duckhts_cgranges_from_query: Execute a SQL query on an extension-owned DuckDB connection, append its interval rows into a session-scoped cgranges registry entry, and leave the populated index ready for explicit finalization with duckhts_cgranges_index(…).
  • duckhts_cgranges_from_table: Reserved convenience constructor for bulk cgranges population from a table name. The current implementation is intentionally deferred and directs callers to duckhts_cgranges_from_query(…).
  • duckhts_cgranges_has_overlap: Vectorized scalar predicate for streaming provider rows through a finalized session-scoped cgranges index. Returns TRUE when the query interval overlaps at least one indexed interval, or when mode = 'contain' and it fully contains at least one indexed interval; NULL inputs return NULL.
  • duckhts_cgranges_count_overlaps: Vectorized scalar overlap counter for streaming provider rows through a finalized session-scoped cgranges index. Returns the number of indexed intervals that overlap the query interval, or with mode = 'contain' the number fully contained by it; NULL inputs return NULL.
  • duckhts_cgranges_overlaps_list: Vectorized scalar overlap expander for streaming provider rows through a finalized session-scoped cgranges index. Returns a LIST of hit STRUCTs that can be expanded with UNNEST, preserving provider columns while emitting one row per matching indexed interval. Because scalar return types are fixed, labels are returned as text with label_type describing the original cgranges label kind; NULL inputs return NULL.
  • duckhts_cgranges_overlaps: Query a finalized session-scoped cgranges registry entry and return one row per overlapping or containing indexed interval, preserving the original label type and interval coordinates.
  • duckhts_cgranges_overlaps_bulk: Run a SQL query that yields overlap probes, stream those rows through a finalized session-scoped cgranges registry entry, and return one row per matching indexed interval. The probe query runs on the extension-owned helper connection, so it must reference regular tables/views rather than connection-local temp tables. When query_row_id_col is omitted, query_row_id defaults to the 1-based probe row ordinal.
  • regionkey: Encode a genomic interval as an official RegionKey-compatible 64-bit unsigned integer. Start and end use 0-based half-open interval semantics, matching BED-style coordinates; strand accepts -1, 0, or 1.
  • regionkey_hex: Render a RegionKey as its lowercase 16-character hexadecimal string representation.
  • parse_regionkey_hex: Parse a 16-character hexadecimal RegionKey string back into its UBIGINT code. Invalid or non-hex strings return NULL.
  • encode_regionkey: Encode the raw upstream RegionKey fields directly: chromosome code, 0-based start, 0-based end, and strand code (0 = unknown, 1 = +, 2 = -).
  • extract_regionkey_chrom: Extract the raw upstream RegionKey chromosome code.
  • extract_regionkey_startpos: Extract the raw upstream RegionKey 0-based start position.
  • extract_regionkey_endpos: Extract the raw upstream RegionKey 0-based end position.
  • extract_regionkey_strand: Extract the raw upstream RegionKey strand code (0 = unknown, 1 = +, 2 = -).
  • decode_regionkey: Decode a RegionKey into its raw upstream numeric fields: chrom_code, start, end, and strand_code.
  • reverse_regionkey: Decode a RegionKey into a STRUCT with chrom, chrom_code, start, end, strand, and strand_code.
  • extend_regionkey: Extend a RegionKey interval by a fixed number of bases on both sides, clamping to the official 28-bit RegionKey position range.
  • are_overlapping_regions: Return TRUE when two explicit 0-based half-open intervals overlap on the same canonical chromosome.
  • are_overlapping_region_regionkey: Return TRUE when a 0-based half-open interval overlaps the supplied RegionKey interval.
  • are_overlapping_regionkeys: Return TRUE when two RegionKeys overlap.

Quality Control

  • duckhts_fastq_qc: Aggregate canonical sequence and Phred+33 quality strings directly into exact read/base/Q20/Q30/Q40, nucleotide, quality-sum, and per-cycle sufficient statistics. The nested cycles list supports mean-quality, nucleotide-content, GC, and read-length curves without expanding one SQL row per base. Rows with any NULL input are ignored. Per-cycle state defaults to at most 1,048,576 cycles; pass a constant max_cycles per aggregate group to choose a larger explicit limit, up to 16,777,216.

Sample Identity

  • duckhts_somalier_import_sites: Import an already selected Somalier sites VCF/BCF as one canonical panel and population-frequency relation.
  • duckhts_somalier_vcf_counts: Extract a complete panel-aligned A/B/other count relation from VCF/BCF FORMAT/AD.
  • duckhts_somalier_bam_counts: Extract complete panel-aligned A/B/other base counts from one indexed BAM or CRAM source.
  • duckhts_somalier_panel_sha256: Derive a stable SHA-256 identity for an ordered biallelic sample-fingerprinting panel.
  • duckhts_somalier_frequency_sha256: Derive a stable identity for panel-aligned population-B allele frequencies.
  • duckhts_somalier_classify: Classify one measured A/B/other count tuple for Somalier-derived autosomal relatedness.
  • duckhts_somalier_prepare_sketches: Build one panel-verified packed relatedness sketch per sample from typed count evidence.
  • duckhts_somalier_verify_sketches: Verify persisted relatedness sketches against their retained raw count evidence.
  • duckhts_somalier_relatedness: Compute fused Somalier-derived relatedness and concordance statistics for two prepared sketches.
  • duckhts_somalier_verify_relatedness: Verify a typed relatedness result against its two sealed sketches.
  • duckhts_somalier_charr: Estimate per-sample contamination with a bounded Somalier-derived CHARR reduction.
  • duckhts_somalier_matched_contamination: Estimate directional contamination for explicitly selected receiver/anchor sample pairs.

Metadata

  • detect_quality_encoding: Inspect a FASTQ file's observed quality ASCII range and report compatible legacy encodings with a heuristic guessed encoding.
  • duckhts_samtools_idxstats: Write samtools idxstats-compatible TAB-delimited output for BAM, CRAM, or SAM input. Indexed BAM uses hts_idx_get_stat(...) for the fast path; CRAM, SAM, and unindexed BAM fall back to a full scan while preserving samtools-style contig rows plus the final * row.
  • read_hts_header: Inspect HTS headers in parsed, raw, or combined form across supported formats. Raw VCF/BCF mode includes the final #CHROM sample header line so the returned text is suitable for Parquet metadata and future VCF/BCF regeneration.
  • read_hts_index: Inspect high-level HTS index metadata such as sequence names and mapped counts.
  • read_hts_index_spans: Expand index metadata into span and chunk rows suitable for low-level index inspection.
  • read_hts_index_raw: Return the raw on-disk HTS index blob together with basic identifying metadata.

Compression

  • bgzip: Compress a plain file to BGZF and return the created output path and byte counts.
  • bgunzip: Decompress a BGZF-compressed file and return the created output path and byte counts.

Indexing

  • bam_index: Build a BAM or CRAM index and report the written index path and format.
  • bcf_index: Build a TBI or CSI index for a VCF or BCF file and report the written index path and format.
  • tabix_index: Build a tabix index for a BGZF-compressed text file using a preset or explicit coordinate columns.

Variants

  • variantkey: Encode a normalized biallelic variant as an official VariantKey-compatible 64-bit unsigned integer. This DuckHTS wrapper accepts 1-based VCF/DuckHTS POS to match bcftools %VKX / +add-variantkey, internally converts to the upstream 0-based field, and preserves the official hashed nonreversible mode for large, ambiguous, and symbolic REF/ALT strings. Only CHROM, POS, REF, and ALT are encoded; END, SVLEN, mate breakend coordinates, and other SV metadata are not.
  • variantkey_hex: Render a VariantKey as its lowercase 16-character hexadecimal string representation.
  • parse_variantkey_hex: Parse a 16-character hexadecimal VariantKey string back into its UBIGINT code. Invalid or non-hex strings return NULL.
  • encode_variantkey: Encode the raw upstream VariantKey fields directly: chromosome code, 0-based position, and 31-bit REF+ALT code.
  • extract_variantkey_chrom: Extract the raw upstream VariantKey chromosome code.
  • extract_variantkey_pos: Extract the raw upstream VariantKey 0-based position field.
  • extract_variantkey_refalt: Extract the raw upstream 31-bit VariantKey REF+ALT code.
  • decode_variantkey: Decode a VariantKey into its raw upstream numeric fields: chrom_code, pos0, and refalt_code.
  • reverse_variantkey: Decode a VariantKey into a STRUCT with chrom, chrom_code, 1-based pos, upstream 0-based pos0, ref, alt, refalt_code, and reversible. For hashed nonreversible keys, reversible is FALSE and ref/alt are returned as NULL because DuckHTS v1 does not ship the optional NRVK lookup sidecar.
  • variantkey_range: Return the inclusive minimum and maximum VariantKey bounds for a chromosome plus 1-based VCF position range, suitable for numeric range filtering on precomputed VariantKeys.
  • duckhts_contig_key: Return a conservative contig join key by removing one non-empty leading chr prefix case-insensitively and normalizing M/MT to MT. X and Y are uppercased; all other suffixes are preserved. This does not map numeric sex chromosomes, accessions, patches, or alternate loci.
  • bcftools_liftover: Row-oriented liftover kernel intended to mirror bcftools +liftover semantics as closely as possible while returning one STRUCT per input row with fields: src_chrom, src_pos, src_ref, src_alt, dest_chrom, dest_pos, dest_end, dest_ref, dest_alt, mapped, reverse_complemented, swap, reject_reason, and note. Set no_left_align := true to skip post-liftover left-alignment of lifted indels (mirrors –no-left-align in bcftools +liftover).
  • duckdb_liftover: DuckDB-specific wrapper over bcftools_liftover that takes either a table name or a derived-table expression plus column-name strings for chrom/pos/ref/alt and returns the lifted table. The no_left_align parameter mirrors –no-left-align in bcftools +liftover.
  • bcftools_norm_row: Normalize one variant against FASTA with bcftools/vt-style left alignment.
  • duckhts_bcftools_norm: Normalize variants from a table or derived-table expression while preserving input columns.
  • bcftools_score: Compute polygenic scores from genotype VCF/BCF and summary statistics using bcftools +score dosage semantics.
  • bcftools_munge_row: Normalize one summary-statistics row into GWAS-VCF-style fields (chrom/pos/ref/alt/effect metrics), resolving REF/ALT orientation against a FASTA reference and applying swap-aware sign/frequency/count transforms. The output flag alleles_swapped means REF/ALT orientation was swapped to match the FASTA reference.
  • duckdb_munge: DuckDB macro wrapper over bcftools_munge_row that maps source columns (via preset or explicit map) and returns normalized GWAS-VCF-style rows with lean outputs and explicit alleles_swapped semantics. Output columns: chrom, pos, id, ref, alt, alleles_swapped, filter, ns, ez, nc, es, se, lp, af, ac, ne (16 columns). For METAL meta-analysis output with SI/I2/CQ/ED columns, use duckdb_munge_metal.
  • duckdb_munge_metal: Extended munge macro with METAL meta-analysis output columns. Same as duckdb_munge but additionally emits: si (imputation info, from INFO input), i2 (Cochran's I² heterogeneity, from HET_I2), cq (Cochran's Q -log10 p, from HET_LP or -log10(HET_P)), and ed (effect direction string, from DIRE; +/- flipped on allele swap). The R wrapper rduckhts_munge() auto-dispatches to this macro when metal keys (INFO, HET_I2, HET_P, HET_LP, DIRE) are present in the resolved column map.

Sequence UDFs

  • seq_revcomp: Compute the reverse complement of a DNA sequence using A, C, G, T, and N bases. Overloaded: accepts either a VARCHAR text sequence (returns VARCHAR) or a UTINYINT[] of htslib nt16 codes as produced by read_bam(sequence_encoding := 'nt16') (returns UTINYINT[]); the nt16 overload is bit-identical to the text path after decoding, so BAM pipelines can reverse-complement without leaving the nt16 encoding.
  • seq_canonical: Return the lexicographically smaller of a sequence and its reverse complement. Overloaded: accepts either a VARCHAR text sequence (returns VARCHAR) or a UTINYINT[] of htslib nt16 codes as produced by read_bam(sequence_encoding := 'nt16') (returns UTINYINT[]); the nt16 overload compares by decoded base order and is bit-identical to the text path after decoding.
  • seq_hash_2bit: Encode a short DNA sequence as a 2-bit unsigned integer hash. Overloaded to also accept a UTINYINT[] of htslib nt16 codes (from read_bam(sequence_encoding := 'nt16')); non-ACGT codes yield NULL, bit-identical to the text path.
  • seq_encode_4bit: Encode an IUPAC DNA sequence as a list of 4-bit base codes, preserving ambiguity symbols including N.
  • seq_decode_4bit: Decode a list of 4-bit IUPAC DNA base codes back into a sequence string.
  • seq_gc_content: Compute GC fraction for a DNA sequence as a value between 0 and 1. Overloaded: accepts either a VARCHAR text sequence or a UTINYINT[] of htslib nt16 codes as produced by read_bam(sequence_encoding := 'nt16'); the nt16 overload classifies codes directly and is bit-identical to the text path, so BAM pipelines can compute GC without decoding sequences back to text.
  • seq_kmers: Expand a sequence into positional k-mers with optional canonicalization.

SAM Flag UDFs

  • sam_flag_bits: Decode a SAM flag into a struct of boolean bit fields using explicit SAM-oriented names such as is_paired, is_proper_pair, is_next_segment_unmapped, and is_supplementary.
  • sam_flag_has: Test whether any bits from the provided SAM flag mask are set in a flag value.
  • is_forward_aligned: Test whether a mapped segment is aligned to the forward strand. Returns NULL for unmapped segments because SAM flag 0x10 does not define genomic strand when 0x4 is set.
  • is_paired: Test whether the SAM flag indicates that the template has multiple segments in sequencing (0x1).
  • is_proper_pair: Test whether the SAM flag indicates that each segment is properly aligned according to the aligner (0x2).
  • is_unmapped: Test whether the read itself is unmapped according to the SAM flag.
  • is_next_segment_unmapped: Test whether the next segment in the template is flagged as unmapped (0x8).
  • is_reverse_complemented: Test whether SEQ is stored reverse complemented (0x10); for mapped reads this corresponds to reverse-strand alignment.
  • is_next_segment_reverse_complemented: Test whether SEQ of the next segment in the template is stored reverse complemented (0x20).
  • is_first_segment: Test whether the read is marked as the first segment in the template.
  • is_last_segment: Test whether the read is marked as the last segment in the template.
  • is_secondary: Test whether the alignment is marked as secondary.
  • is_qc_fail: Test whether the read failed vendor or pipeline quality checks.
  • is_duplicate: Test whether the alignment is flagged as a duplicate.
  • is_supplementary: Test whether the alignment is marked as supplementary.

CIGAR Utils

  • cigar_has_soft_clip: Test whether a CIGAR string contains any soft-clipped segment (S). Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_has_hard_clip: Test whether a CIGAR string contains any hard-clipped segment (H). Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_left_soft_clip: Return the left-end soft-clipped length from a CIGAR string, or zero if the alignment does not start with S. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_right_soft_clip: Return the right-end soft-clipped length from a CIGAR string, or zero if the alignment does not end with S. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_query_length: Return the query-consuming length from a CIGAR string, counting M, I, S, =, and X. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_aligned_query_length: Return the aligned query length from a CIGAR string, counting M, =, and X but excluding clips and insertions. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_reference_length: Return the reference-consuming length from a CIGAR string, counting M, D, N, =, and X. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.
  • cigar_has_op: Test whether a CIGAR string contains at least one instance of the requested operator. Overloaded to also accept a UINTEGER[] binary CIGAR (as produced by read_bam(cigar_representation := 'binary')); the binary overload is bit-identical to the text path.

Operational notes:

  • Reader table functions accept scan_mode := 'sequential' to force full-file streaming/counting instead of index-backed count/parallel paths where applicable; region queries remain index-backed and are incompatible with sequential mode.
  • Paired FASTQ is supported via mate_path or interleaved := true.
  • CRAM reads are supported with an explicit reference file.
  • GTF/GFF attributes can be returned as a parsed MAP with attributes_map := true via read_gff or read_gtf.
  • Optional SAM tag columns and an auxiliary tag map are available through standard_tags and auxiliary_tags.
  • Tabix readers support header, header_names, type inference with auto_detect, and explicit column_types.
  • MSVC builds (windows_amd64/windows_arm64) are not supported; use MinGW/RTools on Windows.

Added Functions

function_name function_type description comment examples
__duckhts_somalier_certify_thresholds scalar NULL NULL  
__duckhts_somalier_charr aggregate NULL NULL  
__duckhts_somalier_contamination_filter scalar NULL NULL  
__duckhts_somalier_matched_profiles scalar NULL NULL  
__duckhts_somalier_matched_settings_valid scalar NULL NULL  
__duckhts_somalier_sketch aggregate NULL NULL  
__duckvep_projection_base macro NULL NULL  
__duckvep_projection_code scalar NULL NULL  
__duckvep_projection_peptide macro NULL NULL  
__duckvep_projection_peptide_finish macro NULL NULL  
__duckvep_projection_residue macro NULL NULL  
__duckvep_projection_span macro NULL NULL  
_duckvep_annotate_breakend_compact scalar NULL NULL  
_duckvep_annotate_breakend_rich scalar NULL NULL  
_duckvep_annotate_small_compact scalar NULL NULL  
_duckvep_annotate_small_hgvs scalar NULL NULL  
_duckvep_annotate_small_projected scalar NULL NULL  
_duckvep_annotate_small_projected_hgvs scalar NULL NULL  
_duckvep_annotate_small_rich scalar NULL NULL  
_duckvep_annotate_small_rich_hgvs scalar NULL NULL  
_duckvep_annotate_structural_compact scalar NULL NULL  
_duckvep_annotate_structural_rich scalar NULL NULL  
_duckvep_phase_call scalar NULL NULL  
_duckvep_raw_gt scalar NULL NULL  
_duckvep_record_order scalar NULL NULL  
are_overlapping_region_regionkey scalar NULL NULL  
are_overlapping_regionkeys scalar NULL NULL  
are_overlapping_regions scalar NULL NULL  
bam_bin_counts table NULL NULL  
bam_index table NULL NULL  
bcf_index table NULL NULL  
bcftools_liftover scalar NULL NULL  
bcftools_munge_row scalar NULL NULL  
bcftools_norm_row scalar NULL NULL  
bcftools_score table NULL NULL  
bgunzip table NULL NULL  
bgzip table NULL NULL  
cigar_aligned_query_length scalar NULL NULL  
cigar_has_hard_clip scalar NULL NULL  
cigar_has_op scalar NULL NULL  
cigar_has_soft_clip scalar NULL NULL  
cigar_left_soft_clip scalar NULL NULL  
cigar_query_length scalar NULL NULL  
cigar_reference_length scalar NULL NULL  
cigar_right_soft_clip scalar NULL NULL  
decode_regionkey scalar NULL NULL  
decode_variantkey scalar NULL NULL  
detect_quality_encoding table NULL NULL  
duckdb_liftover table_macro NULL NULL  
duckdb_munge table_macro NULL NULL  
duckdb_munge_metal table_macro NULL NULL  
duckdb_munge_preset_map macro NULL NULL  
duckdb_munge_resolved_map macro NULL NULL  
duckhts_alt_to_list scalar NULL NULL  
duckhts_bam_bed_coverage table NULL NULL  
duckhts_bam_convert_parquet_sql macro NULL NULL  
duckhts_bcf_convert_parquet_sql macro NULL NULL  
duckhts_bcftools_norm table_macro NULL NULL  
duckhts_cgranges_add scalar NULL NULL  
duckhts_cgranges_count_overlaps scalar NULL NULL  
duckhts_cgranges_create scalar NULL NULL  
duckhts_cgranges_destroy scalar NULL NULL  
duckhts_cgranges_from_query scalar NULL NULL  
duckhts_cgranges_from_table scalar NULL NULL  
duckhts_cgranges_has_overlap scalar NULL NULL  
duckhts_cgranges_index scalar NULL NULL  
duckhts_cgranges_overlaps table NULL NULL  
duckhts_cgranges_overlaps_bulk table NULL NULL  
duckhts_cgranges_overlaps_list scalar NULL NULL  
duckhts_contig_key scalar NULL NULL  
duckhts_duckdb_supports_geometry macro NULL NULL  
duckhts_duckdb_supports_variant macro NULL NULL  
duckhts_duckdb_type_supported macro NULL NULL  
duckhts_fastq_qc aggregate NULL NULL  
duckhts_gff_convert_parquet_sql macro NULL NULL  
duckhts_htslib_feature_string scalar NULL NULL  
duckhts_htslib_features scalar NULL NULL  
duckhts_htslib_version scalar NULL NULL  
duckhts_json_path_key macro NULL NULL  
duckhts_mosdepth table NULL NULL  
duckhts_parquet_copy_sql macro NULL NULL  
duckhts_parquet_ident_list macro NULL NULL  
duckhts_parquet_metadata_args macro NULL NULL  
duckhts_quote_ident macro NULL NULL  
duckhts_quote_string macro NULL NULL  
duckhts_raw_header_text macro NULL NULL  
duckhts_samtools_idxstats table NULL NULL  
duckhts_simd_backend scalar NULL NULL  
duckhts_simd_backend_available scalar NULL NULL  
duckhts_simd_backend_compiled scalar NULL NULL  
duckhts_simd_backend_cpu_supported scalar NULL NULL  
duckhts_simd_info table NULL NULL  
duckhts_simd_kernel_info table NULL NULL  
duckhts_simd_requested_backend scalar NULL NULL  
duckhts_simd_set_backend table NULL NULL  
duckhts_somalier_bam_counts table NULL NULL  
duckhts_somalier_charr table_macro NULL NULL  
duckhts_somalier_classify scalar NULL NULL  
duckhts_somalier_frequency_sha256 macro NULL NULL  
duckhts_somalier_import_sites table_macro NULL NULL  
duckhts_somalier_matched_contamination table_macro NULL NULL  
duckhts_somalier_panel_sha256 macro NULL NULL  
duckhts_somalier_prepare_sketches table_macro NULL NULL  
duckhts_somalier_relatedness scalar NULL NULL  
duckhts_somalier_vcf_counts table_macro NULL NULL  
duckhts_somalier_verify_relatedness scalar NULL NULL  
duckhts_somalier_verify_sketches macro NULL NULL  
duckhts_tabix_convert_parquet_sql macro NULL NULL  
duckvep_allele_geometry scalar NULL NULL  
duckvep_annotate table_macro NULL NULL  
duckvep_breakend_geometry scalar NULL NULL  
duckvep_ensembl_regions table_macro NULL NULL  
duckvep_ensembl_regulation_features table_macro NULL NULL  
duckvep_ensembl_transcripts table_macro NULL NULL  
duckvep_haplotypes table NULL NULL  
duckvep_model_drop scalar NULL NULL  
duckvep_model_load table NULL NULL  
duckvep_model_receipt table_macro NULL NULL  
duckvep_phase_call macro NULL NULL  
duckvep_repeat_alleles macro NULL NULL  
duckvep_so_terms table NULL NULL  
duckvep_transcript_projection table_macro NULL NULL  
encode_regionkey scalar NULL NULL  
encode_variantkey scalar NULL NULL  
extend_regionkey scalar NULL NULL  
extract_regionkey_chrom scalar NULL NULL  
extract_regionkey_endpos scalar NULL NULL  
extract_regionkey_startpos scalar NULL NULL  
extract_regionkey_strand scalar NULL NULL  
extract_variantkey_chrom scalar NULL NULL  
extract_variantkey_pos scalar NULL NULL  
extract_variantkey_refalt scalar NULL NULL  
fasta_index table NULL NULL  
fasta_nuc table NULL NULL  
hts_region_union_query macro NULL NULL  
hts_union_query macro NULL NULL  
is_duplicate scalar NULL NULL  
is_first_segment scalar NULL NULL  
is_forward_aligned scalar NULL NULL  
is_last_segment scalar NULL NULL  
is_next_segment_reverse_complemented scalar NULL NULL  
is_next_segment_unmapped scalar NULL NULL  
is_paired scalar NULL NULL  
is_proper_pair scalar NULL NULL  
is_qc_fail scalar NULL NULL  
is_reverse_complemented scalar NULL NULL  
is_secondary scalar NULL NULL  
is_supplementary scalar NULL NULL  
is_unmapped scalar NULL NULL  
parse_regionkey_hex scalar NULL NULL  
parse_variantkey_hex scalar NULL NULL  
read_bam table NULL NULL  
read_bcf table NULL NULL  
read_bcf_samples table NULL NULL  
read_bed table NULL NULL  
read_bigwig table NULL NULL  
read_fasta table NULL NULL  
read_fastq table NULL NULL  
read_geno table NULL NULL  
read_gff table NULL NULL  
read_gtf table NULL NULL  
read_hts_header table NULL NULL  
read_hts_index table NULL NULL  
read_hts_index_raw table_macro NULL NULL  
read_hts_index_spans table NULL NULL  
read_pileup table NULL NULL  
read_tabix table NULL NULL  
regionkey scalar NULL NULL  
regionkey_hex scalar NULL NULL  
reverse_regionkey scalar NULL NULL  
reverse_variantkey scalar NULL NULL  
sam_flag_bits scalar NULL NULL  
sam_flag_has scalar NULL NULL  
seq_canonical scalar NULL NULL  
seq_decode_4bit scalar NULL NULL  
seq_encode_4bit scalar NULL NULL  
seq_gc_content scalar NULL NULL  
seq_hash_2bit scalar NULL NULL  
seq_kmers table NULL NULL  
seq_revcomp scalar NULL NULL  
tabix_index table NULL NULL  
variantkey scalar NULL NULL  
variantkey_hex scalar NULL NULL  
variantkey_range scalar NULL NULL  

Overloaded Functions

This extension does not add any function overloads.

Added Types

This extension does not add any types.

Added Settings

This extension does not add any settings.