Search Shortcut cmd + k | ctrl + k
anofox_tabfm

Zero-shot tabular machine learning inside DuckDB — classification, regression, synthetic-data generation and imputation with real tabular foundation models (Mitra, TabDPT, TabPFN v2/2.5/2.6/3, TabICL, Orion-BiX/MSP, TabFM) on ONNX Runtime, no training loop

Maintainer(s): sipemu, jrosskopf

Installing and Loading

INSTALL anofox_tabfm FROM community;
LOAD anofox_tabfm;

Example

-- Fetch a real tabular foundation model once (Mitra, Apache-2.0, ~300 MB,
-- no license gate). Cached under ~/.cache/anofox-tabfm and reused after.
INSTALL httpfs; LOAD httpfs;                       -- weights are fetched over HTTPS
CALL tabfm_download('classification', model := 'mitra');

-- Label a few rows; leave the ones you want scored as NULL. The model reads
-- the labelled rows as in-context examples and predicts the rest — no training.
CREATE TABLE iris AS SELECT * FROM VALUES
  (5.1, 3.5, 1.4, 0.2, 'setosa'),
  (4.9, 3.0, 1.4, 0.2, 'setosa'),
  (7.0, 3.2, 4.7, 1.4, 'versicolor'),
  (6.4, 3.2, 4.5, 1.5, 'versicolor'),
  (6.3, 3.3, 6.0, 2.5, 'virginica'),
  (5.8, 2.7, 5.1, 1.9, 'virginica'),
  (5.0, 3.6, 1.4, 0.2, NULL),                      -- predict me
  (6.5, 3.0, 5.8, 2.2, NULL)                       -- and me
  AS t(sepal_len, sepal_wid, petal_len, petal_wid, species);

SELECT petal_len, petal_wid, yhat AS predicted_species, yhat_score
FROM tabfm_classify('iris', 'species', model := 'mitra')
WHERE species IS NULL;

About anofox_tabfm

anofox_tabfm embeds real tabular foundation models — TabPFN-style in-context learners — into DuckDB, so tabular classification and regression become a single SQL statement. There is no training loop, no Python, and no MLOps: the model reads your labelled rows as context and predicts the rest.

Built-in models

Eleven models ship in the extension and are selected with model := (or a SET anofox_tabfm_default_model once per session). The catalog covers the entire top five of the TabArena v0.1.4 leaderboard.

Commercially usable:

  • mitra — AWS AutoGluon (Apache-2.0), no license gate, ~300 MB. A great default.
  • tabdpt — Layer 6 AI (Apache-2.0), no license gate. Classification and regression from one checkpoint; needs no weight conversion step.
  • tabpfn-v2 — Prior Labs (Apache-2.0, attribution).
  • tabicl-v2 — Inria (BSD-3-Clause).
  • orion-bix — Lexsi Labs (MIT), no license gate. Classification only.
  • orion-msp — Lexsi Labs (MIT), no license gate. Classification only.

Non-commercial — the weights and their outputs may not be used for commercial or production purposes:

  • tabpfn-v2-5 — Prior Labs TabPFN 2.5.
  • tabpfn-v2-5-real — RealTabPFN 2.5, the same architecture continued pre-trained on real tabular data.
  • tabpfn-v2-6 — Prior Labs TabPFN 2.6. Handles up to 50k rows / 2000 features.
  • tabpfn-v3 — Prior Labs TabPFN 3. Testing, evaluation and internal benchmarking only.
  • tabfm-v1 — Google TabFM (gated; ~6.6 GB).

SELECT model, license, commercial FROM tabfm_list_models() reports each model's license and whether it is commercially usable — worth checking before shipping anything built on one.

Only weight-free computation graphs are bundled — no model weights are distributed with the extension. You download the weights yourself from Hugging Face into a local cache (~/.cache/anofox-tabfm), and for a gated model you first accept its license (SET anofox_tabfm_accept_hf_license = true). Repositories that Hugging Face itself gates additionally need a token, supplied with a standard DuckDB secret:

CREATE SECRET hf (TYPE http, BEARER_TOKEN 'hf_xxx', SCOPE 'https://huggingface.co');

Bring your own model

Register any compatible model entirely from SQL — no external JSON manifest. Weights can be .safetensors or a native PyTorch .ckpt (read without Python):

CALL tabfm_register_model(
  id := 'my-model',
  classification_graph   := 'model.onnx',
  classification_weights := 'model.safetensors',
  license := 'apache-2.0');

Synthetic data and imputation

The same in-context engine also runs backwards: instead of predicting one column, it models the whole table as a joint distribution and samples from it.

SELECT * FROM tabfm_generate('customers', 500);        -- 500 synthetic rows
CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw');  -- fill every NULL

Generation works column by column, each column sampled conditioned on the ones already generated (the chain rule), so the relationships between columns survive — not just each column's marginal. It needs only classification weights, so it works with every model above, including classification-only ones.

tabfm_impute is the deterministic sibling: it takes the conditional best estimate rather than sampling, so continuous fills keep full precision and non-NULL cells are never modified.

On the Prior Labs breast-cancer benchmark (30 features), a classifier given only synthetic in-context examples scores 97.8% on held-out real rows against 98.3% for the real training data, preserving correlation structure at 0.97 across all 435 feature pairs. Correlations are somewhat attenuated by the quantile binning used for continuous columns — see docs/GENERATE.md in the repository for what the method does and does not preserve. It is not a differential-privacy mechanism.

Inference

Runs on ONNX Runtime, statically linked into the extension (CPU execution provider). CUDA and ROCm/MIGraphX flavors exist in the source tree for self-builds; this community build ships the portable CPU flavor.

Surface

  • tabfm_classify / tabfm_regress — zero-shot predict (a single table with NULL targets, or a separate test := set)
  • tabfm_generate / tabfm_impute — synthesize rows from a table's joint distribution, or fill its NULL cells
  • tabfm_predict, tabfm_predict_by, tabfm_predict_agg, tabfm_predict_win
  • tabfm_register_model / tabfm_unregister_model — pure-SQL model registration
  • tabfm_download / tabfm_models / tabfm_list_models / tabfm_load / tabfm_unload / tabfm_remove — all accept model :=
  • tabfm_devices — discover CPU/GPU execution providers
  • SET anofox_tabfm_* settings (default model, license gate, cache dir, threads, device, tracing)

Full function names are anofox_tabfm_* with short tabfm_* aliases. See the project repository for the full SQL API and examples.

Added Functions

function_name function_type description comment examples
__anofox_tabfm_generate_agg aggregate NULL NULL  
__anofox_tabfm_impute_agg aggregate NULL NULL  
__anofox_tabfm_predict_agg aggregate NULL NULL  
__anofox_tabfm_predict_win aggregate NULL NULL  
anofox_tabfm_accelerate table Set up GPU acceleration in one call: find the accelerator, fetch its backend plugin (and runtime), and verify it loads. Returns one row per step (step, status, detail). Idempotent — re-running re-verifies without re-downloading. Reconnect afterwards for it to take effect. NULL [CALL tabfm_accelerate();]
anofox_tabfm_accuracy aggregate Compute classification accuracy: correct / total as DOUBLE over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.accuracy_score(normalize=True). NULL [SELECT tabfm_accuracy(actual, predicted) FROM predictions;, SELECT round(tabfm_accuracy(actual, predicted), 4) FROM predictions;]
anofox_tabfm_backends table Which discovered device can serve which model, and the reason when one cannot (model, task, device, backend, supported, reason). Built on the same servability predicate dispatch uses, so it cannot promise what a predict would refuse. 'supported' is NULL when the answer is not yet knowable — a bundled GPU graph cannot be matched against weights that have not been downloaded. NULL [SELECT * FROM tabfm_backends() WHERE NOT supported;]
anofox_tabfm_classify table_macro Zero-shot tabular classification with the TabFM foundation model. Uses the labelled rows of data as in-context examples to score the rows whose target is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat, yhat_score, is_training and (detail mode) a proba MAP. Optional features restricts the feature columns; opts is a MAP of options (seed, softmax_temperature, output_mode, …). NULL [SELECT age, plan, yhat, yhat_score FROM tabfm_classify('customers', 'churned') WHERE churned IS NULL;]
anofox_tabfm_confusion_matrix table_macro Compute a tidy confusion matrix from a table, returning (actual, predicted, count) rows for each observed (actual class, predicted class) pair. Column identifiers are safely double-quote escaped (T-01-02-01). Returns one row per distinct (actual, predicted) combination, ordered by actual then predicted. NULL [SELECT * FROM tabfm_confusion_matrix('predictions', 'actual', 'predicted');]
anofox_tabfm_cross_validate table_macro Run leakage-safe k-fold cross-validation in SQL. Assigns rows to k folds via hash(row_key, seed) % k, then for each fold trains tabfm_classify (or tabfm_regress) on the other k-1 folds and predicts the held-out fold via the two-table form. Returns k per-fold rows (fold_id 0..k-1) plus one aggregate row (fold_id=-1) with mean and stddev of the per-fold metric. Default metric: tabfm_accuracy (classification) or tabfm_rmse (regression). Target identifier is safely double-quoted to prevent SQL injection (CV-04). CV-01: hash-based fold assignment (deterministic, order-independent, seed-sensitive). CV-02: leakage-safe — held-out rows arrive as NULL-label test rows, never as training context. CV-03: per-fold rows + aggregate mean±std row. CV-04: safe target quoting. NULL [SELECT * FROM tabfm_cross_validate('patients', 'diagnosis', 'patient_id', k := 5, seed := 42);]
anofox_tabfm_crps aggregate Compute the mean Continuous Ranked Probability Score (CRPS) over regression predictions. CRPS is a proper scoring rule for predictive distributions — lower is better. The second argument must be the yhat_dist STRUCT output of tabfm_regress with output_mode := 'distribution'. Non-uniform bin widths from borders[i+1]-borders[i] are used (PSR-01). NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_crps(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
anofox_tabfm_devices table List the inference devices this build can see (device_id, ep, name, arch, vram, driver, usable). The cpu row always exists; GPU rows appear only in the matching flavor (cuda/rocm) and report usable=false when a device is present but unsupported. NULL [SELECT * FROM tabfm_devices();]
anofox_tabfm_download table Download the TabFM model weights for a task ('classification' or 'regression') from Hugging Face into the local cache. Requires SET anofox_tabfm_accept_hf_license = true. Returns one row per file (file, url, bytes, status). NULL [CALL tabfm_download('classification');]
anofox_tabfm_download_runtime table Download a backend plugin library ('cuda', 'rocm' or 'mlx') into the directory named by SET anofox_tabfm_ep_path (default: the cache dir's 'runtime' subdirectory), so that device can be driven without a matching compile-time flavor build. Returns one row per extracted file (file, bytes, status). NULL [CALL tabfm_download_runtime('cuda');]
anofox_tabfm_ece aggregate Compute Expected Calibration Error (ECE) using 10 equal-width bins over [0,1]. For each row, extracts argmax class and max probability from the proba MAP. ECE = sum_m (|B_m|/N)*|accuracy(B_m) - confidence(B_m)|. NULL rows skipped. Reference: Guo et al. 2017 (On Calibration of Modern Neural Networks). NULL [SELECT tabfm_ece(actual, proba) FROM predictions;]
anofox_tabfm_f1 aggregate Compute F1-score with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined per-class precision/recall returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.f1_score(average=avg, zero_division=0). NULL [SELECT tabfm_f1(actual, predicted, 'macro') FROM predictions;, SELECT tabfm_f1(actual, predicted, 'weighted') FROM predictions;]
anofox_tabfm_fold_assign table_macro Assign each row of data to one of k deterministic folds using (hash(row_key, seed) % k)::INTEGER. Output: all input columns plus fold_id INTEGER in [0, k-1]. Deterministic and order-independent — only the row_key expression determines the fold, not row position. hash() is stable within a DuckDB version; fold assignments change on DuckDB upgrade (Assumption A6). CV-01 compliant: not row_number()-based. NULL [SELECT id, label, fold_id FROM tabfm_fold_assign('my_table', 5, 'id', 42);]
anofox_tabfm_generate table_macro Generate synthetic rows from the joint distribution of data using a tabular foundation model. Factorizes the table column by column (the chain rule) and samples each column conditioned on the ones already generated, so correlations between columns are preserved rather than each column being drawn independently. Returns n rows with the same columns as data plus synthetic_id. Continuous columns are sampled via quantile bins, so values stay inside the observed range. Costs one model call per column, run sequentially. Options: seed, temperature (higher = more diverse), bins, column_order, model. NULL [SELECT * FROM tabfm_generate('customers', 100);]
anofox_tabfm_gpu_precompile table Warm the GPU path for a task by compiling the model for a shape bucket ahead of the first predict (on ROCm this builds and caches the .mxr program; a no-op cost on CPU/CUDA). Returns task, rows, features, device, status. NULL [CALL tabfm_gpu_precompile('classification', 1000, 50);]
anofox_tabfm_impute table_macro Fill the NULL cells of data with a tabular foundation model, conditioning each missing value on the other columns of its row. Returns the same columns as data, with non-NULL cells untouched, so it round-trips: CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw'). Unlike tabfm_generate this does not sample — it takes the conditional best estimate (classification argmax, regression point estimate), so continuous columns keep full precision. Optional columns restricts which columns are filled; opts accepts seed, rounds (MICE-style refinement sweeps), model. NULL [SELECT * FROM tabfm_impute('customers', columns := ['income']);]
anofox_tabfm_interval_score aggregate Compute the mean Gneiting & Raftery (2007) interval score at a given coverage level. IS = (u-l) + (2/(1-coverage)) * underflow + overflow penalties. Lower is better. The coverage parameter (default 0.9) is the nominal coverage — must be strictly between 0 and 1. The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_interval_score(actual, yhat_dist, coverage := 0.9) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
anofox_tabfm_list_models table List every model in the registry (built-ins + user manifests), downloaded or not: model, family, capabilities, license, commercial, size regime (max_rows/features/classes), downloaded. NULL [SELECT * FROM tabfm_list_models();]
anofox_tabfm_load table Eagerly load a downloaded TabFM model for a task into memory so the first predict is warm (otherwise the model loads lazily on first use). NULL [CALL tabfm_load('classification');]
anofox_tabfm_log_loss aggregate Compute log-loss (cross-entropy) from a per-class probability MAP(VARCHAR,DOUBLE). Probabilities are clipped to [1e-15, 1-1e-15] (sklearn eps default) to avoid infinity. Missing MAP keys are treated as p=0 then clipped. NULL rows skipped. Matches sklearn.metrics.log_loss(normalize=True, eps=1e-15). NULL [SELECT tabfm_log_loss(actual, proba) FROM predictions;]
anofox_tabfm_log_score aggregate Compute the mean negative log-likelihood (NLL / log-score) of actual values under the predictive bar distribution. Log-score is a proper scoring rule — lower is better. For y in bin i: NLL = log(w_i) - log(max(1e-10, p_i)). Out-of-support or zero-width bins return the max penalty -log(1e-10) (T-03-04). The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_log_score(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
anofox_tabfm_mae aggregate Compute mean absolute error (MAE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_error. NULL [SELECT tabfm_mae(actual, predicted) FROM predictions;, SELECT round(tabfm_mae(actual, predicted), 4) FROM predictions;]
anofox_tabfm_mape aggregate Compute mean absolute percentage error (MAPE) as a DIMENSIONLESS RATIO over (actual, predicted) pairs. Rows where actual == 0 are SKIPPED to avoid division by zero (documented behavior); returns NULL when all actuals are zero. Rows where actual OR predicted is NULL are also skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_percentage_error (fraction, not percent). NULL [SELECT tabfm_mape(actual, predicted) FROM predictions;, SELECT round(tabfm_mape(actual, predicted) * 100, 2) || '%' FROM predictions;]
anofox_tabfm_medae aggregate Compute median absolute error (MedAE) over (actual, predicted) pairs. Accumulates |actual - predicted| residuals in a heap-allocated vector (O(N) memory). For even N, returns the mean of the two middle residuals. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.median_absolute_error. NULL [SELECT tabfm_medae(actual, predicted) FROM predictions;, SELECT round(tabfm_medae(actual, predicted), 4) FROM predictions;]
anofox_tabfm_models table List the TabFM models known to the local cache (model, task, revision, path, bytes, loaded, license). NULL [SELECT * FROM tabfm_models();]
anofox_tabfm_precision aggregate Compute precision with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined precision (no predicted positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.precision_score(zero_division=0). NULL [SELECT tabfm_precision(actual, predicted, 'macro') FROM predictions;]
anofox_tabfm_r2 aggregate Compute the coefficient of determination R² = 1 - SS_res/SS_tot over (actual, predicted) pairs. Constant-target convention (matches sklearn): when the target variance is 0, returns 1.0 for a perfect prediction and 0.0 otherwise — never NaN or Inf. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.r2_score. NULL [SELECT tabfm_r2(actual, predicted) FROM predictions;, SELECT round(tabfm_r2(actual, predicted), 4) FROM predictions;]
anofox_tabfm_recall aggregate Compute recall with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined recall (no actual positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.recall_score(zero_division=0). NULL [SELECT tabfm_recall(actual, predicted, 'macro') FROM predictions;]
anofox_tabfm_register_model table Register a model in SQL (no manifest file). Named args: id, classification_graph / regression_graph (path or url to the weight-free ONNX graph), classification_weights / regression_weights, tensor_map (or classification_tensor_map / regression_tensor_map), weights_repo, license, commercial, gate_setting, preprocessing_profile, max_rows / max_features / max_classes. Then use model := ''. NULL [CALL tabfm_register_model(id := 'my', classification_graph := '/p/g.onnx', classification_weights := '/p/w.safetensors', tensor_map := '/p/map.json', license := 'apache-2.0');]
anofox_tabfm_regress table_macro Zero-shot tabular regression with the TabFM foundation model. Uses the rows of data with a known numeric target as in-context examples to predict the target for rows where it is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat (yhat_score is NULL for regression). Optional features restricts the feature columns; opts is a MAP of options. NULL [SELECT * FROM tabfm_regress('sold_homes', 'price', test := 'listings');]
anofox_tabfm_remove table Delete a downloaded TabFM model's weights from the local cache (by task, optionally a specific revision). NULL [CALL tabfm_remove('classification');]
anofox_tabfm_rmse aggregate Compute root-mean-squared error (RMSE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.root_mean_squared_error. NULL [SELECT tabfm_rmse(actual, predicted) FROM predictions;, SELECT round(tabfm_rmse(actual, predicted), 4) FROM predictions;]
anofox_tabfm_roc_auc aggregate Compute ROC-AUC with multiclass support and correct rank-sum tie handling. avg is REQUIRED: 'ovr' (one-vs-rest, macro average) or 'ovo' (one-vs-one, macro average). Tied scores handled via the Mann-Whitney U rank-sum form (no optimistic bias). NULL rows skipped. Matches sklearn.metrics.roc_auc_score(multi_class='ovr'/'ovo'). NULL [SELECT tabfm_roc_auc(actual, proba, 'ovr') FROM predictions;]
anofox_tabfm_unload table Unload a loaded TabFM model from memory (all models if no task is given), freeing its RAM/VRAM. NULL [CALL tabfm_unload('classification');]
anofox_tabfm_unregister_model table Remove a model registered with tabfm_register_model. Returns model, status. NULL [CALL tabfm_unregister_model('my_model');]
tabfm_accelerate table Set up GPU acceleration in one call: find the accelerator, fetch its backend plugin (and runtime), and verify it loads. Returns one row per step (step, status, detail). Idempotent — re-running re-verifies without re-downloading. Reconnect afterwards for it to take effect. NULL [CALL tabfm_accelerate();]
tabfm_accuracy aggregate Compute classification accuracy: correct / total as DOUBLE over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.accuracy_score(normalize=True). NULL [SELECT tabfm_accuracy(actual, predicted) FROM predictions;, SELECT round(tabfm_accuracy(actual, predicted), 4) FROM predictions;]
tabfm_backends table Which discovered device can serve which model, and the reason when one cannot (model, task, device, backend, supported, reason). Built on the same servability predicate dispatch uses, so it cannot promise what a predict would refuse. 'supported' is NULL when the answer is not yet knowable — a bundled GPU graph cannot be matched against weights that have not been downloaded. NULL [SELECT * FROM tabfm_backends() WHERE NOT supported;]
tabfm_classify table_macro Zero-shot tabular classification with the TabFM foundation model. Uses the labelled rows of data as in-context examples to score the rows whose target is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat, yhat_score, is_training and (detail mode) a proba MAP. Optional features restricts the feature columns; opts is a MAP of options (seed, softmax_temperature, output_mode, …). NULL [SELECT age, plan, yhat, yhat_score FROM tabfm_classify('customers', 'churned') WHERE churned IS NULL;]
tabfm_confusion_matrix table_macro Compute a tidy confusion matrix from a table, returning (actual, predicted, count) rows for each observed (actual class, predicted class) pair. Column identifiers are safely double-quote escaped (T-01-02-01). Returns one row per distinct (actual, predicted) combination, ordered by actual then predicted. NULL [SELECT * FROM tabfm_confusion_matrix('predictions', 'actual', 'predicted');]
tabfm_cross_validate table_macro Run leakage-safe k-fold cross-validation in SQL. Assigns rows to k folds via hash(row_key, seed) % k, then for each fold trains tabfm_classify (or tabfm_regress) on the other k-1 folds and predicts the held-out fold via the two-table form. Returns k per-fold rows (fold_id 0..k-1) plus one aggregate row (fold_id=-1) with mean and stddev of the per-fold metric. Default metric: tabfm_accuracy (classification) or tabfm_rmse (regression). Target identifier is safely double-quoted to prevent SQL injection (CV-04). CV-01: hash-based fold assignment (deterministic, order-independent, seed-sensitive). CV-02: leakage-safe — held-out rows arrive as NULL-label test rows, never as training context. CV-03: per-fold rows + aggregate mean±std row. CV-04: safe target quoting. NULL [SELECT * FROM tabfm_cross_validate('patients', 'diagnosis', 'patient_id', k := 5, seed := 42);]
tabfm_crps aggregate Compute the mean Continuous Ranked Probability Score (CRPS) over regression predictions. CRPS is a proper scoring rule for predictive distributions — lower is better. The second argument must be the yhat_dist STRUCT output of tabfm_regress with output_mode := 'distribution'. Non-uniform bin widths from borders[i+1]-borders[i] are used (PSR-01). NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_crps(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
tabfm_devices table List the inference devices this build can see (device_id, ep, name, arch, vram, driver, usable). The cpu row always exists; GPU rows appear only in the matching flavor (cuda/rocm) and report usable=false when a device is present but unsupported. NULL [SELECT * FROM tabfm_devices();]
tabfm_download table Download the TabFM model weights for a task ('classification' or 'regression') from Hugging Face into the local cache. Requires SET anofox_tabfm_accept_hf_license = true. Returns one row per file (file, url, bytes, status). NULL [CALL tabfm_download('classification');]
tabfm_download_runtime table Download a backend plugin library ('cuda', 'rocm' or 'mlx') into the directory named by SET anofox_tabfm_ep_path (default: the cache dir's 'runtime' subdirectory), so that device can be driven without a matching compile-time flavor build. Returns one row per extracted file (file, bytes, status). NULL [CALL tabfm_download_runtime('cuda');]
tabfm_ece aggregate Compute Expected Calibration Error (ECE) using 10 equal-width bins over [0,1]. For each row, extracts argmax class and max probability from the proba MAP. ECE = sum_m (|B_m|/N)*|accuracy(B_m) - confidence(B_m)|. NULL rows skipped. Reference: Guo et al. 2017 (On Calibration of Modern Neural Networks). NULL [SELECT tabfm_ece(actual, proba) FROM predictions;]
tabfm_f1 aggregate Compute F1-score with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined per-class precision/recall returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.f1_score(average=avg, zero_division=0). NULL [SELECT tabfm_f1(actual, predicted, 'macro') FROM predictions;, SELECT tabfm_f1(actual, predicted, 'weighted') FROM predictions;]
tabfm_fold_assign table_macro Assign each row of data to one of k deterministic folds using (hash(row_key, seed) % k)::INTEGER. Output: all input columns plus fold_id INTEGER in [0, k-1]. Deterministic and order-independent — only the row_key expression determines the fold, not row position. hash() is stable within a DuckDB version; fold assignments change on DuckDB upgrade (Assumption A6). CV-01 compliant: not row_number()-based. NULL [SELECT id, label, fold_id FROM tabfm_fold_assign('my_table', 5, 'id', 42);]
tabfm_generate table_macro Generate synthetic rows from the joint distribution of data using a tabular foundation model. Factorizes the table column by column (the chain rule) and samples each column conditioned on the ones already generated, so correlations between columns are preserved rather than each column being drawn independently. Returns n rows with the same columns as data plus synthetic_id. Continuous columns are sampled via quantile bins, so values stay inside the observed range. Costs one model call per column, run sequentially. Options: seed, temperature (higher = more diverse), bins, column_order, model. NULL [SELECT * FROM tabfm_generate('customers', 100);]
tabfm_gpu_precompile table Warm the GPU path for a task by compiling the model for a shape bucket ahead of the first predict (on ROCm this builds and caches the .mxr program; a no-op cost on CPU/CUDA). Returns task, rows, features, device, status. NULL [CALL tabfm_gpu_precompile('classification', 1000, 50);]
tabfm_impute table_macro Fill the NULL cells of data with a tabular foundation model, conditioning each missing value on the other columns of its row. Returns the same columns as data, with non-NULL cells untouched, so it round-trips: CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw'). Unlike tabfm_generate this does not sample — it takes the conditional best estimate (classification argmax, regression point estimate), so continuous columns keep full precision. Optional columns restricts which columns are filled; opts accepts seed, rounds (MICE-style refinement sweeps), model. NULL [SELECT * FROM tabfm_impute('customers', columns := ['income']);]
tabfm_interval_score aggregate Compute the mean Gneiting & Raftery (2007) interval score at a given coverage level. IS = (u-l) + (2/(1-coverage)) * underflow + overflow penalties. Lower is better. The coverage parameter (default 0.9) is the nominal coverage — must be strictly between 0 and 1. The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_interval_score(actual, yhat_dist, coverage := 0.9) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
tabfm_list_models table List every model in the registry (built-ins + user manifests), downloaded or not: model, family, capabilities, license, commercial, size regime (max_rows/features/classes), downloaded. NULL [SELECT * FROM tabfm_list_models();]
tabfm_load table Eagerly load a downloaded TabFM model for a task into memory so the first predict is warm (otherwise the model loads lazily on first use). NULL [CALL tabfm_load('classification');]
tabfm_log_loss aggregate Compute log-loss (cross-entropy) from a per-class probability MAP(VARCHAR,DOUBLE). Probabilities are clipped to [1e-15, 1-1e-15] (sklearn eps default) to avoid infinity. Missing MAP keys are treated as p=0 then clipped. NULL rows skipped. Matches sklearn.metrics.log_loss(normalize=True, eps=1e-15). NULL [SELECT tabfm_log_loss(actual, proba) FROM predictions;]
tabfm_log_score aggregate Compute the mean negative log-likelihood (NLL / log-score) of actual values under the predictive bar distribution. Log-score is a proper scoring rule — lower is better. For y in bin i: NLL = log(w_i) - log(max(1e-10, p_i)). Out-of-support or zero-width bins return the max penalty -log(1e-10) (T-03-04). The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. NULL [SELECT tabfm_log_score(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});]
tabfm_mae aggregate Compute mean absolute error (MAE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_error. NULL [SELECT tabfm_mae(actual, predicted) FROM predictions;, SELECT round(tabfm_mae(actual, predicted), 4) FROM predictions;]
tabfm_mape aggregate Compute mean absolute percentage error (MAPE) as a DIMENSIONLESS RATIO over (actual, predicted) pairs. Rows where actual == 0 are SKIPPED to avoid division by zero (documented behavior); returns NULL when all actuals are zero. Rows where actual OR predicted is NULL are also skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_percentage_error (fraction, not percent). NULL [SELECT tabfm_mape(actual, predicted) FROM predictions;, SELECT round(tabfm_mape(actual, predicted) * 100, 2) || '%' FROM predictions;]
tabfm_medae aggregate Compute median absolute error (MedAE) over (actual, predicted) pairs. Accumulates |actual - predicted| residuals in a heap-allocated vector (O(N) memory). For even N, returns the mean of the two middle residuals. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.median_absolute_error. NULL [SELECT tabfm_medae(actual, predicted) FROM predictions;, SELECT round(tabfm_medae(actual, predicted), 4) FROM predictions;]
tabfm_models table List the TabFM models known to the local cache (model, task, revision, path, bytes, loaded, license). NULL [SELECT * FROM tabfm_models();]
tabfm_precision aggregate Compute precision with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined precision (no predicted positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.precision_score(zero_division=0). NULL [SELECT tabfm_precision(actual, predicted, 'macro') FROM predictions;]
tabfm_r2 aggregate Compute the coefficient of determination R² = 1 - SS_res/SS_tot over (actual, predicted) pairs. Constant-target convention (matches sklearn): when the target variance is 0, returns 1.0 for a perfect prediction and 0.0 otherwise — never NaN or Inf. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.r2_score. NULL [SELECT tabfm_r2(actual, predicted) FROM predictions;, SELECT round(tabfm_r2(actual, predicted), 4) FROM predictions;]
tabfm_recall aggregate Compute recall with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined recall (no actual positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.recall_score(zero_division=0). NULL [SELECT tabfm_recall(actual, predicted, 'macro') FROM predictions;]
tabfm_register_model table Register a model in SQL (no manifest file). Named args: id, classification_graph / regression_graph (path or url to the weight-free ONNX graph), classification_weights / regression_weights, tensor_map (or classification_tensor_map / regression_tensor_map), weights_repo, license, commercial, gate_setting, preprocessing_profile, max_rows / max_features / max_classes. Then use model := ''. NULL [CALL tabfm_register_model(id := 'my', classification_graph := '/p/g.onnx', classification_weights := '/p/w.safetensors', tensor_map := '/p/map.json', license := 'apache-2.0');]
tabfm_regress table_macro Zero-shot tabular regression with the TabFM foundation model. Uses the rows of data with a known numeric target as in-context examples to predict the target for rows where it is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat (yhat_score is NULL for regression). Optional features restricts the feature columns; opts is a MAP of options. NULL [SELECT * FROM tabfm_regress('sold_homes', 'price', test := 'listings');]
tabfm_remove table Delete a downloaded TabFM model's weights from the local cache (by task, optionally a specific revision). NULL [CALL tabfm_remove('classification');]
tabfm_rmse aggregate Compute root-mean-squared error (RMSE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.root_mean_squared_error. NULL [SELECT tabfm_rmse(actual, predicted) FROM predictions;, SELECT round(tabfm_rmse(actual, predicted), 4) FROM predictions;]
tabfm_roc_auc aggregate Compute ROC-AUC with multiclass support and correct rank-sum tie handling. avg is REQUIRED: 'ovr' (one-vs-rest, macro average) or 'ovo' (one-vs-one, macro average). Tied scores handled via the Mann-Whitney U rank-sum form (no optimistic bias). NULL rows skipped. Matches sklearn.metrics.roc_auc_score(multi_class='ovr'/'ovo'). NULL [SELECT tabfm_roc_auc(actual, proba, 'ovr') FROM predictions;]
tabfm_unload table Unload a loaded TabFM model from memory (all models if no task is given), freeing its RAM/VRAM. NULL [CALL tabfm_unload('classification');]
tabfm_unregister_model table Remove a model registered with tabfm_register_model. Returns model, status. NULL [CALL tabfm_unregister_model('my_model');]

Overloaded Functions

This extension does not add any function overloads.

Added Types

This extension does not add any types.

Added Settings

name description input_type scope aliases
anofox_tabfm_accept_hf_license Accept the upstream model license (tabfm-non-commercial-v1.0: non-commercial use, no redistribution). Downloads of Google-licensed weights fail without this. BOOLEAN GLOBAL []
anofox_tabfm_cache_dir Weight cache root directory (default ~/.cache/anofox-tabfm) VARCHAR GLOBAL []
anofox_tabfm_context_cache Encode the labelled context once and reuse it across calls, for models that ship a split graph pair (prepare/query). Off by default. It pays off when the same context is scored more than once – chunked scoring, repeated queries against a fixed training table – and costs extra on a single call, which pays for the context it will not reuse. Test-row predictions match the combined graph; the fitted values on CONTEXT rows differ, because the query half has no label path and so no longer sees a context row's own label. Inert for a model that ships no pair. BOOLEAN GLOBAL []
anofox_tabfm_cpu_prepack Enable ONNX Runtime weight prepacking on the CPU EP: faster matmuls at ~+16% resident memory. BOOLEAN GLOBAL []
anofox_tabfm_default_model Default model id for tabfm_classify/regress/download/… when model := is not given. '' = resolve to the single-file manifest model, else the sole registered model. VARCHAR GLOBAL []
anofox_tabfm_device Execution device: auto|cpu|cuda|rocm|coreml|mlx ('migraphx' alias). 'auto' resolves per model: it picks the best device that model can actually be served on (its plugin present and a graph available), else the CPU — so it never promises an accelerator a model cannot use. SELECT * FROM tabfm_backends() shows the choice and the reason. Naming a device explicitly is a hard request: it errors rather than falling back. coreml is flavor-gated and 'auto' does not reach for it. VARCHAR GLOBAL []
anofox_tabfm_ep_path Directory holding the GPU backend plugins (anofox_tabfm_{cuda,migraphx,mlx}_plugin, with this platform's library prefix and extension) and the runtime libraries they load alongside themselves. Defaults to the 'runtime' subdirectory of anofox_tabfm_cache_dir, which is where CALL tabfm_download_runtime(…) and CALL tabfm_accelerate() put them — set this only to point somewhere else. VARCHAR GLOBAL []
anofox_tabfm_gpu_precision GPU numeric mode: fp32|tf32|bf16|fp16. fp32 (default) is strict — a device switch does not change answers, measured exact on both GPUs; on CUDA it disables TF32 tensor-core rounding. tf32 re-enables that rounding (CUDA only; fp32 storage, faster matmuls). bf16/fp16 quantize the MIGraphX program on ROCm (~2x faster on RDNA4, half the VRAM/.mxr; label flips vs fp32 are rare (~2%% measured) but NOT confined to near-ties — measured flips include high-confidence rows, so validate bf16 on your own data) and are rejected on CUDA rather than silently running fp32. VARCHAR GLOBAL []
anofox_tabfm_max_features Maximum feature columns per predict call BIGINT GLOBAL []
anofox_tabfm_max_memory Refuse a predict call when this process's resident memory is already at or above this size (e.g. '16GB') before the call starts, so the failure is a DuckDB exception instead of a cgroup OOM-kill. '' (default) disables the check. Checked against resident memory at call time, not an estimate of the call's own cost – it does not bound how much a single large call can grow memory by itself. VARCHAR GLOBAL []
anofox_tabfm_max_rows Maximum rows per predict call or group BIGINT GLOBAL []
anofox_tabfm_max_sessions Cap on cached model sessions across all models, devices and precisions (default 4); beyond it the oldest-loaded session is evicted. Sessions are cached per (model, device, precision) so switching devices does not rebuild multi-GB sessions, and this cap keeps that from accumulating unbounded memory. BIGINT GLOBAL []
anofox_tabfm_mxr_source Directory holding precompiled MIGraphX .mxr programs (offline/CI/shared cache). Before compiling a shape-bucket (~27 min on ROCm), a matching '____T_H.mxr' here is staged into the cache and reused; empty ('' default) always compiles on-device. Artifacts are arch- and ROCm-version-specific. VARCHAR GLOBAL []
anofox_tabfm_threads ONNX Runtime intra-op thread count for CPU inference BIGINT GLOBAL []
anofox_tabfm_trace_level Diagnostic verbosity: error|warn|info|debug|trace VARCHAR GLOBAL []
anofox_telemetry_enabled Enable or disable anonymous usage telemetry BOOLEAN GLOBAL []
anofox_telemetry_key PostHog API key for telemetry VARCHAR GLOBAL []
datazoo_banner Show the DataZoo feedback banner when an extension is loaded in an interactive terminal (at most once a day per extension). BOOLEAN GLOBAL []