Zero-shot tabular machine learning inside DuckDB — classification, regression, synthetic-data generation and imputation with real tabular foundation models (Mitra, TabDPT, TabPFN v2/2.5/2.6/3, TabICL, Orion-BiX/MSP, TabFM) on ONNX Runtime, no training loop
Installing and Loading
INSTALL anofox_tabfm FROM community;
LOAD anofox_tabfm;
Example
-- Fetch a real tabular foundation model once (Mitra, Apache-2.0, ~300 MB,
-- no license gate). Cached under ~/.cache/anofox-tabfm and reused after.
INSTALL httpfs; LOAD httpfs; -- weights are fetched over HTTPS
CALL tabfm_download('classification', model := 'mitra');
-- Label a few rows; leave the ones you want scored as NULL. The model reads
-- the labelled rows as in-context examples and predicts the rest — no training.
CREATE TABLE iris AS SELECT * FROM VALUES
(5.1, 3.5, 1.4, 0.2, 'setosa'),
(4.9, 3.0, 1.4, 0.2, 'setosa'),
(7.0, 3.2, 4.7, 1.4, 'versicolor'),
(6.4, 3.2, 4.5, 1.5, 'versicolor'),
(6.3, 3.3, 6.0, 2.5, 'virginica'),
(5.8, 2.7, 5.1, 1.9, 'virginica'),
(5.0, 3.6, 1.4, 0.2, NULL), -- predict me
(6.5, 3.0, 5.8, 2.2, NULL) -- and me
AS t(sepal_len, sepal_wid, petal_len, petal_wid, species);
SELECT petal_len, petal_wid, yhat AS predicted_species, yhat_score
FROM tabfm_classify('iris', 'species', model := 'mitra')
WHERE species IS NULL;
About anofox_tabfm
anofox_tabfm embeds real tabular foundation models — TabPFN-style in-context learners — into DuckDB, so tabular classification and regression become a single SQL statement. There is no training loop, no Python, and no MLOps: the model reads your labelled rows as context and predicts the rest.
Built-in models
Eleven models ship in the extension and are selected with model := (or a
SET anofox_tabfm_default_model once per session). The catalog covers the
entire top five of the TabArena
v0.1.4 leaderboard.
Commercially usable:
- mitra — AWS AutoGluon (Apache-2.0), no license gate, ~300 MB. A great default.
- tabdpt — Layer 6 AI (Apache-2.0), no license gate. Classification and regression from one checkpoint; needs no weight conversion step.
- tabpfn-v2 — Prior Labs (Apache-2.0, attribution).
- tabicl-v2 — Inria (BSD-3-Clause).
- orion-bix — Lexsi Labs (MIT), no license gate. Classification only.
- orion-msp — Lexsi Labs (MIT), no license gate. Classification only.
Non-commercial — the weights and their outputs may not be used for commercial or production purposes:
- tabpfn-v2-5 — Prior Labs TabPFN 2.5.
- tabpfn-v2-5-real — RealTabPFN 2.5, the same architecture continued pre-trained on real tabular data.
- tabpfn-v2-6 — Prior Labs TabPFN 2.6. Handles up to 50k rows / 2000 features.
- tabpfn-v3 — Prior Labs TabPFN 3. Testing, evaluation and internal benchmarking only.
- tabfm-v1 — Google TabFM (gated; ~6.6 GB).
SELECT model, license, commercial FROM tabfm_list_models() reports each
model's license and whether it is commercially usable — worth checking
before shipping anything built on one.
Only weight-free computation graphs are bundled — no model weights are
distributed with the extension. You download the weights yourself from
Hugging Face into a local cache (~/.cache/anofox-tabfm), and for a gated
model you first accept its license
(SET anofox_tabfm_accept_hf_license = true). Repositories that Hugging Face
itself gates additionally need a token, supplied with a standard DuckDB
secret:
CREATE SECRET hf (TYPE http, BEARER_TOKEN 'hf_xxx', SCOPE 'https://huggingface.co');
Bring your own model
Register any compatible model entirely from SQL — no external JSON manifest.
Weights can be .safetensors or a native PyTorch .ckpt (read without Python):
CALL tabfm_register_model(
id := 'my-model',
classification_graph := 'model.onnx',
classification_weights := 'model.safetensors',
license := 'apache-2.0');
Synthetic data and imputation
The same in-context engine also runs backwards: instead of predicting one column, it models the whole table as a joint distribution and samples from it.
SELECT * FROM tabfm_generate('customers', 500); -- 500 synthetic rows
CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw'); -- fill every NULL
Generation works column by column, each column sampled conditioned on the ones already generated (the chain rule), so the relationships between columns survive — not just each column's marginal. It needs only classification weights, so it works with every model above, including classification-only ones.
tabfm_impute is the deterministic sibling: it takes the conditional best
estimate rather than sampling, so continuous fills keep full precision and
non-NULL cells are never modified.
On the Prior Labs breast-cancer benchmark (30 features), a classifier given
only synthetic in-context examples scores 97.8% on held-out real rows
against 98.3% for the real training data, preserving correlation structure at
0.97 across all 435 feature pairs. Correlations are somewhat attenuated by the
quantile binning used for continuous columns — see docs/GENERATE.md in the
repository for what the method does and does not preserve. It is not a
differential-privacy mechanism.
Inference
Runs on ONNX Runtime, statically linked into the extension (CPU execution provider). CUDA and ROCm/MIGraphX flavors exist in the source tree for self-builds; this community build ships the portable CPU flavor.
Surface
tabfm_classify/tabfm_regress— zero-shot predict (a single table with NULL targets, or a separatetest :=set)tabfm_generate/tabfm_impute— synthesize rows from a table's joint distribution, or fill its NULL cellstabfm_predict,tabfm_predict_by,tabfm_predict_agg,tabfm_predict_wintabfm_register_model/tabfm_unregister_model— pure-SQL model registrationtabfm_download/tabfm_models/tabfm_list_models/tabfm_load/tabfm_unload/tabfm_remove— all acceptmodel :=tabfm_devices— discover CPU/GPU execution providersSET anofox_tabfm_*settings (default model, license gate, cache dir, threads, device, tracing)
Full function names are anofox_tabfm_* with short tabfm_* aliases. See the
project repository for the full
SQL API and examples.
Added Functions
| function_name | function_type | description | comment | examples |
|---|---|---|---|---|
| __anofox_tabfm_generate_agg | aggregate | NULL | NULL | |
| __anofox_tabfm_impute_agg | aggregate | NULL | NULL | |
| __anofox_tabfm_predict_agg | aggregate | NULL | NULL | |
| __anofox_tabfm_predict_win | aggregate | NULL | NULL | |
| anofox_tabfm_accelerate | table | Set up GPU acceleration in one call: find the accelerator, fetch its backend plugin (and runtime), and verify it loads. Returns one row per step (step, status, detail). Idempotent — re-running re-verifies without re-downloading. Reconnect afterwards for it to take effect. | NULL | [CALL tabfm_accelerate();] |
| anofox_tabfm_accuracy | aggregate | Compute classification accuracy: correct / total as DOUBLE over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.accuracy_score(normalize=True). | NULL | [SELECT tabfm_accuracy(actual, predicted) FROM predictions;, SELECT round(tabfm_accuracy(actual, predicted), 4) FROM predictions;] |
| anofox_tabfm_backends | table | Which discovered device can serve which model, and the reason when one cannot (model, task, device, backend, supported, reason). Built on the same servability predicate dispatch uses, so it cannot promise what a predict would refuse. 'supported' is NULL when the answer is not yet knowable — a bundled GPU graph cannot be matched against weights that have not been downloaded. | NULL | [SELECT * FROM tabfm_backends() WHERE NOT supported;] |
| anofox_tabfm_classify | table_macro | Zero-shot tabular classification with the TabFM foundation model. Uses the labelled rows of data as in-context examples to score the rows whose target is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat, yhat_score, is_training and (detail mode) a proba MAP. Optional features restricts the feature columns; opts is a MAP of options (seed, softmax_temperature, output_mode, …). |
NULL | [SELECT age, plan, yhat, yhat_score FROM tabfm_classify('customers', 'churned') WHERE churned IS NULL;] |
| anofox_tabfm_confusion_matrix | table_macro | Compute a tidy confusion matrix from a table, returning (actual, predicted, count) rows for each observed (actual class, predicted class) pair. Column identifiers are safely double-quote escaped (T-01-02-01). Returns one row per distinct (actual, predicted) combination, ordered by actual then predicted. | NULL | [SELECT * FROM tabfm_confusion_matrix('predictions', 'actual', 'predicted');] |
| anofox_tabfm_cross_validate | table_macro | Run leakage-safe k-fold cross-validation in SQL. Assigns rows to k folds via hash(row_key, seed) % k, then for each fold trains tabfm_classify (or tabfm_regress) on the other k-1 folds and predicts the held-out fold via the two-table form. Returns k per-fold rows (fold_id 0..k-1) plus one aggregate row (fold_id=-1) with mean and stddev of the per-fold metric. Default metric: tabfm_accuracy (classification) or tabfm_rmse (regression). Target identifier is safely double-quoted to prevent SQL injection (CV-04). CV-01: hash-based fold assignment (deterministic, order-independent, seed-sensitive). CV-02: leakage-safe — held-out rows arrive as NULL-label test rows, never as training context. CV-03: per-fold rows + aggregate mean±std row. CV-04: safe target quoting. | NULL | [SELECT * FROM tabfm_cross_validate('patients', 'diagnosis', 'patient_id', k := 5, seed := 42);] |
| anofox_tabfm_crps | aggregate | Compute the mean Continuous Ranked Probability Score (CRPS) over regression predictions. CRPS is a proper scoring rule for predictive distributions — lower is better. The second argument must be the yhat_dist STRUCT output of tabfm_regress with output_mode := 'distribution'. Non-uniform bin widths from borders[i+1]-borders[i] are used (PSR-01). NULL rows are skipped; empty input returns NULL. | NULL | [SELECT tabfm_crps(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| anofox_tabfm_devices | table | List the inference devices this build can see (device_id, ep, name, arch, vram, driver, usable). The cpu row always exists; GPU rows appear only in the matching flavor (cuda/rocm) and report usable=false when a device is present but unsupported. | NULL | [SELECT * FROM tabfm_devices();] |
| anofox_tabfm_download | table | Download the TabFM model weights for a task ('classification' or 'regression') from Hugging Face into the local cache. Requires SET anofox_tabfm_accept_hf_license = true. Returns one row per file (file, url, bytes, status). | NULL | [CALL tabfm_download('classification');] |
| anofox_tabfm_download_runtime | table | Download a backend plugin library ('cuda', 'rocm' or 'mlx') into the directory named by SET anofox_tabfm_ep_path (default: the cache dir's 'runtime' subdirectory), so that device can be driven without a matching compile-time flavor build. Returns one row per extracted file (file, bytes, status). | NULL | [CALL tabfm_download_runtime('cuda');] |
| anofox_tabfm_ece | aggregate | Compute Expected Calibration Error (ECE) using 10 equal-width bins over [0,1]. For each row, extracts argmax class and max probability from the proba MAP. ECE = sum_m (|B_m|/N)*|accuracy(B_m) - confidence(B_m)|. NULL rows skipped. Reference: Guo et al. 2017 (On Calibration of Modern Neural Networks). | NULL | [SELECT tabfm_ece(actual, proba) FROM predictions;] |
| anofox_tabfm_f1 | aggregate | Compute F1-score with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined per-class precision/recall returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.f1_score(average=avg, zero_division=0). | NULL | [SELECT tabfm_f1(actual, predicted, 'macro') FROM predictions;, SELECT tabfm_f1(actual, predicted, 'weighted') FROM predictions;] |
| anofox_tabfm_fold_assign | table_macro | Assign each row of data to one of k deterministic folds using (hash(row_key, seed) % k)::INTEGER. Output: all input columns plus fold_id INTEGER in [0, k-1]. Deterministic and order-independent — only the row_key expression determines the fold, not row position. hash() is stable within a DuckDB version; fold assignments change on DuckDB upgrade (Assumption A6). CV-01 compliant: not row_number()-based. |
NULL | [SELECT id, label, fold_id FROM tabfm_fold_assign('my_table', 5, 'id', 42);] |
| anofox_tabfm_generate | table_macro | Generate synthetic rows from the joint distribution of data using a tabular foundation model. Factorizes the table column by column (the chain rule) and samples each column conditioned on the ones already generated, so correlations between columns are preserved rather than each column being drawn independently. Returns n rows with the same columns as data plus synthetic_id. Continuous columns are sampled via quantile bins, so values stay inside the observed range. Costs one model call per column, run sequentially. Options: seed, temperature (higher = more diverse), bins, column_order, model. |
NULL | [SELECT * FROM tabfm_generate('customers', 100);] |
| anofox_tabfm_gpu_precompile | table | Warm the GPU path for a task by compiling the model for a shape bucket ahead of the first predict (on ROCm this builds and caches the .mxr program; a no-op cost on CPU/CUDA). Returns task, rows, features, device, status. | NULL | [CALL tabfm_gpu_precompile('classification', 1000, 50);] |
| anofox_tabfm_impute | table_macro | Fill the NULL cells of data with a tabular foundation model, conditioning each missing value on the other columns of its row. Returns the same columns as data, with non-NULL cells untouched, so it round-trips: CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw'). Unlike tabfm_generate this does not sample — it takes the conditional best estimate (classification argmax, regression point estimate), so continuous columns keep full precision. Optional columns restricts which columns are filled; opts accepts seed, rounds (MICE-style refinement sweeps), model. |
NULL | [SELECT * FROM tabfm_impute('customers', columns := ['income']);] |
| anofox_tabfm_interval_score | aggregate | Compute the mean Gneiting & Raftery (2007) interval score at a given coverage level. IS = (u-l) + (2/(1-coverage)) * underflow + overflow penalties. Lower is better. The coverage parameter (default 0.9) is the nominal coverage — must be strictly between 0 and 1. The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. |
NULL | [SELECT tabfm_interval_score(actual, yhat_dist, coverage := 0.9) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| anofox_tabfm_list_models | table | List every model in the registry (built-ins + user manifests), downloaded or not: model, family, capabilities, license, commercial, size regime (max_rows/features/classes), downloaded. | NULL | [SELECT * FROM tabfm_list_models();] |
| anofox_tabfm_load | table | Eagerly load a downloaded TabFM model for a task into memory so the first predict is warm (otherwise the model loads lazily on first use). | NULL | [CALL tabfm_load('classification');] |
| anofox_tabfm_log_loss | aggregate | Compute log-loss (cross-entropy) from a per-class probability MAP(VARCHAR,DOUBLE). Probabilities are clipped to [1e-15, 1-1e-15] (sklearn eps default) to avoid infinity. Missing MAP keys are treated as p=0 then clipped. NULL rows skipped. Matches sklearn.metrics.log_loss(normalize=True, eps=1e-15). | NULL | [SELECT tabfm_log_loss(actual, proba) FROM predictions;] |
| anofox_tabfm_log_score | aggregate | Compute the mean negative log-likelihood (NLL / log-score) of actual values under the predictive bar distribution. Log-score is a proper scoring rule — lower is better. For y in bin i: NLL = log(w_i) - log(max(1e-10, p_i)). Out-of-support or zero-width bins return the max penalty -log(1e-10) (T-03-04). The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. | NULL | [SELECT tabfm_log_score(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| anofox_tabfm_mae | aggregate | Compute mean absolute error (MAE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_error. | NULL | [SELECT tabfm_mae(actual, predicted) FROM predictions;, SELECT round(tabfm_mae(actual, predicted), 4) FROM predictions;] |
| anofox_tabfm_mape | aggregate | Compute mean absolute percentage error (MAPE) as a DIMENSIONLESS RATIO over (actual, predicted) pairs. Rows where actual == 0 are SKIPPED to avoid division by zero (documented behavior); returns NULL when all actuals are zero. Rows where actual OR predicted is NULL are also skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_percentage_error (fraction, not percent). | NULL | [SELECT tabfm_mape(actual, predicted) FROM predictions;, SELECT round(tabfm_mape(actual, predicted) * 100, 2) || '%' FROM predictions;] |
| anofox_tabfm_medae | aggregate | Compute median absolute error (MedAE) over (actual, predicted) pairs. Accumulates |actual - predicted| residuals in a heap-allocated vector (O(N) memory). For even N, returns the mean of the two middle residuals. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.median_absolute_error. | NULL | [SELECT tabfm_medae(actual, predicted) FROM predictions;, SELECT round(tabfm_medae(actual, predicted), 4) FROM predictions;] |
| anofox_tabfm_models | table | List the TabFM models known to the local cache (model, task, revision, path, bytes, loaded, license). | NULL | [SELECT * FROM tabfm_models();] |
| anofox_tabfm_precision | aggregate | Compute precision with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined precision (no predicted positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.precision_score(zero_division=0). | NULL | [SELECT tabfm_precision(actual, predicted, 'macro') FROM predictions;] |
| anofox_tabfm_r2 | aggregate | Compute the coefficient of determination R² = 1 - SS_res/SS_tot over (actual, predicted) pairs. Constant-target convention (matches sklearn): when the target variance is 0, returns 1.0 for a perfect prediction and 0.0 otherwise — never NaN or Inf. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.r2_score. | NULL | [SELECT tabfm_r2(actual, predicted) FROM predictions;, SELECT round(tabfm_r2(actual, predicted), 4) FROM predictions;] |
| anofox_tabfm_recall | aggregate | Compute recall with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined recall (no actual positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.recall_score(zero_division=0). | NULL | [SELECT tabfm_recall(actual, predicted, 'macro') FROM predictions;] |
| anofox_tabfm_register_model | table | Register a model in SQL (no manifest file). Named args: id, classification_graph / regression_graph (path or url to the weight-free ONNX graph), classification_weights / regression_weights, tensor_map (or classification_tensor_map / regression_tensor_map), weights_repo, license, commercial, gate_setting, preprocessing_profile, max_rows / max_features / max_classes. Then use model := ' |
NULL | [CALL tabfm_register_model(id := 'my', classification_graph := '/p/g.onnx', classification_weights := '/p/w.safetensors', tensor_map := '/p/map.json', license := 'apache-2.0');] |
| anofox_tabfm_regress | table_macro | Zero-shot tabular regression with the TabFM foundation model. Uses the rows of data with a known numeric target as in-context examples to predict the target for rows where it is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat (yhat_score is NULL for regression). Optional features restricts the feature columns; opts is a MAP of options. |
NULL | [SELECT * FROM tabfm_regress('sold_homes', 'price', test := 'listings');] |
| anofox_tabfm_remove | table | Delete a downloaded TabFM model's weights from the local cache (by task, optionally a specific revision). | NULL | [CALL tabfm_remove('classification');] |
| anofox_tabfm_rmse | aggregate | Compute root-mean-squared error (RMSE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.root_mean_squared_error. | NULL | [SELECT tabfm_rmse(actual, predicted) FROM predictions;, SELECT round(tabfm_rmse(actual, predicted), 4) FROM predictions;] |
| anofox_tabfm_roc_auc | aggregate | Compute ROC-AUC with multiclass support and correct rank-sum tie handling. avg is REQUIRED: 'ovr' (one-vs-rest, macro average) or 'ovo' (one-vs-one, macro average). Tied scores handled via the Mann-Whitney U rank-sum form (no optimistic bias). NULL rows skipped. Matches sklearn.metrics.roc_auc_score(multi_class='ovr'/'ovo'). | NULL | [SELECT tabfm_roc_auc(actual, proba, 'ovr') FROM predictions;] |
| anofox_tabfm_unload | table | Unload a loaded TabFM model from memory (all models if no task is given), freeing its RAM/VRAM. | NULL | [CALL tabfm_unload('classification');] |
| anofox_tabfm_unregister_model | table | Remove a model registered with tabfm_register_model. Returns model, status. | NULL | [CALL tabfm_unregister_model('my_model');] |
| tabfm_accelerate | table | Set up GPU acceleration in one call: find the accelerator, fetch its backend plugin (and runtime), and verify it loads. Returns one row per step (step, status, detail). Idempotent — re-running re-verifies without re-downloading. Reconnect afterwards for it to take effect. | NULL | [CALL tabfm_accelerate();] |
| tabfm_accuracy | aggregate | Compute classification accuracy: correct / total as DOUBLE over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.accuracy_score(normalize=True). | NULL | [SELECT tabfm_accuracy(actual, predicted) FROM predictions;, SELECT round(tabfm_accuracy(actual, predicted), 4) FROM predictions;] |
| tabfm_backends | table | Which discovered device can serve which model, and the reason when one cannot (model, task, device, backend, supported, reason). Built on the same servability predicate dispatch uses, so it cannot promise what a predict would refuse. 'supported' is NULL when the answer is not yet knowable — a bundled GPU graph cannot be matched against weights that have not been downloaded. | NULL | [SELECT * FROM tabfm_backends() WHERE NOT supported;] |
| tabfm_classify | table_macro | Zero-shot tabular classification with the TabFM foundation model. Uses the labelled rows of data as in-context examples to score the rows whose target is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat, yhat_score, is_training and (detail mode) a proba MAP. Optional features restricts the feature columns; opts is a MAP of options (seed, softmax_temperature, output_mode, …). |
NULL | [SELECT age, plan, yhat, yhat_score FROM tabfm_classify('customers', 'churned') WHERE churned IS NULL;] |
| tabfm_confusion_matrix | table_macro | Compute a tidy confusion matrix from a table, returning (actual, predicted, count) rows for each observed (actual class, predicted class) pair. Column identifiers are safely double-quote escaped (T-01-02-01). Returns one row per distinct (actual, predicted) combination, ordered by actual then predicted. | NULL | [SELECT * FROM tabfm_confusion_matrix('predictions', 'actual', 'predicted');] |
| tabfm_cross_validate | table_macro | Run leakage-safe k-fold cross-validation in SQL. Assigns rows to k folds via hash(row_key, seed) % k, then for each fold trains tabfm_classify (or tabfm_regress) on the other k-1 folds and predicts the held-out fold via the two-table form. Returns k per-fold rows (fold_id 0..k-1) plus one aggregate row (fold_id=-1) with mean and stddev of the per-fold metric. Default metric: tabfm_accuracy (classification) or tabfm_rmse (regression). Target identifier is safely double-quoted to prevent SQL injection (CV-04). CV-01: hash-based fold assignment (deterministic, order-independent, seed-sensitive). CV-02: leakage-safe — held-out rows arrive as NULL-label test rows, never as training context. CV-03: per-fold rows + aggregate mean±std row. CV-04: safe target quoting. | NULL | [SELECT * FROM tabfm_cross_validate('patients', 'diagnosis', 'patient_id', k := 5, seed := 42);] |
| tabfm_crps | aggregate | Compute the mean Continuous Ranked Probability Score (CRPS) over regression predictions. CRPS is a proper scoring rule for predictive distributions — lower is better. The second argument must be the yhat_dist STRUCT output of tabfm_regress with output_mode := 'distribution'. Non-uniform bin widths from borders[i+1]-borders[i] are used (PSR-01). NULL rows are skipped; empty input returns NULL. | NULL | [SELECT tabfm_crps(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| tabfm_devices | table | List the inference devices this build can see (device_id, ep, name, arch, vram, driver, usable). The cpu row always exists; GPU rows appear only in the matching flavor (cuda/rocm) and report usable=false when a device is present but unsupported. | NULL | [SELECT * FROM tabfm_devices();] |
| tabfm_download | table | Download the TabFM model weights for a task ('classification' or 'regression') from Hugging Face into the local cache. Requires SET anofox_tabfm_accept_hf_license = true. Returns one row per file (file, url, bytes, status). | NULL | [CALL tabfm_download('classification');] |
| tabfm_download_runtime | table | Download a backend plugin library ('cuda', 'rocm' or 'mlx') into the directory named by SET anofox_tabfm_ep_path (default: the cache dir's 'runtime' subdirectory), so that device can be driven without a matching compile-time flavor build. Returns one row per extracted file (file, bytes, status). | NULL | [CALL tabfm_download_runtime('cuda');] |
| tabfm_ece | aggregate | Compute Expected Calibration Error (ECE) using 10 equal-width bins over [0,1]. For each row, extracts argmax class and max probability from the proba MAP. ECE = sum_m (|B_m|/N)*|accuracy(B_m) - confidence(B_m)|. NULL rows skipped. Reference: Guo et al. 2017 (On Calibration of Modern Neural Networks). | NULL | [SELECT tabfm_ece(actual, proba) FROM predictions;] |
| tabfm_f1 | aggregate | Compute F1-score with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined per-class precision/recall returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.f1_score(average=avg, zero_division=0). | NULL | [SELECT tabfm_f1(actual, predicted, 'macro') FROM predictions;, SELECT tabfm_f1(actual, predicted, 'weighted') FROM predictions;] |
| tabfm_fold_assign | table_macro | Assign each row of data to one of k deterministic folds using (hash(row_key, seed) % k)::INTEGER. Output: all input columns plus fold_id INTEGER in [0, k-1]. Deterministic and order-independent — only the row_key expression determines the fold, not row position. hash() is stable within a DuckDB version; fold assignments change on DuckDB upgrade (Assumption A6). CV-01 compliant: not row_number()-based. |
NULL | [SELECT id, label, fold_id FROM tabfm_fold_assign('my_table', 5, 'id', 42);] |
| tabfm_generate | table_macro | Generate synthetic rows from the joint distribution of data using a tabular foundation model. Factorizes the table column by column (the chain rule) and samples each column conditioned on the ones already generated, so correlations between columns are preserved rather than each column being drawn independently. Returns n rows with the same columns as data plus synthetic_id. Continuous columns are sampled via quantile bins, so values stay inside the observed range. Costs one model call per column, run sequentially. Options: seed, temperature (higher = more diverse), bins, column_order, model. |
NULL | [SELECT * FROM tabfm_generate('customers', 100);] |
| tabfm_gpu_precompile | table | Warm the GPU path for a task by compiling the model for a shape bucket ahead of the first predict (on ROCm this builds and caches the .mxr program; a no-op cost on CPU/CUDA). Returns task, rows, features, device, status. | NULL | [CALL tabfm_gpu_precompile('classification', 1000, 50);] |
| tabfm_impute | table_macro | Fill the NULL cells of data with a tabular foundation model, conditioning each missing value on the other columns of its row. Returns the same columns as data, with non-NULL cells untouched, so it round-trips: CREATE TABLE clean AS SELECT * FROM tabfm_impute('raw'). Unlike tabfm_generate this does not sample — it takes the conditional best estimate (classification argmax, regression point estimate), so continuous columns keep full precision. Optional columns restricts which columns are filled; opts accepts seed, rounds (MICE-style refinement sweeps), model. |
NULL | [SELECT * FROM tabfm_impute('customers', columns := ['income']);] |
| tabfm_interval_score | aggregate | Compute the mean Gneiting & Raftery (2007) interval score at a given coverage level. IS = (u-l) + (2/(1-coverage)) * underflow + overflow penalties. Lower is better. The coverage parameter (default 0.9) is the nominal coverage — must be strictly between 0 and 1. The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. |
NULL | [SELECT tabfm_interval_score(actual, yhat_dist, coverage := 0.9) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| tabfm_list_models | table | List every model in the registry (built-ins + user manifests), downloaded or not: model, family, capabilities, license, commercial, size regime (max_rows/features/classes), downloaded. | NULL | [SELECT * FROM tabfm_list_models();] |
| tabfm_load | table | Eagerly load a downloaded TabFM model for a task into memory so the first predict is warm (otherwise the model loads lazily on first use). | NULL | [CALL tabfm_load('classification');] |
| tabfm_log_loss | aggregate | Compute log-loss (cross-entropy) from a per-class probability MAP(VARCHAR,DOUBLE). Probabilities are clipped to [1e-15, 1-1e-15] (sklearn eps default) to avoid infinity. Missing MAP keys are treated as p=0 then clipped. NULL rows skipped. Matches sklearn.metrics.log_loss(normalize=True, eps=1e-15). | NULL | [SELECT tabfm_log_loss(actual, proba) FROM predictions;] |
| tabfm_log_score | aggregate | Compute the mean negative log-likelihood (NLL / log-score) of actual values under the predictive bar distribution. Log-score is a proper scoring rule — lower is better. For y in bin i: NLL = log(w_i) - log(max(1e-10, p_i)). Out-of-support or zero-width bins return the max penalty -log(1e-10) (T-03-04). The second argument must be the yhat_dist STRUCT from tabfm_regress with output_mode := 'distribution'. NULL rows are skipped; empty input returns NULL. | NULL | [SELECT tabfm_log_score(actual, yhat_dist) FROM tabfm_regress('tbl', 'y', opts := MAP{'output_mode':'distribution'});] |
| tabfm_mae | aggregate | Compute mean absolute error (MAE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_error. | NULL | [SELECT tabfm_mae(actual, predicted) FROM predictions;, SELECT round(tabfm_mae(actual, predicted), 4) FROM predictions;] |
| tabfm_mape | aggregate | Compute mean absolute percentage error (MAPE) as a DIMENSIONLESS RATIO over (actual, predicted) pairs. Rows where actual == 0 are SKIPPED to avoid division by zero (documented behavior); returns NULL when all actuals are zero. Rows where actual OR predicted is NULL are also skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.mean_absolute_percentage_error (fraction, not percent). | NULL | [SELECT tabfm_mape(actual, predicted) FROM predictions;, SELECT round(tabfm_mape(actual, predicted) * 100, 2) || '%' FROM predictions;] |
| tabfm_medae | aggregate | Compute median absolute error (MedAE) over (actual, predicted) pairs. Accumulates |actual - predicted| residuals in a heap-allocated vector (O(N) memory). For even N, returns the mean of the two middle residuals. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.median_absolute_error. | NULL | [SELECT tabfm_medae(actual, predicted) FROM predictions;, SELECT round(tabfm_medae(actual, predicted), 4) FROM predictions;] |
| tabfm_models | table | List the TabFM models known to the local cache (model, task, revision, path, bytes, loaded, license). | NULL | [SELECT * FROM tabfm_models();] |
| tabfm_precision | aggregate | Compute precision with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined precision (no predicted positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.precision_score(zero_division=0). | NULL | [SELECT tabfm_precision(actual, predicted, 'macro') FROM predictions;] |
| tabfm_r2 | aggregate | Compute the coefficient of determination R² = 1 - SS_res/SS_tot over (actual, predicted) pairs. Constant-target convention (matches sklearn): when the target variance is 0, returns 1.0 for a perfect prediction and 0.0 otherwise — never NaN or Inf. Rows where actual OR predicted is NULL are skipped. Returns NULL on empty or all-NULL input. Matches sklearn.metrics.r2_score. | NULL | [SELECT tabfm_r2(actual, predicted) FROM predictions;, SELECT round(tabfm_r2(actual, predicted), 4) FROM predictions;] |
| tabfm_recall | aggregate | Compute recall with the specified averaging mode. avg is REQUIRED: pass 'macro', 'micro', or 'weighted'. Undefined recall (no actual positives for a class) returns 0.0 (zero_division=0). NULL rows skipped. Matches sklearn.metrics.recall_score(zero_division=0). | NULL | [SELECT tabfm_recall(actual, predicted, 'macro') FROM predictions;] |
| tabfm_register_model | table | Register a model in SQL (no manifest file). Named args: id, classification_graph / regression_graph (path or url to the weight-free ONNX graph), classification_weights / regression_weights, tensor_map (or classification_tensor_map / regression_tensor_map), weights_repo, license, commercial, gate_setting, preprocessing_profile, max_rows / max_features / max_classes. Then use model := ' |
NULL | [CALL tabfm_register_model(id := 'my', classification_graph := '/p/g.onnx', classification_weights := '/p/w.safetensors', tensor_map := '/p/map.json', license := 'apache-2.0');] |
| tabfm_regress | table_macro | Zero-shot tabular regression with the TabFM foundation model. Uses the rows of data with a known numeric target as in-context examples to predict the target for rows where it is NULL (single-relation form) or every row of the test relation (train/test form). Returns one row per scored row with yhat (yhat_score is NULL for regression). Optional features restricts the feature columns; opts is a MAP of options. |
NULL | [SELECT * FROM tabfm_regress('sold_homes', 'price', test := 'listings');] |
| tabfm_remove | table | Delete a downloaded TabFM model's weights from the local cache (by task, optionally a specific revision). | NULL | [CALL tabfm_remove('classification');] |
| tabfm_rmse | aggregate | Compute root-mean-squared error (RMSE) over (actual, predicted) pairs. Rows where actual OR predicted is NULL are skipped (SQL aggregate NULL semantics). Returns NULL on empty or all-NULL input. Matches sklearn.metrics.root_mean_squared_error. | NULL | [SELECT tabfm_rmse(actual, predicted) FROM predictions;, SELECT round(tabfm_rmse(actual, predicted), 4) FROM predictions;] |
| tabfm_roc_auc | aggregate | Compute ROC-AUC with multiclass support and correct rank-sum tie handling. avg is REQUIRED: 'ovr' (one-vs-rest, macro average) or 'ovo' (one-vs-one, macro average). Tied scores handled via the Mann-Whitney U rank-sum form (no optimistic bias). NULL rows skipped. Matches sklearn.metrics.roc_auc_score(multi_class='ovr'/'ovo'). | NULL | [SELECT tabfm_roc_auc(actual, proba, 'ovr') FROM predictions;] |
| tabfm_unload | table | Unload a loaded TabFM model from memory (all models if no task is given), freeing its RAM/VRAM. | NULL | [CALL tabfm_unload('classification');] |
| tabfm_unregister_model | table | Remove a model registered with tabfm_register_model. Returns model, status. | NULL | [CALL tabfm_unregister_model('my_model');] |
Overloaded Functions
This extension does not add any function overloads.
Added Types
This extension does not add any types.
Added Settings
| name | description | input_type | scope | aliases |
|---|---|---|---|---|
| anofox_tabfm_accept_hf_license | Accept the upstream model license (tabfm-non-commercial-v1.0: non-commercial use, no redistribution). Downloads of Google-licensed weights fail without this. | BOOLEAN | GLOBAL | [] |
| anofox_tabfm_cache_dir | Weight cache root directory (default ~/.cache/anofox-tabfm) | VARCHAR | GLOBAL | [] |
| anofox_tabfm_context_cache | Encode the labelled context once and reuse it across calls, for models that ship a split graph pair (prepare/query). Off by default. It pays off when the same context is scored more than once – chunked scoring, repeated queries against a fixed training table – and costs extra on a single call, which pays for the context it will not reuse. Test-row predictions match the combined graph; the fitted values on CONTEXT rows differ, because the query half has no label path and so no longer sees a context row's own label. Inert for a model that ships no pair. | BOOLEAN | GLOBAL | [] |
| anofox_tabfm_cpu_prepack | Enable ONNX Runtime weight prepacking on the CPU EP: faster matmuls at ~+16% resident memory. | BOOLEAN | GLOBAL | [] |
| anofox_tabfm_default_model | Default model id for tabfm_classify/regress/download/… when model := is not given. '' = resolve to the single-file manifest model, else the sole registered model. | VARCHAR | GLOBAL | [] |
| anofox_tabfm_device | Execution device: auto|cpu|cuda|rocm|coreml|mlx ('migraphx' alias). 'auto' resolves per model: it picks the best device that model can actually be served on (its plugin present and a graph available), else the CPU — so it never promises an accelerator a model cannot use. SELECT * FROM tabfm_backends() shows the choice and the reason. Naming a device explicitly is a hard request: it errors rather than falling back. coreml is flavor-gated and 'auto' does not reach for it. | VARCHAR | GLOBAL | [] |
| anofox_tabfm_ep_path | Directory holding the GPU backend plugins (anofox_tabfm_{cuda,migraphx,mlx}_plugin, with this platform's library prefix and extension) and the runtime libraries they load alongside themselves. Defaults to the 'runtime' subdirectory of anofox_tabfm_cache_dir, which is where CALL tabfm_download_runtime(…) and CALL tabfm_accelerate() put them — set this only to point somewhere else. | VARCHAR | GLOBAL | [] |
| anofox_tabfm_gpu_precision | GPU numeric mode: fp32|tf32|bf16|fp16. fp32 (default) is strict — a device switch does not change answers, measured exact on both GPUs; on CUDA it disables TF32 tensor-core rounding. tf32 re-enables that rounding (CUDA only; fp32 storage, faster matmuls). bf16/fp16 quantize the MIGraphX program on ROCm (~2x faster on RDNA4, half the VRAM/.mxr; label flips vs fp32 are rare (~2%% measured) but NOT confined to near-ties — measured flips include high-confidence rows, so validate bf16 on your own data) and are rejected on CUDA rather than silently running fp32. | VARCHAR | GLOBAL | [] |
| anofox_tabfm_max_features | Maximum feature columns per predict call | BIGINT | GLOBAL | [] |
| anofox_tabfm_max_memory | Refuse a predict call when this process's resident memory is already at or above this size (e.g. '16GB') before the call starts, so the failure is a DuckDB exception instead of a cgroup OOM-kill. '' (default) disables the check. Checked against resident memory at call time, not an estimate of the call's own cost – it does not bound how much a single large call can grow memory by itself. | VARCHAR | GLOBAL | [] |
| anofox_tabfm_max_rows | Maximum rows per predict call or group | BIGINT | GLOBAL | [] |
| anofox_tabfm_max_sessions | Cap on cached model sessions across all models, devices and precisions (default 4); beyond it the oldest-loaded session is evicted. Sessions are cached per (model, device, precision) so switching devices does not rebuild multi-GB sessions, and this cap keeps that from accumulating unbounded memory. | BIGINT | GLOBAL | [] |
| anofox_tabfm_mxr_source | Directory holding precompiled MIGraphX .mxr programs (offline/CI/shared cache). Before compiling a shape-bucket (~27 min on ROCm), a matching ' |
VARCHAR | GLOBAL | [] |
| anofox_tabfm_threads | ONNX Runtime intra-op thread count for CPU inference | BIGINT | GLOBAL | [] |
| anofox_tabfm_trace_level | Diagnostic verbosity: error|warn|info|debug|trace | VARCHAR | GLOBAL | [] |
| anofox_telemetry_enabled | Enable or disable anonymous usage telemetry | BOOLEAN | GLOBAL | [] |
| anofox_telemetry_key | PostHog API key for telemetry | VARCHAR | GLOBAL | [] |
| datazoo_banner | Show the DataZoo feedback banner when an extension is loaded in an interactive terminal (at most once a day per extension). | BOOLEAN | GLOBAL | [] |