Search Shortcut cmd + k | ctrl + k
anofox_statistics

A DuckDB extension for statistical regression and inference in SQL - OLS, Ridge, Elastic Net, LARS/LassoLars, WLS, recursive least squares, robust estimators (Huber, RANSAC, Theil-Sen), GLMs (Poisson, Negative Binomial, Binomial, Tweedie, Gamma, Logistic) with offset support, mixed-effects/hierarchical GLMs with random slopes and crossed/nested factors, AFT survival regression, explicit priors with Laplace intervals, empirical-Bayes shrinkage, and time-series regression, all with full diagnostics and inference.

Maintainer(s): sipemu

Installing and Loading

INSTALL anofox_statistics FROM community;
LOAD anofox_statistics;

Added Functions

function_name function_type description comment examples
aft_cdf scalar NULL NULL  
aft_fit_agg aggregate NULL NULL  
aft_quantile scalar NULL NULL  
aic scalar Computes Akaike Information Criterion (AIC) from residual sum of squares, number of observations, and number of parameters. NULL [aic(rss, n, k)]
aid_agg aggregate Classifies demand patterns (smooth, intermittent, erratic, lumpy) using Automatic Identification of Demand (AID). NULL [aid_agg(y)]
aid_agg aggregate Classifies demand patterns using AID with a MAP of options (intermittent_threshold, outlier_method). NULL [aid_agg(y, {'intermittent_threshold': 0.3})]
aid_anomaly_agg aggregate Identifies anomalies in demand time series using AID with a MAP of options (intermittent_threshold, outlier_method). NULL [aid_anomaly_agg(y, {'outlier_method': 'iqr'})]
aid_anomaly_agg aggregate Identifies anomalies in demand time series using the AID classification framework. NULL [aid_anomaly_agg(y)]
aid_anomaly_by table_macro NULL NULL  
aid_by table_macro NULL NULL  
alm_fit_agg aggregate Fits an Additive Linear Model (ALM) and returns coefficients and fit statistics. NULL [alm_fit_agg(y, x)]
alm_fit_agg aggregate Fits an Additive Linear Model (ALM) and returns coefficients and fit statistics. NULL [alm_fit_agg(y, x, {'fit_intercept': true})]
alm_fit_predict_agg aggregate Fits an Additive Linear Model on training rows with a MAP of options and predicts all rows. NULL [alm_fit_predict_agg(y, x, split_col, {'distribution': 'laplace'})]
alm_fit_predict_agg aggregate Fits an Additive Linear Model over a partition and returns per-row predictions. NULL [alm_fit_predict_agg(y, x)]
alm_fit_predict_agg aggregate Fits an Additive Linear Model over a partition with a MAP of options and returns per-row predictions. NULL [alm_fit_predict_agg(y, x, {'distribution': 'laplace'})]
alm_fit_predict_agg aggregate Fits an Additive Linear Model using only training rows (split_col='train') and predicts all rows. NULL [alm_fit_predict_agg(y, x, split_col)]
alm_fit_predict_by table_macro NULL NULL  
bic scalar Computes Bayesian Information Criterion (BIC) from residual sum of squares, number of observations, and number of parameters. NULL [bic(rss, n, k)]
binom_test_agg aggregate Performs an exact binomial test comparing an observed success count to a hypothesized probability, using default options. NULL [binom_test_agg(value)]
binom_test_agg aggregate Performs an exact binomial test comparing an observed success count to a hypothesized probability. NULL [binom_test_agg(value, {'p0': 0.5, 'alternative': 'two_sided'})]
binomial_fit_agg aggregate Fits a Binomial GLM (logit link by default) and returns coefficients, deviance, AIC, dispersion, and fit statistics. y is the success rate in [0, 1]. NULL [binomial_fit_agg(y, x)]
binomial_fit_agg aggregate Fits a Binomial GLM (user-selectable link: logit / probit / cloglog; default logit) and returns coefficients, deviance, AIC, dispersion (= 1 for canonical binomial), and fit statistics. y is the success rate in [0, 1]. NULL [binomial_fit_agg(y, x, {'binomial_link': 'logit', 'fit_intercept': true})]
bls_fit_agg aggregate Fits a Bounded Least Squares (BLS) regression with coefficient bounds and returns fit statistics. NULL [bls_fit_agg(y, x)]
bls_fit_agg aggregate Fits a Bounded Least Squares (BLS) regression with coefficient bounds and returns fit statistics. NULL [bls_fit_agg(y, x, {'lower_bound': 0.0, 'upper_bound': 1.0})]
bls_fit_predict_agg aggregate Fits a Bounded Least Squares model on training rows with a MAP of options and predicts all rows. NULL [bls_fit_predict_agg(y, x, split_col, {'null_policy': 'drop'})]
bls_fit_predict_agg aggregate Fits a Bounded Least Squares model over a partition and returns per-row predictions. NULL [bls_fit_predict_agg(y, x)]
bls_fit_predict_agg aggregate Fits a Bounded Least Squares model over a partition with a MAP of options and returns per-row predictions. NULL [bls_fit_predict_agg(y, x, {'null_policy': 'drop'})]
bls_fit_predict_agg aggregate Fits a Bounded Least Squares model using only training rows (split_col='train') and predicts all rows. NULL [bls_fit_predict_agg(y, x, split_col)]
bls_fit_predict_by table_macro NULL NULL  
brown_forsythe_agg aggregate Tests equality of variances across groups using the Brown-Forsythe test. NULL [brown_forsythe_agg(value, group_id)]
brunner_munzel_agg aggregate Performs the Brunner-Munzel test for stochastic equality of two independent samples, using default options. NULL [brunner_munzel_agg(value, group_id)]
brunner_munzel_agg aggregate Performs the Brunner-Munzel test for stochastic equality of two independent samples. NULL [brunner_munzel_agg(value, group_id, {'alternative': 'two_sided'})]
chisq_gof_agg aggregate Performs a chi-squared goodness-of-fit test comparing observed frequencies to expected probabilities. NULL [chisq_gof_agg(observed, expected_prob)]
chisq_test_agg aggregate Performs a chi-squared test of independence on a 2×2 contingency table from two categorical columns. NULL [chisq_test_agg(row_var, col_var)]
chisq_test_agg aggregate Performs a chi-squared test of independence on a 2×2 contingency table from two categorical columns. NULL [chisq_test_agg(row_var, col_var, {'correction': true})]
clark_west_agg aggregate Performs the Clark-West test to compare a nested forecast model against an encompassing model, using default options. NULL [clark_west_agg(actual, forecast_restricted, forecast_unrestricted)]
clark_west_agg aggregate Performs the Clark-West test to compare a nested forecast model against an encompassing model. NULL [clark_west_agg(actual, forecast_restricted, forecast_unrestricted, {'horizon': 1})]
cohen_kappa_agg aggregate Computes Cohen's kappa, a measure of inter-rater agreement for categorical classifications. NULL [cohen_kappa_agg(rater1, rater2)]
cohen_kappa_agg aggregate Computes Cohen's kappa, a measure of inter-rater agreement for categorical classifications. NULL [cohen_kappa_agg(rater1, rater2, {'weighted': false})]
contingency_coef_agg aggregate Computes the contingency coefficient (C), a measure of association for categorical variables. NULL [contingency_coef_agg(row_var, col_var)]
cramers_v_agg aggregate Computes Cramér's V, a measure of association strength for nominal categorical variables. NULL [cramers_v_agg(row_var, col_var)]
dagostino_k2_agg aggregate Performs the D'Agostino-Pearson K² omnibus normality test based on skewness and kurtosis. NULL [dagostino_k2_agg(value)]
diebold_mariano_agg aggregate Performs the Diebold-Mariano test to compare predictive accuracy of two forecast models, using default options. NULL [diebold_mariano_agg(actual, forecast1, forecast2)]
diebold_mariano_agg aggregate Performs the Diebold-Mariano test to compare predictive accuracy of two forecast models. NULL [diebold_mariano_agg(actual, forecast1, forecast2, {'loss': 'squared'})]
distance_cor_agg aggregate Computes the distance correlation between two variables, detecting both linear and nonlinear dependence, using default options. NULL [distance_cor_agg(x, y)]
distance_cor_agg aggregate Computes the distance correlation between two variables, detecting both linear and nonlinear dependence. NULL [distance_cor_agg(x, y, {'n_permutations': 1000})]
eb_shrink_agg aggregate NULL NULL  
eb_shrink_by table_macro NULL NULL  
elasticnet_fit scalar Fits an ElasticNet regression model (L1+L2 regularization) with optional MAP of settings (fit_intercept, alpha, l1_ratio, max_iterations, tolerance). NULL [elasticnet_fit(y, x, {'alpha': 1.0, 'l1_ratio': 0.5})]
elasticnet_fit scalar Fits an ElasticNet regression model combining L1 and L2 regularization to the given response and feature data. NULL [elasticnet_fit(y, x)]
elasticnet_fit_agg aggregate Fits an ElasticNet model combining L1 and L2 regularization and returns coefficients and fit statistics. NULL [elasticnet_fit_agg(y, x)]
elasticnet_fit_agg aggregate Fits an ElasticNet model combining L1 and L2 regularization and returns coefficients and fit statistics. NULL [elasticnet_fit_agg(y, x, {'alpha': 1.0, 'l1_ratio': 0.5})]
elasticnet_fit_predict aggregate Fits an ElasticNet regression model over a window partition and returns predictions with confidence intervals. NULL [elasticnet_fit_predict(y, x)]
elasticnet_fit_predict aggregate Fits an ElasticNet regression model over a window partition and returns predictions with confidence intervals. NULL [elasticnet_fit_predict(y, x, {'null_policy': 'drop'})]
elasticnet_fit_predict_agg aggregate Fits ElasticNet regression on training rows with a MAP of options and predicts all rows. NULL [elasticnet_fit_predict_agg(y, x, split_col, {'null_policy': 'drop'})]
elasticnet_fit_predict_agg aggregate Fits ElasticNet regression over a partition and returns per-row predictions with confidence intervals. NULL [elasticnet_fit_predict_agg(y, x)]
elasticnet_fit_predict_agg aggregate Fits ElasticNet regression over a partition with a MAP of options and returns per-row predictions with confidence intervals. NULL [elasticnet_fit_predict_agg(y, x, {'null_policy': 'drop'})]
elasticnet_fit_predict_agg aggregate Fits ElasticNet regression using only training rows (split_col='train') and predicts all rows. NULL [elasticnet_fit_predict_agg(y, x, split_col)]
elasticnet_fit_predict_by table_macro NULL NULL  
energy_distance_agg aggregate Computes the energy distance between two samples as a measure of distributional difference, using default options. NULL [energy_distance_agg(value, group_id)]
energy_distance_agg aggregate Computes the energy distance between two samples as a measure of distributional difference. NULL [energy_distance_agg(value, group_id, {'n_permutations': 1000})]
fisher_exact_agg aggregate Performs Fisher's exact test for association in a 2×2 contingency table. NULL [fisher_exact_agg(row_var, col_var)]
fisher_exact_agg aggregate Performs Fisher's exact test for association in a 2×2 contingency table. NULL [fisher_exact_agg(row_var, col_var, {'alternative': 'two_sided'})]
g_test_agg aggregate Performs a G-test (log-likelihood ratio test) for goodness of fit or independence. NULL [g_test_agg(row_var, col_var)]
gamma_fit_agg aggregate Fits a Gamma GLM (log link, var_power = 2.0 fixed) and returns coefficients, deviance, AIC, dispersion, and fit statistics. NULL [gamma_fit_agg(y, x)]
gamma_fit_agg aggregate Fits a Gamma GLM (log link, var_power = 2.0 fixed) and returns coefficients, deviance, AIC, dispersion, and fit statistics. y must be strictly positive. NULL [gamma_fit_agg(y, x, {'fit_intercept': true})]
glmm_fit_agg aggregate NULL NULL  
glmm_fit_by table_macro NULL NULL  
huber_fit scalar Fits a Huber M-estimator regression model with optional MAP of settings (epsilon, alpha, fit_intercept, compute_inference, confidence_level, max_iterations, tolerance). NULL [huber_fit(y, x, {'epsilon': 1.35, 'alpha': 0.01})]
huber_fit scalar Fits a Huber M-estimator robust regression model. Returns coefficients, fit statistics, the MAD-based scale, and the outlier count as a struct. NULL [huber_fit(y, x)]
huber_fit_agg aggregate Fits a Huber M-estimator robust regression model and returns coefficients, fit statistics, the MAD-based scale, and the outlier count as a struct. NULL [huber_fit_agg(y, x)]
huber_fit_agg aggregate Fits a Huber M-estimator robust regression model and returns coefficients, fit statistics, the MAD-based scale, and the outlier count as a struct. NULL [huber_fit_agg(y, x, {'epsilon': 1.35, 'fit_intercept': true})]
huber_fit_predict aggregate Fits a Huber M-estimator robust regression over a window partition and returns the prediction for the current row with confidence intervals. NULL [huber_fit_predict(y, x) OVER (PARTITION BY g ORDER BY t)]
huber_fit_predict aggregate Fits a Huber M-estimator robust regression over a window with a MAP of options. NULL [huber_fit_predict(y, x, {'epsilon': 1.5}) OVER (…)]
huber_fit_predict_agg aggregate Fits Huber regression on training rows with a MAP of options and predicts all rows. NULL [huber_fit_predict_agg(y, x, split_col, {'epsilon': 1.5})]
huber_fit_predict_agg aggregate Fits Huber regression over a partition with a MAP of options and returns per-row predictions with confidence intervals. NULL [huber_fit_predict_agg(y, x, {'epsilon': 1.35, 'null_policy': 'drop'})]
huber_fit_predict_agg aggregate Fits Huber regression using only training rows (split_col='train') and predicts all rows. NULL [huber_fit_predict_agg(y, x, split_col)]
huber_fit_predict_agg aggregate Fits a Huber M-estimator robust regression over a partition and returns per-row predictions with confidence intervals. NULL [huber_fit_predict_agg(y, x)]
huber_fit_predict_by table_macro NULL NULL  
icc_agg aggregate Computes the Intraclass Correlation Coefficient (ICC) to measure rater or measurement consistency. NULL [icc_agg(value, subject_id, rater_id)]
icc_agg aggregate Computes the Intraclass Correlation Coefficient (ICC) to measure rater or measurement consistency. NULL [icc_agg(value, subject_id, rater_id, {'type': 'single'})]
isotonic_fit_predict_agg aggregate Fits an isotonic regression model on training rows with a MAP of options and predicts all rows. NULL [isotonic_fit_predict_agg(y, x, split_col, {'increasing': true})]
isotonic_fit_predict_agg aggregate Fits an isotonic regression model over a partition and returns per-row predictions. NULL [isotonic_fit_predict_agg(y, x)]
isotonic_fit_predict_agg aggregate Fits an isotonic regression model over a partition with a MAP of options and returns per-row predictions. NULL [isotonic_fit_predict_agg(y, x, {'increasing': true})]
isotonic_fit_predict_agg aggregate Fits an isotonic regression model using only training rows (split_col='train') and predicts all rows. NULL [isotonic_fit_predict_agg(y, x, split_col)]
isotonic_fit_predict_by table_macro NULL NULL  
jarque_bera scalar Tests whether a sample has skewness and kurtosis consistent with a normal distribution (Jarque-Bera test). NULL [jarque_bera(values)]
jarque_bera_agg aggregate Aggregate version of the Jarque-Bera normality test, applied to a column of values. NULL [jarque_bera_agg(value)]
kendall_agg aggregate Computes Kendall's tau rank correlation coefficient and tests its significance. NULL [kendall_agg(x, y)]
kendall_agg aggregate Computes Kendall's tau rank correlation coefficient and tests its significance. NULL [kendall_agg(x, y, {'alternative': 'two_sided'})]
kruskal_wallis_agg aggregate Performs the Kruskal-Wallis H-test, a nonparametric alternative to one-way ANOVA. NULL [kruskal_wallis_agg(value, group_id)]
lars_fit_agg aggregate Fits a Least Angle Regression (LARS / LassoLars) model and returns coefficients and fit statistics. NULL [lars_fit_agg(y, x, {'fit_intercept': true, 'alpha': 0.0})]
lars_fit_agg aggregate Fits a Least Angle Regression (LARS) model and returns coefficients and fit statistics. NULL [lars_fit_agg(y, x)]
logistic_fit_agg aggregate Fits a binary Logistic regression (binomial GLM with logit link; classifier API). y must be binary (0 or 1). Result struct extends the standard GLM shape with accuracy (on training data, at the configured threshold) and the threshold echo. Optional L2 (ridge) penalty. NULL [logistic_fit_agg(y, x, {'glm_lambda': 0.1, 'threshold': 0.5})]
logistic_fit_agg aggregate Fits a binary Logistic regression (binomial GLM with logit link; classifier API). y must be binary (0 or 1). Result struct extends the standard GLM shape with accuracy and the threshold echo. NULL [logistic_fit_agg(y, x)]
mann_whitney_u_agg aggregate Performs the Mann-Whitney U test (Wilcoxon rank-sum) for two independent samples, using default options. NULL [mann_whitney_u_agg(value, group_id)]
mann_whitney_u_agg aggregate Performs the Mann-Whitney U test (Wilcoxon rank-sum) for two independent samples. NULL [mann_whitney_u_agg(value, group_id, {'alternative': 'two_sided'})]
mcnemar_agg aggregate Performs McNemar's test for marginal homogeneity in paired categorical data. NULL [mcnemar_agg(var1, var2)]
mcnemar_agg aggregate Performs McNemar's test for marginal homogeneity in paired categorical data. NULL [mcnemar_agg(var1, var2, {'correction': true})]
mmd_agg aggregate Computes the Maximum Mean Discrepancy (MMD) between two samples to test distributional similarity, using default options. NULL [mmd_agg(value, group_id)]
mmd_agg aggregate Computes the Maximum Mean Discrepancy (MMD) between two samples to test distributional similarity. NULL [mmd_agg(value, group_id, {'n_permutations': 1000})]
negbinom_fit_agg aggregate Fits a Negative Binomial GLM (log link, overdispersion parameter estimated jointly) and returns coefficients, deviance, AIC, dispersion (= alpha), and fit statistics. NULL [negbinom_fit_agg(y, x)]
negbinom_fit_agg aggregate Fits a Negative Binomial GLM (log link, overdispersion parameter estimated jointly) and returns coefficients, deviance, AIC, dispersion (= alpha), and fit statistics. NULL [negbinom_fit_agg(y, x, {'fit_intercept': true})]
nnls_fit_agg aggregate Fits a Non-Negative Least Squares (NNLS) regression with non-negativity constraints. NULL [nnls_fit_agg(y, x)]
nnls_fit_agg aggregate Fits a Non-Negative Least Squares (NNLS) regression with non-negativity constraints. NULL [nnls_fit_agg(y, x, {'tolerance': 1e-6})]
ols_fit scalar Fits an OLS regression model with optional MAP of settings (fit_intercept, compute_inference, confidence_level, solver, hc_type). NULL [ols_fit(y, x, {'compute_inference': true, 'confidence_level': 0.95})]
ols_fit scalar Fits an Ordinary Least Squares (OLS) regression model to the given response and feature data. NULL [ols_fit(y, x)]
ols_fit_agg aggregate Fits an OLS regression model and returns coefficients and fit statistics as a struct. NULL [ols_fit_agg(y, x)]
ols_fit_agg aggregate Fits an OLS regression model and returns coefficients and fit statistics as a struct. NULL [ols_fit_agg(y, x, {'fit_intercept': true})]
ols_fit_predict aggregate Fits an OLS model over a window partition and returns predictions for each row, including confidence intervals. NULL [ols_fit_predict(y, x)]
ols_fit_predict aggregate Fits an OLS model over a window partition and returns predictions for each row, including confidence intervals. NULL [ols_fit_predict(y, x, {'null_policy': 'drop'})]
ols_fit_predict_agg aggregate Fits OLS regression on training rows with a MAP of options and predicts all rows. NULL [ols_fit_predict_agg(y, x, split_col, {'null_policy': 'drop'})]
ols_fit_predict_agg aggregate Fits OLS regression over a partition and returns per-row predictions with confidence intervals. NULL [ols_fit_predict_agg(y, x)]
ols_fit_predict_agg aggregate Fits OLS regression over a partition with a MAP of options and returns per-row predictions with confidence intervals. NULL [ols_fit_predict_agg(y, x, {'null_policy': 'drop'})]
ols_fit_predict_agg aggregate Fits OLS regression using only training rows (split_col='train') and predicts all rows. NULL [ols_fit_predict_agg(y, x, split_col)]
ols_fit_predict_by table_macro NULL NULL  
one_way_anova_agg aggregate Performs a one-way ANOVA F-test to compare means across multiple groups. NULL [one_way_anova_agg(value, group_id)]
pearson_agg aggregate Computes Pearson's product-moment correlation coefficient and tests its significance. NULL [pearson_agg(x, y)]
pearson_agg aggregate Computes Pearson's product-moment correlation coefficient and tests its significance. NULL [pearson_agg(x, y, {'alternative': 'two_sided'})]
permutation_t_test_agg aggregate Performs a permutation-based two-sample t-test using resampling, using default options. NULL [permutation_t_test_agg(value, group_id)]
permutation_t_test_agg aggregate Performs a permutation-based two-sample t-test using resampling. NULL [permutation_t_test_agg(value, group_id, {'alternative': 'two_sided', 'n_permutations': 10000})]
phi_coefficient_agg aggregate Computes the phi coefficient (φ), a measure of association for 2×2 contingency tables. NULL [phi_coefficient_agg(row_var, col_var)]
pls_fit_predict_agg aggregate Fits a Partial Least Squares model on training rows with a MAP of options and predicts all rows. NULL [pls_fit_predict_agg(y, x, split_col, {'n_components': 2})]
pls_fit_predict_agg aggregate Fits a Partial Least Squares model over a partition and returns per-row predictions. NULL [pls_fit_predict_agg(y, x)]
pls_fit_predict_agg aggregate Fits a Partial Least Squares model over a partition with a MAP of options and returns per-row predictions. NULL [pls_fit_predict_agg(y, x, {'n_components': 2})]
pls_fit_predict_agg aggregate Fits a Partial Least Squares model using only training rows (split_col='train') and predicts all rows. NULL [pls_fit_predict_agg(y, x, split_col)]
pls_fit_predict_by table_macro NULL NULL  
poisson_fit_agg aggregate Fits a Poisson regression (GLM with log link) and returns coefficients, deviance, AIC, and fit statistics. NULL [poisson_fit_agg(y, x)]
poisson_fit_agg aggregate Fits a Poisson regression (GLM with log link) and returns coefficients, deviance, AIC, and fit statistics. NULL [poisson_fit_agg(y, x, {'fit_intercept': true})]
poisson_fit_predict_agg aggregate Fits a Poisson regression on training rows with a MAP of options and predicts all rows. NULL [poisson_fit_predict_agg(y, x, split_col, {'link': 'log'})]
poisson_fit_predict_agg aggregate Fits a Poisson regression over a partition and returns per-row predictions. NULL [poisson_fit_predict_agg(y, x)]
poisson_fit_predict_agg aggregate Fits a Poisson regression over a partition with a MAP of options and returns per-row predictions. NULL [poisson_fit_predict_agg(y, x, {'link': 'log'})]
poisson_fit_predict_agg aggregate Fits a Poisson regression using only training rows (split_col='train') and predicts all rows. NULL [poisson_fit_predict_agg(y, x, split_col)]
poisson_fit_predict_by table_macro NULL NULL  
predict scalar Applies pre-fitted coefficients and intercept to feature data to generate predictions. NULL [predict(x, coefficients, intercept)]
prop_test_one_agg aggregate Tests whether an observed proportion differs from a hypothesized value (one-sample proportion test), using default options. NULL [prop_test_one_agg(value)]
prop_test_one_agg aggregate Tests whether an observed proportion differs from a hypothesized value (one-sample proportion test). NULL [prop_test_one_agg(value, {'p0': 0.5, 'alternative': 'two_sided'})]
prop_test_two_agg aggregate Tests whether two observed proportions are equal (two-sample proportion test), using default options. NULL [prop_test_two_agg(value, group_id)]
prop_test_two_agg aggregate Tests whether two observed proportions are equal (two-sample proportion test). NULL [prop_test_two_agg(value, group_id, {'alternative': 'two_sided'})]
quantile_fit_predict_agg aggregate Fits a quantile regression model on training rows with a MAP of options and predicts all rows. NULL [quantile_fit_predict_agg(y, x, split_col, {'quantile': 0.5})]
quantile_fit_predict_agg aggregate Fits a quantile regression model over a partition and returns per-row predictions. NULL [quantile_fit_predict_agg(y, x)]
quantile_fit_predict_agg aggregate Fits a quantile regression model over a partition with a MAP of options and returns per-row predictions. NULL [quantile_fit_predict_agg(y, x, {'quantile': 0.5})]
quantile_fit_predict_agg aggregate Fits a quantile regression model using only training rows (split_col='train') and predicts all rows. NULL [quantile_fit_predict_agg(y, x, split_col)]
quantile_fit_predict_by table_macro NULL NULL  
ransac_fit scalar Fits a RANSAC regression model with optional MAP of settings (residual_threshold, max_trials, min_samples, stop_probability, stop_n_inliers, random_state, fit_intercept, compute_inference, confidence_level). NULL [ransac_fit(y, x, {'residual_threshold': 0.5, 'random_state': 42})]
ransac_fit scalar Fits a RANSAC robust regression model. Returns coefficients, fit statistics, the residual threshold used, and the inlier / trial counts as a struct. NULL [ransac_fit(y, x)]
ransac_fit_agg aggregate Fits a RANSAC robust regression model and returns coefficients, fit statistics, the residual threshold used, and the inlier / trial counts as a struct. NULL [ransac_fit_agg(y, x)]
ransac_fit_agg aggregate Fits a RANSAC robust regression model and returns coefficients, fit statistics, the residual threshold used, and the inlier / trial counts as a struct. NULL [ransac_fit_agg(y, x, {'residual_threshold': 0.5, 'random_state': 42})]
ransac_fit_predict aggregate Fits a RANSAC regression over a window with a MAP of options. NULL [ransac_fit_predict(y, x, {'residual_threshold': 0.5}) OVER (…)]
ransac_fit_predict aggregate Fits a RANSAC robust regression over a window partition and returns the prediction for the current row. NULL [ransac_fit_predict(y, x) OVER (PARTITION BY g ORDER BY t)]
ransac_fit_predict_agg aggregate Fits RANSAC on training rows with a MAP of options and predicts all rows. NULL [ransac_fit_predict_agg(y, x, split_col, {'residual_threshold': 0.5})]
ransac_fit_predict_agg aggregate Fits RANSAC over a partition with a MAP of options and returns per-row predictions. NULL [ransac_fit_predict_agg(y, x, {'residual_threshold': 0.5, 'random_state': 42})]
ransac_fit_predict_agg aggregate Fits RANSAC using only training rows (split_col='train') and predicts all rows. NULL [ransac_fit_predict_agg(y, x, split_col)]
ransac_fit_predict_agg aggregate Fits a RANSAC robust regression over a partition and returns per-row predictions with confidence intervals. NULL [ransac_fit_predict_agg(y, x)]
ransac_fit_predict_by table_macro NULL NULL  
residuals_diagnostics scalar Computes raw and standardized residuals, leverage, and Cook's distance from actuals and predictions. NULL [residuals_diagnostics(y, y_hat)]
residuals_diagnostics scalar Computes residual diagnostics including studentized residuals when feature matrix and residual standard error are supplied. NULL [residuals_diagnostics(y, y_hat, x, rse, true)]
residuals_diagnostics_agg aggregate Aggregate version of residuals diagnostics with feature matrix: computes raw, standardized, studentized residuals and leverage from predicted and actual values. NULL [residuals_diagnostics_agg(y, y_hat, x)]
residuals_diagnostics_agg aggregate Aggregate version of residuals diagnostics: computes raw, standardized, studentized residuals and leverage from predicted and actual values. NULL [residuals_diagnostics_agg(y, y_hat)]
ridge_fit scalar Fits a Ridge regression model with L2 regularization and optional MAP of settings (fit_intercept, compute_inference, confidence_level, alpha, solver). NULL [ridge_fit(y, x, {'alpha': 1.0, 'compute_inference': true})]
ridge_fit scalar Fits a Ridge regression model with L2 regularization to the given response and feature data. NULL [ridge_fit(y, x)]
ridge_fit_agg aggregate Fits a Ridge regression model with L2 regularization and returns coefficients and fit statistics. NULL [ridge_fit_agg(y, x)]
ridge_fit_agg aggregate Fits a Ridge regression model with L2 regularization and returns coefficients and fit statistics. NULL [ridge_fit_agg(y, x, {'alpha': 1.0})]
ridge_fit_predict aggregate Fits a Ridge regression model over a window partition and returns predictions with confidence intervals. NULL [ridge_fit_predict(y, x)]
ridge_fit_predict aggregate Fits a Ridge regression model over a window partition and returns predictions with confidence intervals. NULL [ridge_fit_predict(y, x, {'null_policy': 'drop'})]
ridge_fit_predict_agg aggregate Fits Ridge regression on training rows with a MAP of options and predicts all rows. NULL [ridge_fit_predict_agg(y, x, split_col, {'null_policy': 'drop'})]
ridge_fit_predict_agg aggregate Fits Ridge regression over a partition and returns per-row predictions with confidence intervals. NULL [ridge_fit_predict_agg(y, x)]
ridge_fit_predict_agg aggregate Fits Ridge regression over a partition with a MAP of options and returns per-row predictions with confidence intervals. NULL [ridge_fit_predict_agg(y, x, {'null_policy': 'drop'})]
ridge_fit_predict_agg aggregate Fits Ridge regression using only training rows (split_col='train') and predicts all rows. NULL [ridge_fit_predict_agg(y, x, split_col)]
ridge_fit_predict_by table_macro NULL NULL  
rls_fit scalar Fits a Recursive Least Squares (RLS) model to the given response and feature data. NULL [rls_fit(y, x)]
rls_fit scalar Fits a Recursive Least Squares (RLS) model with optional MAP of settings (fit_intercept, forgetting_factor, initial_p_diagonal). NULL [rls_fit(y, x, {'forgetting_factor': 0.99})]
rls_fit_agg aggregate Fits a Recursive Least Squares model and returns coefficients and fit statistics. NULL [rls_fit_agg(y, x)]
rls_fit_agg aggregate Fits a Recursive Least Squares model and returns coefficients and fit statistics. NULL [rls_fit_agg(y, x, {'forgetting_factor': 0.99})]
rls_fit_predict aggregate Fits a Robust Least Squares model over a window partition and returns predictions with confidence intervals. NULL [rls_fit_predict(y, x)]
rls_fit_predict aggregate Fits a Robust Least Squares model over a window partition and returns predictions with confidence intervals. NULL [rls_fit_predict(y, x, {'null_policy': 'drop'})]
rls_fit_predict_agg aggregate Fits Robust LS regression on training rows with a MAP of options and predicts all rows. NULL [rls_fit_predict_agg(y, x, split_col, {'null_policy': 'drop'})]
rls_fit_predict_agg aggregate Fits Robust LS regression over a partition and returns per-row predictions. NULL [rls_fit_predict_agg(y, x)]
rls_fit_predict_agg aggregate Fits Robust LS regression over a partition with a MAP of options and returns per-row predictions. NULL [rls_fit_predict_agg(y, x, {'null_policy': 'drop'})]
rls_fit_predict_agg aggregate Fits Robust LS regression using only training rows (split_col='train') and predicts all rows. NULL [rls_fit_predict_agg(y, x, split_col)]
rls_fit_predict_by table_macro NULL NULL  
shapiro_wilk_agg aggregate Performs the Shapiro-Wilk test for normality on a sample. NULL [shapiro_wilk_agg(value)]
spearman_agg aggregate Computes Spearman's rank correlation coefficient and tests its significance. NULL [spearman_agg(x, y)]
spearman_agg aggregate Computes Spearman's rank correlation coefficient and tests its significance. NULL [spearman_agg(x, y, {'alternative': 'two_sided'})]
t_test_agg aggregate Performs a two-sample t-test (Welch or Student) comparing values between two groups, using default options. NULL [t_test_agg(value, group_id)]
t_test_agg aggregate Performs a two-sample t-test (Welch or Student) comparing values between two groups. NULL [t_test_agg(value, group_id, {'alternative': 'two_sided'})]
theil_sen_fit scalar Fits a Theil-Sen regression model with optional MAP of settings (max_subpopulation, n_subsamples, max_iterations, tolerance, random_state, fit_intercept, compute_inference, confidence_level). NULL [theil_sen_fit(y, x, {'random_state': 42, 'max_subpopulation': 5000})]
theil_sen_fit scalar Fits a Theil-Sen robust regression model. Returns coefficients and fit statistics as a struct. NULL [theil_sen_fit(y, x)]
theil_sen_fit_agg aggregate Fits a Theil-Sen robust regression model and returns coefficients and fit statistics as a struct. NULL [theil_sen_fit_agg(y, x)]
theil_sen_fit_agg aggregate Fits a Theil-Sen robust regression model and returns coefficients and fit statistics as a struct. NULL [theil_sen_fit_agg(y, x, {'random_state': 42, 'max_subpopulation': 5000})]
theil_sen_fit_predict aggregate Fits a Theil-Sen regression over a window with a MAP of options. NULL [theil_sen_fit_predict(y, x, {'random_state': 42}) OVER (…)]
theil_sen_fit_predict aggregate Fits a Theil-Sen robust regression over a window partition and returns the prediction for the current row. NULL [theil_sen_fit_predict(y, x) OVER (PARTITION BY g ORDER BY t)]
theil_sen_fit_predict_agg aggregate Fits Theil-Sen on training rows with a MAP of options and predicts all rows. NULL [theil_sen_fit_predict_agg(y, x, split_col, {'random_state': 42})]
theil_sen_fit_predict_agg aggregate Fits Theil-Sen over a partition with a MAP of options and returns per-row predictions. NULL [theil_sen_fit_predict_agg(y, x, {'random_state': 42})]
theil_sen_fit_predict_agg aggregate Fits Theil-Sen using only training rows (split_col='train') and predicts all rows. NULL [theil_sen_fit_predict_agg(y, x, split_col)]
theil_sen_fit_predict_agg aggregate Fits a Theil-Sen robust regression over a partition and returns per-row predictions with confidence intervals. NULL [theil_sen_fit_predict_agg(y, x)]
theil_sen_fit_predict_by table_macro NULL NULL  
tost_correlation_agg aggregate Tests equivalence of a correlation to a reference value using the TOST procedure, using default options. NULL [tost_correlation_agg(x, y)]
tost_correlation_agg aggregate Tests equivalence of a correlation to a reference value using the TOST procedure. NULL [tost_correlation_agg(x, y, {'delta': 0.1})]
tost_paired_agg aggregate Tests equivalence of paired measurements using the TOST procedure, using default options. NULL [tost_paired_agg(x, y)]
tost_paired_agg aggregate Tests equivalence of paired measurements using the TOST procedure. NULL [tost_paired_agg(x, y, {'delta': 0.5})]
tost_t_test_agg aggregate Tests equivalence of two groups using the Two One-Sided Tests (TOST) procedure with a t-test, using default options. NULL [tost_t_test_agg(value, group_id)]
tost_t_test_agg aggregate Tests equivalence of two groups using the Two One-Sided Tests (TOST) procedure with a t-test. NULL [tost_t_test_agg(value, group_id, {'delta': 1.0})]
tweedie_fit_agg aggregate Fits a Tweedie GLM (log link, default power = 1.5 — compound Poisson-Gamma) and returns coefficients, deviance, AIC, dispersion (= phi), and fit statistics. NULL [tweedie_fit_agg(y, x)]
tweedie_fit_agg aggregate Fits a Tweedie GLM (log link, user-specified power 1 < p < 2 for compound Poisson-Gamma) and returns coefficients, deviance, AIC, dispersion (= phi), and fit statistics. Default power is 1.5. NULL [tweedie_fit_agg(y, x, {'power': 1.5, 'fit_intercept': true})]
vif scalar Computes Variance Inflation Factor (VIF) for each column of a feature matrix to detect multicollinearity. NULL [vif(x)]
vif_agg aggregate Aggregate version of VIF: computes Variance Inflation Factor for each feature from a column of feature vectors. NULL [vif_agg(x)]
wilcoxon_signed_rank_agg aggregate Performs the Wilcoxon signed-rank test for paired samples, using default options. NULL [wilcoxon_signed_rank_agg(x, y)]
wilcoxon_signed_rank_agg aggregate Performs the Wilcoxon signed-rank test for paired samples. NULL [wilcoxon_signed_rank_agg(x, y, {'alternative': 'two_sided'})]
wls_fit scalar Fits a WLS regression model with optional MAP of settings (fit_intercept, compute_inference, confidence_level, solver, hc_type). NULL [wls_fit(y, x, weights, {'compute_inference': true})]
wls_fit scalar Fits a Weighted Least Squares (WLS) regression model using per-observation weights. NULL [wls_fit(y, x, weights)]
wls_fit_agg aggregate Fits a Weighted Least Squares regression model and returns coefficients and fit statistics. NULL [wls_fit_agg(y, x, weight)]
wls_fit_agg aggregate Fits a Weighted Least Squares regression model and returns coefficients and fit statistics. NULL [wls_fit_agg(y, x, weight, {'fit_intercept': true})]
wls_fit_predict aggregate Fits a WLS regression model over a window partition using per-row weights and returns predictions. NULL [wls_fit_predict(y, x, weight)]
wls_fit_predict aggregate Fits a WLS regression model over a window partition using per-row weights and returns predictions. NULL [wls_fit_predict(y, x, weight, {'null_policy': 'drop'})]
wls_fit_predict_agg aggregate Fits WLS regression on training rows with weights and a MAP of options and predicts all rows. NULL [wls_fit_predict_agg(y, x, weights, split_col, {'null_policy': 'drop'})]
wls_fit_predict_agg aggregate Fits WLS regression over a partition using weights and returns per-row predictions. NULL [wls_fit_predict_agg(y, x, weights)]
wls_fit_predict_agg aggregate Fits WLS regression over a partition using weights with a MAP of options and returns per-row predictions. NULL [wls_fit_predict_agg(y, x, weights, {'null_policy': 'drop'})]
wls_fit_predict_agg aggregate Fits WLS regression using only training rows (split_col='train') with weights and predicts all rows. NULL [wls_fit_predict_agg(y, x, weights, split_col)]
wls_fit_predict_by table_macro NULL NULL  
yuen_agg aggregate Performs Yuen's trimmed-means t-test, robust to outliers and non-normality, using default options. NULL [yuen_agg(value, group_id)]
yuen_agg aggregate Performs Yuen's trimmed-means t-test, robust to outliers and non-normality. NULL [yuen_agg(value, group_id, {'trim': 0.2})]

Overloaded Functions

This extension does not add any function overloads.

Added Types

This extension does not add any types.

Added Settings

name description input_type scope aliases
anofox_telemetry_enabled Enable or disable anonymous usage telemetry BOOLEAN GLOBAL []
anofox_telemetry_key PostHog API key for telemetry VARCHAR GLOBAL []
datazoo_banner Show the DataZoo feedback banner when an extension is loaded in an interactive terminal (at most once a day per extension). BOOLEAN GLOBAL []