Distributional and varying-coefficient models¶
OpenBoost estimates the parameters of a statistical model as functions of covariates, by gradient boosting.
That is distributional regression when the model is a conditional
distribution F(y | x): each parameter (location, scale, shape, ...)
gets its own additive predictor. The statistics literature calls the
GAM version GAMLSS;
the boosting version is gamboostLSS
/ NGBoost. NaturalBoost and WeibullAFT live here.
It is a varying-coefficient model
(Hastie & Tibshirani, 1993)
when you supply a structural formula y ≈ f(θ(z), x). The trees learn
θ(z); x enters only through f. That is FormulaBoost.
z are the covariates the trees split on. θ = (θ₁, …, θₖ) is the
parameter vector of the statistical model. Each θᵢ is its own ensemble.
The per-round update is a (damped) generalized Gauss-Newton / natural gradient step:
Models are configurations of f, the loss, and the metric M.
| Model | Statistical object | Loss | Metric M |
K |
|---|---|---|---|---|
| NaturalBoost | Conditional distribution | NLL | analytic Fisher | # distribution params |
| WeibullAFT | Censored Weibull AFT | censored NLL | expected Fisher | 2 (scale, shape) |
| FormulaBoost | Varying-coefficient formula | MSE | GGN (plain / diag / full) |
# formula params |
GradientBoosting |
Conditional mean | MSE / logloss / … | scalar Hessian | 1 |
K = 1 is ordinary mean regression. K > 1 is the point of this library.
What this is not¶
Not a faster drop-in for XGBoost/LightGBM mean regression. Those are optimized C++ and should stay the default for MSE/logloss.
Not only "faster NGBoost." NGBoost is distributional regression for a fixed set of families, with exact-split trees, on CPU. OpenBoost does the same job with histogram trees and a GPU path, and also fits varying-coefficient formulas and a censored Weibull whose shape depends on covariates, neither of which NGBoost can write down.
On the overlap (Normal NLL on UCI) quality is tied-or-better and GPU training is much faster: Benchmarks.
Why the metric matters¶
Raw gradients of a multi-parameter likelihood are often badly scaled.
On the sales-curve formula, precond="plain" diverges (extrapolation RMSE
~3.9). Diagonal GGN is usable. Full GGN recovers the hard coefficient
(b(z) corr 0.877 vs 0.730 diag vs 0.599 for an XGBoost custom objective,
which cannot represent off-diagonals).
Same story for Weibull AFT: the observed Hessian's scale term explodes
when λ is wrong. The expected Fisher does not, and is what
WeibullAFT uses.
precond="full" is the default for FormulaBoost. Do not turn it off
unless you are debugging.
Shared training API¶
NaturalBoost, FormulaBoost, and WeibullAFT share one trainer:
model.fit(
X, y, # FormulaBoost also needs model_input=x
eval_set=[...], # early stopping / logging
callbacks=[ob.EarlyStopping(patience=20)],
early_stopping_rounds=20,
)
params = model.predict_params(X) # dict of per-row parameter arrays
- NaturalBoost / WeibullAFT:
eval_setis(X_val, y_val)(pluseventfor AFT). - FormulaBoost:
eval_setis(X_val, y_val, model_input_val).
GPU: histogram trees run on CUDA when openboost[cuda] is installed.
NaturalBoost Normal/Poisson have device gradient kernels. FormulaBoost GGN
is currently host-side (still faster than XGBoost-diag at 200K on A100).
Choosing a model¶
| You have | Use |
|---|---|
y should be a distribution given X |
NaturalBoost* |
A known curve y = f(θ, x) with θ depending on other features |
FormulaBoost |
| Right-censored times, Weibull, shape should depend on covariates | WeibullAFT |
| A custom NLL that is still a distribution | NaturalBoost + custom distribution |
| Ordinary mean regression | GradientBoosting (or XGBoost / LightGBM) |