Model Training¶
You want to train a machine learning model from your pipeline data. The Model Training node supports three gradient-boosted tree families (CatBoost, XGBoost and LightGBM), the Explainable Boosting Machine (EBM, from InterpretML), and the GLM (generalised linear model, via RustyStats). You choose the family when you add the node, and it stays fixed: a different family is a new node. Every run records reproducible evaluation evidence and saves a native model plus its feature contract. A completed result can then be logged to MLflow as a candidate run and scored by a Model Scoring node.
What is MLflow?
MLflow is an open-source platform for tracking experiments and storing models. If you're new to MLflow, the key concepts are: an experiment groups related training runs, a run is a single training attempt with its metrics and parameters, and the model registry stores production-ready models by name and version.
This node takes a single input and produces no downstream data: it is a terminal node. Training in the editor keeps the model and its evaluation artifacts with that training result; logging it to MLflow and saving it to a project file are explicit actions in the node's EXPORT pane.
Choosing the model family¶
A new Model Training node opens on Select algorithm, with one button per family:
| Button | Description on the button |
|---|---|
| CatBoost | Gradient boosting - handles categoricals natively, fast GPU training |
| EBM | Explainable boosting - additive shape functions and pairwise interactions you can read directly |
| GLM | Generalised linear model - interpretable coefficients, regulatory-friendly |
| LightGBM | Gradient boosting - fast leaf-wise trees with native categoricals and early stopping on CPU |
| XGBoost | Gradient boosting - histogram trees with native categoricals and early stopping on CPU |
Choosing a tree or EBM family fills in its starting parameters (see The PARAMETERS pane), and every family starts from a default split: a random split with seed 42, a 20% holdout validation set and no test set. The panel then shows six panes, TARGET, FEATURES, PARAMETERS, SPLIT, TRAIN and EXPORT, described below in that order. A pane whose settings stop the model training shows a warning mark on its tab, and the TRAIN tab shows a running indicator while a model trains. The node has no COLUMNS tab.
Model families¶
| CatBoost | XGBoost | LightGBM | EBM | GLM | |
|---|---|---|---|---|---|
| Tasks | Regression, binary classification | Regression, binary classification | Regression, binary classification | Regression, binary classification | Regression |
| Losses | RMSE, MAE, Poisson, Tweedie, Logloss, CrossEntropy | RMSE, MAE, Poisson, Gamma, Tweedie, Logloss | RMSE, MAE, Poisson, Gamma, Tweedie, Logloss | RMSE, Poisson, Gamma, Tweedie, Logloss | Family and Link Function |
| Round budget (in Parameters JSON) | iterations |
num_boost_round |
num_iterations |
max_rounds (required) |
— |
| Early stopping | early_stopping_rounds |
early_stopping_rounds |
early_stopping_round |
Never | — |
| Final refit (when refitting) | Validation-weighted round count | Validation-weighted round count | Validation-weighted round count | Winning max_rounds, unchanged |
Same settings |
| Tuning | Yes | Yes | Yes | Yes (max_rounds searchable) |
No |
| Monotone constraints | Yes | Yes, except with MAE | Yes, except with MAE | Yes, not on an interaction | Per term |
| Feature weights | Yes (pipeline file only) | No | No | No | No |
| Interactions | Learned by trees | Learned by trees | Learned by trees | Chosen count or explicit pairs | Explicit cards |
| Offset | Baseline | base_margin |
init_score (added at prediction) |
init_score |
Offset term |
| Trace explanation | SHAP values | Native contributions | Native contributions | Term scores (an interaction is one term) | Term contributions |
| Model file | .cbm |
.ubj |
.lgbm |
.ebm |
.rsglm |
| Compute | CPU, optional GPU | CPU, optional CUDA GPU | CPU | CPU, one thread | CPU |
Tree families train with HAUTE_TRAINING_THREADS threads (default: every logical CPU).
On macOS, XGBoost and LightGBM need Homebrew's libomp (brew install libomp).
The TARGET pane¶
The pane opens with the chosen family (for example Algorithm: CatBoost). For CatBoost, XGBoost, LightGBM and EBM it has three sections; the GLM's version is described in GLM target below.
Target and objective
| Field | What it does |
|---|---|
| Target column | The column the model predicts. Pick it from the searchable list (Select target… until you do). Training is blocked with Select a target column. until it is set. |
| Objective | The training loss, one button per loss the family supports (see the table above). Choosing a loss also sets the task (regression or classification) and switches on that loss's usual metrics; clicking the selected loss again clears it. Training is blocked until a loss is chosen, because an unset loss would silently train under the library default. MAE is not offered for an EBM. |
| Variance power | Shown for the Tweedie loss: a slider and a number between 1 (Poisson) and 2 (Gamma), exclusive. A new Tweedie selection starts at 1.5. A saved value outside that range shows Saved variance power must be greater than 1 and less than 2. |
| Positive class | Shown for a classification loss when the target is not Boolean: the label the model predicts the probability of. Predictions above 0.5 are labelled with it. For a numeric target the label reads Positive class (only if the labels are not 0/1); for text labels it is required and shows Choose which label is the positive class. until set. |
Binary classification needs exactly two target classes. A Boolean or 0/1 target is
positive at True/1; any other pair needs a Positive class. Every family labels a
prediction positive when its positive-class probability is above 0.5.
Weight and offset
| Field | What it does |
|---|---|
| Weight column (optional) | A numeric column that weights each row in the loss (for example exposure). None leaves rows unweighted. |
| Offset column (optional) | A numeric column the model folds in through its link function. Under a log link (a log-link GLM, or a Poisson, Gamma or Tweedie loss) it is a strictly positive exposure multiplier: 2× exposure gives 2× the expected count, and null, zero, or negative values are refused when training and when scoring. Under any other link it is added to the prediction. A constant column of 1 is the unit basis under a log link. The offset column must be present when the model scores. Different from the weight, which weights the loss. The info icon beside the label repeats this. |
One column plays one role: the target is never offered as the weight or offset, and the weight and offset lists hold numeric columns only.
Metrics
The evaluation metrics, as toggle buttons: Gini, RMSE, MAE, MSE, R², Poisson Deviance, Tweedie Deviance, Gamma Deviance, AUC and Logloss. Only the metrics that suit the chosen loss can be switched on: regression losses offer the first eight, classification losses AUC and Logloss. Hovering over a disabled metric says why (for example AUC is not available with Poisson, or Choose an objective to enable compatible metrics).
Choosing a metric
For frequency models (Poisson), use Poisson deviance. For severity models, use Gamma deviance with a Gamma loss or Tweedie deviance with a Tweedie loss. For general regression, RMSE or Gini are common choices. For classification, use AUC or Logloss.
GLM target¶
A GLM's TARGET pane reads Algorithm Rustystats and has these controls:
| Field | What it does |
|---|---|
| Target column | As above. |
| Family | The distribution: Poisson (claim frequency), Gamma (claim severity), Tweedie (pure premium), Gaussian (linear regression), Binomial (binary outcomes), Quasi-Poisson (overdispersed counts), Quasi-Binomial (overdispersed binary outcomes) or Neg. Binomial (overdispersed counts with an explicit theta). Hover over a button for its use. Choosing a family resets the link to the family's canonical one and switches on its usual metrics. Training is blocked until a family is chosen, because an unset family would silently train a Gaussian model. |
| Link Function | auto (...) uses the family's canonical link: log for Poisson, Quasi-Poisson, Gamma, Tweedie and Negative Binomial; logit for Binomial and Quasi-Binomial; identity for Gaussian. The other buttons are the links RustyStats supports for the family: Gaussian, Poisson, Quasi-Poisson, Gamma, Tweedie and Negative Binomial accept log or identity, and Binomial and Quasi-Binomial accept logit, log or identity. |
| Variance power (1.0=Poisson, 2.0=Gamma) | Tweedie only, and required: click Set variance power (required for Tweedie) to start at 1.5, then adjust the slider between 1 and 2. Estimate from data profiles it on the node's training data. |
| Dispersion theta (required for Neg. Binomial) | Negative Binomial only: a positive dispersion (variance = mean + mean²/theta). RustyStats refuses to fit without it. Type a value or click Estimate from data, which profiles the likelihood on the node's training data. A value of 0 or less shows Theta must be greater than 0. |
| Weight column (optional), Offset column (optional) | As above. |
| Intercept | Whether to fit an intercept. Ticked by default. |
| Metrics | Gini, RMSE, MAE, Poisson Dev., Tweedie Dev., R², AUC and Logloss. |
If an estimate fails, the pane shows Estimation failed: and the reason.
The FEATURES pane¶
For CatBoost, XGBoost, LightGBM and EBM the pane lists every column that could be a feature, with a count of how many are included and excluded. Columns that already have a role (target, weight, offset, and the group or date column the split uses) are listed under Excluded from predictors and are never features.
- Search features narrows the list by name.
- All, Included and Excluded filter the list by membership; each shows its count.
- Include all and Exclude all tick or clear every listed feature.
- The tick box beside each feature includes or excludes it.
- Monotonicity sets a constraint on an included numeric feature: ↓ decreasing, − no constraint (the default) and ↑ increasing. The buttons are disabled for a non-numeric feature (Monotonicity is only available for numeric features.) and for an excluded one (Include this feature to set monotonicity.).
- An excluded column that is no longer upstream shows as
- not found with a button to remove the exclusion.
Feature selection is explicit. Haute uses every available column except the target, weight, offset, any fold or identifier columns, the group or date column the split uses, and the columns you exclude here. If your data contains ID columns, dates, or columns derived from the target, exclude them to prevent data leakage. The target and weight are kept out because they have training roles; exclusions are for identifiers, dates, leakage-prone fields, and any other columns you deliberately do not want the model to learn from. The training result records the final ordered feature set, the retained metadata, and every excluded column with its reason.
Haute validates feature selection before collecting training data, so a missing, invalid, or unsuitable feature fails clearly before a large collection begins. See Execution Strategy for the all-except strategy and for reading execution diagnostics.
XGBoost and LightGBM refuse monotone constraints with the MAE loss: training is blocked with a message asking you to remove them here or choose another loss. XGBoost's absolute-error objective re-fits each leaf after the tree is built, which breaks the constraint.
Pairwise interactions (EBM)
An EBM adds a Pairwise interactions section. Every included feature is a main effect; interactions add two-feature terms, each kept as one term in results and explanations. A monotone-constrained feature cannot take part in an interaction.
- Let EBM choose (the default): Up to a number of interactions (10 to start), the strongest pairs EBM finds on the training rows. 0 means main effects only.
- Choose pairs: fix the pairs yourself. Add interaction adds a row with two feature lists; a monotone feature is listed but disabled, marked (monotone). A pair needs two different features, and a pair listed twice is refused.
GLM terms and interactions¶
A GLM's features are its terms and interaction cards; the exclusions and monotonicity buttons above do not apply to it.
The FEATURES list shows how many columns are in the model. Above it:
- Search features narrows the list by name.
- Fit all with defaults adds a term with its default fit for every column that has none.
- Remove all terms clears every term and interaction, after asking Remove every term and interaction from this model?
- In model only hides columns with no term.
- Builder and JSON switch between the cards and a JSON editor for the whole terms dictionary (saved when you leave the box).
Each column's dtype decides which fits it offers:
| Column dtype | Fits | Default |
|---|---|---|
| Float, Decimal | Linear, B-spline, Nat. spline, Monotone spline; numeric expressions | Linear |
| Integer | The float fits plus Categorical, Target enc. and Frequency enc. | Linear |
| Boolean | Categorical, Target enc., Frequency enc. | Categorical |
| String, Categorical, Enum | Categorical, Target enc., Frequency enc. | Categorical |
Dates, times, durations, lists, structs, and binary columns cannot be fitted; the pane hides them and says how many it hid. Target, weight, offset, fold, identifier, and split-key columns cannot be terms, expression columns, or interaction factors.
Add term on a column adds its default fit. Clicking it again on a numeric column adds an
Expression term (starting at column ** 2); on an integer or categorical column it adds
a second encoding. When a column has every fit it supports, the button is disabled and says
why. Each term card has:
| Field | What it does |
|---|---|
| Fit type | How the term is fitted (the fits in the table above). |
| df mode, df | Spline fits: Auto leaves the degrees of freedom to RustyStats, which chooses the smoothing (the Summary lists each smooth term's effective degrees of freedom); Fixed takes a df, an integer up to 20 and at least degree + 1 for B-splines or at least 2 for natural and monotone splines. |
| Degree | B-spline and monotone spline: the polynomial degree (3 unless set). |
| Advanced | Spline fits: Basis size (k) (Auto only), Interior knots (Fixed only; they replace df) and Boundary knots. Set only one of df, k and knots. Categorical fits: Reference level and Levels (below). Target encoding: Permutations, from 1 to 100 (default 4). |
| Monotonicity | Linear, B-spline, monotone spline and expression terms: increasing or decreasing. |
| Reference level, Levels | A categorical fit shares the intercept with its first level in sorted order. Reference level chooses that baseline level; Levels fits indicators only for the listed levels. They replace each other, and the rest (and levels unseen at scoring) share the intercept. Training refuses a reference or listed level the data does not contain and lists the observed labels. Quote numeric labels exactly as the column holds them, for example "2" for an integer column (or "2.0" if it also has nulls). |
| Prior weight, Weight | Target encoding: Auto leaves the prior weight to RustyStats; Fixed takes a non-negative Weight. |
| Name, Expression | Expression terms: the term's name and its expression. Expressions have one of the forms x, x ** n, x + y, x - y, x * y, or x / y, where y is a column or a number. Columns in an expression must be numeric and named with letters, digits, and underscores, starting with a letter or underscore; compute logs and other transforms in an upstream Polars node. |
Terms that cannot be fitted (a role column, a column no longer upstream, an unsupported dtype, or an entry with no fit type) are listed under Unresolved terms with the reason. Fix or remove them before training. Training is also blocked until at least one feature has a term (Add a term to at least one feature.).
INTERACTIONS
Add interaction adds a card. Pick at least two features with the feature lists and + feature, then choose the card's Fit type:
- Product fits the features' product. Each feature fits as its override on the card (Linear, Categorical, B-spline, Nat. spline or Target enc.) if you set one; otherwise as its main-effect term, when it has exactly one that interactions honour; otherwise as its dtype default.
- Target enc. and Frequency enc. encode the combination of the raw feature values (for example brand and region). Joint encodings accept integer, boolean, and categorical features, and target encoding takes Prior weight and Permutations.
A monotone spline, a monotone linear or B-spline term, a categorical term with Levels or a Reference level, a frequency encoding, or several main-effect terms for one feature cannot be used inside an interaction, so the card asks for an explicit fit. Categorical fits apply only over a categorical main effect, and linear, spline, and target-encoding fits never do. A categorical feature may use Target enc. when every other feature is Linear; this always adds the target-encoded main effect, even with Include main effects off. Different cards may fit the same feature with different local splines.
Include main effects adds a main effect for each feature that has none, using the fit its cards agree on; the feature list tags that feature "Main effect from Interaction N". Cards that would add different main effects for one feature are refused together: add a main term for it, or give the cards the same fit. Features with a term keep it. Product, target-encoded, and frequency-encoded interactions can coexist over the same factors; a duplicate within one mode shows Duplicate interaction: another card fits these features the same way.
The PARAMETERS pane¶
For CatBoost, XGBoost, LightGBM and EBM, Parameter strategy chooses between Fixed parameters and Tune parameters. The GLM's version is described in GLM penalty and solver below.
Fixed parameters
Parameters JSON holds the family's fixed parameters as a JSON object. A new node starts from:
| Family | Starting Parameters JSON |
|---|---|
| CatBoost | iterations 1000, learning_rate 0.05, depth 6, l2_leaf_reg 3, early_stopping_rounds 50, one_hot_max_size 10 |
| XGBoost | num_boost_round 1000, eta 0.1, max_depth 6, early_stopping_rounds 50 |
| LightGBM | num_iterations 1000, learning_rate 0.05, num_leaves 31, early_stopping_round 50 |
| EBM | max_rounds 2000, learning_rate 0.02, interactions 10 |
Invalid JSON, or a key Haute sets itself, shows Parameters JSON: with the reason under the box and blocks training. Haute sets the objective, threads, seed, offset and categorical handling: choose the loss in the TARGET pane and monotonicity in the FEATURES pane. For CatBoost, GPU training is set in the TRAIN pane rather than here.
CatBoost parameters
Any CatBoost constructor parameter can go in Parameters JSON, except the ones Haute sets. The starting keys are:
| Key | Description |
|---|---|
iterations |
Maximum boosting rounds; tuning uses this as its ceiling |
depth |
Tree depth |
learning_rate |
Step-size shrinkage: smaller values are slower but often more accurate |
l2_leaf_reg |
L2 regularisation of leaf values |
early_stopping_rounds |
Stop a validation fit when its metric stops improving |
one_hot_max_size |
Categorical columns with up to this many levels are one-hot encoded; above it CatBoost uses target statistics, which are much slower to train. The pane repeats this under the box, and the TRAIN pane's run summary shows the resulting Categorical encoding. |
XGBoost parameters
XGBoost trains CPU histogram trees. Parameters JSON accepts num_boost_round,
early_stopping_rounds, eta, max_depth, max_leaves, grow_policy,
min_child_weight, gamma, max_delta_step, subsample, colsample_bytree,
colsample_bylevel, colsample_bynode, lambda, alpha, max_bin,
max_cat_to_onehot and max_cat_threshold. Aliases such as learning_rate or
n_estimators are refused with the canonical name; Haute sets the objective, threads,
seed, categorical handling and offset. Categories are encoded against the levels the
model was trained on, so a value the model never saw fails instead of scoring as
missing.
LightGBM parameters
LightGBM trains CPU leaf-wise trees. Parameters JSON accepts num_iterations,
early_stopping_round, learning_rate, num_leaves, max_depth, min_data_in_leaf,
min_sum_hessian_in_leaf, feature_fraction, bagging_fraction, bagging_freq,
lambda_l1, lambda_l2, min_gain_to_split, max_bin, max_cat_to_onehot,
max_cat_threshold, cat_smooth, cat_l2 and min_data_per_group. LightGBM itself
silently accepts conflicting aliases, so Haute refuses every alias (for example
num_boost_round or eta). A fit that runs out of useful splits before its round
budget records native_exhaustion as its stopping reason. LightGBM cannot apply
monotone constraints with the MAE loss.
EBM parameters
An Explainable Boosting Machine is a sum of one learned shape per feature plus chosen
pairwise interactions, so the model is readable term by term. EBM never stops early:
every fit trains on its own training rows only (validation rows never enter an EBM
fit), for exactly max_rounds rounds, which is required (training is blocked until it
is a positive whole number). Parameters JSON also accepts learning_rate,
interactions, max_bins, max_interaction_bins, min_samples_leaf, min_hessian,
max_leaves, smoothing_rounds, interaction_smoothing_rounds, greedy_ratio,
cyclic_progress, reg_alpha, reg_lambda, max_delta_step, gain_scale,
min_cat_samples, cat_smooth and missing.
interactions is either a count (EBM picks up to that many of the strongest pairs) or
a list of feature pairs; the FEATURES pane's Pairwise interactions section sets
it for you. A monotone-constrained feature cannot take part in an interaction. MAE is
not available.
An .ebm file is loaded only with the feature contract saved beside it, and only
under the exact interpret-core version that contract records; upgrading
interpret-core past a minor version means retraining.
Tune parameters
Every family except the GLM can tune a bounded search over the validation plan set in the SPLIT pane. Choosing Tune parameters hides Parameters JSON, ticks Refit on training + validation in the SPLIT pane (tuning requires the refit) and shows:
- A note saying whether a test set is held out during tuning (No test set is reserved. Reserve one in Split for an independent evaluation. otherwise), with Review split →, and the fit budget, for example 101 total fits: 20 trials × 5 validation fits + 1 final fit.
- Trial count: how many trials to run, from 5 to 50 (20 to start). Trial zero is always your current fixed parameters, so the search must beat the model you would otherwise train.
- Seed: the sampler's seed (42 to start).
- Selection metric: the metric that picks the winner, from the metrics switched on in the TARGET pane (the first of them to start).
- Search space JSON: the values to search, keyed by the family's own parameter names.
It starts from a small family-specific search: CatBoost
depth,learning_rateandl2_leaf_reg; XGBoostmax_depth,etaandlambda; LightGBMnum_leaves,learning_rateandmin_data_in_leaf; EBMmax_rounds,learning_rateandinteractions.
For an ordinary search entry, list every candidate value directly. Values keep their JSON
type, so the engine receives numbers, strings, or Booleans exactly as written. A conditional
entry uses {"choices": [...], "when": {...}} instead (see In the pipeline file below).
Haute rejects unknown fields, lists outside two through fifty distinct finite values, invalid
or cyclic conditions, and keys Haute owns, such as a tree family's round budget,
loss/objective, device, threads, callbacks, write directories, or random seed. Training is
blocked until tuning is complete: 5–50 trials, a selection metric, a non-empty search space,
and at most 200 trial-validation fits in total.
Tuning needs holdout validation or cross-validation and evaluates every trial in turn with a
deterministic seeded sampler. The selected parameters are refitted once on all development
data; the test set, when there is one, is then evaluated once. A tree family refits with the
winner's validation-weighted round count (the final tree count); an EBM refits with the
winner's max_rounds unchanged. Choosing Fixed parameters again turns tuning off.
GLM penalty and solver¶
A GLM has no tuning. Its PARAMETERS pane holds the penalty and the solver:
| Field | What it does |
|---|---|
| Regularization | None (the default), Ridge, Lasso or Elastic Net. It cannot be combined with automatically smoothed splines: the pane says which splines conflict and asks you to set Fixed df on them or turn regularisation off. |
| Penalty | With regularisation: Cross-validated chooses the penalty strength on held-out folds of the training data; Fixed uses the alpha you enter. |
| Folds | Cross-validated penalty: the number of folds, from 2 to 20 (5 to start). |
| Selection rule | Cross-validated penalty: Minimum deviance (the default) or One standard error (the largest penalty within one standard error). |
| Alpha | Fixed penalty: a positive number (1 to start). |
| L1 ratio | Elastic Net only, and required: a slider from 0 (ridge) to 1 (lasso), 0.5 to start. |
| Solver | Opens the solver settings: Maximum iterations (1 to 10000) and Tolerance (between 0 and 1); leave them blank for RustyStats' defaults. Robust standard errors is Off or HC0, HC1, HC2 or HC3 heteroskedasticity-robust standard errors. They cannot be combined with regularisation, monotonicity constraints, or automatically smoothed splines, whose standard errors are not valid; the pane names the conflict. |
A cross-validated penalty also records a seed (42) so the same data selects the same penalty; it has no field in the pane.
The SPLIT pane¶
The SPLIT pane decides which rows the model is trained, validated and tested on. It separates three roles:
- Development data is available for model selection and the final refit.
- Validation data estimates candidate settings: a holdout validation set, cross-validation, or no validation.
- An optional test set remains unseen until model selection is complete and is evaluated exactly once.
| Field | What it does |
|---|---|
| Row limit | Trains on a seeded random sample of at most this many rows. Blank (All rows) uses every row. |
| Split strategy | Random split (the default), Group split or Time-based split. |
| Group column | Group split: the column whose rows must stay together, such as a customer, policyholder, household, or claim identifier. |
| Date column | Time-based split: the date the split boundaries apply to. |
| Validation strategy | Holdout validation (the default), Cross-validation or No validation. |
| Validation set (%) | Holdout validation with a random or group split: the share of source rows held out, above 0 and below 100 (20 to start). |
| Validation starts | Holdout validation with a time-based split: rows from this date on form the validation set. |
| Fold count | Cross-validation: from 2 to 10 folds (5 to start). With a time-based split the folds use an expanding window. |
| Refit on training + validation | Holdout validation only, ticked by default: refits the final model on all development rows. Untick it to keep the one model trained on the training rows during validation, with no second fit; without a test set, its diagnostics are then labelled as validation diagnostics. Cross-validation and no validation always perform their final fit, and tuning requires the refit, so the box is ticked and locked while Tune parameters is on. |
| Test set (%) | Random or group split: the share of source rows reserved as the test set, from 0 (no test set) to below 100. |
| Test starts | Time-based split: rows from this date on form the test set; leave it blank for none. |
Above the strategies, a bar shows the Target allocation of training, validation and test rows. Once the split settings are complete and the pipeline's rows can be counted, it becomes the Exact allocation: the planned development and test row counts and the range of training and validation rows per fit. A time-based split shows Awaiting preview until the counts arrive.
How the strategies split the rows:
- Random split: classification is stratified by the target; regression is seeded but unstratified. Every requested partition must contain rows.
- Group split: the planner assigns complete groups while balancing row counts, so rows for the same entity never appear in different partitions. The group column is kept as split metadata and is not offered as a feature.
- Time-based split: explicit date boundaries, and a later date never appears in a validation fit's training data. Equal dates always stay together; null or invalid dates fail with an actionable error.
With No validation, Haute performs one final fit and tuning is unavailable. The pane shows the reason when the split cannot be used: Validation and test must total below 100% so training retains some rows., Validation must start before the test set., or Complete the split settings: split strategy, validation strategy, and required group/date fields.
The TRAIN pane¶
Run summary lists what will train: Model (family and loss), Target, Inputs (the feature count), Categorical encoding (CatBoost with categorical features), Evaluation (the validation strategy and the allocation), Fit budget (the total fits, and for a cross-validated GLM penalty the internal folds each fit uses) and Compute (CPU or GPU (CUDA)).
GPU training
CatBoost trains on a GPU when its GPU training box is ticked. XGBoost trains on an NVIDIA GPU when its GPU training box is ticked; the box names the device once the server has found one. LightGBM, EBM and the GLM train on the CPU only.
Haute installs XGBoost's CPU-only build, so XGBoost trains on the CPU with no extra
setup. GPU training is an optional extra: if you want it, run haute gpu-setup once in the
project (see
Installing Haute) and
restart haute serve. Until the server can train on a CUDA GPU, the XGBoost box is
disabled and says why.
XGBoost never falls back to the CPU silently. A GPU training request fails before the fit
when no GPU can be used, and fails afterwards if XGBoost trained somewhere else; nothing is
saved in either case. The run summary shows GPU (CUDA) and the result's fit evidence
records the device used (for example cuda:0). Before launch, Haute estimates the GPU memory
the data needs and refuses a job that does not fit.
A GPU model is saved for CPU scoring, so it scores the same in every deployment, including ones with only the CPU build. A GPU fit is a different model from a CPU fit with the same settings: XGBoost's device algorithm finds different splits, so the predictions differ. Switching the device marks the last result as out of date.
Memory estimate
Before training, the pane estimates the RAM the data needs (Estimating dataset size... while it counts):
- Dataset fits in memory shows the Source rows, the Est. training RAM and the Available RAM (and, for a GPU fit, Est. GPU VRAM against the GPU VRAM free).
- Will downsample means the data would not fit: training uses a seeded random sample of rows capped at a limit that fits, shown as Training rows, and the result records the downsampling. A Row limit you set in the SPLIT pane still applies when it is lower.
- Row count not proven appears when the data passes through a join that has no key contract: the pane shows the upper bound instead of a verdict, and an Open "…" button for each such join. Declaring the join many-to-one bounds the rows by its base input.
- Memory estimate unavailable appears when the rows or columns reaching the node cannot be known before it runs.
When a GPU fit would need more memory than the GPU has free, the pane says so and suggests selecting the CPU or reducing rows or features.
Training
- Train Model (or Tune & Train while tuning is on) trains the model. While Haute collects the data the button reads Preparing training data...; then it shows the progress message. If anything blocks training, a Complete before training list names each problem with a Go to … link to the pane that fixes it.
- Cancel training stops a run in progress.
- While a CatBoost, XGBoost or LightGBM model trains, the pane draws its loss curve live; a tuned run also shows the trial and fold it is on and the best objective so far.
- Config changed since last training appears when you change a training setting after training, with Re-train to train again.
A finished run shows Model trained - results in preview panel below; a failed or cancelled run shows Training failed or Training cancelled with the reason.
The EXPORT pane¶
The EXPORT pane acts on the node's last completed training result. Nothing is logged or saved automatically. Both actions stay disabled until the model has been trained and while a new training run is in progress (Training is running - export is available when it completes.). If you change training settings after training, the pane warns Training settings changed since this model was trained. Exports use the last trained model.
MLFLOW LOGGING
| Field | What it does |
|---|---|
| MLflow destination | Databricks, MLflow server or Local folder (the default, the project's local MLflow folder). A remote that is not configured is greyed out, and clicking it opens MLflow settings instead. A node keeps its choice even when other destinations are configured later. |
| Experiment path | The MLflow experiment the run is logged into. On Databricks it is a workspace folder path; on an MLflow server or local folder it is a plain name. Leave it blank to use the default shown in the box: the node's name, or /Shared/haute/<node name> on Databricks. |
- Log run to MLflow logs the run. If the chosen destination is unavailable, the pane says why with a Configure MLflow link. If the destination cannot be reached or rejects its credentials, the pane shows why with Test connection in MLflow settings, and nothing is logged anywhere else instead.
- After a log, the pane shows Last logged to … with Open run, and the button becomes Log again. Logging a result that is already logged asks first (Log as a new run or Cancel) and creates a new run.
- If a log fails because the response never arrived, Retry repeats the same log rather than risking a duplicate run.
A logged run follows Haute's candidate-run contract so a separate promotion process can
find and compare it: tags such as haute.node_id, haute.trained_at and
haute.evaluation_plan_sha256, metrics named by evaluation set (final_test_gini,
development_gini, selection_gini_mean), and the model, feature contract and evaluation
evidence as artifacts.
Registering and promoting models
Haute logs candidate runs but never registers a trained model. Registering a run in the MLflow model registry, and promoting it (for example after comparing it with the current champion and moving an alias), is a separate process outside Haute. A Model Scoring node can then load the registered version or alias, or score a logged run directly.
MODEL FILE
FILENAME OR PATH chooses where Save model to file writes a copy of the trained model
and its feature contract, the same way a Data Output writes a file: type a name or pick a
file. A bare filename saves in the project's models/ folder, paths are relative to the
project root, and the model's extension (.cbm, .ubj, .lgbm, .ebm or .rsglm, by
family) is added if you leave it off. The pane shows the Destination: before you save,
refuses an extension that does not match the model format, and asks before replacing a file
that already exists (Replace existing file). After saving it shows where the model and
its Feature contract went. Keep an .ebm file's feature contract beside it: the model
cannot load without it.
The pane remembers where the result was last logged and saved, including after you reload the page while the server still holds the training result. After a server restart the pane says the result is no longer available and asks you to train again.
Reading the result¶
The finished result opens in a results panel under the canvas, with these tabs; a tab appears only when the result has something to show in it:
- Summary: the metrics, evaluation and tuning evidence described below.
- Coefficients and Relativities (GLM): each term's estimate and uncertainty, and its effect relative to a baseline of 1.
- Terms (EBM): each main effect's shape and each interaction's score table.
- Loss: training and validation loss across iterations. After a holdout validation fit and a refit, it draws the validation fit and says which fit it is.
- Lift: how well predictions separate lower and higher outcomes.
- Residuals: prediction errors and actual against predicted.
- Features: feature importance.
- AvE: actual against expected across each feature's groups, with exposure.
- PDP: partial dependence, how predictions change as one feature varies.
The Summary keeps model-selection evidence distinct from final performance:
- Test metrics are the performance on the untouched test set, when one was reserved.
- Test diagnostics, Validation diagnostics or Training diagnostics are named after the rows they were evaluated on. Without a test set they are usually training diagnostics, which are in-sample performance and are marked as such.
- When a validation fit ran and no test set was reserved, the Summary leads with a Validation, N rows card (Validation (K-fold mean), N rows under cross-validation): the out-of-sample metrics used to select the model. The collapsed results bar shows the same metrics.
- Candidate selection holds the holdout or cross-validation metrics used to compare fixed and tuned candidates. Cross-validation summaries are weighted by the number of validation rows in each fit.
- A tuned run shows the baseline, winning trial, improvement, selected parameters, final tree count (tree families) or the winning round budget (EBM), and exact total fit count. Use best as fixed parameters copies the winning parameters into the node's Parameters JSON and turns tuning off, after asking.
- An EBM's Terms view shows each main effect's shape (including the score for missing values) and each interaction's score table, as additive scores on the model's link scale. They are the model itself, not SHAP values.
- A GLM shows its fit statistics, the penalty actually applied (with the folds, rule, and seed when it was cross-validated), and each automatic spline's effective degrees of freedom. Standard errors and p-values are valid only for an unpenalised, unconstrained fit without automatic splines; otherwise the Coefficients view shows dashes and says why. Relativities (exponentiated coefficients) exist only for log-link models, and a relativity too large to compute is reported as a diagnostic error naming the terms.
The model, feature contract, evaluation plan, results and report, and the optional tuning plan, trials and report are published together as that training result's own files, so a later training run never changes what an earlier result exports. MLflow logging attaches the same evidence and selected final parameters to one final run.
Example¶
A Poisson claim-frequency GLM with an exposure offset:
- Add a Model Training node, connect your training data to it and click GLM.
- In the TARGET pane, set Target column to
claim_count, click Poisson under Family (the link stays auto (log)), and set Offset column (optional) toexposure. - In the FEATURES pane, click Add term on
driver_ageand set its Fit type to B-spline with df mode Fixed and df 5. Addvehicle_ageas Linear with decreasing monotonicity, andareaas Categorical with Reference levelurban(under Advanced). - Under INTERACTIONS, click Add interaction, pick
driver_ageandarea, leave Fit type on Product, setdriver_age's fit on the card to Linear andarea's to Categorical, and keep Include main effects ticked.areaneeds its own fit because its main effect's reference level does not apply inside an interaction. - In the PARAMETERS pane, choose Ridge with a Cross-validated penalty, Folds 5 and Selection rule Minimum deviance.
- In the TRAIN pane, check the run summary and click Train Model. The results panel opens with the Summary, Coefficients and Relativities tabs.
In the pipeline file
The node's settings are stored in a JSON sidecar, config/model_training/<node name>.json,
which the pipeline's .py file names in the node's decorator:
@pipeline.modelling(config="config/model_training/<node name>.json").
| Setting in the editor | Stored as |
|---|---|
| Select algorithm | algorithm: "catboost", "xgboost", "lightgbm", "ebm" or "glm" |
| Target column | target |
| Objective | loss_function (the loss name, for example "Poisson"), and task: "regression" or "classification" |
| Variance power (tree and EBM Tweedie) | variance_power |
| Positive class | positive_class (a Boolean, an integer or a string label) |
| Weight column (optional) | weight |
| Offset column (optional) | offset |
| Metrics | metrics: "gini", "rmse", "mae", "mse", "r2", "auc", "logloss", "poisson_deviance", "tweedie_deviance", "gamma_deviance" |
| Feature tick boxes | exclude (the columns left out; tree and EBM families) |
| Monotonicity | monotone_constraints, a map of column to -1 or 1 (tree and EBM families) |
| Pairwise interactions | params.interactions (EBM) |
| Parameters JSON | params |
| Tune parameters | tuning (absent while Fixed parameters is chosen) |
| Row limit | row_limit |
| SPLIT pane | evaluation (see below) |
| Refit on training + validation | refit_on_development (absent means ticked) |
| GPU training (CatBoost) | params.task_type: "GPU" |
| GPU training (XGBoost) | device: "gpu" (absent means CPU) |
| MLflow destination | mlflow_destination: "databricks" or "server"; absent for the local folder |
| Experiment path | mlflow_experiment (absent or blank uses the default) |
| FILENAME OR PATH | model_export_path |
| Family (GLM) | family: "gaussian", "poisson", "quasipoisson", "binomial", "quasibinomial", "gamma", "tweedie" or "negbinomial" |
| Link Function (GLM) | link: "identity", "log" or "logit"; blank for the canonical link |
| Variance power (GLM Tweedie) | var_power |
| Dispersion theta (GLM) | theta |
| Intercept (GLM) | intercept (defaults to true) |
| Term cards (GLM) | terms, a map of term name to term spec |
| Interaction cards (GLM) | interactions, a list of cards |
| Regularization (GLM) | regularization: "ridge", "lasso" or "elastic_net"; absent for None |
| Penalty, Alpha (GLM) | alpha: a positive number fixes the penalty; absent or 0 cross-validates it |
| Folds, Selection rule (GLM) | cv_folds (2 to 20), cv_selection ("min" or "1se") |
| L1 ratio (GLM) | l1_ratio |
| Maximum iterations, Tolerance (GLM) | max_iter, tol |
| Robust standard errors (GLM) | robust_standard_errors: "HC0", "HC1", "HC2" or "HC3" |
These keys have no editor control:
| Key | What it does |
|---|---|
name |
Names the training artifacts; training uses the node's id when it is absent. |
feature_columns |
An explicit feature list (tree and EBM families): Haute uses exactly those columns instead of the all-except set. |
feature_weights |
CatBoost feature weights. |
fold_column, id_columns |
Columns kept as training metadata rather than features. |
output_dir |
Where training writes its working files (outputs by default). |
cv_seed |
The GLM penalty's cross-validation seed; the editor writes 42 when you choose a regularisation. |
evaluation.seed |
The split seed; the editor writes 42. |
evaluation.validation.window |
"expanding" for time-based cross-validation; the editor sets it. |
categorical_levels |
Declared category levels for categorical feature columns. |
exclude, feature_columns and monotone_constraints apply to the tree and EBM
families (feature_weights to CatBoost only) and are ignored for a GLM, whose features
are its terms and interaction factors.
Evaluation. Every node has one versioned evaluation object. The retired top-level
split and cross_validation fields are not accepted. A random split reserving 20% of
source rows as a test set, with five validation folds, then one refit on all development
rows:
{
"evaluation": {
"schema_version": 1,
"strategy": "random",
"seed": 42,
"test": {"size": 0.2},
"validation": {"method": "cross_validation", "fold_count": 5}
}
}
A group split uses "strategy": "group" with a group_column, and a time-based split
uses "strategy": "temporal" with a date_column and date boundaries, for example
"test": {"start": "2025-01-01"}. Fractions are source-relative numbers from 0
(inclusive) to 1 (exclusive). The validation shapes are:
| Shape | Behaviour |
|---|---|
{"method": "none"} |
No candidate validation; perform one final fit. Tuning is unavailable. |
{"method": "single", "size": 0.2} |
One random or group validation fit using a source-relative fraction. |
{"method": "single", "start": "2024-07-01"} |
One time-based validation fit at an explicit boundary. |
{"method": "cross_validation", "fold_count": 5} |
Two to ten random or group validation fits. |
{"method": "cross_validation", "fold_count": 5, "window": "expanding"} |
Two to ten expanding time-based validation fits. |
Tuning. trial_count includes the baseline and must be 5–50; total trial-validation
fits may not exceed 200. A search entry is either a list of values or a conditional
entry:
{
"tuning": {
"schema_version": 1,
"trial_count": 20,
"seed": 42,
"metric": "gini",
"search_space": {
"depth": [4, 6, 8, 10],
"learning_rate": [0.01, 0.03, 0.05, 0.1, 0.2],
"grow_policy": ["SymmetricTree", "Depthwise"],
"min_data_in_leaf": {
"choices": [10, 25, 50, 100],
"when": {"grow_policy": ["Depthwise"]}
}
}
}
}
GLM terms and interactions. GLM settings sit directly on the node config, not inside
params. The example above is stored as:
{
"algorithm": "glm",
"task": "regression",
"target": "claim_count",
"offset": "exposure",
"family": "poisson",
"terms": {
"driver_age": { "type": "bs", "df": 5 },
"vehicle_age": { "type": "linear", "monotonicity": "decreasing" },
"area": { "type": "categorical", "reference": "urban" }
},
"interactions": [
{ "factors": ["driver_age", "area"], "specs": {"driver_age": {"type": "linear"}, "area": {"type": "categorical"}}, "include_main": true }
],
"intercept": true,
"regularization": "ridge",
"cv_folds": 5,
"cv_selection": "min",
"cv_seed": 42
}
A native term is keyed by the column it fits; an "expression" term
({"type": "expression", "expr": "age ** 2"}) or a second encoding
({"type": "frequency_encoding", "variable": "area"}) is keyed by a name that is not a
column. Term types are linear, categorical, bs, ns, ms, target_encoding,
frequency_encoding and expression; spline settings are df, degree, k, knots
and boundary_knots, categorical settings reference and levels, target-encoding
settings prior_weight and n_permutations, and monotonicity is "increasing" or
"decreasing". An interaction has factors (two or more columns), include_main,
optional per-feature specs, and encoding: absent for Product, or
"target_encoding" or "frequency_encoding":
See also:
- Model Scoring - to score data with your trained model
- Execution Strategy - how training data is collected and diagnosed