Quick start¶
Use one of the following examples after installing the Python package to get started:
TBoostRegressor¶
import numpy as np
import polars as pl
from t_boost import TBoostRegressor
# initialize data
rng = np.random.default_rng(0)
n = 10_000
data = pl.DataFrame({
"DriverAge": rng.integers(18, 80, n).astype(np.float32),
"VehicleAge": rng.integers(0, 20, n).astype(np.float32),
"Region": rng.choice(["North", "South", "East", "West"], n),
"Exposure": rng.uniform(0.1, 1.0, n),
})
rate = np.exp(-2.0 + 0.5 * (data["DriverAge"].to_numpy() < 25))
data = data.with_columns(ClaimCount=rng.poisson(rate * data["Exposure"].to_numpy()))
train_data, test_data = data[:8_000], data[8_000:]
# specify the training parameters
model = TBoostRegressor(objective="poisson")
# train the model
model.fit(train_data, "ClaimCount", exposure="Exposure")
# make the prediction using the resulting model
preds = model.predict(test_data)
print(preds)
The target and the exposure are given as column names of the training frame, and those
columns are not used as features. Region is a String column, so it is treated as a
categorical feature automatically. At prediction time,
columns are matched by name and extra columns are ignored.
For a Poisson, Gamma or Tweedie model predict returns the rate per unit of exposure. The
expected number of claims for a row is predict(X) * exposure.
TBoostClassifier¶
import numpy as np
import polars as pl
from t_boost import TBoostClassifier
# initialize data
rng = np.random.default_rng(0)
train_data = pl.DataFrame({
"Age": rng.integers(18, 80, 1000).astype(np.float32),
"Channel": rng.choice(["Broker", "Direct", "Online"], 1000),
"Lapsed": rng.integers(0, 2, 1000),
})
test_data = train_data.drop("Lapsed").head(5)
model = TBoostClassifier()
# train the model
model.fit(train_data, "Lapsed")
# make the prediction using the resulting model
preds_class = model.predict(test_data)
preds_proba = model.predict_proba(test_data)
print("class = ", preds_class)
print("proba = ", preds_proba)
A target with two classes trains a logistic model. A target with three or more classes trains a softmax model, with one set of rating tables per class.
Rating tables¶
The fitted model is a set of rating tables. Continue the TBoostRegressor example to look at them:
import json
tables = json.loads(model.tables(train_data)) # the rating tables
for table in tables["tables"]:
print(table["feature_names"], table["shape"])
contributions = model.predict_contributions(test_data.head(5)) # per-prediction breakdown
importances = model.feature_importances_ # share of variance per feature
For each prediction, the base value plus the sum of the contributions equals the raw score on the link scale, so the explanation is exact rather than estimated. See Model analysis for details.
Note
t-boost computes in 32-bit floating point. Numeric feature columns are converted to
float32 before fitting and scoring, and fit issues a PrecisionWarning once per
estimator if any numeric feature column has another dtype. Cast the columns to
pl.Float32 to skip the conversion.