Node Types¶
Every step in a Haute pipeline is a node. You connect nodes on the canvas to define how data flows from source to output. Each node type is described on its own page.
First pipeline?
If you're building your first pipeline, a common path is: Quote Input or Data Input → Polars (clean your data) → Banding and Rating Step (build your rating structure) → Quote Response. You don't need every node type to get started.
How these pages are written
Each page walks through the node's panel in the editor: what each tab, field and button is for. How the same settings are stored in the pipeline file is in a closed In the pipeline file section at the end of each page, for when you read the generated code or review a change.
Quick reference¶
| I want to... | Use this node |
|---|---|
| Bring in quote data for live pricing | Quote Input |
| Load a parquet or CSV file, a database, lakehouse or Databricks table | Data Input |
| Store fixed parameters (tax rate, loadings) | Constant |
| Join, filter, or calculate new columns | Polars |
| Join another dataframe into an existing connection | Edge Join |
| Convert ages or values into bands | Banding |
| Look up rating factors from a table | Rating Step |
| Score data with a trained model | Model Scoring or Load File |
| Train a new model | Model Training |
| Optimise prices subject to constraints | Expander + Optimisation |
| Apply saved optimisation results | Apply Optimisation |
| Switch between live and batch data | Source Switch |
| Choose which columns to return from the API | Quote Response |
| Profile a dataset, build pivot tables and charts | Explore |
| Save results to a file or table | Data Output |
| Group nodes into a reusable block | Submodel |
| Reuse a node's logic with different inputs | Instances |
Working with any node¶
- Adding and connecting nodes. Drag a node from the Nodes palette on the left onto the canvas, then drag a connection from one node to the next. A connection carries the upstream node's data, under the upstream node's name.
-
The node panel. Click a node to open its panel on the right. Most nodes have tabs along its top:
- CONFIG holds the node's own settings, described on its page. On a Polars node this tab is called POLARS, because the node's settings are its steps.
- POLARS, on Data Input, Load File, Expander, Rating Step and Model Scoring nodes, adds optional steps that run on the node's result, built the same way as a Polars node's (see Building the node from steps).
- COLUMNS lists the node's Output Columns: untick a column to stop the node passing it on. Filter columns... finds a column by name, and All and None tick or clear every box. The list appears once the node has been previewed. Quote Input, Quote Response, Model Training, Optimisation, Explore and Submodel nodes have no Columns tab.
Model Training, Optimisation and Explore nodes split their settings into panes instead, described on their pages. - The preview. Under the canvas, the preview shows the selected node's output. With Calculation set to Automatic in Pipeline settings (the toolbar's Pipeline button), clicking a node calculates its preview; set to Manual, a node shows its last result until you click Refresh (Ctrl+Enter). Preview rows in Pipeline settings sets how many rows a preview shows (0 means no limit), and Search columns... narrows the columns shown. Click a cell to trace how its value was calculated (see Price tracing). - Renaming, copying and deleting. Right-click a node for Rename, Duplicate and Delete. A node's name is also the name downstream nodes use for its data.
Example: your first pricing pipeline¶
Below is a simple motor pricing pipeline that takes in quote data, enriches it, applies rating factors, and returns a premium. Each box is a node on the canvas, and the arrows show the direction data flows.
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Quote Input │────▶│ Polars │────▶│ Banding │
│ (load quotes) │ │ (vehicle_age) │ │ (driver_age) │
└──────────────┘ └──────────────┘ └──────────────┘
│
▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│Quote Response │◀────│ Polars │◀────│ Rating Step │
│ (return cols) │ │ (premium) │ │ (factors) │
└──────────────┘ └──────────────┘ └──────────────┘
Step by step:
-
Quote Input - Loads sample quotes containing
driver_age,area,vehicle_value, andyear_of_manufacture. This is the entry point of the pipeline. -
Polars - Calculates
vehicle_agefromyear_of_manufacture(e.g. current year minus manufacture year). This adds a new column to the table. -
Banding - Bands
driver_ageinto groups: 18–25, 26–65, and 65+. The output is a new column (driver_age_band) that the rating step can look up against. -
Rating Step - Looks up
area_factorandage_factorfrom rating tables usingareaanddriver_age_band, then multiplies them into acombined_factor. -
Polars - Calculates the final premium:
final_premium = base_rate * combined_factor. This is a single expression that produces the price. -
Quote Response - Selects the columns to return from the API:
quote_id,final_premium,area_factor, andage_factor.
Try it yourself
You can recreate this pipeline on the canvas in a few minutes. Start with a Quote Input node, then chain each step by dragging a connection from one node's output to the next node's input.
Inputs¶
Nodes that bring data into your pipeline. They have no upstream connections.
- Quote Input - entry point for live pricing; reads a preview file during development
- Data Input - reads files (parquet, CSV and more), lakehouse, database or Databricks tables, or inline records
- Constant - stores fixed values like expense loadings or tax rates
Transforms¶
- Polars - general-purpose node for joins, filters, and calculations
- Edge Join - compact join node created from canvas connections
- Banding - bands numeric or date values by breakpoints, or categorical values by mapping
- Rating Step - looks up rating factors from tables and combines them
- Expander - generates a range of candidate values for each row (used with Optimisation)
- Source Switch - toggles between live and batch data sources
Models¶
- Model Training - trains a CatBoost, XGBoost, LightGBM, EBM or GLM model
- Model Scoring - scores data with an MLflow-managed model
- Load File - loads and scores a standalone model file
Optimisation¶
- Optimisation - find the best price per quote (or the best factor table) subject to your constraints
- Apply Optimisation - apply the saved results to new data at deployment time
Pipeline outputs¶
- Quote Response - chooses which columns to return in the API response
- Data Output - saves results to a file, lakehouse or database table
Analysis¶
- Explore - profiles the data at any step and builds pivot tables and charts, without feeding the pipeline
Organisation¶
- Submodel - groups nodes into a collapsible, reusable block
- Instances - reuse a node's logic with different inputs
Key terms¶
A quick glossary for terms you'll see throughout these docs.
| Term | What it means |
|---|---|
| Pipeline | A chain of connected steps that transforms data into a price. |
| Node | A single step in the pipeline. It might be a data source, a calculation, a model score, or an output. Each node type has its own page in this section. |
| Canvas | The visual editor workspace where you drag, drop, and connect nodes to build your pipeline. |
| DataFrame / df | A table of data (rows and columns). In code, df is shorthand for this. |
| Polars | The data engine Haute uses under the hood. When you see pl.col("x") in code, it means "the column called x." Think of it as a formula language for tables. |
| Parquet | A file format for tabular data, like CSV but faster and smaller. You don't need to understand the internals - just know it's a data file. |
| MLflow | An open-source platform Haute uses to track model training experiments and store trained models. Think of it as version control for models. |
| JSON | A text-based data format. Haute uses it for configuration files and preview data. You won't usually write it by hand - the UI generates it. |
| Terminal node | A node that produces a result (a trained model, an optimisation output) but doesn't pass data to the next node in the pipeline. |