Data Output¶
You want to save results - scoring a full dataset and writing the output to parquet, CSV or a database table for downstream analysis.
When to use
Use this for batch scoring - processing a full dataset and saving the results. For live API responses, use Quote Response instead.
This node takes a single input and passes it on unchanged: running or previewing the pipeline passes data through a Data Output without writing anything. It writes only when you click Write.
The CONFIG tab¶
| Field | What it does |
|---|---|
| PROVIDER | Where the data goes: File (the default), Database or Lakehouse. Changing it keeps only the settings the new provider uses. |
| FORMAT | The format to write, from the formats the provider supports (see the table below). A format whose writer needs a package that is not installed is listed with needs one of: and the package, and a format Polars marks as unstable is listed with (unstable). |
| MODE | sink (streaming) or write (in memory). Shown when the format supports more than one mode, or once a mode is set (a new node starts on sink). Each format offers only the modes Polars supports for it. |
| FILENAME OR PATH | File provider: the file to write. Type a name or pick one from the file browser. Paths are resolved from your project folder, not from the pipeline file's folder. A bare file name goes in the project's outputs/ folder, and the format's extension is added when the name has none: scored_policies with Parquet selected writes outputs/scored_policies.parquet. A relative path such as exports/scored.parquet is taken from the project folder, and an absolute path must stay inside the project. Missing folders are created. |
| TABLE LOCATOR | Lakehouse provider: the table to write. |
| CONNECTION ENVIRONMENT REFERENCE | Database provider: the name of an environment variable that holds the database URI. |
| CREDENTIAL-FREE URI | Database provider: a database URI without credentials, used instead of the connection reference. |
| TABLE | Database provider: the table to write. |
| ARGUMENTS | Extra keyword arguments for the Polars writer. Add argument adds a row with a name (suggested from the arguments the writer accepts) and a JSON value (for example "," or true). Only arguments the writer accepts are allowed; an unknown name or a value that is not valid JSON is flagged. |
The formats each provider offers:
| PROVIDER | FORMAT |
|---|---|
| File | Parquet, CSV, NDJSON, JSON, Arrow IPC / Feather, Arrow IPC stream, Avro, Excel |
| Lakehouse | Delta Lake, Iceberg |
| Database | Database (URI) |
Under the fields, the panel says how the format is written (a streaming sink or an eager writer) and how it is published.
Writing¶
When the settings are complete, the panel shows the resolved Destination: before you write, and warns if the destination's extension does not match the selected format.
- Write runs the pipeline up to the node and writes the full dataset to the destination. It is disabled until the settings are complete, and reads Writing... while it runs. When it finishes, the panel reports the result with the number of rows and the path.
- If the destination file already exists, the panel asks you to confirm with Replace existing file before overwriting it. For a lakehouse or database destination, what happens when the table already exists is the Polars writer's own behaviour, which you can set with ARGUMENTS (for example
if_table_existsfor a database table).
The COLUMNS tab¶
By default every input column is written; to write only some of them, untick the others on the COLUMNS tab, with no upstream node needed (see Working with any node).
Example¶
Writing a scored dataset to parquet:
- Connect the scored data to a Data Output node.
- Leave PROVIDER on File, FORMAT on Parquet and MODE on sink.
- Type
scored_policiesin FILENAME OR PATH. The panel shows the destination,outputs/scored_policies.parquet. - Click Write.
Not part of a deployed API
A deployed pricing API never writes: its Data Output nodes pass data through without invoking the writer.
Multiple outputs
You can have multiple Data Output nodes in a pipeline - for example, to write both a parquet file and a CSV, or to save results at different stages.
In the pipeline file
The node's settings are stored in a JSON sidecar, config/data_output/<node name>.json, which the pipeline's .py file names in the node's decorator: @pipeline.data_output(config="config/data_output/<node name>.json"). Every configuration chooses exactly one destination with outputType.
| Setting in the editor | Stored as |
|---|---|
| PROVIDER | outputType: "file", "lakehouse" or "database" |
| FORMAT | format: "parquet", "csv", "ndjson", "json", "ipc", "ipc_stream", "avro" or "excel" for a file; "delta" or "iceberg" for a lakehouse; "database" for a database |
| MODE | mode: "sink" or "write"; absent uses the format's usual mode |
| FILENAME OR PATH, TABLE LOCATOR | path |
| CONNECTION ENVIRONMENT REFERENCE | connection |
| CREDENTIAL-FREE URI | uri |
| TABLE | table |
| ARGUMENTS | arguments, a map of argument name to value |
| COLUMNS tab | selected_columns |
The column metadata any node may carry, column_renames and categorical_levels, has no
editor control on this node.
The example above is stored as:
See also:
- Quote Response - define the API response for live pricing