Skip to main content
A dataset publishes one of your tables to an Apache Iceberg table in your own catalog and object storage. Bijection keeps it current, records every version, and retains exact snapshots for as long as a reader needs them. Your own services read those tables for bulk work that does not belong in your app: training, simulation, warehouse analytics, large spatial joins, raster processing. Your app keeps the operational database. A dataset is a copy that other systems read; results come back through customer compute, never by writing to your tables directly.

Declare a dataset

Install the datasets component once:
bijection/bijection.config.ts
Then declare which columns of a table to publish:
bijection/datasets.ts
Each column keeps a stable Iceberg field ID across deployments. Exact integers, decimals, timestamps and geometry keep their exact values: a geometry column is ISO WKB in OGC:CRS84 with GeoParquet metadata, so DuckDB, Spark, Sedona and GeoPandas read it directly.

Bind it to your catalog

Register your catalog once with a stored credential, then bind the dataset to a table. Bijection creates the table if it does not exist:
--check validates a binding against the live catalog without submitting it.

Read an exact snapshot

A checkpoint retains exact versions as a named Iceberg tag on every selected table, so a job reads a fixed snapshot even while publication continues:
The manifest names each table’s snapshot and frozen metadata file; it carries no storage credentials. Your job reads with its own catalog identity:
Release the checkpoint when every reader has finished:
With --lease, the checkpoint releases itself when the lease ends unless you retain it again, so a job that dies does not hold its snapshots forever.