> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bijection.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Datasets

> Publish your tables to your own Apache Iceberg lake and read them at exact, retained snapshots

A dataset publishes one of your tables to an Apache Iceberg table in your own
catalog and object storage. Bijection keeps it current, records every version,
and retains exact snapshots for as long as a reader needs them. Your own services
read those tables for bulk work that does not belong in your app: training,
simulation, warehouse analytics, large spatial joins, raster processing.

Your app keeps the operational database. A dataset is a copy that other systems
read; results come back through [customer compute](/datasets/customer-compute),
never by writing to your tables directly.

## Declare a dataset

Install the datasets component once:

```ts bijection/bijection.config.ts theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import { defineApp } from "bijection/server";
import datasets from "@bijection/datasets/bijection.config.js";

const app = defineApp();
app.use(datasets);
export default app;
```

Then declare which columns of a table to publish:

```ts bijection/datasets.ts theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
import { defineDataset } from "@bijection/datasets";
import { tables } from "./_generated/tables";

export const orders = defineDataset({
  key: "orders",
  source: tables.orders,
  select: ["location", "quantity"],
});
```

Each column keeps a stable Iceberg field ID across deployments. Exact integers,
decimals, timestamps and geometry keep their exact values: a geometry column is
ISO WKB in OGC:CRS84 with GeoParquet metadata, so DuckDB, Spark, Sedona and
GeoPandas read it directly.

## Bind it to your catalog

Register your catalog once with a stored credential, then bind the dataset to a
table. Bijection creates the table if it does not exist:

```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
bijection dataset catalog register analytics --credential-ref ANALYTICS_CATALOG
bijection dataset bind orders --catalog analytics --table reporting.orders --wait
```

`--check` validates a binding against the live catalog without submitting it.

## Read an exact snapshot

A checkpoint retains exact versions as a named Iceberg tag on every selected table,
so a job reads a fixed snapshot even while publication continues:

```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
bijection dataset checkpoint retain nightly --datasets orders --lease 2h --wait --output manifest.json
```

The manifest names each table's snapshot and frozen metadata file; it carries no
storage credentials. Your job reads with its own catalog identity:

```sql theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
SELECT * FROM reporting.orders VERSION AS OF 'nightly'
```

Release the checkpoint when every reader has finished:

```bash theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
bijection dataset checkpoint release nightly
```

With `--lease`, the checkpoint releases itself when the lease ends unless you
retain it again, so a job that dies does not hold its snapshots forever.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.