Apache Iceberg tables
Transactions, schema evolution and rollback to an earlier state, on your S3, in a format the whole ecosystem can read.
Lakehouse and SQL
Your data stays in your object storage, as Apache Iceberg tables. Trino queries it in SQL, together with your existing databases, without copying. A shared catalog tells what each table holds, where it comes from and who may read it.
Key capabilities
Transactions, schema evolution and rollback to an earlier state, on your S3, in a format the whole ecosystem can read.
One query joins your Iceberg tables and your PostgreSQL, Oracle or SQL Server databases, without moving the data.
Read-only queries, paginated results, bounded time and volume: exploration never brings the cluster to its knees.
Layers, databases, tables and fields, with their descriptions and full-text search. It updates at the end of every run.
For each column, the source columns and the processing that produced it, captured automatically at run time.
graal profiles a sample and proposes a description for a table and its fields. A human validates it before it is published.
How it works
Step 01
Pipelines and jobs write Iceberg, layer by layer: raw, cleaned, ready to use.
Step 02
Every table written joins the catalog with its schema, owner and lineage.
Step 03
Analysts, notebooks, applications and agents query the same tables in SQL, under the same permissions.
The data is Parquet files in your object storage; Iceberg adds the metadata that turns them into transactional tables. Compute comes on demand: Trino for interactive SQL, Spark for heavy processing. You size one without touching the other.
The graal catalog is the one read by the SQL editor, notebooks, pipelines and agents. A table described once is described for everyone, and its permissions apply whichever way it is accessed.
Standards and integrations
Governance
No. Tables live in your S3-compatible storage, and metadata in the PostgreSQL database of your installation.
Yes. Iceberg is an open format, and graal exposes an Iceberg REST catalog: Spark, PyIceberg or another compatible engine reads them directly.
No. Trino queries them where they are. You move to Iceberg what benefits from it, at your own pace.
A query joining regional consumption and temperature: it is one of the acts of the demonstration.