Data and SQL
The principle fits in one sentence: your files stay on your object storage. graal catalogs them and makes them queryable; it does not host them and does not copy them elsewhere.
Buckets and explorer
Section titled “Buckets and explorer”Each project has buckets on the S3-compatible storage you provide. The explorer serves the two everyday gestures: dropping a dataset, and looking at what a run has just produced.
Jobs write there in Parquet by default. It is an open format, readable by Spark, Pandas, DuckDB and Trino without conversion.
Catalog: layers, databases, tables, fields
Section titled “Catalog: layers, databases, tables, fields”The catalog is organised in four levels:
- layer — the refining stage (raw, cleaned, exposed, per your own convention);
- database — the logical grouping inside a layer;
- table — the dataset itself, backed by files in the bucket;
- field — the columns, their type and their description.
It is not just a directory: it is what makes it possible to know what you have and who may touch it, project by project. That is the substantive difference with a datalab, where knowledge of the data lives in the head of whoever dropped it.
Rights are set per layer, per database and per table, with the same roles and assignments as the rest of the platform (see roles and permissions).
Apache Iceberg tables
Section titled “Apache Iceberg tables”Tables use the Apache Iceberg format, on your object storage: transactions, schema evolution and time travel, in a format the whole ecosystem reads.
The Iceberg catalog is a JDBC catalog that lives in graal’s PostgreSQL database, in a dedicated schema per tenant: it adds no service to host. Trino, Spark and PyIceberg all read it. When a job writes a table, its fields appear in the graal catalog with no manual entry.
SQL warehouse: Trino
Section titled “SQL warehouse: Trino”The warehouse queries your tables in SQL with Trino, deployed by the platform chart in the application namespace. A catalog table becomes queryable with no copy and no prior loading, and Trino joins your Iceberg tables with your existing databases.
Queries run from the console’s SQL editor, from a notebook, or from your BI tool through Trino’s standard JDBC access. Existing Hive-format Parquet or ORC tables are read on the same catalog.
Lineage
Section titled “Lineage”The catalog tells where each table comes from: lineage is followed at column level, from the job that produced it to the sources it read. The site’s Lakehouse and SQL page also presents AI-generated dataset documentation.