Capabilities overview
graal covers the whole data & AI cycle, from the raw source to the agent in production. This page groups its capabilities into ten families and points, for each one, to the site page that presents it and to the documentation page that details it.
The ten families
Section titled “The ten families”| Family | What it covers | On the site | In the documentation |
|---|---|---|---|
| 1. Ingestion and connectors | Databases, files and S3, SFTP, REST APIs, Hadoop, SAS, change data capture | Connectors | Secrets and libraries |
| 2. Low-code pipelines | Visual editor, Pandas and PySpark export, per-step preview, quality checks | Low-code pipelines | Jobs and workflows |
| 3. Orchestration | Jobs, graph workflows, cron, events, backfills, retries, alerts | Orchestration | Jobs and workflows |
| 4. Compute | Distributed Spark, instance types, GPUs, restricted pods with no cluster rights | Orchestration and compute | Runtime images |
| 5. Lakehouse and SQL | Apache Iceberg tables, federated SQL with Trino, catalog, SQL editor, lineage | Lakehouse and SQL | Data and SQL |
| 6. Notebooks | Jupyter and VS Code in the browser, SSO, code assistant | Notebooks | Workspaces |
| 7. ML and generative AI | Experiments, model registry, serving, drift, LLM gateway, RAG | Machine learning · Generative AI | Machine learning |
| 8. AI agents and governance | MCP server, service accounts, human approval, SSO, RBAC, audit, costs | AI agents · Governance | AI agents (MCP) · Roles and permissions |
| 9. Deployment and operations | Helm chart, high availability, offline, OpenShift, GitOps, observability, multi-tenancy | Deployment · Architecture | Installation · Architecture |
| 10. Openness and reversibility | Exported code, open formats, OpenAPI REST API, MLflow API, MCP | Sovereignty | REST API · Sovereignty and architecture |
Family by family
Section titled “Family by family”1. Ingestion and connectors
Section titled “1. Ingestion and connectors”PostgreSQL, Oracle, SQL Server and MySQL databases; CSV and Parquet files on S3; SFTP; REST APIs as
a source; a Hadoop and HDFS bridge; reading SAS data (.sas7bdat); change data capture (CDC). The
credentials of a source are project secrets, never fields of the pipeline.
2. Low-code pipelines to code
Section titled “2. Low-code pipelines to code”A pipeline is drawn by drag and drop and exported as Pandas or PySpark: the code can be read, versioned and run without graal. Each step can be previewed on a sample, quality checks (not null, uniqueness, range, freshness) are placed as steps, and a pipeline can be described in natural language before it is reviewed.
3. Orchestration
Section titled “3. Orchestration”Python, Spark, SQL, bash and notebook jobs, SAS programs and specialised types (dbt and others). Graph workflows, cron scheduling, event triggers (a file arriving, another workflow finishing), backfills over a date range, automatic retries and alerts.
4. Compute
Section titled “4. Compute”Distributed Spark on your Kubernetes, instance types per project, GPU scheduling. Every pod follows the PSA restricted profile, and the platform needs no cluster administration right.
5. Lakehouse and SQL
Section titled “5. Lakehouse and SQL”Apache Iceberg tables on your S3 storage, federated SQL with Trino, a catalog of layers, databases, tables and fields, a SQL editor in the console, column-level lineage and AI-generated dataset documentation.
6. Notebooks and workspaces
Section titled “6. Notebooks and workspaces”Jupyter and VS Code in the browser, behind the same authentication as the console, with access to the project’s tables and buckets and a code assistant connected to a model hosted on your side.
7. ML and generative AI
Section titled “7. ML and generative AI”Experiment tracking and model registry, compatible with the MLflow client; models served over REST behind your SSO, drift monitoring; an LLM gateway to local models or to providers you choose; RAG on your documents.
8. AI agents and governance
Section titled “8. AI agents and governance”An MCP server your agents drive under their own service account, a live activity feed, revoking an agent in one gesture, human approval of sensitive actions and per-agent policies. OIDC, SAML and LDAP/AD SSO, RBAC per project and per group, an audit log of human and agent actions, encrypted secrets, quotas and costs per project.
9. Deployment and operations
Section titled “9. Deployment and operations”A Helm chart, installed on your Kubernetes or your OpenShift, on premises, in a private cloud or with a SecNumCloud-qualified or HDS-certified hosting provider, including offline. High availability, GitOps, Prometheus metrics, OpenTelemetry, backup and restore, one realm per tenant, white labelling.
10. Openness and reversibility
Section titled “10. Openness and reversibility”Pipeline code is exported, data stays in open formats (Parquet, Iceberg, S3, SQL), and everything the console does goes through a REST API described in OpenAPI. On top of that: the MLflow API, the MCP protocol, Git versioning of projects, a Terraform provider and a command line.
If you arrive here from an old link
Section titled “If you arrive here from an old link”The former graal site published tutorials and per-tool integrations. They have been replaced by the families above:
- dbt, Dask, Beam, Airflow: the jobs and workflows of family 3, Orchestration;
- Debezium: the change data capture of family 1, Connectors;
- Jenkins, Azure DevOps, Azure Data Factory: the REST API and the GitOps of family 9;
- Migrating from Databricks: the platform migration use case.