Skip to content

Capabilities overview

graal covers the whole data & AI cycle, from the raw source to the agent in production. This page groups its capabilities into ten families and points, for each one, to the site page that presents it and to the documentation page that details it.

Family What it covers On the site In the documentation
1. Ingestion and connectors Databases, files and S3, SFTP, REST APIs, Hadoop, SAS, change data capture Connectors Secrets and libraries
2. Low-code pipelines Visual editor, Pandas and PySpark export, per-step preview, quality checks Low-code pipelines Jobs and workflows
3. Orchestration Jobs, graph workflows, cron, events, backfills, retries, alerts Orchestration Jobs and workflows
4. Compute Distributed Spark, instance types, GPUs, restricted pods with no cluster rights Orchestration and compute Runtime images
5. Lakehouse and SQL Apache Iceberg tables, federated SQL with Trino, catalog, SQL editor, lineage Lakehouse and SQL Data and SQL
6. Notebooks Jupyter and VS Code in the browser, SSO, code assistant Notebooks Workspaces
7. ML and generative AI Experiments, model registry, serving, drift, LLM gateway, RAG Machine learning · Generative AI Machine learning
8. AI agents and governance MCP server, service accounts, human approval, SSO, RBAC, audit, costs AI agents · Governance AI agents (MCP) · Roles and permissions
9. Deployment and operations Helm chart, high availability, offline, OpenShift, GitOps, observability, multi-tenancy Deployment · Architecture Installation · Architecture
10. Openness and reversibility Exported code, open formats, OpenAPI REST API, MLflow API, MCP Sovereignty REST API · Sovereignty and architecture

PostgreSQL, Oracle, SQL Server and MySQL databases; CSV and Parquet files on S3; SFTP; REST APIs as a source; a Hadoop and HDFS bridge; reading SAS data (.sas7bdat); change data capture (CDC). The credentials of a source are project secrets, never fields of the pipeline.

A pipeline is drawn by drag and drop and exported as Pandas or PySpark: the code can be read, versioned and run without graal. Each step can be previewed on a sample, quality checks (not null, uniqueness, range, freshness) are placed as steps, and a pipeline can be described in natural language before it is reviewed.

Python, Spark, SQL, bash and notebook jobs, SAS programs and specialised types (dbt and others). Graph workflows, cron scheduling, event triggers (a file arriving, another workflow finishing), backfills over a date range, automatic retries and alerts.

Distributed Spark on your Kubernetes, instance types per project, GPU scheduling. Every pod follows the PSA restricted profile, and the platform needs no cluster administration right.

Apache Iceberg tables on your S3 storage, federated SQL with Trino, a catalog of layers, databases, tables and fields, a SQL editor in the console, column-level lineage and AI-generated dataset documentation.

Jupyter and VS Code in the browser, behind the same authentication as the console, with access to the project’s tables and buckets and a code assistant connected to a model hosted on your side.

Experiment tracking and model registry, compatible with the MLflow client; models served over REST behind your SSO, drift monitoring; an LLM gateway to local models or to providers you choose; RAG on your documents.

An MCP server your agents drive under their own service account, a live activity feed, revoking an agent in one gesture, human approval of sensitive actions and per-agent policies. OIDC, SAML and LDAP/AD SSO, RBAC per project and per group, an audit log of human and agent actions, encrypted secrets, quotas and costs per project.

A Helm chart, installed on your Kubernetes or your OpenShift, on premises, in a private cloud or with a SecNumCloud-qualified or HDS-certified hosting provider, including offline. High availability, GitOps, Prometheus metrics, OpenTelemetry, backup and restore, one realm per tenant, white labelling.

Pipeline code is exported, data stays in open formats (Parquet, Iceberg, S3, SQL), and everything the console does goes through a REST API described in OpenAPI. On top of that: the MLflow API, the MCP protocol, Git versioning of projects, a Terraform provider and a command line.

The former graal site published tutorials and per-tool integrations. They have been replaced by the families above:

  • dbt, Dask, Beam, Airflow: the jobs and workflows of family 3, Orchestration;
  • Debezium: the change data capture of family 1, Connectors;
  • Jenkins, Azure DevOps, Azure Data Factory: the REST API and the GitOps of family 9;
  • Migrating from Databricks: the platform migration use case.