Architecture
The specification, for your platform team
What graal installs, what it creates at runtime, what it never creates, and what you provide once: everything you need to validate an installation before the first command.
The result
Two namespaces, and nothing above them
graal installs into an existing Kubernetes or OpenShift cluster, with no cluster-scoped object.
- namespaces, provided by your platform team
- 2
- CRDs: no custom resources
- 0
- cluster-admin rights, no cluster-wide role
- 0
- of pods under PSA restricted, workloads included
- 100 %
Components
What runs, and where
Every image is pinned by digest and pulled from the registry you designate.
| Component | Role | Namespace |
|---|---|---|
| Console | Web interface for your teams: projects, pipelines, jobs, catalog, models, agents. | graal |
| graal API | REST contract described in OpenAPI; permissions, audit, secrets and workload orchestration. | graal |
| LLM gateway | Built into the API: local models or providers you choose, token quotas, logs, personal data masking. | graal |
| MCP server | Entry point for AI agents, under a service account, with the same permissions as the API. | graal |
| Keycloak | Identity: one realm per organization, OIDC, SAML or LDAP federation. | graal |
| PostgreSQL | Metadata, catalog, encrypted secrets, scheduling. | graal |
| Low-code engine | Turns a drawn pipeline into Pandas or PySpark code. | graal |
| Trino | Federated SQL over your Apache Iceberg tables. | graal |
| Application gateway | SSO access to Jupyter and VS Code notebooks and to served models. | graal |
| Workloads | Jobs, workflows, Spark, notebooks and served models, started by the API. | graal-run |
Flows
Who talks to whom
Two ways in, your teams and your agents, and a single set of permissions.
01
Your teams
The browser reaches the console and the API behind your SSO (OIDC or SAML). Every request carries the person’s identity and their permissions on the project.
02
Your AI agents
The MCP client talks to the MCP server, which calls the API under a service account scoped to one project. Sensitive actions wait for human approval.
03
Workloads
The API creates pods in graal-run. Each run gets a short-lived token and a network policy that limits its egress to your storage, the API and Trino.
04
Data
Read and written on your S3 storage, as Parquet files and Iceberg tables. It is never copied outside your cluster.
05
Secrets
Encrypted in the database (AES-256-GCM) and delivered to the workload in memory at startup. They appear neither in pod specs nor in logs.
06
Observability
Prometheus metrics and OpenTelemetry traces to your own tools. The trace context follows the request all the way to the run.
Kubernetes objects
What graal creates, and what it never creates
Every pod, workloads included, runs under PSA restricted: non-privileged user, all capabilities dropped, RuntimeDefault seccomp profile.
At runtime, in graal-run
- Pods and Jobs for workloads, in graal-run
- Services and ConfigMaps belonging to a run
- One NetworkPolicy per run, created before its pods
- Deployments for notebooks and served models
- A common owner per run: a single delete cleans everything up
Never
- A namespace
- A CRD (CustomResourceDefinition)
- A ClusterRole or ClusterRoleBinding
- A ServiceAccount, Role or RoleBinding: you provide them
- Any object outside its two namespaces
Your platform team
What you provide, once
It is the first question a platform lead asks; the answer fits in five lines.
- Two namespaces
- graal for the platform, graal-run for workloads, with your quotas and your baseline network policy.
- A service account and a namespaced role
- The role grants nothing outside the two namespaces; the exact list of its verbs is in the installation documentation.
- A PostgreSQL database
- Yours, or the one shipped with the chart.
- S3 storage and an image registry
- graal hosts neither: it connects to them.
- A domain name and a TLS certificate
- For the console, the API, the MCP server and the application gateway.
Variables, probes and chart values: installation documentation.
Operations
What your platform team operates
- High availability
- The API runs as several replicas; scheduling and run recovery are coordinated by a database lease, with no duplicate and no lost run.
- Backup
- PostgreSQL archived continuously to a dedicated bucket; restoration is documented and timed.
- Offline
- OCI chart, image list derived from the rendered manifests, copy with digests and signatures; no network installation at runtime.
- OpenShift
- User and group IDs left to OpenShift; images compatible with an arbitrary UID.
Have your platform team review the architecture
The detailed architecture pack and the rendered chart are sent to you on request, directly.