A sovereign LLM gateway
An endpoint compatible with the OpenAI API, in front of local models on your GPUs or the providers you allow for your organization.
Generative AI
An LLM gateway serves the models you choose, hosted on your GPUs or with the provider of your choice, with quotas, a log and personal data masking. RAG answers over your documents, cites its sources and respects everyone’s permissions.
Key capabilities
An endpoint compatible with the OpenAI API, in front of local models on your GPUs or the providers you allow for your organization.
A token budget per project, a clear refusal beyond it, and a log of calls, metadata only by default.
French social security numbers, IBANs, phone numbers, emails, company IDs: personal data is masked before anything is sent to an external provider.
Your PDFs and tables feed a vector index in your PostgreSQL. Every answer cites its sources, and every index has its permissions.
An assistant in Jupyter and VS Code, and AI-proposed documentation for catalog tables, on the model you chose.
An operations agent that diagnoses a failed run, a data engineering agent that proposes a pipeline, each under its own account and rules.
How it works
Step 01
You declare the local models and allowed providers, then each project’s quotas.
Step 02
A job splits your documents, computes their embeddings through the gateway and stores them in an index, with the permissions of their source.
Step 03
Your applications, notebooks and agents call the gateway or the index, under quotas and with a log.
The LLM gateway lives inside graal, next to permissions, the secret vault and the audit trail. Your applications hold no provider key: they call the gateway with their graal identity, and the gateway applies the quota, masks personal data, picks the model and logs the call. Changing models does not change a line of your applications.
Standards and integrations
Governance
Not with a local model: everything stays in your cluster. With an external provider, only the necessary excerpts go out, after personal data masking.
The open models you host on your GPUs, or on CPUs for the smallest ones, and the providers you allow through their URL and key.
Yes. Each project has a token quota; beyond it, the call is refused with an explicit message, and consumption is read per project.
RAG over a corpus of your choice, served by a model hosted in your infrastructure.