← Overview

Generative AI

Generative AI, on your models and your documents

An LLM gateway serves the models you choose, hosted on your GPUs or with the provider of your choice, with quotas, a log and personal data masking. RAG answers over your documents, cites its sources and respects everyone’s permissions.

Interface illustration

Key capabilities

Your models, your rules

A sovereign LLM gateway

An endpoint compatible with the OpenAI API, in front of local models on your GPUs or the providers you allow for your organization.

Quotas and a log

A token budget per project, a clear refusal beyond it, and a log of calls, metadata only by default.

Personal data masking

French social security numbers, IBANs, phone numbers, emails, company IDs: personal data is masked before anything is sent to an external provider.

RAG over your documents

Your PDFs and tables feed a vector index in your PostgreSQL. Every answer cites its sources, and every index has its permissions.

Coding assistant and documentation

An assistant in Jupyter and VS Code, and AI-proposed documentation for catalog tables, on the model you chose.

Built-in agents

An operations agent that diagnoses a failed run, a data engineering agent that proposes a pipeline, each under its own account and rules.

How it works

From model to application

  1. Step 01

    Choose the models

    You declare the local models and allowed providers, then each project’s quotas.

  2. Step 02

    Index

    A job splits your documents, computes their embeddings through the gateway and stores them in an index, with the permissions of their source.

  3. Step 03

    Connect

    Your applications, notebooks and agents call the gateway or the index, under quotas and with a log.

One door to the models

The LLM gateway lives inside graal, next to permissions, the secret vault and the audit trail. Your applications hold no provider key: they call the gateway with their graal identity, and the gateway applies the quota, masks personal data, picks the model and logs the call. Changing models does not change a line of your applications.

Standards and integrations

Standard interfaces

  • OpenAI-compatible API
  • MCP
  • vLLM
  • llama.cpp
  • pgvector
  • PostgreSQL
  • GPU
  • OIDC

Governance

Every call is counted and traced

  • Token quotas per project and per application
  • A log of calls, available to your CISO, with encrypted content as an option
  • An index only answers the people allowed on its source

Frequently asked questions

Do our documents go to a provider?

Not with a local model: everything stays in your cluster. With an external provider, only the necessary excerpts go out, after personal data masking.

Which models are supported?

The open models you host on your GPUs, or on CPUs for the smallest ones, and the providers you allow through their URL and key.

Can we cap spending?

Yes. Each project has a token quota; beyond it, the call is refused with an explicit message, and consumption is read per project.

An assistant over your documents

RAG over a corpus of your choice, served by a model hosted in your infrastructure.