Use case · Public sector
Your documents queried in plain language, with a model hosted on your premises
Procedures, memos, contracts, guidance: graal indexes your documents, and your teams and agents get answers that cite their sources. The model runs on your GPUs, access follows each project's permissions, and no request goes to a provider you have not chosen.
The challenge
Valuable documents, scattered and sensitive
- Thousands of documents spread across file shares, business tools and mailboxes.
- Sensitive content that cannot be sent to an online AI service.
- Answers that must cite their source to be usable.
- Access rights that differ from one department to the next.
How graal handles it
From document to cited answer
Step 01
Drop
Your documents are dropped on S3 or read from an SFTP server. An indexing job splits them into passages and keeps their provenance.
Step 02
Index
Passages are embedded by a model served through the LLM gateway, then stored in a vector index, within your infrastructure.
Step 03
Protect
Each index carries its permissions: a user or an agent only queries the indexes of their projects. Personal data is masked before it reaches the model.
Step 04
Ask
Your applications call the gateway's OpenAI-compatible API, and your agents go through MCP. Every answer cites the passages it draws on.
Step 05
Track
Token quotas per project, call logs and audit: you know who asked what, and what it consumed.
The typical scenario starts from a well-defined document collection — a department’s procedures, a body of contracts, a directorate’s guidance — and an open model served on your GPUs. If you prefer an external provider for some uses, the gateway connects it too, project by project: you decide where requests go.
What you get
- Sourced answers: every cited passage links back to the original document
- An assistant your applications and agents call through a standard API
- Access rights enforced in the search itself
- Token consumption measured and capped per project
What stays with you
- Your documents and their vector index
- The model, served on your GPUs
- The logs of questions and answers
Capabilities involved
- Generative AI
LLM gateway, local models, RAG and cited answers.
- Connectors
Documents on S3 or on an SFTP server.
- Orchestration and compute
Scheduled re-indexing of new documents.
- Governance
Per-index permissions, token quotas and audit.
- AI agents
The assistant, available to your agents through MCP.
Frequently asked questions
Which model should we use?
An open model served on your GPUs by the gateway, or the provider of your choice, project by project. Changing the generation model does not require rebuilding the index.
Do we need GPUs?
To serve a model on your premises in production, yes: the gateway assigns them by instance type. A small model on CPU is enough for a first trial.
Are our documents used to train the model?
No. They are indexed, not learned: the model reads the passages retrieved at question time, and nothing goes into its weights.
Try it on your document collection
In 8 weeks, an assistant on a first corpus, with a model served on your premises.