← All use cases

Use case · Public sector

Your documents queried in plain language, with a model hosted on your premises

Procedures, memos, contracts, guidance: graal indexes your documents, and your teams and agents get answers that cite their sources. The model runs on your GPUs, access follows each project's permissions, and no request goes to a provider you have not chosen.

Interface illustration

The challenge

Valuable documents, scattered and sensitive

  • Thousands of documents spread across file shares, business tools and mailboxes.
  • Sensitive content that cannot be sent to an online AI service.
  • Answers that must cite their source to be usable.
  • Access rights that differ from one department to the next.

How graal handles it

From document to cited answer

  1. Step 01

    Drop

    Your documents are dropped on S3 or read from an SFTP server. An indexing job splits them into passages and keeps their provenance.

  2. Step 02

    Index

    Passages are embedded by a model served through the LLM gateway, then stored in a vector index, within your infrastructure.

  3. Step 03

    Protect

    Each index carries its permissions: a user or an agent only queries the indexes of their projects. Personal data is masked before it reaches the model.

  4. Step 04

    Ask

    Your applications call the gateway's OpenAI-compatible API, and your agents go through MCP. Every answer cites the passages it draws on.

  5. Step 05

    Track

    Token quotas per project, call logs and audit: you know who asked what, and what it consumed.

The typical scenario starts from a well-defined document collection — a department’s procedures, a body of contracts, a directorate’s guidance — and an open model served on your GPUs. If you prefer an external provider for some uses, the gateway connects it too, project by project: you decide where requests go.

What you get

  • Sourced answers: every cited passage links back to the original document
  • An assistant your applications and agents call through a standard API
  • Access rights enforced in the search itself
  • Token consumption measured and capped per project

What stays with you

  • Your documents and their vector index
  • The model, served on your GPUs
  • The logs of questions and answers

Capabilities involved

Frequently asked questions

Which model should we use?

An open model served on your GPUs by the gateway, or the provider of your choice, project by project. Changing the generation model does not require rebuilding the index.

Do we need GPUs?

To serve a model on your premises in production, yes: the gateway assigns them by instance type. A small model on CPU is enough for a first trial.

Are our documents used to train the model?

No. They are indexed, not learned: the model reads the passages retrieved at question time, and nothing goes into its weights.

Try it on your document collection

In 8 weeks, an assistant on a first corpus, with a model served on your premises.