← Overview

Machine learning

From the first training run to the served model

Every training run leaves a trace: parameters, metrics, artifacts and code. The version you keep enters the registry, then becomes a REST endpoint behind your SSO, with its drift monitored. Your data scientists keep their MLflow client.

Interface illustration

Key capabilities

The whole model lifecycle

Comparable experiments

Parameters, metrics and artifacts of every run, linked to the job or notebook that produced them. You compare, you choose.

A versioned model registry

Each version keeps its origin: the experiment, the code and the training data. Validation steps are read in the same place.

Works with the MLflow client

Your existing scripts log their runs and models into graal without changing a line, through the MLflow API.

REST serving

A registry version becomes a REST endpoint, with as many replicas as needed, behind your SSO or a service token.

Drift monitoring

A sample of requests is kept, then compared with the training data by a scheduled job. Drift triggers an alert.

GPUs on demand

Training and inference on your cluster’s GPUs, reserved run by run and capped by project quotas.

How it works

From notebook to endpoint

  1. Step 01

    Train

    In a notebook or a job, your code logs its runs to an experiment, through the MLflow client.

  2. Step 02

    Register

    You promote the best run to the registry. Going to production can require a second person’s approval.

  3. Step 03

    Serve and monitor

    The version becomes a REST endpoint; drift is measured at regular intervals and alerts you.

A served model is a run that does not stop

An endpoint is a deployment that graal manages like everything else: its containers run without privileges, its network rules are in place before them, and every call goes through an authentication check. The model is loaded from the registry at startup: what is served is exactly the registered version.

  • The European regulation on artificial intelligence, (EU) 2024/1689, requires high-risk AI systems to technically allow the automatic recording of events over their lifetime (Article 12). graal helps you trace and document: runs, versions, training data and deployments.

Standards and integrations

Interfaces you know

  • MLflow
  • Python
  • XGBoost
  • PyTorch
  • REST
  • MLServer
  • GPU

Governance

Every model has an owner

  • Per-project permissions on experiments, registry and endpoints
  • Every version and every deployment is attributed to a person or an agent
  • Endpoints only accept authenticated calls

Frequently asked questions

Do we have to change our MLflow scripts?

No. You point the MLflow client at graal; your tracking and model registration calls work as they are.

Where are models stored?

In your S3-compatible storage, with their artifacts. Nothing leaves your infrastructure.

Who can call a served model?

The people and applications allowed on the project, through your SSO or a revocable service token.

Follow a model all the way to production

An experiment, a registry version, an endpoint: shown in the demonstration on the energy scenario.