Use case · Energy
A regional demand forecast, ready every morning
Every day, regional consumption and temperatures land in your lakehouse. A thermosensitivity model trains, the forecast is published at 6 a.m., and an AI agent watches over the pipeline. Everything runs on your infrastructure.
The challenge
Two series to join every day, a forecast expected on time
- Half-hourly consumption and daily temperature have neither the same rhythm nor the same format: they must be aligned before modelling.
- Raw data holds surprises: an "ND" value in a numeric column, UTC timestamps next to local dates.
- A forecast is only worth something if it is published on time, every day, even when a source arrives late.
- Metering data stays within your infrastructure.
How graal handles it
The pipeline, from source to forecast
Step 01
Ingest
Two connectors load regional consumption and daily temperatures from S3, then write them as Iceberg tables.
Step 02
Prepare
In the visual editor, a pipeline casts types, joins both series on the local date and aggregates by region. The generated PySpark code exports and belongs to you.
Step 03
Explore
In Jupyter or the SQL editor, your analysts check the relationship between temperature and consumption, on the same tables.
Step 04
Train
A job trains the thermosensitivity model. Parameters, metrics and artifacts are recorded as experiments; the chosen version enters the registry.
Step 05
Schedule and delegate
A workflow recomputes the forecast every morning at 6 a.m., with retries and alerts. An AI agent creates, runs and schedules the jobs through a service account scoped to the project.
This scenario is the one graal’s demo plays. It relies on two open datasets from ODRÉ (Open Data Réseaux Énergies): consolidated and final regional consumption, and daily regional temperature. In a pilot, your own series — metering, production, weather history — take their place, and the pipeline does not change.
Source: RTE and Weathernews France via ODRÉ — Licence Ouverte v2.0 (Etalab).
Consumption: consolidated and final regional éCO2mix data, produced by RTE and published on ODRÉ under the Licence Ouverte v2.0 (Etalab).
Temperatures: daily regional temperature since January 2016, produced by Weathernews France and published on ODRÉ under the Licence Ouverte v2.0 (Etalab).
What you get
- A regional forecast published every morning, at a fixed time
- Every model version linked to the data, parameters and run that produced it
- A pipeline readable as code, which your teams review, version and run elsewhere if needed
- A pipeline watched by an agent whose every action is traced
What stays with you
- The consumption and temperature series, on your S3 storage
- The pipeline and job code, exportable as PySpark
- The model, its versions and its metrics
- The log of human and agent actions
Capabilities involved
- Connectors
Open and internal series, loaded from S3 or your databases.
- Low-code pipelines
Casting, joining and aggregation, exported as PySpark.
- Orchestration and compute
Daily workflow, cron, retries and alerts.
- Lakehouse and SQL
Iceberg tables, catalog and SQL queries on the series.
- Machine learning
Experiments, registry and forecast versions.
- AI agents
An agent that runs the pipeline, under a service account.
Frequently asked questions
Can we use our own metering data?
Yes. The demo relies on open data; in a pilot, your internal series replace or complement it, and the pipeline stays the same.
Can we forecast production instead of consumption?
Yes. The same sequence applies to wind or solar output, with the relevant weather variables.
Does the model need GPUs?
No. A thermosensitivity model trains on CPU. GPUs remain available, by instance type, for heavier models.
What exactly does the AI agent do?
It creates, runs and schedules the project's jobs through a service account scoped to that project. No delete tool is enabled by default, and you can revoke its access at any time.
Prepare this use case on your data
In 8 weeks, the same pipeline on your series, installed in your infrastructure.