Use case · Research
A notebook that becomes a scheduled, governed and monitored job
Your data scientists explore in Jupyter or VS Code. When a notebook produces a useful result, it becomes a scheduled job versioned in Git, with its permissions, costs and monitoring, without being rewritten. The datalab remains the place to explore; graal takes over for production.
The challenge
Exploration works, production waits
- Promising notebooks that only run in their author's workspace.
- Going to production means a rewrite, and often another team.
- Permissions, costs and traces to justify project by project.
- A datalab that must stay free for exploration.
How graal handles it
From notebook to scheduled job
Step 01
Explore
Jupyter and VS Code open in the browser, behind your SSO, with access to the lakehouse tables and the project's S3 storage.
Step 02
Version
The notebook is saved in the project's Git repository. Every change is a commit, reviewable and reversible.
Step 03
Schedule
The same notebook runs as is, as a job. It chains with other steps in a scheduled workflow, with retries and alerts.
Step 04
Track
On every run, the executed notebook is kept with its outputs; the models it produces enter the registry.
Step 05
Govern
Per-project permissions and costs, human approval before going to production: every release has an owner.
The typical scenario: a research or analytics team exploring in a datalab, some of whose jobs now have to run every week, with an owner and a budget. graal does not replace the datalab, it complements it: what was explored on one side is industrialized on the other, without changing the tools people work with.
Onyxia is an open source datalab (MIT licence), developed by INSEE and supported by DINUM. graal complements it: Onyxia to explore, graal to industrialize.
What you get
- Notebooks in production without rewriting
- A Git history of every job
- Scheduled runs, monitored and retried automatically on failure
- Costs and permissions you can read, project by project
What stays with you
- The notebooks and their Git history
- The data, on your S3 storage
- The executed notebooks and the models
Capabilities involved
- Notebooks
Jupyter and VS Code behind your SSO.
- Orchestration and compute
Notebook jobs, workflows, cron, retries and alerts.
- Machine learning
A registry for the models your runs produce.
- Governance
Permissions, costs and approval per project.
- Openness
Git, REST API and open formats.
Frequently asked questions
Do we have to give up our datalab?
No. Both coexist: the datalab to explore, graal to schedule, govern and monitor what goes into production. They can share the same S3 storage.
Do our notebooks have to be rewritten?
No. A notebook runs as is, as a job; its parameters are passed at launch, and the executed notebook is kept with the run.
Who can put a job into production?
The people the project grants that right to. You can require human approval before every release.
Industrialize your first notebooks
In 8 weeks, a first explored job goes into production, in your infrastructure.