Skip to content

Jobs and workflows

This is the line between a platform and a datalab: a datalab executes while you sit in front of the screen, a platform executes while you are away.

Every job has a type, which decides what it executes and which image it runs in.

Type What it executes Image
python a Python script or package, dependencies installed at startup python
bash a shell command, for operations tasks toolbox
spark a Spark workload (spark-submit), sized by instance type spark-py
lowcode a graph from the low-code editor, translated to code python

The contract also carries specialised types — docker, dbt, dask, flink, ray —, each served by its own runtime image. A type whose image is not configured on your installation answers 501 feature_disabled (see runtime images).

Creating a job, step by step and as commands, is in the quickstart.

A job that is started produces a run. While it executes, four things are readable:

  • the logs, cursor-paginated — the console follows them live, an API client replays them from any point;
  • the events, timestamped: queued, started, step, finished;
  • the container metrics;
  • the files produced, in the project’s bucket.

A run carries a status (queued, running, succeeded, failed, cancelled, retrying, skipped) and a duration. That is what the job list shows.

A job also carries its maximum duration (timeout_seconds) and its number of retries (max_retries): a run that fails is restarted automatically up to that number.

A workflow chains steps in a graph: each step is a job, dependencies are edges. A failed step stops the branch that depends on it without taking the others down.

A workflow or a job is triggered:

  • by hand, from the console, the API or an agent;
  • by a standard cron schedule — on a job, it is set by JSON Patch on the schedule field (example on the REST API page);
  • on an event: a file arriving, another workflow finishing.

Backfills replay a scheduled job over a range of past dates, and each run receives the logical date it processes.