Automate your own development process with Flyte, using event-driven runs, review and release gates, human approvals, and evaluation gates for systems whose behavior depends on data and models.
Software development lifecycle
The process around your workloads can run on Flyte too. Triage an issue when it is filed, rebuild a dataset when a schema changes, evaluate a model before it is promoted, or wait days for a person to approve a release. As Flyte runs, these chores are typed, retried, observable, and auditable.
Why data and AI systems need more than CI
A conventional software project has one input that changes: code. Everything downstream of a commit is reproducible from that commit, so a pull request fully describes a change and a green build fully answers whether it is safe.
A data or AI system has four inputs that change: code, data, models, and prompts. Three of them change without a commit, which breaks the assumptions CI is built on:
| CI assumes | For a data or AI system |
|---|---|
| A change arrives as a diff | A retrained model, a drifted upstream table, or a provider’s model update arrives with no diff |
| The same input gives the same output | Generative steps are non-deterministic, so two runs of one suite can disagree |
| Tests pass or fail | Quality is a distribution. The question is whether the change is worse than what is live by more than the noise |
| A build takes minutes and is cheap to repeat | An evaluation can need GPUs, paid API calls, and an hour |
| The artifact under test is in the repository | A 40 GB checkpoint is not, and the comparison point is whatever is in production |
So the useful gate is rarely “do the tests pass”. It is “is this model better than the one serving, on the slice that matters, by more than the run-to-run variance”. Answering that takes a pipeline: fan-out over a dataset, GPUs, caching, retries, a readable report, and often a human decision. Running that pipeline on the same orchestrator as production gives you one system, one set of credentials, and one place to look when something goes wrong.
What Flyte provides
The section uses no special lifecycle features. Each need maps to an ordinary part of Flyte:
| Need | Use |
|---|---|
| Start a run from GitHub, Slack, Linear, ClickUp, or Jira | Software development tools |
| Start a run on a schedule | Triggers |
| Launch once per event when a delivery is retried or resent | run_once, keyed on the event. See Event-driven automation for the concurrent case. |
| Wait for a person, possibly for days | External conditions |
| Keep an expensive gate fast | Caching, fan-out, and per-task resources |
| Survive a flaky provider | Retries and timeouts |
| Trace an output back to its inputs and commit | Run history, task reports, and lineage |
| Tell a person the result | Notifications, or a Slack message from inside the run |
Integrations by stage
Most teams use a handful of these. None is required.
| Stage | Integrations |
|---|---|
| Trigger a run from an event | GitHub, Slack, Linear, ClickUp, Jira |
| Configure the run | Hydra, OmegaConf |
| Validate inputs | Pandera for dataframe contracts, TypeSafe AI for typed, confidence-scored model answers |
| Test generated code | Code generation, which runs generated code in a sandbox |
| Compute | Ray, Spark, Dask, PyTorch |
| Read from a warehouse | Snowflake, BigQuery, Databricks |
| Move data between steps | Polars, Lance, JSONL, Hugging Face |
| Record results | MLflow, Weights & Biases |
| Report to a reviewer | Papermill, for a parameterized notebook as the gate’s output |
| Run agents | Agent frameworks, with ten SDKs that run as durable tasks |
| Observe production | OpenTelemetry, Grafana Agent |
Scope of this section
- Deploying from CI is covered in CI/CD deployments: API keys,
flyte deploy, and commit-pinned versions. This section assumes you have that in place. - Your CI system stays. Keep linting, unit tests, and type checking in CI. They are fast and need no cluster. Use Flyte for checks that need real data, real compute, retries, or a person.