ESC
Type to search guides, tutorials, and reference documentation.
← AI and Machine Learning
⚙️

AI Operations (MLOps)

Machine learning systems fail silently. MLOps is the discipline of making that failure visible, reversible, and reproducible.

MLOps is the operational discipline for systems whose behavior is determined by data as well as by code. It covers how models are built reproducibly, promoted to production safely, observed once there, and retired or retrained when they stop working. The reason it needs its own name is simple: conventional software fails loudly and deterministically, whereas a machine learning system usually fails silently, keeps returning well-formed answers, and degrades over weeks while every dashboard stays green.

Why it is not just DevOps

A deployed model is the product of three versioned artifacts, not one: the code that defines the pipeline, the data it was trained on, and the resulting model weights. Reproducing a result requires all three, and only the first is naturally handled by source control. On top of that, the same code can produce different weights across runs, correctness is statistical rather than binary, and there is a hard boundary — the moment features are computed at serving time rather than training time — where the two halves of the system can silently disagree.

What to version and track

  • Data snapshots. An immutable, addressable reference to the exact rows used, not a query against a mutable table.
  • Feature definitions. The transformation logic itself, versioned, so training and serving can be proven to share it.
  • Experiments. Parameters, metrics, environment and artifacts for every run, recorded automatically rather than in a notebook someone will close.
  • Model registry. Named, versioned model artifacts with a lifecycle stage, the evaluation that justified promotion, and a pointer back to the run that produced them.
  • Lineage. The chain from a prediction in production back to the model, the code commit and the data. Without it, incident response is guesswork.

Full bit-exact reproducibility is often unattainable — non-deterministic accelerator kernels and unpinned dependencies see to that. The achievable and useful goal is provenance: knowing exactly what went in, even when the output cannot be recreated byte for byte.

Train/serve skew

This is the canonical MLOps failure. A feature is implemented once in a batch pipeline for training and again in application code for serving, and the two drift apart — a different default for missing values, a subtly different time window, a lookup that includes data unavailable at prediction time. The model receives inputs that do not match what it learned from, and quality drops with no error raised anywhere.

The structural fixes are to compute features once and share the definition between paths (the core purpose of a feature store), and to log the exact feature vector used for each prediction so training and serving distributions can be compared directly rather than argued about.

Deployment patterns

Choose the serving shape from the latency the decision actually needs. Batch scoring writes predictions on a schedule and is the cheapest and most operationally boring option — prefer it whenever the consumer can tolerate staleness. Online serving computes on request and introduces availability, latency and autoscaling concerns. Streaming scores events as they arrive. Embedded or edge inference runs inside the client, which removes network cost and adds a fleet-update problem.

Promotion should be gradual regardless of shape. Shadow deployment sends real traffic to the new model without using its output, which validates the plumbing and the input distribution with zero blast radius. Canary routing exposes a small share of traffic. Champion/challenger keeps the incumbent live as the comparison baseline. In every case the rollback path should be a routing change, not a rebuild — and it should have been tested.

Monitoring on three levels

Operational monitoring — latency, error rate, saturation, queue depth — is ordinary service monitoring and is necessary but tells you nothing about whether predictions are any good.

Data monitoring watches the inputs: schema conformance, null and cardinality rates, range violations, and distribution shift in the features and in the predicted-value distribution. It is the earliest available signal because it needs no labels.

Model monitoring watches quality itself, which requires ground truth. The difficulty is label delay: the outcome may arrive hours, months or never after the prediction. Where labels eventually arrive, join them back and track performance over the aligned window. Where they do not, fall back on proxies — downstream acceptance rates, override rates, complaint volume — and be explicit that they are proxies.

One rule saves more incidents than any tool: a monitor over an empty table must report "no data", never "healthy". A pipeline that quietly stopped writing and a system that is genuinely fine look identical to a naive aggregate, and the failure mode is a green dashboard over a dead process.

Retraining and feedback loops

Retraining can be scheduled or triggered by a monitored condition. Scheduled retraining is predictable and easy to reason about; triggered retraining reacts faster but needs a threshold that is neither noisy nor asleep. Either way, a retrained model is a candidate, not a release — it goes through the same evaluation and promotion gates, and automatic retraining without automatic evaluation is a mechanism for deploying regressions on a timer.

Watch for feedback loops. When a model's predictions influence the behavior that generates its next training set — recommendations shaping what gets clicked, risk scores shaping who is approved — the system trains on the consequences of its own decisions and narrows over time. Holding out a small randomized slice of traffic is the standard defense, and it is a decision to make deliberately rather than discover later.

Governance, cost and failure modes

For regulated or high-impact use, document the intended use, training data, evaluation and known limitations of each model version, and keep an approval record tied to the registry entry. Cost control is mostly right-sizing: batch where possible, share accelerators, measure utilization rather than allocation, and check whether a smaller or compressed model meets the requirement — see model compression.

The failures worth designing against: no tested rollback; a pipeline reporting success while the model rots; drift alerts nobody owns or acts on; monitors computed over empty inputs; training data that cannot be reconstructed; and a model in production that no living engineer can rebuild. See also machine learning for the modeling groundwork and MLOps pipeline architecture for a fuller treatment of pipeline design.

RELATED GUIDES