MLOps

·

5 min read

Our pre-production MLOps checklist

Six items, none of them exotic, that decide whether a model degrades gracefully and visibly — or silently, until a business result makes it someone else's problem to discover.

Priya Nandakumar

By

Priya Nandakumar

,

Head of Data & MLOps

5 min read

·

Our pre-production MLOps checklist

In this article

Filed under

MLOps

Published

Reading time

5 min read

Written by

Priya Nandakumar

A model that’s accurate in a training notebook and a model that’s accurate on a Tuesday six months from now, after three data-schema changes and a seasonal demand shift, are not the same model — even if the weights never change. The gap between them is what MLOps is actually for, and it’s the part most AI projects underbuild, because none of it shows up in a launch demo.

Before any model we build takes production traffic, it goes through the same checklist, regardless of the client or the use case.

Versioned data, not just versioned code. If you can’t reproduce the exact dataset a model was trained and evaluated on, you can’t debug it when it starts behaving differently, and you definitely can’t prove to a risk or compliance team what it was and wasn’t trained on. Every feature we build goes into a versioned feature store, not a one-off script.

A held-out evaluation set that’s re-run on a schedule, not just at launch. We set a fixed cadence — weekly, for most forecasting models — where the model gets scored against fresh held-out data and the result gets logged. This is how we caught degrading accuracy at Ferrovia Logistics before it showed up as a stockout, instead of after.

Drift monitoring with a defined threshold, not a dashboard nobody checks. A dashboard is necessary but not sufficient. We set specific thresholds — a MAPE increase past a defined point, a shift in input feature distributions past a defined point — that trigger an alert to a named person, not a chart that quietly turns orange.

A rollback path that’s been tested, not just documented. If deploying a new model version requires more than a single command to reverse, it will not get reversed quickly enough during an actual incident. We test the rollback path before launch, the same way we test the forward deployment path.

Cost and latency budgets set before launch, not discovered after the first invoice. A model that’s accurate but too slow or too expensive to run at your actual query volume isn’t production-ready either. We size infrastructure against realistic load before go-live, not against the load in a demo environment.

Clear ownership of what happens when the model is wrong. Every production model will eventually make a wrong call. The question is whether there’s a defined path for what happens next — a human review step, a fallback rule, an escalation — or whether the wrong call just propagates downstream silently. We define this before launch, as part of the architecture, not as an incident-response afterthought.

None of these six items is exotic. All six are the difference between a model that degrades gracefully and visibly, and one that degrades silently until a business result — a stockout, a missed fraud flag, a bad forecast — makes the degradation someone else’s problem to discover. We’ve built this checklist the hard way, watching models we didn’t build fail in exactly these predictable spots. Running it before launch costs a few extra weeks. Skipping it costs a lot more, later, when nobody’s watching.

Priya Nandakumar

Written by

Priya Nandakumar

Head of Data & MLOps

View profile →

03 · Start

Have a use case

like this one?

Most of what’s in this article came out of a real engagement. Tell us about yours.

Create a free website with Framer, the website builder loved by startups, designers and agencies.