01 · Northwind / Approach
Most AI failures are
process failures.
A model that scores well in a notebook and a system that holds up in production are separated by work most teams skip: scoping, evaluation, monitoring, ownership. We run every engagement through the same five stages, in the same order.
5
Stages, none skipped
11 wks
Kickoff to production
64
Systems through it
Assess
Design
Build
Deploy
Operate
02 · The five stages
Five stages.
In this order, for a reason.
Each stage ends at a written gate. Nothing moves forward until the gate is signed by the people who will own the system.
i
Stage 01 · Weeks 1–3
Assess.
Find out what’s worth building.
AI Strategy & Readiness
We audit the data and infrastructure behind each candidate use case and score it on effort, risk and payback, using the same matrix every time so three departments’ pet projects compete on equal terms. It is also where we find out how far from AI-ready your data really is. That changes sequencing, not whether the roadmap gets built.
Data & infrastructure audit
Week 1
Interviews with the teams closest to the work
6–10
Use cases scored: effort, risk, payback
All
12-month roadmap with budget bands
Gate
ii
Stage 02 · 1–2 weeks
Design.
Fix the evaluation plan first.
Agents
RAG
Document AI
An agent’s scope gets bounded here: what it decides alone, what it hands to a person and where the guardrails sit. It is also where we define how we will know it works: a gold question set, a confidence threshold per document field, an accuracy bar for a forecast. Skipping Design doesn’t save time; it moves the same decisions into production.
Workflow mapped, exception paths included
All
Guardrails and human checkpoints
Written
Gold set or evaluation harness
200+ cases
Confidence thresholds and routing rules
Gate
iii
Stage 03 · 4–8 weeks
Build.
Against the messy version.
Shadow mode
We build against real inputs from week one: real documents, real edge cases, real exceptions. A system tuned on the tidy sample breaks the first week it meets production traffic. Every build runs in shadow mode against live data before it touches a real decision.
Core system: extraction, retrieval, agent or model
Weekly demos
Integrated into the tools your team uses
In place
Evaluation harness running continuously
Every commit
Shadow run on live traffic
Gate
iv
Stage 04 · 1–2 weeks
Deploy.
With monitoring already on.
Data & MLOps
Rollback tooling, drift and performance monitoring and access logging are live on day one. A system is not “in production” until it runs on real traffic with someone, us or your team, watching the dashboard rather than the deployment checklist.
Staged rollout where the risk calls for it
Canary
Rollback verified before go-live
< 5 min
Access logging and audit trail
Confirmed
Runbook handed to on-call
Gate
v
Stage 05 · Ongoing
Operate.
Keep it accurate, keep it used.
Enablement
Retainer
A system that silently degrades six months after launch has failed as completely as one that never shipped, just more expensively. Operate is monitoring against agreed thresholds, retraining triggered and logged (never silent), and the adoption work that turns a deployed system into one people rely on.
Drift reviewed against thresholds
Weekly
Retraining cycles, triggered and logged
On signal
Adoption and usage dashboards
Monthly
Ownership handoff or Retainer
Your call
04 · The gates
Every stage ends
at a written gate.
Stage
Gate artefact
Signed by
Typical length
Maps to
01
Assess
Scored use-case matrix and a 12-month roadmap with budget bands
Sponsor + finance
2–5 weeks
Strategy & Readiness
02
Design
Architecture, guardrails and an evaluation set with a pass bar
Process owner + risk
1–2 weeks
Agents · RAG · Document AI
03
Build
Shadow-mode report: accuracy, cost per run, exceptions routed
Process owner
4–8 weeks
Agents · RAG · Document AI
04
Deploy
Runbook, verified rollback, monitoring and access logging live
On-call owner + security
1–2 weeks
Data & MLOps
05
Operate
Monthly drift and adoption review against agreed thresholds
System owner
Ongoing
MLOps · Enablement
05 · Why this order matters
Skip a stage and
you pay for it later.
01
Skip Assess
Teams build the best version of the wrong use case: a well-engineered system nobody prioritised, competing for adoption with the workflow people actually needed fixed.
02
Skip Design
Guardrails get bolted on after an incident instead of scoped before one, because there was never an evaluation harness to catch the failure before a customer did.
03
Skip Build discipline
The system is tuned on a clean sample and meets the malformed PDF, the missing field and the holiday-weekend backlog for the first time in production.
04
Skip Deploy monitoring
A forecast that was accurate at launch drifts for months before anyone notices. We found exactly that at a logistics client before they engaged us.
05
Skip Operate
A strong research copilot stalls under 20% usage, not because the tool is wrong but because nobody trained people on when to rely on it.
→
Run all five
Eleven weeks, on average, from kickoff to a monitored system in production.
06 · All five stages, in practice
07 · Questions
Method,
answered.
Every stage applies to every engagement, scaled to its size. A Retainer is mostly Operate; a Pilot moves through Assess quickly.
01
Can we skip the Assessment stage if we already know what we want to build?
Sometimes. If the use case is narrow and the data is understood, we fold a lightweight Assess into Design. We never drop it entirely: it is usually where a scoping problem gets caught before it becomes a Build problem.
02
How do you scope Design versus just starting to Build?
03
What does Operate actually involve, month to month?
04
Do all five stages apply to every engagement?
05
How long does the full cycle take?

08 · Start
Five stages.
Eleven weeks.
Tell us the workflow. In 30 minutes we’ll tell you which stage it is really at, and what the next gate looks like.





