01 · Northwind / Approach

How we work

Most AI failures are

process failures.

A model that scores well in a notebook and a system that holds up in production are separated by work most teams skip: scoping, evaluation, monitoring, ownership. We run every engagement through the same five stages, in the same order.

5

Stages, none skipped

11 wks

Kickoff to production

64

Systems through it

Assess

Design

Build

Deploy

Operate

02 · The five stages

Same order, every engagement

Five stages.

In this order, for a reason.

Each stage ends at a written gate. Nothing moves forward until the gate is signed by the people who will own the system.

i

Stage 01 · Weeks 1–3

Assess.

Find out what’s worth building.

AI Strategy & Readiness

We audit the data and infrastructure behind each candidate use case and score it on effort, risk and payback, using the same matrix every time so three departments’ pet projects compete on equal terms. It is also where we find out how far from AI-ready your data really is. That changes sequencing, not whether the roadmap gets built.

Data & infrastructure audit

Week 1

Interviews with the teams closest to the work

6–10

Use cases scored: effort, risk, payback

All

12-month roadmap with budget bands

Gate

ii

Stage 02 · 1–2 weeks

Design.

Fix the evaluation plan first.

Agents

RAG

Document AI

An agent’s scope gets bounded here: what it decides alone, what it hands to a person and where the guardrails sit. It is also where we define how we will know it works: a gold question set, a confidence threshold per document field, an accuracy bar for a forecast. Skipping Design doesn’t save time; it moves the same decisions into production.

Workflow mapped, exception paths included

All

Guardrails and human checkpoints

Written

Gold set or evaluation harness

200+ cases

Confidence thresholds and routing rules

Gate

iii

Stage 03 · 4–8 weeks

Build.

Against the messy version.

Shadow mode

We build against real inputs from week one: real documents, real edge cases, real exceptions. A system tuned on the tidy sample breaks the first week it meets production traffic. Every build runs in shadow mode against live data before it touches a real decision.

Core system: extraction, retrieval, agent or model

Weekly demos

Integrated into the tools your team uses

In place

Evaluation harness running continuously

Every commit

Shadow run on live traffic

Gate

iv

Stage 04 · 1–2 weeks

Deploy.

With monitoring already on.

Data & MLOps

Rollback tooling, drift and performance monitoring and access logging are live on day one. A system is not “in production” until it runs on real traffic with someone, us or your team, watching the dashboard rather than the deployment checklist.

Staged rollout where the risk calls for it

Canary

Rollback verified before go-live

< 5 min

Access logging and audit trail

Confirmed

Runbook handed to on-call

Gate

v

Stage 05 · Ongoing

Operate.

Keep it accurate, keep it used.

Enablement

Retainer

A system that silently degrades six months after launch has failed as completely as one that never shipped, just more expensively. Operate is monitoring against agreed thresholds, retraining triggered and logged (never silent), and the adoption work that turns a deployed system into one people rely on.

Drift reviewed against thresholds

Weekly

Retraining cycles, triggered and logged

On signal

Adoption and usage dashboards

Monthly

Ownership handoff or Retainer

Your call

The platform

Six layers.
One platform.

Everything an AI product needs, built to work together.

  1. Connect every source in minutes, with lineage on every field.

04 · The gates

What gets signed, by whom

Every stage ends

at a written gate.

Stage

Gate artefact

Signed by

Typical length

Maps to

01

Assess

Scored use-case matrix and a 12-month roadmap with budget bands

Sponsor + finance

2–5 weeks

Strategy & Readiness

02

Design

Architecture, guardrails and an evaluation set with a pass bar

Process owner + risk

1–2 weeks

Agents · RAG · Document AI

03

Build

Shadow-mode report: accuracy, cost per run, exceptions routed

Process owner

4–8 weeks

Agents · RAG · Document AI

04

Deploy

Runbook, verified rollback, monitoring and access logging live

On-call owner + security

1–2 weeks

Data & MLOps

05

Operate

Monthly drift and adoption review against agreed thresholds

System owner

Ongoing

MLOps · Enablement

05 · Why this order matters

What skipping looks like

Skip a stage and

you pay for it later.

01

Skip Assess

Teams build the best version of the wrong use case: a well-engineered system nobody prioritised, competing for adoption with the workflow people actually needed fixed.

02

Skip Design

Guardrails get bolted on after an incident instead of scoped before one, because there was never an evaluation harness to catch the failure before a customer did.

03

Skip Build discipline

The system is tuned on a clean sample and meets the malformed PDF, the missing field and the holiday-weekend backlog for the first time in production.

04

Skip Deploy monitoring

A forecast that was accurate at launch drifts for months before anyone notices. We found exactly that at a logistics client before they engaged us.

05

Skip Operate

A strong research copilot stalls under 20% usage, not because the tool is wrong but because nobody trained people on when to rely on it.

→

Run all five

Eleven weeks, on average, from kickoff to a monitored system in production.

06 · All five stages, in practice

01

Meridian Trust Bank

Financial Services

Avg. loan-file review

1.5 h

Was

4.1 h

−63%

Loan-file review cut 63% with agents

that know when to stop.

Underwriters were re-keying the same fields across four systems. An agent now does the cross-checking and flags only what needs a human decision.

$2.1M

Annual cost avoided

0.62

Review threshold, conf.

9 wks

Kickoff to production

Read the case

01

Meridian Trust Bank

Avg. loan-file review

1.5 h

Was

4.1 h

−63%

Loan-file review cut 63% with agents

that know when to stop.

Underwriters were re-keying the same fields across four systems. An agent now does the cross-checking and flags only what needs a human decision.

$2.1M

Annual cost avoided

0.62

Review threshold, conf.

9 wks

Kickoff to production

Read the case

02

Ferrovia Logistics

Manufacturing & Logistics

Forecast error (MAPE)

12.5%

Was

18.9%

−6.4 pts

Forecasts that say when

they are losing accuracy.

Demand models across 22 distribution centres, monitored for drift and retrained on a schedule the planners can see.

−22%

Safety-stock cost

22

Warehouses

16 wks

Kickoff to production

Read the case

02

Ferrovia Logistics

Forecast error (MAPE)

12.5%

Was

18.9%

−6.4 pts

Forecasts that say when

they are losing accuracy.

Demand models across 22 distribution centres, monitored for drift and retrained on a schedule the planners can see.

−22%

Safety-stock cost

22

Warehouses

16 wks

Kickoff to production

Read the case

07 · Questions

About the method

Method,

answered.

Every stage applies to every engagement, scaled to its size. A Retainer is mostly Operate; a Pilot moves through Assess quickly.

01

Can we skip the Assessment stage if we already know what we want to build?

Sometimes. If the use case is narrow and the data is understood, we fold a lightweight Assess into Design. We never drop it entirely: it is usually where a scoping problem gets caught before it becomes a Build problem.

02

How do you scope Design versus just starting to Build?

03

What does Operate actually involve, month to month?

04

Do all five stages apply to every engagement?

05

How long does the full cycle take?

08 · Start

Five stages.

Eleven weeks.

Tell us the workflow. In 30 minutes we’ll tell you which stage it is really at, and what the next gate looks like.

Create a free website with Framer, the website builder loved by startups, designers and agencies.