The five-stage AI rollout that survives procurement
A playbook for buying, piloting, and scaling AI agents inside a 500+ person org.
A playbook for buying, piloting, and scaling AI agents inside a 500+ person org.
Every successful enterprise AI deployment we've seen follows roughly the same arc: Shadow, Pilot, Limited, Broad, and Enterprise-wide. The names are obvious. The pacing and the success criteria at each gate are not.
Shadow, in weeks 1 to 2, means the agent runs in your environment, observes every conversation, but takes no action. This stage exists to build your golden conversation dataset and to give security, legal and IT a chance to review data flows before anything is customer-facing. Don't skip it. The shadow stage is your fastest path through a security review because you can show the data flows before you're arguing about them under time pressure.
Pilot, in weeks 3 to 6, means the agent handles a defined subset of conversations with human review on every response. The goal isn't deflection, it's calibration. You're learning what the agent gets right, what it gets wrong, and what your escalation criteria should be. Define measurable success criteria before you start, because you'll need those numbers to get budget for the next stage.
Limited rollout, in weeks 7 to 10, expands to 20 to 30% of traffic with reduced human oversight. Broad rollout, in months 3 to 4, goes to full traffic. Enterprise-wide is when you replicate the pattern across other teams and use cases, often the most valuable stage because the organisation has now built the muscle for agent deployment.
The single biggest accelerant for enterprise procurement is a pre-filled security questionnaire. Most large organisations use the same 200 to 400 question security assessment. If your vendor can hand you a completed questionnaire on day one, procurement time drops from months to weeks.
The second biggest blocker is data residency. Know where your data lives, where it transits, and where any model calls go before the first security meeting. Saying the model call goes to a third-party API is a showstopper in some regulated industries, not because it's wrong but because it wasn't disclosed upfront and now legal has questions. Map it out in a one-page data flow diagram.
Legal reviews almost always flag the same three things: data retention policy, model training opt-out, and indemnification for AI-generated content. Get written answers to all three from your vendor before legal gets involved. Walking in with a pre-answered FAQ saves two to four weeks.
A pilot without success criteria is just a demo. Before you launch, get written agreement from your key stakeholders on the three numbers that will determine whether the pilot succeeded: deflection rate target, quality floor, and escalation precision floor. Write them down. Get sign-off. You'll need them when the pilot ends and someone moves the goalposts.
Run the pilot on a use case where the stakes of failure are low and the volume is high. Billing FAQs, onboarding questions, password resets: high frequency, low risk. Don't pilot on exception handling or sensitive escalations. You want the pilot to succeed, and success requires picking a slice where the agent can shine.
Instrument everything before you start. You want to know, in real time, how many conversations the agent handled, how many escalated, what the escalation reasons were, and what users said in follow-up satisfaction surveys. Pilots that can't produce this data on demand during stakeholder reviews tend not to get funded for the next stage.
The jump from pilot to org-wide deployment is primarily an organisational challenge, not a technical one. You need an owner, not a committee or a working group, but one person whose job it is to make this succeed. You need a documented runbook that a new team member can follow without asking the original implementer. And you need a feedback loop: a regular meeting where the team reviews quality metrics and decides what to improve.
The teams that scale fastest treat their agent deployment as a product, not a project. Products have roadmaps, owners, and user feedback loops. Projects have end dates. If your agent initiative is structured as a project, it will end, usually at the point where the initial enthusiasm fades and the hard operational work begins.
Replication across teams gets easier with each deployment because the organisation builds institutional knowledge. The second deployment takes half as long as the first. The third takes a third of the time. Invest in documenting what you learn at each stage, because it compounds.
Start free, or book 30 minutes and we will walk through it against your stack.