Smruti Patel

Smruti Patel

I make enterprise software just work, and help organizations go AI-native.
Most recently SVP Engineering at Apollo GraphQL; previously Stripe and VMware.

From Pilots to Production: The Rollout

Updated September 2026.

BLUF: Most AI pilots stall between demo and production. This is the rollout playbook: five phases, paced crawl-walk-run, with the red flags that tell you you’re not ready to scale.

“Fast forward one year. You have two engineers. One delivered significant impact using agentic AI. The other delivered comparable impact without it. How do you assess their performance?”

A leader on my team asked me that, and I paused, because every answer reveals a different failure mode. Reward the one who used it and you’re incentivizing tool adoption. Reward the other and you’re taxing innovation. Reward them equally and you’ve admitted AI created no leverage. Reward neither and you’ve told your team outcomes don’t matter.

It turned out to be a systems question, and one most organizations can’t answer yet.

I’ve spent the past year using coding assistants and talking to engineering leaders at companies of every size. The rollouts that stall look alike: impressive demo, enthusiastic leadership, then nothing.

A good strategy has three parts: an honest diagnosis of what’s actually true about your organization, guiding principles for the constraints you’re designing around, and coherent actions that move you systematically. This post is the third part. Skip the first two and you’ll execute a rollout plan beautifully in service of the wrong thing.

The diagnosis, in short

Code got cheap, and judgment became the scarce thing: deciding what’s worth building, and knowing whether what shipped actually landed with customers. Most dashboards are well-lit in the middle of the delivery pipeline, where speed and quality live, and dark at both ends, where those calls get made and felt.

Judgment Is the New Bottleneck makes that argument in full, and Instrumenting the Dark Ends covers what to measure at each end. Start there if you haven’t: this rollout depends on those metrics.

The principles, in short

Diagnosis tells you what to optimize for. Principles tell you how not to break things on the way. Five, and every one of them became a hard design constraint when we built Apollo’s MCP server.

Clarity. Every AI-generated output needs to carry its reasoning: why this code, what alternatives were considered, what constraints applied. Without it, reviewers rubber-stamp work they can’t evaluate.

Context. Context management is the difference between an agent that matches your patterns and one that violates your architecture. If your data is scattered, poorly documented, or hard to access, AI workflows break down immediately. Models hallucinate, and teams stop trusting the output.

Coupling. Loosely coupled systems are easier to evolve with AI. Put abstractions in front of vendors and keep contracts explicit at every interface. Design for graceful degradation, so that when an agent generates something wrong the blast radius is contained.

Controls. Guardrails are what make experimentation survivable. Rate-limit what agents can do, scope what they can touch, and put a circuit breaker on them that trips when error rates spike.

Confidence. Engineers need to trust that AI is helping them, and to find out early when it’s quietly accruing debt. That takes feedback loops that show when it’s working and when it isn’t, and a way to override it without friction.

The sharpest of those was Coupling. Building the MCP server, we were split on how. One camp wanted it merged into the runtime codebase, with the same modularity and release rigor we’d spent years getting right. The other wanted it standalone, because standalone was faster. Standalone won, and at ship time that looked like the wrong call to a lot of people: we’d taken something strategically important and put it outside the thing we were most disciplined about.

Then usage climbed. And because it was out there on its own, we could see how people were actually using it, which was nothing like what we’d assumed when we designed it. That’s what reset our thinking on where the whole agentic access platform needed to go.

Across all five, clear guardrails make teams faster, because they let people move with high agency and clear accountability at the same time. (The platform thinking underneath this is in Building AI Platforms That Scale Human Agency.)

Now the rollout.

The Five E’s

For every tool you’ll have early adopters, early majority, and late majority. Your strategy depends on where your organization sits on that curve. The timeline depends on your size, your executive sponsorship, and how change-ready the org actually is.

The Five E's: Crawl-Walk-Run Timeline

1. Experimentation — let early adopters explore (weeks 1–4)

Your discovery phase, driven by the engineers already using AI whether you know it or not.

The cultural principle: if engineers feel they have to hide their AI usage, you lose the most useful data you have. Some shadow AI is inevitable, and the only productive response is to bring it into the open.

  • Give early adopters a few select tools and a shared space to document what works and what doesn’t
  • Map engineering workflows to find the high-friction points
  • Identify where AI creates leverage versus where it creates overhead

Done when: you have use cases ranked by value and feasibility, baseline metrics established, and engineers sharing experiments openly.

2. Evaluation — build the business case (weeks 5–8)

Now get honest about readiness and ROI.

  • Assess infrastructure gaps: data quality, API consistency, observability
  • Calculate the full cost: tooling plus review time, debugging time, maintenance overhead
  • Define north-star metrics using Precision, Speed, Quality, and Impact (what goes in each)
  • Decide what’s allowed, what requires review, what’s prohibited

In most orgs I’ve seen, infrastructure becomes the limiting factor before engineering talent does.

Red flags that mean you are not ready to scale:

  • Undocumented or inconsistent APIs
  • No secrets management or access control
  • Manual deployment processes
  • Code review already a bottleneck
  • No baseline engineering metrics
  • Unclear data classification policies

Done when: you have executive buy-in, approved budget, a go/no-go framework, and a named list of infrastructure investments.

3. Education & enablement — build the foundation (month 3)

This is where you move from early adopters to early majority. Do not skip the first two phases.

  • Deploy to a pilot group (10–20% of engineering)
  • Run monthly show-and-tells where engineers demo real workflows
  • Share rulesets, prompts, successful patterns, and failures
  • Teach where not to use AI, which matters as much as where to use it
  • Create an AI runbook: documented prompts, workflows, anti-patterns

Almost every codebase has a comment like “call Bob, only he knows the right sequence to upgrade our SPOF database.” We had a Bob. He was at a music concert, entirely legitimately, and we waited a week for him to get back. That was always a risk. Now it’s a bottleneck for your agents too, because an agent can’t ask Bob. It can only work from what’s written down. Everything we’ve been telling ourselves to do for decades, codify the knowledge, write the spec, document the review standards, just went from hygiene to leverage. The runbook that saves a new engineer their first week is the same runbook your agent can’t work without.

Explicit human sponsors. Every AI-generated change needs a human owner. Early on, that means approving the merge. Once a path has earned sample-after, it means owning the outcome. No orphaned AI code. While humans are still approving merges, ownership also means shepherding code through review: the Slack pings, the standup mentions, knowing when to nudge. Agents don’t do that. Humans must.

Quality gates. Train reviewers to ask: Why this architectural choice? What edge cases does it miss? What are you optimizing for? How would you debug this at 2am?

Guardrails. Flag PRs over 200 lines. Require decomposition for complex tasks. Time-box review SLAs: review promptly or close. Don’t let AI PRs linger.

Done when: pilot teams show measurable gains on at least one Precision or Impact signal, and best practices are written down.

4. Expansion — scale what works (months 4–6)

Scale only what’s been validated.

  • Roll out proven tools organization-wide
  • Standardize on one primary tool for common use cases, to reduce fragmentation
  • Build custom solutions only for high-value, company-specific problems
  • Create AI champions in each team to coach others

When I set the technical strategy at Apollo, this was one of the calls: pick a single tool as the core agentic platform, build shared configuration around it, and stay ready to switch as the landscape moved. We’d done the same for Kotlin with IntelliJ: strongly encouraged, with tooling built around it so the default was also the easiest path. The bet was on lower COGS from fewer tools, and on every shared config, prompt and runbook compounding across teams instead of splintering.

Be surgical about where you apply it:

AI Use Case Risk Matrix

Done when: delivery improvements are measurable, at least one Precision or Impact signal has moved, and AI-related incidents are declining.

Watch for: inconsistent quality across teams; engineers who can’t explain their own code; fragmented tooling creating support overhead; and code slop: verbose, poorly structured output that technically passes.

5. Enforcement & evolution — make it normal (ongoing)

Enforcement, kept lightweight: team norms for PR size, review quality, and ownership. Documentation required before a new experimental tool spreads. Make it easy to do the right thing and hard to do the wrong one.

Evolution: measure quarterly against your north-star metrics, sunset tools that aren’t delivering, invest in advanced capabilities only where ROI is proven, and feed production learnings back into strategy.

Done when: AI is just part of how you build software.

Pacing: crawl, walk, run

Crawl (Experimentation + Evaluation). Go slow. Understand before committing. The slowness is the work.

Walk (Education & Enablement). Measured progress with tight feedback loops. Expect surprises and high output variance. This is where you learn what works in your specific context.

Run (Expansion + Enforcement & Evolution). Move fast on a solid foundation. Infrastructure is ready, processes are proven, the organization is aligned.

Don’t try to transform everything at once. Small, proven improvements compound quickly once the foundation holds.

If you don’t have six months

Most people reading this don’t have executive cover for a six-month program. The compressed version, in order:

  1. Stop hiding shadow AI. One week. Ask openly what people are already using. You cannot plan a rollout around a baseline you refuse to measure.
  2. Run the red-flag list. One afternoon. If you fail three or more, your next ninety days belong to infrastructure work. Treat that as a finding you needed.
  3. Pick one workflow before you pick one tool. Instrument it end to end, idea to production, including review and rework, then compare against your baseline.
  4. Set the two guardrails that do the most work. PR size limits and review SLAs. Most of the downstream problems I’ve seen trace back to one of those being absent.
  5. Write down what would make you stop. Do it before you scale.

That’s enough to know whether you have a rollout or a science project.

Nothing in this post is AI-specific. It’s ordinary engineering leadership: diagnosis, principles, disciplined execution, applied to a change that happens to be moving unusually fast. The organizations that win here will be the ones that stayed honest about what was working, retired what wasn’t, and kept shipping value through the hype.

So, the two engineers. You don’t assess them on whether they used AI. You assess the judgment in what they chose to build, the systems they improved along the way, the teams they upskilled, and the value they delivered to customers. That’s how you amplify human impact: with AI in service of the craft.


I run Judgment Is the New Bottleneck as a 90-minute working session with engineering leadership teams: your staff brings the dashboard they actually use and leaves with one metric retired and one signal added. If your org is drowning in AI output and can’t tell motion from progress, get in touch.