Smruti Patel

Smruti Patel

I make enterprise software just work, and help organizations go AI-native.
Most recently SVP Engineering at Apollo GraphQL; previously Stripe and VMware.

Instrumenting the Dark Ends

What actually goes in the Precision and Impact columns

BLUF: Precision and Impact are where the high-leverage judgment now sits, and the metrics haven’t settled because the workflows haven’t.

In the last post I shared four lenses for engineering velocity, mapped against the software delivery lifecycle: Precision, Speed, Quality, Impact. This post is about the two ends of that lifecycle nobody instruments.

I’m not going to hand you four metrics for the AI era. I don’t know your business, and anyone offering a clean dashboard for this without knowing it is guessing. What I have is where to look, and the instruments I’ve actually used there.

If judgment is the scarce resource now, what exactly are we measuring? I found that it depends. It depends on how much of your SDLC already runs without a human in it.

Think of it as a spectrum. On one end, humans review everything: agents write, but a person looks at every change before it moves. In the middle, agents run multi-step: they write, test, open the PR, maybe fix their own build, and humans approve at checkpoints. On the far end, agents ship and you sample after: some path to production for your prime business use cases where no human looks at every change.

The autonomy spectrum: as checkpoints disappear, less of the outcome is visible to your metrics

Every step along that spectrum, a checkpoint disappears. The further right you are, the more of your outcome is decided before any human gets to look, and the less helpful your existing metrics get. Speed and Quality are well-lit, but Precision and Impact sit at the opposite, dark ends of the lifecycle.

And the more autonomy agents have, the earlier decisions are locked in, so Precision and Impact matter more.

Precision

Precision is judgment at the planning end of your lifecycle, and it’s where I think most flat Impact actually comes from. When Impact stays flat, the coding is rarely the reason. Usually the other systems never got built: the right problem wasn’t picked, the review didn’t catch anything, nobody knew if the customer was better off. Precision is the first of those systems.

Planning used to mean months of customer research, discovery, validation and design partnership, much of it outside Engineering. That was a defensible trade: coding was the expensive bottleneck downstream, so we paid for high fidelity in which problems to solve. The weight was insurance against a cost that doesn’t exist the same way now.

Planning and engineering are both cheaper now. If you can prototype in an afternoon, validation shifts left and gets cheap. You stop trying to be right in the planning doc. This holds for the work you can afford to throw away. Where you can’t, in regulated or safety-critical surfaces, codify the constraints upstream for the known knowns, and enforce the right human checkpoints downstream so validation happens before anything ships.

Two things to instrument: how you align, and what you prioritize.

Alignment

Start with knowing your customer. Be intentional about which personas you’re optimizing the product for: an enterprise platform team looks very different from a frontend engineer leveraging your open source distribution. This new world lets you build much shorter feedback loops with each of them.

The charter that fixed the “body shop” team in the last post is a Precision instrument: it makes “are we building the right thing” legible in a way no velocity chart can. It worked pre-AI. It isn’t enough now.

What’s different is how quickly you can test an alignment question instead of arguing it.

We were honing the problem definition for a new product. There were two camps: one believed we should build for our existing customer persona, the other believed we should expand our TAM. Our Sr. Staff engineer prototyped both, far enough to show the user experience, the tradeoffs, and a realistic time to production for each. They demoed it to a few key enterprise customers and got feedback that reshaped our thinking and our design.

That argument would have taken a quarter. It took days, and it ended with customer evidence rather than the most senior opinion in the room.

And the further right you sit on the autonomy spectrum, the more this matters: alignment questions that go unanswered get encoded into shipped software before anyone reviews them.

Prototyping, evaluation frameworks, and the discipline to kill things early also let you hold team structures and role definitions more loosely than you used to. That can be uncomfortable and generate a lot of FUD. It also means the charter you wrote last year may be describing a team that no longer exists. Write it again.

Prioritization

I’ve never been excited about wiring up Jira. But agent orchestrators, harnesses and software factories now make it easier to connect product needs, customer advisory board notes, support tickets and GTM calls into one plane.

We used Glean to analyze the Gong calls our GTM teams had with customers over the previous three to four quarters, summarized the key areas for focus, and tuned our backlog for a mix of big-ticket features, papercuts and operational needs. We did this in hours, without being blocked on multiple cross-functional teams.

That corpus always existed. Reading it cost more than it was worth. Now the same throughput that created the firehose can mine it.

What fraction of what you shipped is actually being used — not adopted, used. This matters especially in enterprise software, where customer integrations take a long time, platform upgrade windows are narrow and mission-critical, and every unused feature still has to be maintained.

How often you killed something after starting it. It sounds like failure, because we’ve always been rewarded for shipping. It’s usually good precision. Deprecating or killing something you thought was useful, early, actually saves you compounding costs: less product surface bloat, lower engineering carrying costs, less customer confusion. We wound down a managed hosting offering we had already launched when we found customers preferred to run that critical piece of infrastructure themselves, without another network hop on the request path.

This is only effective when you, as leaders, normalize saying no or pulling the plug, and revisit how you’ve historically perceived throw-away work. To make that stick at Apollo, I hosted an annual “burnpit party” across Engineering to celebrate everything we’d eliminated that year.

What you decided not to build, and whether that decision held. This is the hardest of the three, because the evidence is an absence. It is also the one that tells you whether the alignment work was real or ceremonial. The agent identity call was ours. For login and auth we integrated with an open source IAM library. For caller identity, we leveraged what our customers already had in place. We built only the layer specific to agents: brokering a separate credential for each service an agent calls. The test of that split is whether customers ever needed more from the parts we didn’t build.

Impact

You stop asking “is it shipped?” — a question with an expiry date — and start asking whether you’re building for compounding value. Three places to look.

Customers

Are your customers happy?

How’s the day 0 journey? Was onboarding smooth? What does the last mile of customer integration look like, and how is your forward-deployed engineering supporting it? Could they just plug and play?

How’s day 100? Are they using it actively, and the way you designed it to be used? Our MCP server wasn’t, and that reset our roadmap.

How’s day 1000? Are your systems reliable enough and scaling as your customers’ needs are scaling? How do you show up for your customers when shit hits the fan? Have they stopped worrying about you?

Revenue

Is your revenue growing? You need a sustainable business model to keep the innovation going.

Revenue is a lagging indicator, and in this new world the silos are breaking down. If Engineering used to hear about revenue secondhand, through Product or a quarterly support-ticket review, you can go to the source directly now.

When I’ve drunk all my engineering kool-aid and need a reality check, I go talk to Sales and Customer Success. Three things to look at.

What you charge. How are you pricing and packaging what engineering ships, and is that helping or hurting the business? In one of our products, the pricing structure was disincentivizing the right behaviors, which led our customers down rabbit holes of redesigning much later. So we introduced a new pricing tier that mapped more cleanly to customer use, maturity and complexity.

What it costs. What do the COGS look like? Anchor that against whatever your key north-star business metric is: rides, payment volume, query volume. If the rate of growth in COGS is higher than the rate of growth in the business metric, you have a bottom-line problem. Whether to prioritize it is a conversation with your CFO and the rest of the executive team.

Who stays. Renewal is the customer’s decision, which makes it the honest signal. Churn rate and renewal rate are the obvious measures, and the account-level questions underneath them are the interesting ones. Happy customers are healthy accounts. Which accounts have more consumption than they purchased? Which are far under-utilized against their commits? Where are your customers struggling, and are you solving the right problems? That’s Precision showing up again.

Organization

How are your teams doing? Customers and revenue tell you whether the work landed. This tells you whether the org is getting better at sustainably producing work that lands.

Paved paths. How long until a new engineer ships a meaningful change to production? Now ask the same question about an agent working in that codebase. It’s the same measure.

Software engineering is a team sport, and now that individuals have their own agents, the question is how the org collectively builds its knowledge. I’ve written before about the Bob problem: the runbook that saves a new engineer their first week is the same runbook your agent can’t work without.

What’s good for humans turns out to be doubly good for the agents.

The best paths maintain themselves: one engineer wrote an agent that routinely checks for unused feature flags and deletes them. Work that used to sit on a backlog forever now gets done, and the system gets more resilient for it.

The counterweight is real. LeadDev’s AI Impact Report 2026 found 42% believe over-reliance on AI tools is already eroding engineers’ ability to debug without them. Capability can atrophy while output goes up, and none of your delivery metrics will tell you. Your on-call rotation will.

Unblock paths. How long does work wait for someone with the authority to say yes, no, or not now? At Stripe we had go/unblock: when one team was blocked on another, both first tried to solve it locally, then co-wrote a short escalation doc from a template. What it made visible was how long a team sat blocked before escalating, and how often it had to. Agents speed up everything inside a team; the handoffs between teams are where the time goes now.

If judgment is the bottleneck, this is where you can watch it queue.

Explorers and scalers. Are the right people on each bet for its stage? Zero-to-one work needs explorers who are comfortable throwing things away. Durable systems need scalers who make them boring and reliable. Each has a failure mode: explorers drift toward shiny work, scalers toward repetitive work. AI can take the repetitive work off scalers and make throwaway work cheap for explorers, so both spend more of their time on work that compounds.

And the further right you sit on the spectrum, the longer the gap between a decision and the signal that tells you whether it was right. That gap is what you’re instrumenting against.

The outer loop

The last post opened with three calls my velocity dashboard couldn’t help with. Hire another frontend engineer or burn more tokens. Ship the feature or invest in CI/CD. Build the coding agent or buy one.

I still can’t tell you what to decide. But I can tell you what each one is actually asking.

The hiring call is a COGS question in disguise: if token spend is growing faster than the business metric it serves, decide whether fixing unit economics is the right priority for your business before you add engineers or tokens. The CI/CD call is a spectrum question: the more autonomous your pipeline, the more it is your review, and the less optional it gets. And build-versus-buy is a precision question, the same one we asked about agent identity: which part of the thing is specific enough to your customers that only you should build it.

All three become answerable once you’re looking at the right ends of the pipeline. That’s the outer loop: idea to customer value, with judgment at every handoff.

None of this gives you a clean dashboard. Precision and Impact resist being turned into a number on a slide, which is most of why they stayed dark so long. So don’t wait for the dashboard. Pick one signal from this post, put it in whichever column is emptiest, and start tracking it next quarter with data you already have.


I run Judgment Is the New Bottleneck as a 90-minute working session with engineering leadership teams: your staff brings the dashboard they actually use and leaves with one metric retired and one signal added. If your org is drowning in AI output and can’t tell motion from progress, get in touch.

This is the second half of an argument. The first, Judgment Is the New Bottleneck, is the diagnosis. For the operational side, see From Pilots to Production.