AI Just Broke Your Engineering Metrics. Now What?

Tokenmaxxing

Editor’s Note: This article first appeared in Jellyfish’s LinkedIn newsletter, The Current. You can read and subscribe here.

For years, engineering relied on traditional activity metrics to understand how work was flowing through the organization. These metrics were never perfect, but they at least correlated with engineering effort. Now with AI, that correlation breaks down.

AI coding tools mean software engineers can generate massive amounts of code faster than ever. But that extra code doesn’t necessarily increase impact, and measuring output alone can create a misleading picture of engineering performance. Metrics still matter, but only if they offer a reliable reflection of the value developers are creating.

As Head of AI & Research at Jellyfish, I have a front-row seat to the metric distortions. We have observability into how hundreds of thousands of engineers across ~1000 companies actually adopt AI and how that connects to the code they ship, its quality, and downstream business outcomes. Here’s what the data is showing us:

The Broken Metrics

The two engineering metrics most affected by AI are lines of code and merged pull requests (PRs).

Lines of code

Lines of code was never a great proxy for engineering productivity. AI adoption means it’s now all but meaningless.

Jellyfish data shows that AI-generated code is 20% larger on average than human-generated code. That ability to create huge amounts of code in minutes breaks any remaining connections between volume and value. AI has essentially turned lines of code into an input metric that correlates closely to token spend rather than effort or productivity.

Merged PRs

Merged PRs is still a valuable metric, but it has been distorted by AI acceleration. It’s also not necessarily as indicative of downstream impact as you might expect.

While we continue to see ~2x average PR throughput for companies with a high level of AI adoption over the baseline, we also find that more PR throughput doesn’t translate cleanly into more downstream value. In other words, more merged pull requests doesn’t always equal more features shipped or more revenue. The reason for this is that other bottlenecks, such as antiquated roadmap processes, quality issues, or production review bottlenecks, cause extra capacity to be poured into things like internal tooling and prototypes that don’t have a direct line to building the business.

Pushing for more merged PRs doesn’t make sense once you reach the highest tiers of token consumption. The heaviest-spending developers burn ~10x as many tokens as the median developer but ship just ~2x as many PRs. If you compare the folks at the bottom of the curve to those at the top, the cost-per-merged-PR climbs from $0.26 in the bottom 10% to ~$25 in the top 10% – roughly 100x more expensive.

Value is Still the Most Important Metric

AI hasn’t broken or distorted every engineering metric. Measuring value, whether that’s token spend per shipped feature or another unit of meaningful business outcome, is still extremely important. You need to know that you’re enabling your agent stack in the right way for your unique technology, products, and business. Most importantly, you need to know that you’re spending tokens on high-value outcomes.

Our aggregate data shows that broad, moderate AI adoption is what really drives value. The best ROI comes from moving most people from low to moderate usage, not pushing heavy users even higher. You can only take token consumption so far before it hits what we call the “agentic barrier”: the point at which you need extra infrastructure to capture more gains from your autonomous agents. Whether extreme spending pays off comes down to the ultimate business value of shipped code.

What It Means for Engineering

Our understanding of pull requests as a unit of value is changing. Human PRs merge ~80% of the time; for autonomous-agent PRs, it’s only ~60%. With agentic workflows, PR count blends shipped work with throwaway work: exploratory candidates, proactive fixes no one committed to, and vibe-coded attempts that may ultimately get killed.

Teams that have historically overemphasized coding speed will have an even harder time. With 2x (or more) speedup, bottlenecks become even worse in the review, design, and planning stages. If you have 2x the code capacity but haven’t adapted your roadmap process, delivery mechanisms, and GTM enablement, that capacity pools in side projects and backlog trimming instead of moving the business forward.

Measuring What Matters

Engineering leaders need to move on from metrics that no longer make sense in the AI engineering era. Measuring throughput and cycle times of scoped deliverables and outcomes is particularly valuable in this new environment.

The most useful metrics now tie spend to outcomes. Cost per shipped feature – or even cost per merged PR – tells you far more than token volume alone, because it asks whether the spend is driving business value, not just activity. Throughput measured in scoped deliverables, tracked alongside that cost, is what shows you whether broad adoption is actually paying off.

Cycle time is another metric worth tracking, but it needs a fresh read in the AI era. A shorter cycle time still generally signals a more efficient process. A rising one, though, often isn’t a coding problem at all. When you double output but cycle time creeps up, it usually means the work is now piling up downstream, in review, design, and planning. That’s exactly the bottleneck shift we talked about above, and it’s why measuring cycle time across the whole delivery process matters more than ever. Jellyfish compares key metrics with and without AI so you can quantify improvements in speed, quality, and throughput.

As companies adapt their engineering workflows to integrate AI, they need to make sure their metrics keep pace. To learn more about how Jellyfish can help engineering leaders measure the AI metrics that matter, request an AI Impact demo.

About the author

Nicholas Arcolano

Nicholas Arcolano, Ph.D. is Head of Research at Jellyfish where he leads Jellyfish Research, a multidisciplinary department that focuses on new product concepts, advanced ML and AI algorithms, and analytics and data science support across the company.