In this article
AI Impact Week Day 1
Catch up on the product keynotes, AI research trends and AI impact talks with industry leaders.
Watch the Day 1 RecapDay 1 asked the questions. Day 2 answered them.
I’ve been a huge advocate for the power of Jellyfish to give engineering leaders the best tools to communicate engineering impact for years. So I’m thrilled to recap the tremendous launches our engineering team shared today. We plan to support you through the biggest transformation most of us have ever experienced.
Yesterday’s sessions hit on the gaps between AI adoption and AI impact. Almost every developer is using AI now, and very few engineering leaders can say what it’s worth. At the end of the day, Andrew promised the Jellyfish R&D team would show how we’re closing that gap, live. Today they did. I know, because I got to cheer on so many of my teammates sharing the fruits of their hard work to help make sense of all the change!
Adam Ferrari opened with a stat that sums up the problem:
In our State of Engineering Management report, most orgs said they’re seeing real gains from AI. Less than half said they can measure those gains in any real way.
The rest are going on gut feel. And the more your process changes, the harder that gets to justify.
So the new Jellyfish is built around the three questions every engineering leader gets asked today: Where do we stand? Are we transforming? What is it worth? Here’s what we showed for each one.
Meet the new Jellyfish, built for the AI-DLC
Meet the new Jellyfish, built for the AI-DLC

"It's no longer the software development lifecycle as we've known it. It's something completely new. It's the AI-native software development lifecycle, or AI-DLC."
Adam set the scene. A year or two ago, few were seriously debating whether humans should still review code. Now it’s a live conversation. Teams are rebuilding their workflows, adding agents and software factories, and rethinking the role people play. For most orgs, that has meant a year or more of near-constant change.
While Jellyfish has led engineering analytics for close to a decade, Adam was clear that the old approaches can’t keep up with this. So we’ve rebuilt the product from the ground up. This goes far deeper than a new interface. Under the hood, we rebuilt the data platform too, with faster analytics, more scalable storage, and better data quality checks and lineage. That’s what lets Jellyfish take in new signals as fast as they appear, and what makes custom metrics, custom cohorts and flexible views possible. If you’ve used Jellyfish before, Adam said, “I think you’re in for some big surprises today.”
Measure AI Impact: Where do I stand?
2. Measure AI Impact: Where do I stand?

"It's literally next-level engineering management."
Tricia started with why this is so hard today. Traditional engineering tools understand Git commits and tickets. AI assistance, prompts, and agent output don’t exist in their data model. AI vendor dashboards show usage, and never connect it to team throughput or business value. So leaders stitch it together by hand, “tab-hopping across multiple vendor portals, GitHub logs, spreadsheets.” Jellyfish brings it all into one view, with people using AI and fully autonomous agents side by side.
Then AJ played a VP of Engineering whose execs had just asked for proof: “Is AI actually driving higher velocity? Or is it just inflating PR volume without adding real value?” Everything he showed is available now to all Jellyfish customers.
- Metrics Explorer with AI Assistance: Break down any metric, like PRs merged, by human contributors and autonomous agents. You can show leadership how AI is changing throughput using metrics they already trust.
- Custom metrics agent: AJ’s org measures productivity with weighted PR throughput. He asked the agent, in plain English, to break it down by team. It wrote the query, validated it, and ran it. Power users can drop into the SQL to fine-tune it.
- Blueprints: Ready-made dashboards based on Jellyfish research and best practices, so every team works from the same metrics.
- Dashboard agent: Ask for a new chart, like AI-assisted PR throughput by team over the last three months, and it’s added to your dashboard.
- Flexible Lifecycle Explorer: Change the phases to match how your teams actually work, so you can see where AI is speeding things up and where new bottlenecks are forming.
- Jellyfish Assistant: AJ spotted a spike in PR cycle time in July. Instead of clicking through charts and digging through logs in other tools, he asked the Assistant why. It came back with a likely root cause, “saving us hours of manual effort.” It also posts insights and recommendations right at the top of your dashboards.
Transform with AI-DLC
3. Transform with AI-DLC: Am I transforming?

"Quite bluntly, just using AI isn't sufficient anymore."
"Why did the developer break up with the LLM? It kept saying 'I think this should work,' but never committed."
Ryan opened with a confession. Like a lot of companies, Jellyfish runs an AI Champions group. Early on, it relied on grassroots sharing and anecdotes. We only heard from people it was working for, and we heard slowly. Whether a practice was actually productive was “more of a gut and vibe check than anything else.” That made it hard to change our AI practices quickly. So the team built what we needed, drawing on Jellyfish research across over 90 million prompts and 150 trillion tokens.
Available now:
- Behavioral metrics: Leading indicators from session data, like skill usage, compaction rates, and session agentic ratio (how interactively engineers work with their agents). You can see good practices taking hold before they show up in output. Available for Claude Code.
- Skills metrics: Skill adoption by team and over time, plus every skill ranked by usage. Jellyfish research across 700,000+ developer weeks found that consistent skill use (30+ invocations) can lead to 27% more PRs shipped per month. Available for Claude Code.
- Cohorts tied to outcomes: Compare heavy and light adopters of a practice against outcomes like PR rework rate and cycle time.
Greta showed what this looks like with Shipmate, Jellyfish’s own “super skill” for creating and reviewing PRs. Adoption was trending up, with draft PR as “the hottest skill on the block.” Then she checked impact. The top 25% of Shipmate users had less PR rework. Cycle time was more mixed: faster at first, then not. Her advice was to keep an eye on it. These are early signals, not proof, and that’s the point. You see them in weeks instead of waiting for word of mouth. The data also surfaced skills that were spreading on their own, like one for reducing verbosity.
Coming soon:
- Task classification: What engineers actually use agents for, like code generation, code review and planning, by team or by week. At Jellyfish, code generation is the top task, and our Data Engineering team does far more AI-assisted code review than our Data UX team. That’s a good reason for those two managers to compare notes. Available in alpha on request.
- SDLC phase classification: AI usage and token spend broken out by lifecycle phase, right next to Lifecycle Explorer. It helps answer questions like “We spend a lot of tokens in review. Is it paying off?”
Ryan closed with an open invitation to anyone working on the same problems: “Don’t be a stranger. Let’s collaborate.”
Prove AI ROI
4. Prove AI ROI: What is it worth?

"AI cost has really just exploded over the past few months."
Taylor opened with the number that brought finance into the conversation. Jellyfish research found a 60x increase in token spend per merged PR over the past 12 months. And it’s sneaking up on people. One company Taylor spoke with was already spending 5x what it budgeted for AI this year.
Teams have moved on from “Is everybody using it?” to “Are we using it well?” Taylor sees two reasons to measure AI ROI. The first is to evaluate the investment alongside your CFO and business leaders. The second gets overlooked: ROI data gives R&D teams a feedback loop to learn how to use these tools better. The session walked through three questions.
What are we spending?
- AI spend dashboard (available now): AI spend across Claude, Cursor, and Copilot in one place, by tool, model and team, over time. Every vendor reports spend differently. Jellyfish puts it all in one consistent view so you can see what you’re spending and how it’s tracking.
- Total R&D Cost (available now): A new metric that shows AI spend in the context of everything you invest in R&D. Of course AI spend is material and growing, but most of your investment is still in people. This puts both side by side.
What are we spending it on?
Noah explained why this one is hard. Vendor APIs tell you what people spent. Jellyfish knows what work those same people produced. Nothing connects the two. And the way we’ve always tracked human effort doesn’t work for agents, because a developer can now be focused on one thing while an agent works on something else.
So the team built a hand-labeled dataset covering how developers really work with AI, from going deep with a single agent to juggling a dozen tickets at once, and used it to test and refine the approach. Jellyfish looks at the work that follows the spend, mostly commits, and splits the cost when someone is running several sessions at once.
- AI spend allocation (coming soon): See AI spend with the same groupings you already use for people’s effort, like investment categories and custom fields. For example, if you let agents take on simpler support tickets, you can check whether AI spend actually moved to support and your people’s time moved to roadmap work.
What are we getting for it?
- R&D Cost per Merged PR (available now): A simple unit-economics view of what each dollar buys. Track how your cost to build software changes over time, split by human and AI. R&D Cost per Issue Resolved and per Deliverable Resolved are coming soon, because different teams plan in different units.
- AI capacity metrics (coming soon): Metrics like merged PRs per engineer, which normalize output to headcount so you can tell an AI capacity bump apart from hiring. Taylor took on the obvious objection: aren’t PRs changing? They are getting more verbose. But PRs per Linear project have stayed steady, so PRs are still a reliable unit of work.
Next up for the team: AI cost benchmarks, so you can compare your spend and outcomes with peers, and then velocity metrics.
How the best engineering teams win with AI
5. How the best engineering teams win with AI

"There is no silver bullet. There is no singular answer to it."
Emily has watched the questions from prospective customers change over the past year. First it was which tool to use and whether one was faster than another. Then it was how quality was holding up and where bottlenecks were moving. Now it’s “our spend is out of control” and “what are we getting for it?”
Ryan’s answer: it depends, and it keeps changing. A month ago he would have pointed to skills. With the newest models, some people now argue that skills restrict the model. What works depends on where you are. Earlier in the journey, the best teams make room for engineers to learn the tools, even if that slows things down for a while, and focus on adoption. Further along, the focus shifts to efficacy: the right model for the right work, no duplicated effort, and skills managed like a product. Most orgs are both at once, with one pocket doing something remarkable and most of the org still early. The job is to learn from that pocket and spread what works, without holding everyone to the same bar. And through all of it, don’t let quality slip.
Jellyfish helps in two ways. In the platform, benchmarks and research show you where you stand, and you can ask Jellyfish Assistant what your next best opportunity is. And with people: Ryan introduced the new Jellyfish Advisory Group, or JAG. It’s staffed by practitioners who have been engineers and engineering leaders themselves. They work across our customer base to bring back what’s working, from rolling out AI code review to handling compliance when AI is doing your reviews.
Tomorrow on Day 3, you’ll hear from our partners, including AWS and Linear, about how Jellyfish fits into your stack.
Request a Demo
Know what your AI is worth
Your CFO wants to know what you’re getting for your AI spend. Your teams want to know what’s working. The new Jellyfish helps you answer both: where you stand today, whether you’re actually transforming, and what all of it is worth.
Request a demo and we’ll show you how it works.
About the author
Marilyn is an Engineering Manager at Jellyfish. Her specialities include communicating, programming, researching, working well with others and with the command line, leading, being independently motivated to learn and to create.
She primarily works in Java, JS (mostly React/Redux, some NodeJS, MongoDB, and Angular), and python, with previous experience in PHP, Haskell, C, Perl, and Ruby. Diverse interests along the full stack have led to a plethora of other familiarities, including git, AWS (mostly ECS, S3, RDS, Dynamo, and Cost Explorer), Jenkins, TravisCI, PostgreSQL, Spring, mvn, svn, MySQL, XML, XSLT, CSS, sed, yacc, x86 assembly, and many other acronyms.