AI Impact Week Day 3: Why Scaling AI Takes More Than Tools

Proving AI’s impact takes an ecosystem

Day 1 named the problem: Adoption is solved, Impact isn’t. Day 2 showed the product that answers the three questions every engineering leader gets asked: Where do I stand? Am I transforming? Am I thriving?

Day 3 tackled what comes next. Knowing where you stand is only half the job. You still have to change how your teams work, convince finance the spend is worth it, and keep up as the tools change every quarter. No single platform does all of that alone.

And the pressure is rising. As I shared in my opening, the majority of companies expect to speed up AI spending, with AI’s share of IT budgets growing. That puts real pressure on margins, and on the CTO who has to explain it.

So on Day 3, we brought in the partners who help close the gap: AWS and Linear, where software gets built; IBM and Slalom, who run the change programs; and CFGI, who sits between finance and engineering. Here’s what you missed.

 

Opening: Why the ecosystem matters

1. Opening: Why the ecosystem matters

"Adopting the fast-moving AI-DLC takes more than understanding it. It requires new engineering best practices and a FinOps approach built for dynamic, token-based billing."
Andrew Lau, Co-Founder and CEO, Jellyfish

I kicked off Day 3 by reminding everyone that the three big questions don’t come from Jellyfish. They come from the CFO, the board and the CEO, and they all land on the CTO. The CTO’s job hasn’t changed: build great software, predictably, at a reasonable cost. What’s changed is that people and agents now build it together, and the board wants answers fast.

Then our CEO, Andrew Lau, explained why we built Jellyfish to plug into an ecosystem. Our customers rely on us to answer three questions: where they stand with AI adoption, whether they’re turning their SDLC into an AI-DLC, and whether they’re thriving. Those answers come from data collected across 50+ integrations, from AI tools like Claude Code and GitHub Copilot to issue and incident tracking like Linear and Jira. And because the landscape keeps shifting, we keep adding more. We recently integrated with Kiro and now give visibility into development workflows powered by Amazon Bedrock. Our Linear integration helps teams plan and allocate work across people and agents.

The AI-DLC: From workshop to scaled impact

2. The AI-DLC: From workshop to scaled impact

"When you see that the work which I was going to spend three months on got completed in three weeks, I'm yet to see a better way to develop more conviction."
Anupam Mishra, Director of AI Engineering, AWS

Anupam’s team has worked with almost 1,000 AWS customers this year, and his message was that fully autonomous software is real but still rare: maybe a single-digit percentage of companies. The ones getting there did the groundwork first: high engineering standards, clear documentation, and a way to sort which work is safe to hand to agents. His playbook is to skip the toy demo and run a real project, mixing greenfield work, brownfield work and bug fixes. To prove ROI, start with team-level gains in productivity and quality (one Amazon team cut on-call interruptions from 15 to 7 per engineer), then graduate to business results, like a product launched in two months instead of a year. He also warned against defaulting to the most expensive model, telling the story of a leaderboard whose top AI user was burning tokens on grep searches.

My takeaway: the AI-DLC doesn’t scale on tools alone. It scales on guardrails and a data baseline you can prove your progress against.

How Jellyfish helps (Where do I stand? Am I thriving?): Metrics Explorer shows human, AI-assisted and fully autonomous agent work side by side, giving you the baseline Anupam described. Lifecycle Explorer shows where time actually goes across the AI-DLC. Token Usage and Spend tracks usage by tool and by model, so you can catch costly habits like that grep-search leaderboard early. See everything we launched.

The triad in the agentic era

3. The triad in the agentic era

"Over 50% of the work that's created is created by agents using our MCP server."
Cristina Cordova, COO, Linear

Linear made waves with its “issue tracking is dead” manifesto, and Cristina explained why. Agents now create most of the work in Linear, so the point of a ticket is no longer handing work off. It’s keeping the context of why a customer cares, what was decided and what was rejected, because an agent is like a brand-new hire who needs that context. Roles are blurring too. Designers hand designs to agents and engineers review the result, marketers ship their own website updates, and founders are getting back into the code. Team sizes look the same from the outside, but inside, work happens in two- or three-person pods, and some Linear managers run 20 direct reports with tech leads taking on more.

Her take on efficiency: a small team isn’t the goal. As she put it, efficiency “should give you this ability to invest in opportunities that are working. It shouldn’t be an excuse to, like, under-invest.”

How Jellyfish helps (Am I transforming?): Skill Adoption shows which AI skills and practices are spreading across your org, and which teams are falling behind. Behavioral metrics show how well people are working with agents, so you can see which teams are compounding and which are spinning their wheels. And our Linear integration helps you plan and allocate work across people and agents. See everything we launched.

Tokens on the balance sheet

4. Tokens on the balance sheet: A new playbook for finance and R&D

"Your capitalization of token spend is really only as good as your tagging. So if you can't show what the agent was actually working on, right, then the safe answer is usually to expense it."
Josiah Tubbs, Partner, CFGI

R&D budgets used to be mostly fixed costs like headcount and licenses. Tokens are different: they’re usage-driven and can double in a quarter. Josiah suggested treating AI like infrastructure, with a fixed floor for tools plus tripwires. He warned that per-engineer caps are “kind of like the new time sheet”: they measure activity, not value, and can punish your best people. Classify spend by what it does. Tokens that build the product are R&D, tokens used every time a customer uses your product are cost of revenue, and internal use goes to G&A or sales and marketing. Chad’s reassurance for finance: the rules for capitalizing software haven’t changed, but you now need to tag tokens to the work they did. Each month, the CFO should see spend by team and project, how much is untagged, and cost per outcome.

Their one piece of advice before 2027 budgets lock: CFOs and CTOs should agree on what counts as AI spend. As Josiah said, that costs nothing.

How Jellyfish helps (Am I thriving?): Token Usage and Spend puts all your AI spend in one place, including reconciliation between API-reported and telemetry-reported cost. Spend-to-work attribution does the tagging Josiah described, showing which initiatives, deliverables and roadmap items your AI spend goes to. AI cost benchmarks compare your spend efficiency with hundreds of peers. See everything we launched.

Driving change to climb the AI maturity curve

5. Less art, more science: Driving change to climb the AI maturity curve

"When you have the token leaderboards, you're teaching developers just to burn tokens."
Karl Schwarz, Go-to-Market Lead, Slalom

I opened with a reality check: according to Steve Yegge, only about 20% of developers have really embraced agentic development, while 60% are still on tab-complete. Karl sees most clients stuck between basic AI use and agentic coding, held back by token budgets, security worries and fear of the unknown. Both agreed the bottleneck has moved to process: if you couldn’t ship quickly before AI, AI just makes the leak worse. Facundo suggested budgeting tokens the way you budget cloud, with enterprise, team and individual budgets, and tailoring agentic workflows to each company’s own systems. Karl urged leaders to stop counting tokens and measure outcomes instead. What’s working: Slalom hit 95% adoption at one client by creating an AI transformation group that focused on the middle of the pack, not just the champions. IBM ran a two-day hackathon for 400 to 500 people to help teams rethink their day jobs with AI. Their one metric to add in Q4? Karl said time from idea to production. Facundo said cost per PR paired with PR size.

How Jellyfish helps (Am I transforming?): Lifecycle Explorer pinpoints the process bottlenecks Karl and Facundo described, so you fix the pipe before buying more code generation. AI Cohorts show who your power users are and who’s in the middle of the pack and needs support. Pattern detection automatically surfaces what your high-leverage teams do differently, so you can spread it. See everything we launched.

Closing: Three days, one answer

Closing: Three days, one answer

My biggest takeaway from the partner sessions: the engineering organization that excel treat AI adoption as a science, with a real measure of value and outcomes, and humans taking the lead. So the next time someone asks what you got from your AI investment, you shouldn’t have to say “I think it’s working.” You can show them.

If you are a Software company focused on developer tooling then integrate with Jellyfish to show your customers the impact of your tools. Services firms can use Jellyfish as a diagnostic and measurement layer. You can reach us at partners@jellyfish.co for a conversation.

Request a Demo

Stop saying “I think it’s working.” Start showing it.

Your board and your CFO want to know what AI is worth, and your teams need help turning that answer into change. Jellyfish gives you the data: where your teams stand today, whether they’re actually transforming, and whether your AI spend is paying off, compared against 1,300+ other companies. Our partners help you act on it.

Request a demo and we’ll show you how it works.

About the author

Billy Robins

Billy Robins is Head of Partnerships at Jellyfish.

Read more by this author