How to Build a Secure AI SDLC Framework for Scaling Agentic Development

AI coding tools have spread through engineering teams faster than most governance models can keep up with. Developers adopt AI tools and agents independently, AI-generated pull requests pile up in review queues, and engineering leaders have limited visibility into how those tools affect delivery.

A Harvard study of Jellyfish data covering 100,000 engineers found that AI makes coding faster with no clear drop in code quality, but the gains haven’t yet led to a significant increase in features shipped. Jellyfish points to bottlenecks outside coding, such as review, testing, and cross-team coordination, as one likely reason.

Engineering leaders need an AI SDLC framework to correct that imbalance. This operating and technical blueprint connects AI tools to delivery pipelines, governance policies, and business goals.

The sections below cover the four pillars, a maturity model to assess your team, the tool stack, and a phased plan to scale agentic development.

What Is an AI SDLC Framework?

What Is an AI SDLC Framework?

An AI SDLC framework is the operating model an engineering organization uses to govern AI across the software development life cycle. It sets the rules for which AI tools and agents teams can use, what code and data those tools can access, how humans review AI-generated work, and how leaders measure the return on AI spend.

A written AI usage policy covers part of this work. The framework builds those rules into the delivery pipeline through model gateways, CI/CD quality gates, review standards, and shared telemetry.

Ownership usually spans four groups:

  • Engineering leadership sets adoption goals, owns workflow standards, and reports AI impact to the business.
  • Platform engineering manages model access, repository context, and the pipeline integrations that connect AI tools to delivery.
  • Security and compliance define data handling rules, agent permissions, and audit requirements.
  • Development teams apply the standards day to day and point out where the workflow slows them down.

Example → Under an AI SDLC framework, a developer using an agentic coding tool works through an approved model gateway that keeps secrets and proprietary code within company boundaries. The agent’s pull request passes automated security scans and a size check before a human reviewer opens it. The PR is also tagged as AI-assisted, so leaders can compare its cycle time and rework rate with human-written work.

Diagram of an AI SDLC workflow where a developer or agent works through a model gateway, opens a pull request tagged as AI-assisted, passes automated security and size checks and human review for high-risk changes, then merges, with telemetry tracking adoption, cycle time, rework and AI spend

Core Components of a Modern AI SDLC Framework

Core Components of a Modern AI SDLC Framework

A working AI SDLC framework rests on four pillars. Each one covers a different kind of risk, from leaked source code to AI spend that doesn’t show up in delivery results, and together they give leaders a consistent view of how AI moves through the pipeline.

Pillar What it protects Primary owner Key signal
Governance, IP, and security guardrails Source code, customer data, and production access Security and compliance Unapproved tool usage and policy exceptions
Centralized AI infrastructure and context Code quality and consistency across teams Platform engineering Weekly active usage of approved tools
Human-in-the-loop workflow standards Reviewer capacity and release stability Engineering managers and tech leads PR size and review time on AI-assisted PRs
Intelligence and telemetry The business case for AI investment Engineering leadership Cycle time, change failure rate, and AI spend per developer

Pillar 1: Governance, IP, and Security Guardrails

This pillar sets the boundaries for how AI tools handle source code, customer data, and system access. The need for it grows as AI writes more production code.

Veracode’s 2026 GenAI Code Security Report found that roughly 44% of tested AI code-generation tasks resulted in code with a known security vulnerability.

Key controls:

  • Enterprise agreements with appropriate data-retention and training restrictions can prevent proprietary code from being retained or used for model training.
  • Each agent gets its own service identity, with least-privilege permissions limited to approved repositories and actions.
  • Automated scans can catch unverified packages and exposed credentials before they reach a prompt or a commit.
  • Audit logs record which tool, model, and agent produced each change, supporting internal audits and standards such as NIST SP 800-218A.

Example → A coding agent that reads Jira tickets could pick up hidden instructions in a ticket description and try to push changes to a protected branch. Scoped permissions block the push, and the audit log shows which agent attempted it and when.

Pillar 2: Centralized AI Infrastructure and Context

AI tools produce better code when they can see how the rest of the system fits together. Jellyfish’s State of AI report found that higher AI adoption correlated with roughly 4× greater PR throughput in centralized and balanced codebases.

Bar chart from the Jellyfish State of AI report showing AI productivity gain in PRs per engineer of 3.9x for centralized codebases, 4.1x for balanced, 2.1x for distributed and 0.9x for highly distributed

In the most fragmented setups, the gain dropped to 0.9x, meaning AI usage correlated with a slight decrease in throughput. Centralized infrastructure and shared context give every team the same starting point.

Key controls:

  • A central model gateway routes AI traffic through approved models, applies usage policies, and tracks token spend by team.
  • Repository-level instruction files such as AGENTS.md or CLAUDE.md give every tool the same coding standards, architecture rules, and test commands.
  • Model Context Protocol (MCP) servers connect agents to approved internal sources like documentation, tickets, and service catalogs.
  • Platform teams maintain a short list of vetted assistants and AI agents, with a clear process for requesting new ones.

Example → A platform team adds an AGENTS.md file to its payments repository that names the approved logging library, the test command, and the rule that all database calls go through the shared data client. Every agent working in that repo reads the same guidance, so AI-generated PRs follow house patterns before a reviewer opens them.

Pillar 3: Human-in-the-loop (HITL) workflow standards

Every AI-generated change needs a defined approval path before it can be merged. Depending on the risk, that may involve a human reviewer, automated checks, or both. Teams should define those thresholds in advance, so reviewers know which changes require human sign-off and which can follow a lighter automated path.

Jellyfish’s AI Engineering Trends data shows that lead adopters now merge 9% or more of their PRs with AI as the only reviewer, compared with 0.6% at the median company. Each team needs a clear policy on which changes can take that path and which still need human oversight.

Jellyfish AI Engineering Trends chart showing the percentage of merged PRs that were agentic-merged by company percentile from January 2025 to August 2026, reaching about 9% at the 90th percentile

Key controls:

  • Size limits for AI-generated PRs keep each change small enough for a reviewer to understand in one sitting.
  • Changes to authentication, payments, or infrastructure code go to senior reviewers, while low-risk changes follow a lighter path.
  • Automated tests, static analysis, and AI review agents check each PR before a human opens it, so reviewers can focus on design and logic.
  • PR labels mark AI-assisted and agent-authored work, which tells reviewers how much scrutiny a change needs and gives leaders clean data for measurement.

Example → When Jellyfish enabled AI code review for 18 of its own engineers, their PR output more than doubled and fewer PRs needed changes after review. Jellyfish attributed part of the shift to faster AI feedback, which encouraged engineers to break work into smaller units. Nine months later, their PRs averaged 268 lines, 82% smaller than those of teammates without the tool.

PRO TIP 💡 → Check merge outcomes before approving an AI-only review path for any change type. Jellyfish AI Workflow Insights shows agent-authored PRs, contribution splits, and merge outcomes side by side with human work, so teams can see which kinds of changes already hold up under lighter review.

Jellyfish AI Workflow Insights showing agent-assisted PRs with their collaborators, the human share of commits and each PR's merge outcome

Pillar 4: Intelligence and Telemetry

AI costs are now a budget decision for many engineering organizations. Jellyfish’s State of AI report found that the typical user at a median company went from about $5 a month to $81 in ten months.

Line chart from the Jellyfish State of AI report showing monthly AI cost per user by company percentile from August 2025 to May 2026, with the median company rising from $5 to $81

Intelligence and telemetry give leaders one view across Git, issue tracking, CI/CD, and AI tools, so they can trace AI spend to the work it produced.

Core metrics:

  • Adoption metrics such as weekly active users and the share of AI-assisted and agent-authored PRs show how deeply AI is embedded in daily work.
  • Cycle time, PR pickup time, and review time show whether AI output moves through the pipeline or waits in queues.
  • Change failure rate and 30-day rework rate reveal whether faster output holds up in production.
  • AI spend per developer and per merged PR connects tool and token costs to the work they produce.

Example → A company pays for two AI coding tools and wants to consolidate. Telemetry shows that teams using one tool merge more PRs with lower rework, while teams on the other produce more code that comes back for fixes within 30 days. Leadership keeps the stronger tool and moves the second license budget into enablement for the teams that need it.

The AI SDLC Maturity Model: Where Does Your Team Sit?

The AI SDLC Maturity Model: Where Does Your Team Sit?

Engineering organizations usually adopt AI in stages, and each stage brings its own risks. The model below breaks that path into four levels, with the signs to look for at each one and the step that moves a team forward.

Level What it looks like Signs your team is here Next step
Level 1: Ad hoc and shadow AI Developers use personal accounts and browser-based chatbots with no approved tools or policies. Code gets pasted into public tools, no one knows which teams use AI, and AI costs are scattered across expense reports. Audit current usage, approve a small set of enterprise tools, and set data-handling rules.
Level 2: Standardized tools The company licenses approved AI tools and sets basic security policies, but workflows stay the same. Adoption is high, and PR volume is growing, and review queues keep getting longer. Set PR size limits, add automated checks to CI/CD, and start measuring cycle time and review time.
Level 3: Integrated workflows AI is built into CI/CD, with review standards, quality gates, and delivery metrics that include artificial intelligence activity. AI-assisted PRs are labeled, automated checks catch issues before human review, and leaders track AI’s effect on cycle time and quality. Give agents well-scoped tasks like test generation and dependency updates, with clear approval paths.
Level 4: Agentic SDLC Agents handle routine coding, testing, and maintenance, and engineers focus on architecture, review, and final approval. Agents open a meaningful share of PRs, AI-only review follows a defined policy, and AI spend is budgeted by the value of the work. Expand agent autonomy step by step, and keep measuring outcomes against spend and quality.

Many organizations today are somewhere between Levels 2 and 3. Jellyfish’s State of AI report found that autonomous agents open about 2% of PRs at the median company, compared with 35% at the 90th percentile as of April 2026. Level 4 is still limited to a small group of teams, and the distance between them and everyone else keeps growing.

The move from Level 2 to Level 3 is usually the hardest. Licenses and policies can get a team to Level 2 fairly quickly, and the next step depends on changes to review practices, CI/CD gates, and measurement.

Maturity can also vary across the framework. A team might have Level 3 review workflows but Level 1 measurement, for example. Assessing each of the four pillars separately can expose those gaps before you decide what to improve next.

Keep in mind → Level 4 isn’t the right target for every team. For regulated codebases and safety-critical systems, Level 3 with agents limited to low-risk work may give the best return.

The Modern AI SDLC Tool Stack

The Modern AI SDLC Tool Stack

An AI SDLC framework typically depends on several categories of tools working together. Individual platforms may cover multiple functions, but engineering teams still need capabilities for generation, context, validation, workflow automation, governance, and measurement.

Category What it does Examples
AI coding assistants and agents Generate code, tests, and documentation, and complete multi-step tasks on their own GitHub Copilot, Cursor, Claude Code, Gemini Code Assist, Kiro, Devin
Model access and gateways Centralize access to approved models, enforce usage policies, and monitor consumption LiteLLM, Portkey, Amazon Bedrock, Microsoft Foundry
Context and knowledge Give AI tools access to codebase knowledge, documentation, and internal systems MCP servers, Unblocked, Augment
AI code review Review PRs automatically and identify potential bugs, security issues, and risky changes before or alongside human review CodeRabbit, Graphite, Greptile, Cursor Bugbot
Validation and security Scan code and dependencies for vulnerabilities, exposed secrets, and unverified packages Snyk, SonarQube, Semgrep, Socket
CI/CD and workflow automation Automate builds, tests, policy checks, and deployments, enforce quality gates, and apply rules for PR routing and approvals GitHub Actions, GitLab CI/CD, LinearB (gitStream)
Engineering intelligence (SEI) Connect Git, issue tracking, CI/CD, and AI tool data to measure adoption, delivery, and ROI Jellyfish

Many engineering organizations already own tools in several of these categories. The challenge is getting them to work as one system, with shared policies, shared context, and data that flows into a single place for measurement. A stack assembled one purchase at a time can leave each team with a different picture of how AI is performing.

What to look for when choosing tools:

  • Usage and activity data should flow from each tool into your measurement platform, so AI impact can be tracked across the whole pipeline.
  • Enterprise controls vary by vendor and plan, so confirm what each tool offers for SSO, audit logging, data retention, and use of your code for model training.
  • Agent access should follow least privilege, with permissions limited to each task and actions traceable to the specific agent where the platform supports it.

Future-proofing tip → AI tooling changes quickly. AWS, for example, will end support for Amazon Q Developer IDE plugins and paid subscriptions on April 30, 2027, and recommends Kiro as the replacement for IDE users. Centralized model access, portable context files, and tool-agnostic policies make it easier to switch tools without redesigning workflows from scratch.

A Three-Phase Plan to Implement Your AI SDLC Framework

A Three-Phase Plan to Implement Your AI SDLC Framework

The three phases below move a team through the maturity model, from uncovering shadow AI in Phase 1 to scaling agentic workflows in Phase 3.

The timelines assume a single business unit or pilot group, and larger organizations should expect each phase to take longer.

Phase 1: Baseline and Audit (Weeks 1–2)

Phase 1 documents current AI usage and current delivery performance. The audit lists every AI tool in use, who uses it, whether it is licensed through the company or a personal account, and which repositories and systems it can access.

The baseline uses historical Git, issue tracker, and CI/CD data to record cycle time, review time, change failure rate, and rework rate before any changes are made.

Key actions:

  • An audit of AI tool usage across teams, including personal accounts and browser-based chatbots, uncovers shadow AI and any code or data it may have exposed.
  • Baseline metrics for cycle time, review time, change failure rate, and 30-day rework rate give later phases a point of comparison.
  • Security and platform leads review which repositories, secrets, and systems AI tools can currently reach.
  • Leadership names an owner for each pillar and agrees on the metrics the framework will be judged by.

Ready for the next phase when: The team has a documented inventory of AI tools, baseline delivery metrics covering at least the past quarter, and a named owner for each pillar.

Common pitfall ⚠️: Auditing only the tools the company pays for. Developers often use personal accounts or browser-based chatbots alongside approved tools, and those usually carry the most data exposure risk.

PRO TIP 💡 → The Phase 1 baseline doesn’t require new instrumentation. Jellyfish derives delivery and AI usage signals from the Git, issue tracking, and workflow data teams already have, so the baseline includes historical cycle time and review time from day one. Impact Insights then uses that starting point for before-and-after comparisons in Phases 2 and 3.

Jellyfish Impact Insights chart showing monthly pull requests split into AI-assisted and unassisted work, with automated insights on throughput changes

Phase 2: Standardize and Gate (Weeks 3–6)

Phase 2 puts the controls from the four pillars in place for the pilot group. Approved tools go through a central model gateway, AI-generated PRs follow size and review rules, and automated checks in CI/CD scan each change before a human reviewer sees it.

The target for this phase is Level 3 on the maturity model, with AI built into the workflow and measured against the Phase 1 baseline.

Key actions:

  • Approved AI tools move behind a central model gateway, and personal accounts found in the Phase 1 audit are replaced with enterprise licenses.
  • Repository instruction files such as AGENTS.md or CLAUDE.md document coding standards, test commands, and architecture rules for AI tools to follow.
  • CI/CD pipelines bring static analysis, dependency scanning, and secret detection as required checks on every PR.
  • Review rules set size limits for AI-generated PRs, label AI-assisted work, and send high-risk changes to senior reviewers.

Ready for the next phase when: AI usage in the pilot group goes through approved tools, required CI/CD checks apply to every PR, and review time and change failure rate hold steady or improve against the Phase 1 baseline.

Common pitfall ⚠️: Setting PR size limits without changing how work gets planned. If tickets stay large, developers split code into arbitrary chunks to meet the limit, and reviewers lose the context they need. Breaking work into smaller tickets during planning keeps PRs small for the right reasons.

Phase 3: Scale and Optimize (Week 7 Onward)

Phase 3 brings the framework to the rest of the organization and gives agents a larger share of routine work. Agents start with well-scoped tasks such as test generation, dependency updates, and code migrations, following the approval paths set in Phase 2.

Leaders compare delivery metrics and AI spend with the Phase 1 baseline to decide which workflows to expand and which to adjust.

Level 4 is the goal for teams and codebases where agents produce measurable gains. Higher-risk systems often do better at Level 3, with agents limited to low-risk work.

Key actions:

  • Other teams adopt the framework in waves, using the pilot group’s instruction files, CI/CD checks, and review rules as templates.
  • Low-risk change types, such as documentation updates or minor dependency bumps, can merge after AI review alone under a written policy approved by security.
  • Token budgets are set per team based on the value of the work, with spend tracked alongside throughput and rework.
  • Quarterly reviews compare each team’s metrics with the Phase 1 baseline and reassess its maturity level for each pillar.

Signs it’s working: Delivery metrics improve against the Phase 1 baseline across several teams, AI spend per merged PR stays stable or declines, and change failure rate holds steady as agent-authored work grows.

Common pitfall ⚠️: Treating agent PR volume as the measure of success. More agent-authored PRs only count as progress if cycle time, change failure rate, and rework hold steady or improve at the same time.

Scale Your AI SDLC Framework Safely with Jellyfish

Scale Your AI SDLC Framework Safely with Jellyfish

Most of the work in this guide depends on measurement. The Phase 1 baseline, the maturity assessments, and decisions about where to expand agent use all rely on data that connects AI activity to delivery outcomes. Vendor dashboards report usage for their own tools, so leaders need a separate view that covers the whole pipeline.

That’s the role Jellyfish fills. As a software engineering intelligence platform, Jellyfish brings data from Git, issue trackers, CI/CD pipelines, and AI coding tools into one place. Our AI Impact product measures adoption, spend, and delivery outcomes across assistants, review agents, and autonomous agents, using the same model for every vendor.

Here’s exactly how Jellyfish supports each part of the framework:

  • Adoption Insights: System signals show AI usage across tools, teams, and roles automatically, with no manual reporting. Adoption trends point to where usage grows or plateaus, which teams need more enablement, and where autonomous agents contribute. Phase 1 audits and later expansion both work from the same usage data.

Jellyfish Adoption Insights showing pull requests merged each month assisted by AI, with insights that merged PRs increased 46% and 20% remain unassisted

  • Impact Insights: Before-and-after comparisons measure how delivery speed and PR cycle time change with and without AI, using SDLC data. Engineering leaders can break results down by individual, team, repo, or code area, and see how AI changes the split between roadmap work and keep-the-lights-on (KTLO) work. These comparisons show whether the framework improves on the Phase 1 baseline.
  • AI Workflow Insights: This view tracks agent-authored PRs, contribution splits, and merge outcomes to show how humans and agents work together. It also measures the balance of AI-generated and human-written code and monitors code quality. Engineering managers can use these signals to set review standards and adjust them as agent work grows.

Jellyfish AI Code Ratio chart comparing lines merged with AI lines accepted and the resulting AI code ratio over time

  • AI Token Cost Management: The dashboard tracks token usage and spend by tool, team, and model, and connects that spend to PR throughput over time. Year-to-date and projected spend figures support budget forecasting. Finance and engineering teams get a clear picture of which teams and tools produce the most output for their spend.
  • Vendor Comparison: One vendor-neutral model, built on normalized SDLC signals, benchmarks assistants, code review agents, and autonomous agents against each other. Cycle time, throughput, and quality data identify top and underperforming tools, and task-level comparisons show which ones work best for writing, reviewing, debugging, or documenting. Platform teams can base renewal and consolidation decisions on that evidence.

Jellyfish Vendor Comparison tool breakdown comparing adoption, allocation and issue cycle time for GitHub Copilot, Google Gemini, Cursor, Amazon Q and Windsurf

  • Auto Report Builder: Auto-generated executive reports summarize what’s working, what isn’t, and where to invest next. Quarterly maturity reviews and board updates can start from these reports, so no one has to assemble the data by hand.

The framework in this guide gives AI a clear path through the SDLC, with guardrails, review standards, and a plan for scaling agent work. Jellyfish supplies the data to confirm each step is working, from the Phase 1 baseline to the decision to expand agent autonomy.

Book a demo to see how Jellyfish measures AI’s impact across your engineering organization.

FAQs

FAQs

How does an AI SDLC framework address regulatory compliance and risk?

An AI SDLC framework gives organizations a structured way to practice responsible AI across software development. Many teams align their controls with the NIST AI RMF, a voluntary U.S. framework for managing AI risk, and with NIST’s companion profile for generative AI.

Organizations with EU operations also track the EU AI Act. Some of its obligations already apply, including transparency rules from August 2026, and the Digital Omnibus amendments moved high-risk system requirements to December 2027 for stand-alone systems and August 2028 for AI embedded in regulated products.

Traceability is a common requirement across all of these. Audit logs that record which large language models, tools, and agents produced each change let reviewers and auditors trace AI-generated code back to its source.

How does AI change the early stages of the software development lifecycle?

AI now affects requirements gathering as well as coding. In an AI software development lifecycle, more teams use spec-driven development, where product and engineering write structured, machine-readable specifications before any code exists.

Agentic AI tools can help draft those specs from user interviews, tickets, and design documents, and coding agents then work from them with clear acceptance criteria.

Detailed specs give developers and agents the same context, and they make quality assurance easier because teams can generate tests directly from the requirements. AWS’s Kiro, for example, is built around this spec-first approach.

How does an AI SDLC framework integrate with DevOps and infrastructure?

An AI SDLC framework works inside existing DevOps pipelines. Agents connect to infrastructure through the Model Context Protocol (MCP) or vendor SDKs, and they can help engineers write Terraform modules or update Kubernetes manifests.

Because these changes affect production systems, they need the same CI/CD gates as application code, plus policy checks and human approval for high-impact changes.

A current dependency graph helps reviewers see which services an AI-proposed change will affect, and observability tools confirm how those changes behave after deployment. The same controls apply to agentic systems that act on infrastructure without a developer in the loop.

How do you measure whether an AI SDLC framework is working?

Measurement combines software delivery metrics with signals about developer experience. Cycle time, review time, change failure rate, and rework rate show whether AI speeds up work without hurting stability.

Developer surveys bring context that pipeline data can’t capture, such as trust in AI output and time spent reviewing it. Comparing both sets of data against a pre-framework baseline shows which parts of the framework improve results.

Software engineering intelligence platforms like Jellyfish connect these data sources so leaders can track them in one place.

About the author

Lauren Hamberg

Lauren is Senior Product Marketing Director at Jellyfish where she works closely with the product team to bring software engineering intelligence solutions to market. Prior to Jellyfish, Lauren served as Director of Product Marketing at Pluralsight.

Read more by this author