Writing · Dr. Kai Stalmann

Beyond Agile: Project Management for AI-Driven Development

6 February 2026EnglishFirst published on LinkedIn

This is the first in a two-part series. This article examines why agile's assumptions break down in AI-driven development. The follow-up will present a concrete methodology — built around goals, decisions, and sessions — along with tooling that puts it into practice.

The Shift

Software development is undergoing a fundamental change. AI coding assistants can implement features, write tests, refactor code, and generate documentation — not as a future promise, but today. This changes who does what, how fast things move, and what actually matters for project success.

Agile was designed for a world where humans write all the code. That world is shifting. The practices that made agile revolutionary in 2001 now create friction in a workflow where AI handles much of the implementation and humans handle direction and review.

This isn't about agile being wrong. It's about the assumptions underneath it no longer holding.

What Agile Assumes

Human velocity is the bottleneck

Agile's measurement system — story points, velocity, burndown charts — exists because implementation is slow and unpredictable. A developer might estimate a feature at 5 points, take 3 days, hit an unexpected issue, and need 5 days. Sprint planning exists to manage this uncertainty.

When AI implements a feature in minutes, these measurements lose their meaning. The estimation game collapses because the thing being estimated — human coding time — is no longer the dominant constraint. If you want to measure progress, measure decisions made and goals completed. Those are the outputs that actually advance a project.

Example: A team estimates a REST API endpoint at 3 story points, roughly a day's work. With AI, the implementation takes one prompt. The actual time is spent deciding the API contract, authentication scheme, and error handling strategy. The decisions took two hours. The coding took two minutes. Story points measured the wrong thing.

The NoEstimates movement has made a similar observation — that story points and velocity often create overhead without improving predictability. AI-driven development makes this critique unavoidable: when implementation cost approaches zero, estimating it is pure waste.

Coordination is between implementers

Standups, sprint planning, and retrospectives exist because multiple humans need to stay aligned while building different parts of a system. "I'm working on the payment module, so don't touch the order service until I'm done" — this coordination prevents conflicts and ensures coherent architecture.

In AI-driven development, the implementer is increasingly the AI itself. The humans aren't coordinating implementation; they're coordinating decisions. "We need to decide on the payment provider before either of us can define the integration interface" is a fundamentally different coordination problem than "don't edit the same files I'm editing."

Daily standups asking "what did you code yesterday?" lose their purpose when the answer is always "I told the AI what to build." The useful question becomes "what decisions did you make yesterday, and what decisions are blocking you?"

Iterations discover requirements

Agile's iterative approach — build a little, show it, get feedback, adjust — assumes that building is the expensive part and that seeing working software reveals hidden requirements. Sprint demos exist because stakeholders need to see the software to react to it.

This remains partially true. Seeing working software does surface requirements. But when building is cheap, the iteration cycle compresses dramatically. For well-scoped features, you can build multiple prototypes in an afternoon and discuss them. The iteration overhead — sprint ceremonies, backlog grooming, story writing for each cycle — risks becoming the bottleneck rather than the solution.

What Actually Matters Now

Decisions, not tasks

In traditional development, the workflow is: break work into tasks → assign tasks → implement → review → integrate. The task is the unit of work.

With AI, implementation is not the hard part. The hard part is knowing what to implement. Every feature involves dozens of decisions:

  • What's the API contract? (REST vs GraphQL, field naming, pagination strategy)
  • What authentication method? (JWT vs sessions)
  • How should errors be handled? (exceptions vs result types)
  • How do we handle edge cases? (concurrent writes, network failures, partial success)

Each decision unblocks the AI to implement. Without them, the AI either guesses (risky) or asks (slow). The pace of a project is set by how quickly decisions are made, not how quickly code is written.

A backlog item like "Implement pagination for user list endpoint" is useless in this world — AI can implement pagination in six different styles in seconds. The useful backlog item is the decision: "Choose pagination strategy: cursor-based, offset-based, or keyset? Consider our mobile client's needs."

Example: A team creates 15 user stories with acceptance criteria and assigns them across three developers. By end of sprint, 10 are done, 5 carry over. With AI, those 15 stories could be implemented in a day — but only if all the implicit decisions embedded in the acceptance criteria are actually made. "As a user, I want to reset my password" sounds simple, but hiding inside it: email or SMS? Token expiry? Rate limiting? Account lockout? The user story format obscures these decisions. The decisions are the actual work.

Goals, not stories

User stories ("As a user, I want X so that Y") were a useful format for capturing intent and keeping developers focused on user value. But they operate at the wrong level for AI-driven development.

AI needs clear outcomes and constraints, not reminders about user value. And humans directing AI don't need granular stories — they need goals with enough context to make decisions.

A goal like "Users can authenticate securely with email and OAuth, with proper session management and password recovery" is more useful than 12 separate user stories. It gives the AI full context for coherent implementation choices, and it gives the human a clear criterion for "done."

The granular breakdown happens during implementation: the AI figures out the steps, asks for decisions when needed, and delivers against the goal. The human reviews the result against the goal, not against individual story acceptance criteria.

Review gets harder, not easier

What changes about code review is not whether to do it, but what to look for. AI-generated code has no engineering intent behind its micro-decisions — it chose an approach based on patterns in its training data, not architectural rationale. This has specific consequences.

AI bypasses architecture. This is the most common and damaging pattern in AI-generated code. A team carefully builds a repository layer to encapsulate database access; the AI writes a direct SQL query in a controller because it "works." The project has centralized error handling middleware; the AI adds a try-catch inline. There's a service layer for business logic; the AI puts it straight into the API handler.

Each instance looks harmless. The code is correct, tests pass, the feature works. But cumulatively, these shortcuts erode the architecture. After enough AI sessions, the codebase has two ways of doing everything: the intended way through abstractions, and the direct way the AI chose. This creates maintenance nightmares and defeats the purpose of the architecture.

It happens because AI optimizes for the immediate problem with the context it has. It doesn't respect architectural boundaries unless explicitly told about them. A human developer who participated in the architecture discussions instinctively routes database access through the repository. The AI, seeing a table name and a query need, writes the query. Both produce working code. Only one respects the architecture.

Defects shift. AI doesn't make typos or forget semicolons. It makes structural errors: wrong algorithm for the data characteristics, missed edge cases it wasn't told about, code that works for the happy path but fails under concurrency, security patterns that look correct but aren't. Linters and tests catch the shallow issues. Reviewers must focus on architectural coherence, security, and domain correctness — the things that require genuine understanding.

The review question expands. Traditional review asks: "Is this code correct and clean?" AI-driven review adds: Does this follow our architectural decisions? Are there patterns the AI introduced that contradict our conventions? Would a human engineer have made a different structural choice, and if so, is the AI's choice acceptable?

This is why architectural decisions must be explicitly documented where the AI can read them — and why reviewers must check not just "does it work?" but "does it work the way our architecture intends?"

When AI writes the code, understanding no longer comes through authorship. It must come through deliberate review. The review becomes the primary mechanism for the team to understand the codebase, not a secondary check.

What Breaks

Sprint cadence becomes artificial

When implementation takes minutes instead of days, fixed two-week sprints lose their rationale. A "sprint" might contain the equivalent of months of traditional work, or it might stall for days waiting on a single decision that requires stakeholder input.

The natural rhythm of AI-driven work is the session — a block of time where a human directs AI toward a goal. Sessions might last 30 minutes or 8 hours. They don't align with sprint boundaries. Forcing them into sprints adds ceremony without value.

Roles blur

Scrum defines roles: Product Owner, Scrum Master, Developer. These assume clear boundaries between "deciding what to build," "facilitating the process," and "building it."

With AI, everyone who interacts with it is simultaneously directing, reviewing, and building. Role boundaries dissolve. What matters is who has the authority and context to make specific decisions.

The Tooling and Workflow Gap

Most development tools were built for the old model. Issue trackers manage task lifecycles (to-do → in progress → done) that map to human implementation. Project boards visualize work flowing through coding stages. Standups coordinate who's working on what code. Documentation lives separate from the codebase.

In AI-driven development, each of these assumptions breaks. The lifecycle becomes: decision needed → decided → implemented → reviewed. The board should reflect decision flow, not coding flow. Coordination is about who's making which decisions. And documentation — architecture docs, decision records, CLAUDE.md-style context files — isn't a secondary artifact but the primary interface between human intent and AI execution.

Similarly, git workflows change. Branches become short-lived (hours, not weeks). Their purpose shifts from isolating long-running work to containing a goal-directed session for review. Pull requests shift from reviewing a human's implementation choices to verifying that a goal was met and architectural decisions were followed. The project's memory — architecture, decisions, context — and its code have different ownership patterns and different merge characteristics. Current workflows treat them identically, and this creates friction.

There is nothing today that treats decisions as the unit of progress, goals as the organizing structure, and sessions as the natural work rhythm. This gap is real, and closing it requires not just new tools but a coherent methodology that connects these concepts.

Prior Art

This analysis doesn't exist in a vacuum. Several approaches have pushed in related directions:

Shape Up (Basecamp/37signals) eliminated backlogs, replaced estimation with fixed-time "bets," and gave teams autonomy to figure out implementation within a shaped scope. It recognized that ceremony and granular task tracking create drag. But it predates the AI shift and still assumes human implementation as the core activity.

Spec-Driven Development treats specifications as first-class artifacts that drive AI code generation — "version control for your thinking." This aligns with the insight that documentation becomes the primary medium of direction, but focuses on the specification layer rather than the full project methodology.

AWS's AI-Driven Development Lifecycle (AI-DLC) proposes adaptive workflows where AI handles execution under human oversight, with humans retaining decision authority. It shares the premise that rigid ceremonies don't fit AI-driven work, but focuses on process orchestration rather than the underlying shift in what constitutes work.

Each of these addresses part of the picture. None yet offers a complete methodology built around the primitives that AI-driven development actually requires: goals as the organizing structure, decisions as the unit of progress, sessions as the natural rhythm, and review as the primary point of human engineering judgment.

What Comes Next

Agile was a response to waterfall. It recognized that software development is iterative and that rigid upfront planning fails. AI-driven development is the next shift: when implementation is cheap, the human contribution is direction, decisions, and rigorous review.

The methodology that fits this reality will share agile's core value — responding to change — but organize around different primitives. Goals instead of stories. Decisions instead of tasks. Sessions instead of sprints. Review that focuses on architectural compliance and intent rather than implementation style.

This article has outlined the problem. A follow-up will present a concrete methodology built on these primitives, along with templates, tooling, and workflows designed for AI-driven development environments. The goal is not another framework to adopt ceremonially, but a practical system that treats human decision-making as the central activity it has become.

Dr. Kai Stalmann · qantr GmbH All writing →