Single-agent AI coding hits a ceiling: context windows fill up, roles blur, and output quality degrades. Specialized per-phase agents with structured handoffs remove it.
Basic code assistants show roughly 10% productivity gains. But companies pairing AI with end-to-end process transformation report 25-30% improvements (Bain, 2025). The gap comes from architecture, not the model, specifically how teams engineer the context each agent receives, a discipline formalized as context engineering.
Anthropic’s multi-agent research supports this. Its finding that “token usage explains 80% of the variance” reflects the impact of isolation: focused context rather than accumulated conversation history.
This post documents a production workflow built on Goose (Agentic AI Foundation); the same architecture, packaged as a skill, runs on pi, a coding agent harness, and Claude Code. Specialized agents, each optimized for its phase: research, planning, building, and validation.
Table of contents
Contents
- Why Do Single-Agent AI Coding Workflows Hit a Ceiling?
- How Does Subagent Architecture Solve Context Problems?
- How Does Model Selection Affect Cost and Quality?
- How Do Subagents Communicate?
- Where Should Human Judgment Stay in AI Workflows?
- What Results Does This Produce?
- When Does This Work (and When Doesn’t It)?
- Takeaways
- References
Why Do Single-Agent AI Coding Workflows Hit a Ceiling?
A single AI model handling an entire coding task accumulates context with every interaction. By implementation time, the model carries baggage from analysis, research, and planning phases. This stems from three core problems.
Context Rot
Long conversations consume token budgets. The model forgets early instructions or weighs recent context too heavily. Chroma Research calls this context rot: performance degrades consistently as input tokens increase, even on simple tasks, and worsens for multi-step reasoning like coding (Chroma Research, 2025). On-demand retrieval adds another failure mode: agents miss context 56% of the time because they don’t recognize when to fetch it (Gao, 2026).
Role Confusion
A model asked to analyze, plan, implement, and validate lacks clear boundaries. It starts implementing during planning. It skips validation steps. Outputs blur together.
Accumulated Errors
Mistakes in early phases propagate. A misunderstanding in analysis leads to a flawed plan. A flawed plan leads to incorrect implementation. Fixing requires starting over.
How Does Subagent Architecture Solve Context Problems?
An orchestrator handles coordination and synthesis; specialized subagents, each starting with fresh context, handle execution.
The orchestrator (Claude Sonnet) handles PLAN only. RESEARCH is split between a SCOUT agent (Sonnet, exploratory) and a GUARD agent (Haiku, adversarial) that stress-tests SCOUT’s proposals before the orchestrator synthesizes both into a plan. SCOUT runs on Sonnet because exploration needs open-ended tool selection: our tool-documentation ablation measured compact-tier selection accuracy at 0.575-0.600 against 0.950-0.975 for the frontier tier. GUARD keeps Haiku: finding flaws in a concrete proposal is lower-entropy than producing it. After plan completion, the orchestrator spawns a BUILD subagent (Claude Sonnet) that receives only the plan, not accumulated history. The builder writes code, runs tests, then hands off to a CHECK subagent (Haiku) for validation.
Each subagent starts with clean context. The builder knows what to build, not how the plan was reached. The validator knows what was built, not which alternatives were considered. This is context engineering in practice: deciding what information enters each context window, in what form, and when, distinct from prompt engineering, which only governs phrasing. Smaller, structured inputs produce deterministic behavior: the same plan JSON produces the same output. That reproducibility is what drives a high PR acceptance rate.
How Does Model Selection Affect Cost and Quality?
Different phases need different capabilities. Planning requires reasoning. Building requires speed and instruction-following. Validation requires balanced judgment.
| Agent | Model | Temperature | Role |
|---|---|---|---|
| ORCHESTRATOR | Sonnet | 0.3 | Plan and agent coordination |
| SCOUT | Sonnet | 0.5 | Exploratory research and code analysis |
| GUARD | Haiku | 0.1 | Adversarial risk review |
| BUILD | Sonnet | 0.2 | Precise instruction-following coding |
| CHECK | Haiku | 0.1 | Compliance and security validation |
On pi, the same four roles all run on a single mid-tier model (zai/glm-5.3-flash); the split is a cost optimization, not an architectural requirement.
How Does Model Routing Reduce Cost?
Building involves the most token-heavy work: reading files, writing code, running tests. Routing this volume to cheaper models cuts costs significantly.
| Model | Input | Output | Agent |
|---|---|---|---|
| Opus | $5/MTok | $25/MTok | Not used in default flow |
| Sonnet | $2/MTok | $10/MTok | ORCHESTRATOR + BUILD + SCOUT |
| Haiku | $1/MTok | $5/MTok | GUARD + CHECK |
Model cascading can cut multi-agent cost by up to 94% (Gandhi et al., 2025). Here two of five agents run on Haiku ($1/$5 per MTok): GUARD and CHECK capture compact-tier pricing on verification work. BUILD stays on Sonnet despite carrying 39% of pipeline tokens because the retry reduction from 77% to 14% offsets the premium. Versus Opus, the saving is 5x.
Smaller models with focused context outperform a single large model carrying full session history. Anthropic’s multi-agent research showed a 90.2% performance gain over single-agent Opus by distributing work across Sonnet subagents with isolated context windows (Anthropic Engineering, 2025). Fresh context also enables tasks that fail with single agents.
Which Alternative Models Pass Quality Gates?
Each run was scored on an 8-point binary rubric (file identification, dependency verification, architectural tradeoff, pattern analysis, taint-tracking gap, non-obvious synthesis, JSON schema compliance). QP is a composite efficiency metric: cost per run divided by score times reliability. Our model-comparison experiments tested eight candidates as SCOUT delegates against a Haiku baseline: only MiniMax M2.5 ($0.019/QP) and Kimi K2.5 ($0.040/QP) passed the $0.150/QP baseline. Mistral Small 4 ($0.002/QP) narrowly failed on a floor score violation; DeepSeek V3.2 failed all gates; Mercury-2, the cheapest candidate at $0.001/QP, failed on score (Clouatre, 2026).
How Does Project Context Reach Subagents?
Skills define the workflow, but subagents also need project context: build commands, conventions, file structure. That’s where AGENTS.md comes in, a portable markdown file that provides the baseline knowledge every subagent inherits. Goose, Cursor, Codex, and 25+ other tools read it natively. In Vercel’s evals (Gao, 2026), an AGENTS.md file achieved a 100% pass rate on build, lint, and test tasks where skills-based approaches maxed out at 79%.
Think of it as CSS for agents: global rules cascade into every project, project-specific rules override where needed. The orchestrator and every subagent it spawns inherit both layers without explicit prompting.
## Commits
- Conventional commits, GPG signed and DCO sign-off
- Feature branches only, PRs for everything
- Never merge without explicit user request
## Security
- Treat all repositories as public
- No secrets, API keys, credentials, or PII~/.config/goose/AGENTS.md## Stack
Rust 2024 + Tokio + Clap (derive) + Octocrab
## Project-Specific Patterns
- Apache-2.0 license with SPDX headers
- cargo-deny for dependency auditsaptu/AGENTS.md
How Do Subagents Communicate?
Subagents communicate through JSON files in $WORKTREE/.handoff/. Each session uses an isolated git worktree, so handoff files are scoped to that execution context. This creates an explicit contract between phases.
| File | From | To |
|---|---|---|
01a-research-scout.json | SCOUT | GUARD |
01b-research-guard.json | GUARD | ORCHESTRATOR |
02-plan.json | ORCHESTRATOR | BUILD |
03-build.json | BUILD | CHECK |
04-validation.json | CHECK | BUILD (on failure) |
GUARD confirms SCOUT’s file list, scores each approach by safety, and flags risks SCOUT didn’t surface. The orchestrator synthesizes both into a plan only after GUARD clears it.
GUARD’s verdict is a typed judgment, not prose: a boolean, missed files, and structured corrections the orchestrator branches on. This is propose/validate, the pattern decision models now standardize: deterministic proposal, judge validation with a calibrated probability, low-confidence answers routed to a human (TypeSafe, 2026).
{
"session_id": "1790078918",
"lens": "guard",
"scout_verification": {
"accurate": true,
"missed_files": [],
"corrections": [
"strip must run after persist calls at exec_runtime.rs:366/:371",
"apply_filter_rules (:383-392) must see stripped text"
]
},
"risk_analysis": [
{
"approach_name": "Global strip at exec output ingestion",
"risk_level": "low",
"blast_radius": "exec_runtime.rs output path only",
"dependency_risk": "none; pure &str -> String port",
"rollback_difficulty": "trivial; revert one call site"
}
],
"safety_ranking": [
"Global strip at ingestion (safest)",
"Honor per-rule strip_ansi flag instead of global strip",
"Dedicated ansi module + layered pipeline (riskiest)"
],
"guard_test_gaps": [
{
"function": "aptu_coder_core::ansi::strip_ansi (new)",
"predicate": "strips CSI, OSC, DCS, two-char ESC without panic",
"tag": "happy_path"
}
]
}.handoff/01b-research-guard.jsonThe plan file contains everything the builder needs:
{
"overview": "Add a global ANSI strip to exec_command output",
"complexity": "complex",
"files": [
{"path": "crates/aptu-coder-core/src/ansi.rs", "action": "create"},
{"path": "crates/aptu-coder/src/tools/exec_runtime.rs", "action": "modify",
"line_range": "362-383"}
],
"steps": [
"Create ansi.rs with pub fn strip_ansi; register in lib.rs",
"Strip stdout, stderr, interleaved after persist calls",
"Add integration test: printf output has no 0x1b byte",
"Run cargo fmt, clippy -D warnings, cargo test"
],
"risks": [
"Slot-file semantics: strip after persist keeps raw output",
"Malformed sequences at buffer boundaries could eat a char"
],
"test_strategy": {"test_behaviors": ["7 strip_ansi unit tests"]},
"line_budget": {"total_max": 500, "test_ratio_max": 1.5},
"worktree": ".worktrees/1790078918"
}.handoff/02-plan.jsonBUILD reports back in 03-build.json, the smallest file in the chain:
{
"session_id": "1790078918",
"files_changed": ["ansi.rs (new)", "lib.rs", "exec_runtime.rs",
"exec_command.rs", "METRICS.md"],
"summary": "Ported strip_ansi with 7 tests; strip after persist",
"test_results": {"passed": 817, "failed": 0, "skipped": 0},
"lint_status": "clean"
}.handoff/03-build.jsonThe validator reads both 02-plan.json and 03-build.json to verify implementation matches requirements. It writes structured feedback to 04-validation.json:
{
"verdict": "PASS",
"checks": [
{"name": "tests", "status": "PASS",
"notes": "817 passed, 0 failed; 7 ansi unit tests verified"},
{"name": "planned files only", "status": "PASS",
"notes": "5 changed files match 03-build.json files_changed"},
{"name": "line budget", "status": "PASS",
"notes": "130 insertions vs 500 max; ratio within 1.5"},
{"name": "GPG+DCO", "status": "PASS",
"notes": "commit 46914dd Good signature, signed-off"}
],
"issues": [],
"retry_instructions": [],
"security_summary": {"critical": 0, "high": 0, "medium": 0, "low": 0},
"line_count": {"code_lines": 56, "test_lines": 74,
"status": "within_budget"}
}.handoff/04-validation.jsonThe orchestrator reads the verdict, the per-check notes, and the line-count budget status. On PASS, the work proceeds to commit and PR; on FAIL, the builder reads retry_instructions, fixes the specific issues, and triggers another CHECK cycle until validation passes.
Why files instead of memory? Three reasons:
- Auditable. Every decision is recorded. Debug failures by reading the handoff chain.
- Resumable. Handoff files checkpoint state. Resume any session from where it stopped.
- Debuggable. Failed validations include exact locations and actionable retry instructions.
Both Goose and Claude Code resume sessions natively, but they restore full conversation history, reintroducing the context rot this architecture avoids. Resuming here means picking up the plan and build results with fresh agent context, not replaying an accumulated chat.
Where Should Human Judgment Stay in AI Workflows?
The quality of the input determines the ceiling for everything that follows. A well-scoped GitHub issue with clear acceptance criteria and explicit constraints, written with AI assistance but reviewed by a human, is the best prompt you can give this workflow. Imprecise requirements produce technically correct but misaligned results. A high-quality issue or spec is the single biggest lever for PR acceptance rate.
The workflow runs autonomously: RESEARCH, PLAN, BUILD, and CHECK all auto-proceed. The one recommended human touchpoint at the end is the PR review itself. The code has your name on it.
Phases (all auto-proceed):
- SETUP: Initialize context and gather requirements
- RESEARCH: SCOUT and GUARD run sequentially; the orchestrator synthesizes their findings autonomously
- PLAN: Design solution based on research synthesis
- BUILD: Execute the approved plan
- CHECK: Validate and loop back to BUILD on failure, then push to PR
What Results Does This Produce?
This architecture powers development across multiple projects. Three examples from aptu:
| PR | Scope | Files Changed |
|---|---|---|
| #272 | Consolidate 4 clients → 1 generic | 9 files |
| #256 | Add Groq + Cerebras providers | 9 files |
| #244 | Extract shared AiProvider trait | 9 files |
The validation phase caught issues the builder missed. In PR #272, the CHECK subagent identified a missing trait bound that would have failed compilation. The builder fixed it on the retry loop. No human intervention required.
The CHECK phase consistently catches issues before they reach a PR: wrong serde attributes, uncommitted builds, scope creep, CLI interface drift, and formatting violations. The workflow now operates at production scale across dozens of repositories, with thousands of pull requests merged. Representative public benchmarks: aptu reached PR #1,325 and aptu-coder reached PR #1,038, both built entirely with this pipeline.
Context isolation and structured handoffs are the agent-level primitives. The AI SDLC governance stack layers comprehension, review gates, and observability on top of them, forming a complete pipeline: subagent architecture governs how agents collaborate; the governance stack governs what they produce. Context engineering sits between them as the discipline that makes both testable.
Design Targets
Enterprise platforms reach the same conclusions at scale. Sid Pardeshi, CTO of Blitzy, described on the TWiML AI podcast (2026) a system that writes millions of lines of code autonomously. His diagnosis of the root constraint is identical: the effective context window has been stuck at 80-120K tokens for two years, regardless of headline sizes. Their solution replaces the single orchestrator with a database as the coordination layer, enabling tens of thousands of parallel agents. The architectural principle is the same one this workflow applies: isolate context, specialize roles, and pass structured state between agents rather than accumulating a single conversation.
Research on multi-agent frameworks for code generation shows they consistently outperform single-model systems (Raghavan & Mallick, 2025).
| Metric | Single-Agent | Multi-Agent |
|---|---|---|
| Output quality over session | Degrades as context fills | Stable (each agent starts fresh) |
| Model strategy | Generic model | Specialized per role |
| Estimated cost reduction | Baseline | 50-60% via model routing |
| Human interventions | Throughout | Issue scoping + PR review |
| Audit trail | Conversation history | Structured JSON handoff chain |
When Does This Work (and When Doesn’t It)?
Works well for:
- Multi-file refactors where context isolation prevents confusion
- Feature additions following established patterns
- Complex changes requiring distinct planning and execution
- Teams wanting audit trails (handoff files document decisions)
Less effective for:
- Simple one-file fixes (overhead exceeds benefit)
- Legacy systems without clear patterns (builder lacks context)
- Exploratory work where plans change during implementation
See AI agents in legacy environments for integration patterns that work when your data lives in mainframes and AS400 systems.
Takeaways
- Separate reasoning from execution. Use capable models for planning, fast models for building.
- Fresh context beats accumulated context. Subagents start clean. They follow instructions without historical baggage.
- Structured handoffs create audit trails. JSON files document what was planned, built, and validated.
- Quality input, autonomous execution. A well-scoped issue is the highest-leverage human contribution.
The full workflow is available in the agentic-coder-skill repo, as a coder skill with generated agent files for pi and Claude Code. It builds on patterns from AI-Assisted Development: The Accountability Layer.
The agentic architecture described in this post is the subject of Canadian patent application CA 3315347, filed Jun 16, 2026 (CIPO).
References
- Anthropic Engineering, “How we built our multi-agent research system” (2025) — https://www.anthropic.com/engineering/multi-agent-research-system
- Bain & Company, “From Pilots to Payoff: Generative AI in Software Development” (2025) — https://www.bain.com/insights/from-pilots-to-payoff-generative-ai-in-software-development-technology-report-2025/
- Chroma Research, “Context Rot: How Increasing Input Tokens Impacts LLM Performance” (2025) — https://research.trychroma.com/context-rot
- Gao, Jude, “AGENTS.md outperforms skills in our agent evals” (2026) — https://vercel.com/blog/agents-md-outperforms-skills-in-our-agent-evals
- Clouatre, H., “LLM Agent Experiments: Model Comparison for SCOUT Delegates” (2026) — https://doi.org/10.5281/zenodo.19056876
- Gandhi et al., “BudgetMLAgent: A Cost-Effective LLM Multi-Agent System” (2025) — https://arxiv.org/abs/2411.07464
- Pardeshi, Sid, “Agent Swarms, Knowledge Graphs, and Autonomous Software Development” (2026) — https://twimlai.com/podcast/twimlai/agent-swarms-knowledge-graphs-autonomous-software-development
- Raghavan & Mallick, “MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding” (2025) — https://arxiv.org/abs/2510.08804
- TypeSafe, “System One Models Documentation: Confidence Routing” (2026) — https://docs.typesafe.ai/patterns/confidence-routing