Agent Reliability
Moving autonomous LLM systems into production requires moving past anecdotal prompt engineering toward deterministic validation. This track covers empirical evaluation harnesses, context management patterns, and multi-agent coordination architectures that ensure predictable, testable behavior.
Start here
Reading order
-
AI Delivery Decision Frameworks: Type 1, Type 2, DACI
• UpdatedMisclassifying reversible decisions costs more than the decision itself. Four frameworks unblock AI delivery: Type 1/Type 2, Eisenhower, DACI, and PMBOK.
-
What a Null Result Taught Us About AI Agent Evaluation
• UpdatedWe tested prompt repetition on 20 parallel AI agents. Ceiling effects dominated both experiments. The null result is a finding about evaluation design.
-
Open-Weight LLMs Reach the Structured Output Quality Ceiling
• UpdatedOpen-weight models now match closed-source on structured output at 95x lower cost. Pre-registered blind eval, 30 samples, zero quality delta.
-
How Does Context Engineering Drive Multi-Agent Reliability?
• UpdatedContext engineering turns AI context into infrastructure, preventing the accumulation, starvation, and leakage failures that undermine multi-agent reliability.
-
AI SDLC Governance: Three Layers for Engineering Leaders
• UpdatedA three-layer governance stack (comprehension, review gate, and observability) gives engineering leaders measurable control over AI-generated code volume.
-
AI Agents in Legacy Systems: ROI Without Modernization
• UpdatedLayer AI agents over legacy systems without modernization. 30-80% productivity gains in 3-6 months. Patterns that bypass technical debt.
-
AI-Augmented CI/CD: Shift Left Security With Bounded Risk
• UpdatedAI code review in CI/CD with prompt injection containment. Defensive patterns: review-context profiles, managed review gates, capability isolation.
-
How Much Tool Documentation Do AI Agents Actually Need?
• UpdatedTrimming MCP tool-definition prose cut cold-call input tokens 12-14% with no statistically detectable accuracy loss across two model tiers.
-
Why Your AI Agent Failed in Production
• UpdatedWhy your AI agent failed: missing decision provenance, not metrics. The 3 observability gaps traditional monitoring won't catch.
-
SRE for AI Agents: Error Budgets, Trust, and 90 Trials
• UpdatedCan an AI agent predict scope without hallucinating? We ran 90 trials. It added 1.7 phantom files per change. Error budgets and trust ladders are the gate.
Browse all tracks on the Topics index or filter by tag on the Tags page.