Skip to content

Agent Reliability

Moving autonomous LLM systems into production requires moving past anecdotal prompt engineering toward deterministic validation. This track covers empirical evaluation harnesses, context management patterns, and multi-agent coordination architectures that ensure predictable, testable behavior.

Start here

Reading order

  1. AI Delivery Decision Frameworks: Type 1, Type 2, DACI

    • Updated

    Misclassifying reversible decisions costs more than the decision itself. Four frameworks unblock AI delivery: Type 1/Type 2, Eisenhower, DACI, and PMBOK.

  2. What a Null Result Taught Us About AI Agent Evaluation

    • Updated

    We tested prompt repetition on 20 parallel AI agents. Ceiling effects dominated both experiments. The null result is a finding about evaluation design.

  3. Open-Weight LLMs Reach the Structured Output Quality Ceiling

    • Updated

    Open-weight models now match closed-source on structured output at 95x lower cost. Pre-registered blind eval, 30 samples, zero quality delta.

  4. How Does Context Engineering Drive Multi-Agent Reliability?

    • Updated

    Context engineering turns AI context into infrastructure, preventing the accumulation, starvation, and leakage failures that undermine multi-agent reliability.

  5. AI SDLC Governance: Three Layers for Engineering Leaders

    • Updated

    A three-layer governance stack (comprehension, review gate, and observability) gives engineering leaders measurable control over AI-generated code volume.

  6. AI Agents in Legacy Systems: ROI Without Modernization

    • Updated

    Layer AI agents over legacy systems without modernization. 30-80% productivity gains in 3-6 months. Patterns that bypass technical debt.

  7. AI-Augmented CI/CD: Shift Left Security With Bounded Risk

    • Updated

    AI code review in CI/CD with prompt injection containment. Defensive patterns: review-context profiles, managed review gates, capability isolation.

  8. How Much Tool Documentation Do AI Agents Actually Need?

    • Updated

    Trimming MCP tool-definition prose cut cold-call input tokens 12-14% with no statistically detectable accuracy loss across two model tiers.

  9. Why Your AI Agent Failed in Production

    • Updated

    Why your AI agent failed: missing decision provenance, not metrics. The 3 observability gaps traditional monitoring won't catch.

  10. SRE for AI Agents: Error Budgets, Trust, and 90 Trials

    • Updated

    Can an AI agent predict scope without hallucinating? We ran 90 trials. It added 1.7 phantom files per change. Error budgets and trust ladders are the gate.

Browse all tracks on the Topics index or filter by tag on the Tags page.