Tag: case-studies
All the articles with the tag "case-studies".
-
Does WebMCP Pay Off for Content Sites?
• UpdatedAcross 135 agent runs on a production blog, shrinking a four-tool WebMCP surface yielded little median-token benefit; relevance ordering mattered more.
-
How Much Tool Documentation Do AI Agents Actually Need?
• UpdatedTrimming MCP tool-definition prose cut cold-call input tokens 12-14% with no statistically detectable accuracy loss across two model tiers.
-
Open-Weight LLMs Reach the Structured Output Quality Ceiling
• UpdatedOpen-weight models now match closed-source on structured output at 95x lower cost. Pre-registered blind eval, 30 samples, zero quality delta.
-
AI Adoption in Engineering: Breaking the 50% Plateau
• UpdatedPurpose-built AI tooling cuts per-task cost 21-68%. Three-cohort model and four-phase operating framework for engineering leaders past the 50% adoption plateau.
-
SRE for AI Agents: Error Budgets, Trust, and 90 Trials
• UpdatedCan an AI agent predict scope without hallucinating? We ran 90 trials. It added 1.7 phantom files per change. Error budgets and trust ladders are the gate.
-
What a Null Result Taught Us About AI Agent Evaluation
• UpdatedWe tested prompt repetition on 20 parallel AI agents. Ceiling effects dominated both experiments. The null result is a finding about evaluation design.