---
title: "Does WebMCP Pay Off for Content Sites?"
description: "Across 135 agent runs on a production blog, shrinking a four-tool WebMCP surface yielded little median-token benefit; relevance ordering mattered more."
author: "Hugues Clouâtre"
pubDatetime: 2026-10-04T21:05:56.000Z
modDatetime: 2026-10-05T17:00:00.000Z
url: https://clouatre.ca/posts/webmcp-beyond-checkout/
tags:
  - "agentic-ai"
  - "architecture"
  - "case-studies"
  - "mcp"
---

Stripe has published a controlled benchmark for WebMCP, the browser API that lets a web page hand AI agents callable tools. In 60 checkout tests across six models, agents using WebMCP consumed 42% fewer tokens, made 38% fewer tool calls, and finished 39% faster, while both arms completed every purchase ([Mecozzi & Kaliski, 2026](https://stripe.dev/blog/how-stripe-is-designing-checkout-for-ai-agents)).

Checkout is a narrow slice of what agents do. [HUMAN Security, 2026](https://www.humansecurity.com/learn/blog/state-of-agentic-traffic-august-2026-agentic-traffic-grows-27-reaches-new-high-as-codex-debuts-strongly/) found that in August 2026, product and search routes accounted for 80% of agentic activity, while checkout and payment stood at 2.1%. The published record holds no controlled measurement of progressive disclosure on a site that agents read instead of buy from.

This post runs that test on this blog's four read-only WebMCP tools, built in [Agent-Ready Websites: The 5-Layer Implementation Stack](/posts/agent-ready-website/). Three 45-run batches (baseline, first disclosure build, and rerun) put 135 runs through a fixed rubric, and 134 passed. The rerun held completion, and seven of nine token medians stayed within 4% of baseline; the experiment compares two configurations of an existing WebMCP implementation, not WebMCP against HTML or Markdown access. For a CTO funding agent readiness on a content site, response shape is the lever, and tool count barely registers.

*Table 1: The checkout benchmark against this blog's read-only eval.*

| Dimension | Stripe checkout | This blog, read-only |
|---|---|---|
| Workload | Purchase; each step drops prior tools | Read-only; tools stay valid |
| Tool surface | Disclosure follows checkout state | Four tools cut to two or three |
| Completion | 100% in both arms | 45/45 rerun vs 44/45 baseline |
| Tokens | 42% fewer | Within 4% in seven of nine cells |
| Tool calls | 38% fewer | Flat, one to six per run |
| Duration | 39% faster | Mixed; 6.5 s vs 7.6 s medians |

## Table of contents

## Why Does Stripe's Checkout Benchmark Matter Beyond Commerce?

Progressive disclosure, the design principle behind the benchmark, applies to any site, while the headline number comes from a workload most sites do not run. The tests ran on a fictional outdoor retailer's checkout, counted total tokens including cache, and kept the six models undisclosed.

The mechanism is the transferable part. Stripe built its implementation so that a page exposes only the tools and parameters relevant to its current state. The team also refused an "agent-only" code path that could drift from the human checkout or add engineering toil. Both choices suit a content site: disclosure maps naturally onto page types, and a shared code path keeps tools in sync with the HTML.

The workload may not transfer. A checkout is a multi-step state machine where each step invalidates the previous step's tools, so disclosure removes real clutter; a blog post page has no such sequence. Agent traffic still makes the test worthwhile. The latest monthly edition puts media at 36.2% of agentic traffic ([HUMAN Security, 2026](https://www.humansecurity.com/learn/blog/state-of-agentic-traffic-august-2026-agentic-traffic-grows-27-reaches-new-high-as-codex-debuts-strongly/)), so content sites carry a large share of agent visits, and non-checkout surfaces deserve the same scrutiny. Referrals from AI search engines grew 16x between 2024 and 2026 ([Deda, 2026](https://seranking.com/blog/ai-traffic-research-study/)).

## What Do Independent Studies Say About Tool Surfaces?

Independent studies agree that smaller tool surfaces cut cost at scale, and they also predict little gain when a catalog is already compact. The WindTunnel benchmark ran 21 configurations against 49 tasks on eight sites. WebMCP posted 3-47x lower median cost than the median of other methods, yet raw task-solve rate "does not separate WebMCP from the best screen-driving configuration" ([nekuda, 2026](https://github.com/nekuda-ai/WindTunnel)). One disclosure matters: nekuda sells WebMCP tooling, and its test sites ship purpose-built tools. No published study in this set measured disclosure on a read-only site with fewer than ten tools.

[Gan & Sun (2025)](https://arxiv.org/abs/2505.03275) found that retrieving a small tool subset from a large Model Context Protocol (MCP) pool cut prompt tokens by over 50% and lifted selection accuracy from 13.62% to 43.13%. Anthropic's tool search reached an 85% token reduction, but its guidance lists a "Small tool library (<10 tools)" and compact definitions as cases where it is "less beneficial" ([Anthropic, 2025](https://www.anthropic.com/engineering/advanced-tool-use)). This blog exposes four compact tools, so that guidance predicted small gains before any test ran. Field data is thinner still: our search found no published, measured WebMCP adoption number, and rendered crawls put registered tools near zero outside demos ([Spronta, 2026](https://www.spronta.com/blog/state-of-webmcp-july-2026/)).

*Table 2: Published evidence on WebMCP and tool-surface size, descending from controlled benchmarks to crawl-based field data.*

| Source (Year) | Evidence base | Key finding |
|---|---|---|
| Mecozzi & Kaliski (2026) | 60 checkout tests, six undisclosed models | -42% tokens, -38% calls, -39% time; 100% success in both arms |
| nekuda (2026) | 49 tasks, 8 sites, 21 configurations | 3-47x lower median cost; solve rate does not separate it from best screen-driving setup |
| Gan & Sun (2025) | Large MCP tool pools | Over 50% fewer prompt tokens; selection accuracy 13.62% to 43.13% |
| Anthropic (2025) | Tool search on Claude platform | 85% token cut; less beneficial below 10 tools |
| Spronta (2026) | Rendered crawls, multi-site | No published adoption number; registered tools round to zero outside demos |

None measured disclosure on a read-only site with fewer than ten tools.

## How Was Disclosure Tested on a Read-Only Site?

The test ran one fixed evaluation across three batches: a four-tool baseline, a first disclosure build, and a rerun after a relevance-ordering fix. Every batch crossed the same grid, five iterations per task-model cell, for 135 runs in total, and 134 passed the rubric. The harness, run JSONs, and raw data are archived on Zenodo ([Clouatre, 2026](https://doi.org/10.5281/zenodo.23148368)).

### The Four Read-Only Tools

The four tools search posts, return a post as Markdown, list related posts, and list posts by concept cluster. Each registers on `document.modelContext` and only reads data. Disclosure derives a page kind from the URL and registers the matching subset (Figure 1). The home page offers search and concept lookup; a post page offers its full text, related posts, and concept lookup; other pages, including the archive, offer search alone. The archive keeps its own page kind as future-proofing for listing surfaces, though it maps to the same single tool today.

```mermaid
%% Diagram of WebMCP progressive disclosure: the page URL resolves to a page kind, and each kind registers its tool subset, two on home, three on post pages, one on archive and other pages
graph TD
    U[Page URL] --> K{Page kind?}
    K --> H[Home: 2 tools]
    K --> P[Post: 3 tools]
    K --> A[Archive: 1 tool]
    K --> O[Other: 1 tool]
    H --> S[search_posts]
    H --> C[get_posts_by_concept]
    P --> M[get_post_markdown]
    P --> R[get_related_posts]
    P --> C
    A --> S
    O --> S

    classDef primary fill:#075985,stroke:#282728,stroke-width:2px,color:#fff
    classDef secondary fill:#e6e6e6,stroke:#ece9e9,stroke-width:2px,color:#282728
    classDef accent fill:#38bdf8,stroke:#075985,stroke-width:2px,color:#fff
    classDef muted fill:#fdfdfd,stroke:#ece9e9,stroke-width:1px,color:#282728
    linkStyle default stroke-width:3px

    class U,K accent
    class H,P,A,O primary
    class S,C,M,R secondary
```

*Figure 1: Page kind decides which read-only tools register. Only post pages expose the full-text tool.*

One mapping table holds the whole policy, so adding a page kind means editing a single file.

*Code Snippet 1: The page-kind to tool mapping that drives disclosure.*

```typescript file="src/utils/webmcp-page-kind.ts"
export type PageKind = "home" | "post" | "archive" | "other";

const ALLOWED_TOOLS: Record<PageKind, readonly string[]> = {
  home: ["search_posts", "get_posts_by_concept"],
  post: [
    "get_post_markdown", // [!code highlight]
    "get_related_posts",
    "get_posts_by_concept",
  ],
  archive: ["search_posts"],
  other: ["search_posts"],
};
```

### Eval Design

Discover asks the agent to search for a post and return its slug. Extract asks it to quote one sentence verbatim from a paginated post. Traverse asks it to follow a relationship edge, then a concept cluster, to a hub post.

The models were claude-haiku-4-5, glm-5.3-flash, and gpt-6-luna, driven by [goose](https://github.com/aaif-goose/goose) (Agentic AI Foundation) in headless Chromium. Each run used temperature 0.3, seed 42, and a 12-turn cap, and a deterministic rubric graded the transcript; the fixed seed makes repeat runs comparable, so cell-level differences track the configuration change rather than harness or prompt drift. Wall clock was recorded per run but embeds provider API latency, so cross-batch duration reads are directional. The design follows the fixed-seed, repetition-based pattern argued for in [What a Null Result Taught Us About AI Agent Evaluation](/posts/prompt-repetition-agent-evaluation/).

### A Confound That Favors the Rerun

The baseline started all three tasks on the home page with all four tools registered. The rerun started each task on its natural page with two or three tools. Starting on the task's own page could also help the rerun, so any token gain would overstate disclosure's effect. Even with the rerun spared those navigation tokens, seven of nine cells stayed within 4% of baseline, so the flat result is not an artifact of the head start. Five runs per cell, the same depth Stripe used per model, limits sensitivity to small shifts, but exact Mann-Whitney tests detected the extract/Haiku increase (+25.6% median, p = 0.032), well outside the ±4% band every other cell occupied. The tests found no composite-score difference in any cell; six token cells differed at alpha 0.05, five by less than 4%, exploratory with no correction for nine parallel tests ([Clouatre, 2026](https://doi.org/10.5281/zenodo.23148368)).

## What Did the Agent Runs Show?

The rerun held completion and left tokens flat almost everywhere.

### Completion and Tokens

The rerun completed 45 of 45 runs against 44 of 45 at baseline; the token medians appear in Table 3. A composite score, the fraction of rubric checks each run passed, matched the rubric verdict in 89 of 90 runs with no significant baseline-to-rerun difference in any cell. The single baseline failure came from claude-haiku-4-5 on extract, which did not quote the target sentence verbatim; that cell passed all five times in the rerun. Table 3 reports medians, while Figure 2 plots batch means with bootstrap CIs because the extract spread is skewed. Medians describe the typical passing run; means are more sensitive to unusually expensive runs, so extract/Luna and extract/GLM differ materially between the two summaries (+2.7% median vs +9.8% mean, and -10.0% median vs -0.3% mean, respectively). Median call counts matched across batches almost everywhere: one per discover run, two per traverse run, and five to six per extract run. Table 3 lists the medians over rubric-passing runs.

*Table 3: Median total tokens per run, baseline vs rerun, over rubric-passing runs (five runs per cell, four in baseline extract/Haiku); Change reads baseline to rerun, Passes read baseline / rerun.*

| Task | Model | Tokens (baseline → rerun) | Change | Passes |
|---|---|---|---|---|
| discover | claude-haiku-4-5 | 14,264 → 14,185 | -0.6%, flat | 5 / 5 |
| discover | glm-5.3-flash | 11,755 → 11,704 | -0.4%, flat | 5 / 5 |
| discover | gpt-6-luna | 11,139 → 11,259 | +1.1%, flat | 5 / 5 |
| extract | claude-haiku-4-5 | 64,676 → 81,244 | +25.6%, up | 4 / 5 |
| extract | glm-5.3-flash | 54,121 → 48,710 | -10.0%, down | 5 / 5 |
| extract | gpt-6-luna | 24,381 → 25,037 | +2.7%, flat | 5 / 5 |
| traverse | claude-haiku-4-5 | 22,268 → 22,178 | -0.4%, flat | 5 / 5 |
| traverse | glm-5.3-flash | 18,017 → 18,027 | +0.1%, flat | 5 / 5 |
| traverse | gpt-6-luna | 16,999 → 17,659 | +3.9%, flat | 5 / 5 |

### Why Extract Moved

Extract is the only task that moved, and it moved both ways. claude-haiku-4-5 paged further through the post, with one run reaching 113,806 tokens over eight calls, while glm-5.3-flash spent 10% less. Figure 2 plots every passing run, so the spread stays visible behind the medians.

**Tokens stayed within 4% of baseline in 7 of 9 cells**

| Series | D-Haiku | D-GLM | D-Luna | E-Haiku | E-GLM | E-Luna | T-Haiku | T-GLM | T-Luna |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Baseline mean | 14264 (95% CI 14264-14264) | 11745.2 (95% CI 11723.8-11756.8) | 11128.2 (95% CI 11081.4-11175) | 64682 (95% CI 64651-64713) | 48743.2 (95% CI 40362.2-56063) | 27494 (95% CI 24371.8-33726.6) | 22277.6 (95% CI 22268-22296.8) | 18017 (95% CI 18017-18017) | 16999 (95% CI 16999-16999) |
| Rerun mean | 14191 (95% CI 14182.6-14205.4) | 11703.4 (95% CI 11702.2-11704) | 11257.4 (95% CI 11254.2-11259) | 81640.2 (95% CI 68958.4-97735.4) | 48586.4 (95% CI 38025.2-59147.6) | 30186.2 (95% CI 25027-36599) | 22368 (95% CI 22170.8-22565.4) | 18027 (95% CI 18027-18027) | 16431.4 (95% CI 13969.6-17665.6) |

- Baseline observations 1/9 (n=5): 14264, 14264, 14264, 14264, 14264
- Baseline observations 2/9 (n=5): 11755, 11703, 11758, 11755, 11755
- Baseline observations 3/9 (n=5): 11162, 11139, 11206, 11067, 11067
- Baseline observations 4/9 (n=4): 64695, 64645, 64731, 64657
- Baseline observations 5/9 (n=5): 58892, 41771, 54121, 34727, 54205
- Baseline observations 6/9 (n=5): 39957, 24393, 24374, 24381, 24365
- Baseline observations 7/9 (n=5): 22316, 22268, 22268, 22268, 22268
- Baseline observations 8/9 (n=5): 18017, 18017, 18017, 18017, 18017
- Baseline observations 9/9 (n=5): 16999, 16999, 16999, 16999, 16999
- Rerun observations 1/9 (n=5): 14185, 14185, 14181, 14185, 14219
- Rerun observations 2/9 (n=5): 11704, 11704, 11704, 11704, 11701
- Rerun observations 3/9 (n=5): 11259, 11259, 11251, 11259, 11259
- Rerun observations 4/9 (n=5): 113806, 65882, 65892, 81244, 81377
- Rerun observations 5/9 (n=5): 35375, 68403, 35333, 48710, 55111
- Rerun observations 6/9 (n=5): 25037, 25018, 34779, 41066, 25031
- Rerun observations 7/9 (n=5): 22657, 22169, 22169, 22667, 22178
- Rerun observations 8/9 (n=5): 18027, 18027, 18027, 18027, 18027
- Rerun observations 9/9 (n=5): 11510, 17659, 17670, 17659, 17659

*Figure 2: Total tokens per run by task (D = discover, E = extract, T = traverse) and model; dots are individual passing runs, lines join batch means with bootstrap 95% CIs, and identical runs under the fixed seed collapse the CI to a point in zero-variance cells.*

The chart excludes one run: the failed baseline extract run from claude-haiku-4-5, which stopped at 80,292 tokens. Reading it: the significance test flags two traverse cells whose medians moved +3.9% and +0.1%, so detectability is not magnitude; on extract the Haiku increase is significant (p = 0.032) while Luna's is not, and no cell approaches the 42% cut Stripe reported. Even with the harness context favoring the rerun, the rerun changed behavior nowhere except extract, and there the direction was not consistent; these results show no broad median-token benefit from the revised configuration, and do not establish behavioral equivalence.

These tasks need one to six median calls, the full catalog was four compact tools, and no step invalidated another step's tools. Disclosure trimmed a list that was already short, so the agent's context barely changed.

## What Mattered More Than WebMCP Tool Count?

Payload caps and relevance ordering changed outcomes, and lean descriptions kept the catalog small. Removing one or two tools per page left tokens flat. A similar result surfaced in the server-side ablation behind [How Much Tool Documentation Do AI Agents Actually Need?](/posts/mcp-tool-docs/#what-mattered-more-than-documentation-richness), where the gap between model tiers outweighed documentation richness.

### Payload Caps

Caps bounded the worst case. Before pagination, the Markdown tool returned a whole post in one response: 6,764 tokens for the audited post, `ai-sdlc-governance-stack`, and 7,493 for `mcp-tool-docs`, the longest post at the time. After pagination, the audited post spans six pages that peaked at 1,415 tokens before the payload-cap change ([#1653](https://github.com/clouatre-labs/clouatre.ca/pull/1653)) and 1,422 after, 71% of a 2,000-token budget (Figure 3). The three discovery tools return 338 to 513 tokens against a 600-token budget.

**Pagination and caps held the worst case to 71% of budget**

| Series | Value |
| --- | --- |
| One response, unbounded | 6764 tokens |
| Largest page, after pagination | 1415 tokens |
| Largest page, after caps | 1422 tokens |
| Budget | 2000 tokens |

Audited post: ai-sdlc-governance-stack, 6 pages at both paged commits

*Figure 3: `get_post_markdown` worst case for the audited post, in three steps: the unbounded single response, the largest page right after pagination, and the largest page under the payload caps, against the 2,000-token budget.*

### Relevance Ordering

The most informative result came from ordering. The first run after the caps shipped showed a traverse regression: claude-haiku-4-5 and glm-5.3-flash each added a retry call in every run, and claude-haiku-4-5's median rose from 22,268 to 31,827 tokens. The cause was a windowing bug: related posts were cut to a page in file order, which pushed the highest-weight edge, the one the task needed, off the first page. Sorting edges by weight before windowing returned all 15 traverse runs to two calls. A cap without relevance ordering hides the answer the agent came for.

*Code Snippet 2: `get_related_posts` sorts edges by weight before windowing; the regression shipped without the sort.*

```typescript file="src/utils/webmcp-tools.ts"
const sorted = all.toSorted(
  (a, b) => (Number(b.weight) || 0) - (Number(a.weight) || 0), // [!code highlight]
);
const [offset, limit] = clampWindow(
  input,
  5,
  clampTokenBudget(undefined),
);
const page = sorted.slice(offset, offset + limit);
```

### Lean Descriptions

Lean descriptions followed the earlier post's trim, applied to the browser tools in the same change as the caps. With four compact tools, the site already sat in the small-library case where Anthropic calls tool search less beneficial.

## Where Does WebMCP Stop Reaching, and What Should Leaders Fund?

WebMCP reaches only agents running inside an open browser tab, so content sites should fund Markdown endpoints first and treat it as a bounded, read-only layer. Chrome describes the API as "primarily designed for local browser workflows with a human in the loop," and notes that clients "must visit a site directly to know if it has callable tools" ([Chrome for Developers, 2026](https://developer.chrome.com/docs/ai/webmcp)), and tools exist only while the page stays open. Training crawlers and API-based agents never see them.

Markdown reaches those agents at lower cost. Cloudflare's own announcement post took 16,180 tokens as HTML and 3,150 as Markdown, an 80% reduction for that single page ([Martinho & Allen, 2026](https://blog.cloudflare.com/markdown-for-agents/)). The earlier [What Does Leaner Tool Documentation Unlock Next?](/posts/mcp-tool-docs/#what-does-leaner-tool-documentation-unlock-next) section asked whether lean surfaces carry to the public web; for reading workloads, plain Markdown carries furthest.

The platform is still moving. The origin trial now spans Chrome 149 to 162, ending 2027-03-30 ([Chrome Platform Status, 2026](https://chromestatus.com/feature/5117755740913664)). The specification is a Draft Community Group Report dated 2026-10-02 and sits outside the W3C Standards Track ([W3C Web Machine Learning Community Group, 2026](https://webmachinelearning.github.io/webmcp/)). Sites that ship tools should set `readOnlyHint` on tools that change no state and `untrustedContentHint` on tools that return user-generated or external content ([Chrome for Developers, 2026](https://developer.chrome.com/docs/ai/webmcp/secure-tools)). Table 4 orders the spend.

*Table 4: Agent surfaces for a content site, ordered by when to fund them.*

| Surface | Audience | Runs where | Cost | When to fund |
|---|---|---|---|---|
| Markdown endpoints, llms.txt | Crawlers, answer engines, API agents | Server-side: cacheable at the CDN edge, no client execution | Low: static build output | First, on any content site |
| Structured graph data | Agents traversing related posts | Server-side: static JSON, edge-cached | Medium: curated edges and clusters | Once posts cross-link |
| Read-only WebMCP tools | In-browser agents | Ephemeral: client-side in an open tab, origin trial | Medium: tool code, caps, eval runs | After Markdown, with caps and an eval |
| Mutating WebMCP tools | In-browser agents that transact | Ephemeral: client-side, session-bound | High: state, confirmation, security review | Only when the site sells or books |

That keeps the agent-ready surface small, cheap to defend, and measurable. Row four raises the stakes. Payment protocols such as x402, which builds on the HTTP 402 status code, attach money to tool calls, so a regression surfaces as a charge and the harness becomes part of the correctness case.

## Takeaways

1. **Shrinking the tool surface showed no broad median-token benefit.** Seven of nine cells stayed within 4% of baseline, an order of magnitude below Stripe's reported 42%, despite the harness favoring the rerun; the comparison covers two configurations of the same implementation, not WebMCP versus Markdown.
2. **Response shape outweighed tool count.** A relevance-ordering bug added a retry to every traverse run for two models; sorting by edge weight removed it, and payload caps have held the largest Markdown page at 1,422 tokens.
3. **Fund Markdown before WebMCP.** WebMCP reaches only agents running inside an open browser tab, while Markdown endpoints serve every agent that fetches the page; on read-only surfaces the tools also inherit that pipeline, so the funding order is a dependency, not just a preference.

## References

- Anthropic, "Introducing advanced tool use on the Claude Developer Platform" (2025) -- https://www.anthropic.com/engineering/advanced-tool-use
- Chrome for Developers, "WebMCP" (2026) -- https://developer.chrome.com/docs/ai/webmcp
- Chrome for Developers, "WebMCP tool security" (2026) -- https://developer.chrome.com/docs/ai/webmcp/secure-tools
- Chrome Platform Status, "WebMCP" feature entry (2026) -- https://chromestatus.com/feature/5117755740913664
- Clouatre, H., "WebMCP Description Benchmarks: Behavioral Eval of In-Browser Tool Surfaces" (2026) -- https://doi.org/10.5281/zenodo.23148368
- Deda, Y., "AI traffic grew 16x from 2024 to 2026: Here are the top AI search engines driving visits to websites" (2026) -- https://seranking.com/blog/ai-traffic-research-study/
- Gan, T. & Sun, Q., "RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation" (2025) -- https://arxiv.org/abs/2505.03275
- HUMAN Security, "State of Agentic Traffic - August 2026: Agentic traffic grows 27%, reaches new high as Codex debuts strongly" (2026) -- https://www.humansecurity.com/learn/blog/state-of-agentic-traffic-august-2026-agentic-traffic-grows-27-reaches-new-high-as-codex-debuts-strongly/
- Martinho, C. & Allen, W., "Introducing Markdown for Agents" (2026) -- https://blog.cloudflare.com/markdown-for-agents/
- Mecozzi, C. & Kaliski, S., "How Stripe is designing Checkout for AI agents" (2026) -- https://stripe.dev/blog/how-stripe-is-designing-checkout-for-ai-agents
- nekuda, "WindTunnel: Benchmark WebMCP against other methods browser agents use to interact with websites" (2026) -- https://github.com/nekuda-ai/WindTunnel
- Spronta, "The State of WebMCP: July 2026" (2026) -- https://www.spronta.com/blog/state-of-webmcp-july-2026/
- W3C Web Machine Learning Community Group, "WebMCP" Draft Community Group Report (2026) -- https://webmachinelearning.github.io/webmcp/