Recorded run · 20 September 2026

What this measurement establishes

EffectProbe reached a conclusion on 240 of 580 declared arguments across 10 real MCP servers: 221 were measured to control something, and 19 to control nothing. This measures coverage of argument analysis, not the percentage of attacks blocked. The remaining arguments are unknown, never safe by default.

These are the developer’s recorded runs, not customer results or an independent certification. This summary derives from results/b-w1-2026-09-20/RESULT.md in the EffectProbe repository. Raw evidence is not available on this page; request repository access to review it.

Coverage on real servers

One probe per server, on Docker Desktop, using the same W1 compiler, observer and generator. Images and server configuration were held at the pilot’s versions. No response or state fixtures were used to extend reach.

Arguments that reached a conclusion, out of all declared arguments in each configured server.
MCP server packageConcludedDeclared
@playwright/mcp1269
chrome-devtools-mcp7107
@upstash/context7-mcp44
@modelcontextprotocol/server-filesystem525
exa-mcp-server44
@notionhq/notion-mcp-server6771
@wonderwhy-er/desktop-commander2074
@supabase/mcp-server-supabase3350
github/github-mcp-server87127
@sentry/mcp-server149
Total240580

Supabase was configured with its docs tools excluded, which is why its denominator is 50. Counts describe this configuration and what the probe could reach. They do not rank the servers for safety.

A separate synthetic benchmark

The benchmark uses a synthetic server with known ground truth: 47 tools, including 9 deceptive tools, with three replicas on the local backend. These scores are not measurements of the real servers listed here.

Benchmark scores recorded with the same build.
MeasurementPrecisionRecall
Tool effects1.0000.944
Argument effects1.0000.980
Authority targets1.0000.966
Deception detected1.0000.889

Precision measures how often a finding is right; recall measures how much of the known truth is found. A precision of 1.000 in this fixture does not establish perfect precision on other servers or protection against arbitrary attacks.

What the audit can and cannot say

All 240 conclusions passed the run’s audit for three known failure classes: short-literal matches, read queries mistaken for destinations, and negative conclusions whose requests changed.

The audit re-implements the checks against stored rows. It can catch incorrect application of those rules; it cannot establish that the rules themselves are correct. One probe per server also leaves run-to-run variation unmeasured. Missing state, API responses and observation limits still constrain coverage.