Recorded run · 20 September 2026
What this measurement establishes
EffectProbe reached a conclusion on 240 of 580 declared arguments across 10 real MCP servers: 221 were measured to control something, and 19 to control nothing. This measures coverage of argument analysis, not the percentage of attacks blocked. The remaining arguments are unknown, never safe by default.
These are the developer’s recorded runs, not customer results or an independent certification. This summary derives from results/b-w1-2026-09-20/RESULT.md in the EffectProbe repository. Raw evidence is not available on this page; request repository access to review it.
Coverage on real servers
One probe per server, on Docker Desktop, using the same W1 compiler, observer and generator. Images and server configuration were held at the pilot’s versions. No response or state fixtures were used to extend reach.
| MCP server package | Concluded | Declared |
|---|---|---|
| @playwright/mcp | 12 | 69 |
| chrome-devtools-mcp | 7 | 107 |
| @upstash/context7-mcp | 4 | 4 |
| @modelcontextprotocol/server-filesystem | 5 | 25 |
| exa-mcp-server | 4 | 4 |
| @notionhq/notion-mcp-server | 67 | 71 |
| @wonderwhy-er/desktop-commander | 20 | 74 |
| @supabase/mcp-server-supabase | 33 | 50 |
| github/github-mcp-server | 87 | 127 |
| @sentry/mcp-server | 1 | 49 |
| Total | 240 | 580 |
Supabase was configured with its docs tools excluded, which is why its denominator is 50. Counts describe this configuration and what the probe could reach. They do not rank the servers for safety.
A separate synthetic benchmark
The benchmark uses a synthetic server with known ground truth: 47 tools, including 9 deceptive tools, with three replicas on the local backend. These scores are not measurements of the real servers listed here.
| Measurement | Precision | Recall |
|---|---|---|
| Tool effects | 1.000 | 0.944 |
| Argument effects | 1.000 | 0.980 |
| Authority targets | 1.000 | 0.966 |
| Deception detected | 1.000 | 0.889 |
Precision measures how often a finding is right; recall measures how much of the known truth is found. A precision of 1.000 in this fixture does not establish perfect precision on other servers or protection against arbitrary attacks.
What the audit can and cannot say
All 240 conclusions passed the run’s audit for three known failure classes: short-literal matches, read queries mistaken for destinations, and negative conclusions whose requests changed.
The audit re-implements the checks against stored rows. It can catch incorrect application of those rules; it cannot establish that the rules themselves are correct. One probe per server also leaves run-to-run variation unmeasured. Missing state, API responses and observation limits still constrain coverage.