Testing MCP servers in 2026: 31 tools compared and the gaps between them
A Model Context Protocol (MCP) server is used by a language model, not by a person, so testing it means more than calling each tool by hand. Does it follow the specification? Can an agent pick the right tool from its descriptions? Will it pass review for the Claude connector directory or the ChatGPT app store, and work inside Cursor, VS Code and the command-line agents? How does it behave under load? On 1 October 2026 I reviewed about eighty tools that inspect, test, evaluate, audit, secure or load test an MCP server, open source and paid, and compared thirty-one of them capability by capability. This is what the landscape looks like, and where it is thin.
In brief
- Every capability exists somewhere, but no open, local tool combines them. Inspection, evals, audits, host checks and load testing are spread over about a dozen products.
- MCPJam is the closest thing to a complete workbench. It ships saved eval suites, AI-generated test cases that the user keeps or discards, an LLM judge, client emulation, directory-readiness checks and CI gates. Its AI features depend on its hosted backend and credits.
- The official MCP Inspector is no longer a slow target. Version 2 has been generally available since July 2026, ships about weekly, and has saved collections, assertions, CI flows and a conformance runner on its roadmap for November 2026 to January 2027. It has no LLM or eval work planned.
- Three areas are saturated: manual inspection, security scanning, and the basic “YAML evals plus LLM judge plus GitHub Action”. Protocol conformance is owned by the official suite.
- The gaps are real: a fully local, bring-your-own-model evaluation loop; an open, versioned rule pack for many hosts backed by live behavioural checks; fix suggestions proven by re-scoring; load testing tied to functional scenarios; one combined report; and anything .NET-native.
- The protocol changed fundamentally two months before this review. Specification 2026-07-28 is a stateless redesign, and many small tools are stuck on older versions.
How the ground moved in 2026
The protocol
The current MCP specification is 2026-07-28. Per its changelog it removes the initialize handshake, protocol-level sessions (Mcp-Session-Id) and ping, and makes a new server/discover call mandatory. Elicitation, sampling and roots become multi-round-trip results instead of server-initiated requests. Dynamic Client Registration is deprecated in favour of Client ID Metadata Documents, and Roots, Sampling, Logging and HTTP+SSE are in a deprecation window. Tasks moved to an extension.
The consequence for anyone testing a server: a tester must speak both the legacy era (2025-11-25 and earlier) and the modern one. Hosts lag the specification, so the two will coexist for a long time. The specification has had five revisions in about twenty months, the latest of them breaking.
Official tooling
| Official asset | State today | Planned |
|---|---|---|
| MCP Inspector v2 | Web, CLI and terminal UI; both protocol eras; CI-friendly CLI | Record and replay, diff, latency view (October to November 2026); saved calls, assertions, CI, conformance runner, OTLP export (November 2026 to January 2027); plugins, multi-server (January to February 2027), per its roadmap |
| Conformance suite | Tests any server URL against per-revision requirement sets; GitHub Action; baselines | New specification proposals must ship conformance scenarios (SEP-2484) |
| MCP Registry | Preview; validates namespace and manifest only | Delegates scanning and curation to others |
Nothing on the official roadmap covers evals, LLM-as-judge, tool-design quality, host-guideline validation or load testing.
Adoption, money and churn
The MCP blog reported “close to half-a-billion downloads a month” across Tier 1 SDKs in July 2026, and Glama’s registry listed about 94,600 servers on the day of this review. The money followed: Alpic raised a $6M pre-seed in September 2025, Manufact (the company behind mcp-use) a $6.3M seed in February 2026, and MCPJam is an Open Core Ventures company. Anthropic acquired Stainless in May 2026, Arcade acquired Smithery in August 2026, and Promptfoo’s README says it is now part of OpenAI.
Churn is the other signal. Many small tools went quiet within three to six months of launch (lastmile’s mcp-eval, mclenhard/mcp-evals, steviec’s tester, MCPSpec, mcp-shield, destilabs/mcp-doctor), and a few vanished (Bellwether, Janix-ai’s validator).
.NET is a gap of its own. The C# SDK is Tier 1 and reached version 2.0 in July 2026, yet no .NET-native inspector, test or load tool was found; Aspire users wrap the Node-based official Inspector.
Six jobs, and who does each
A tool in this space does one or more of six jobs: inspect a server by hand, run functional tests and evals against it, audit its design and protocol conformance, check it against the rules of the hosts that will run it, scan it for security problems, and load test it. Observability sits beside them on the production side.
Inspectors and interactive clients
The official MCP Inspector is open source (MIT, moving to Apache-2.0), has about 10.7k stars, and reached version 2 on 28 July 2026, with 2.9.0 on npm on 30 September. It offers a web UI, a CLI and a terminal UI on one shared core, a Docker image, and needs Node 22.19 or later. It speaks stdio, SSE and Streamable HTTP in both protocol eras, and covers tools, resources, prompts, elicitation, sampling with a stubbed response, roots, logging, completions, pagination, Tasks, Skills and MCP Apps. It debugs OAuth, including Client ID Metadata Documents and enterprise-managed authorization, has protocol, network and console monitors, and its CLI gives CI a JSON output, exit codes and a --strict schema portability lint. It is the reference client that tracks the specification first, it is free, and it ships fast. It has no LLM, chat, evals, AI test generation or judge, and none planned; the roadmap says “the Inspector is not a host”. Saved tests, assertions and conformance are planned but not shipped; there are no host rules, no load testing and no reports. It is Node-only, and version 1 had a critical remote code execution flaw (CVE-2025-49596).
MCPJam Inspector (docs) is open core: Apache-2.0 with the eval server code under a separate licence, about 2.2k stars, releases almost daily. It does manual debugging of tools, resources, prompts and elicitation with JSON-RPC logs and a guided OAuth debugger; an LLM playground with up to three models side by side and many providers including Ollama and custom endpoints; a ChatGPT Apps and MCP Apps widget emulator; eval suites with expected tool calls, argument matching, negative cases, deterministic assertions and token budgets; AI test generation, where a Generate button drafts positive and negative cases from the tool catalog and each draft is saved or discarded (docs); an LLM judge with a threshold (default 0.7) and up to ten rubric checks, plus a multi-model matrix with latency, tokens and cost; baselines, quality gates, schedules, pull request comments and JUnit, HTML and JSON reports; static host compatibility verdicts for Claude, ChatGPT, Cursor, Copilot and Codex Desktop, and directory-readiness runs against the Anthropic and OpenAI directory requirements; and protocol, OAuth, MCP Apps and Tasks conformance in the CLI and SDK. Pricing: Free with 200 credits a day, Pro at $24 a month on annual billing, Team at $199 a month, Enterprise custom; evals cost 3 credits per prompt (a credit is $0.01), and runs on your own model key are not billed. It has the broadest feature set in the category, a funded team and a real CI story, and it already ships the generate, review, store and regress loop. But the judge, rubric checks and generation need the hosted backend: the local CLI runs deterministic assertions with your own key but not the judge, and hosted judge traffic goes through OpenRouter. Host compatibility is labelled “prototype · static checks” and omits transport and auth blockers. There is no load or stress testing, only per-run latency, and the hosted app has no stdio. Early versions had an unauthenticated RCE (CVE-2026-23744).
mcp-use Inspector and the Manufact cloud: the inspector is MIT, the cloud is paid, and the framework repository as a whole has about 10.6k stars. It runs hosted, through npx, in Docker, or auto-mounted in the mcp-use dev server, and can be embedded as a React client. Streamable HTTP only; multi-server; tools, resources, prompts and elicitation; bring-your-own-key chat with OpenAI, Anthropic, Google and Ollama; saved tool calls with replay; MCP-UI and OpenAI Apps SDK widget debugging. The cloud (pricing) adds test suites and a publishing checklist on Free, end-to-end checks for ChatGPT and Claude and a submission pack from Hobby ($25 a month), Startup at $250, Enterprise from $1,000. The UI is polished and the community large, but there is no stdio, the open-source inspector has no assertions, AI generation, judge or CI mode (that value sits in the paid cloud), it is tied to the mcp-use framework funnel, and telemetry is on by default.
Postman (proprietary; Free, Solo $9, Team $19 per user a month, Enterprise custom) has an MCP request type over stdio and Streamable HTTP with tools, resources and prompts, an elicitation form, mocked sampling, an MCP Apps preview, OAuth 2.1 with a debugger, and saving to collections shared in workspaces. It is familiar to every API team and strong on collaboration and auth. It has no MCP evals, AI test generation or judge, whether MCP requests work with test scripts, the collection runner or the CLI is not documented, and it is closed source with an account required.
Apidog (proprietary freemium) covers stdio and Streamable HTTP, tools, prompts and resources, many auth types, environment variables, a raw JSON-RPC view and saving to a project. Rich auth options in a familiar workflow; no sampling or elicitation documented, no MCP test scenarios, assertions or CI.
Insomnia (Kong; Apache-2.0 core; MCP client since version 12 in November 2025) covers HTTP transport, tools, prompts and resources, automatic OAuth discovery, roots, events, notifications, elicitation, and sampling through its own LLM settings. Open-source core with Git-based project storage; stdio not mentioned; no MCP test scripts or evals.
Hosted playgrounds:
| Tool | What it does | Limits |
|---|---|---|
| Glama MCP Inspector | Free browser inspector: tools, resources, prompts, elicitation, sampling, roots, tasks, OAuth 2.1; shareable state in the URL | Remote servers only (stdio through a tunnel); no chat, no tests |
| MCP Playground Online | Free inspector with JSON-RPC logs; “Agent Studio” across up to four models; an “AI Readiness Review” that generates user questions, rates tool descriptions and proposes rewrites; security scanner | Closed source; remote servers only; no saved suites or CI |
| Cloudflare AI Playground | Free chat client for remote servers with OAuth and a debug log | No manual tool forms, saving, tests or CI |
| MCPize, Smithery playgrounds | Basic inspector and chat | Not independently verified |
Command-line clients:
| Tool | Licence, stars | What it is | Gap |
|---|---|---|---|
| f/mcptools | MIT, ~1.6k | list and call, shell, mock server, proxy | Stale since May 2025; no auth |
| apify/mcpc | Apache-2.0, ~945 | Persistent sessions, OAuth 2.1, JSON mode | No LLM, no tests |
| IBM/mcp-cli | Apache-2.0, ~2.0k | LLM chat host with many providers | A host, not a tester |
| philschmid/mcp-cli | MIT, ~1.2k | Single binary for shell chaining | No tests |
| wong2/mcp-cli | GPL-3.0, ~440 | Interactive inspector with OAuth | No tests |
| reloaderoo | MIT, ~125 | CLI inspection plus hot-reload proxy | Slow-moving |
| mcp-probe | MIT, ~135 | Rust terminal debugger with a validation suite | Early (v0.1.0) |
Also in this space: FastMCP’s fastmcp dev inspector simply launches the official Inspector; VS Code and Cursor offer MCP logs and debug attach but are hosts, not test tools; Hoppscotch and Bruno have no MCP client. The name “MCP Workbench” is already used by Orkes and at least two other projects.
Functional tests and AI evaluation
MCPJam’s evals, described above, are the benchmark for this category.
testmcpy (Preset; Apache-2.0; 4 stars; active, with PyPI releases through September 2026) is a Python CLI with an optional React UI, SQLite storage, a GitHub Action and tests in YAML. It has more than forty evaluators (tool called, parameters, sequence, call count, answer contents, time and token limits), AI test generation from the CLI and the UI with a generation history, eleven model providers including Ollama, model comparison and a leaderboard, baselines, mutation and metamorphic testing, flaky detection, a load_test step, a 0 to 100 “LLM usability” score, a SARIF scan, a wrapper for the official conformance suite and JUnit output. It is the closest open-source match to a full workbench, and it runs locally. It has almost no adoption, the repository still describes itself as built for testing Superset, generated tests are written as YAML with no accept or reject queue, and LLM-judge and stdio support could not be confirmed.
MCPSpec (MIT; 7 stars; last commit March 2026, which looks stalled) is a TypeScript CLI with a React dashboard and YAML test collections: ten assertion types, variable extraction for chained calls, record and replay with diff, mock-server generation, an eight-rule security audit, a 0 to 100 “MCP Score”, a latency benchmark with P95 and P99, baselines and several reporters. Broad deterministic coverage in a single tool; nothing LLM-driven, by design; one burst of development and no traction.
lastmile-ai/mcp-eval (Apache-2.0; 25 stars; dormant since September 2025) is a Python library and CLI with tests as decorators, pytest or datasets; assertions on content, tool calls, sequence, performance and path efficiency; an LLM judge with a rubric and a minimum score; OpenTelemetry-based latency, token and cost metrics; mcp-eval generate, which writes test files from the tool list; HTML, JSON and Markdown reports; and a GitHub Action. Clean code-first design with a judge and reports, but no UI, no review step for generated tests and no run history. Users report that generated tests are generic and do not exercise the server (#36, #40), and it is unmaintained.
gleanwork/mcp-server-tester (MIT; 19 stars; active, backed by Glean) is a Playwright fixture with matchers for schema, snapshot, response size, tool calls and a judge; an “LLM host mode” with iterations and pass rates, including Claude Code and Codex as hosts; an HTML reporter with a pass-rate trend; and a non-LLM wizard that suggests expectations to accept. Solid engineering that compares servers and hosts side by side. It requires Playwright and TypeScript knowledge, ignores resources, prompts and notifications, and delegates AI authoring to coding-agent skills.
Smaller eval tools:
| Tool | Licence, stars, status | What it does | Main limits |
|---|---|---|---|
| steviec/mcp-server-tester | MIT, 35, dormant since Sept 2025 | YAML direct-call tests and evals with required tools and an LLM judge | Anthropic only; no reports or UI |
| mclenhard/mcp-evals | MIT, 132, unmaintained since mid-2025 | LLM grades answers 1 to 5 on five criteria; GitHub Action | No tool-call assertions; no model matrix |
| mcp-use/eval-action | No licence file, 3 | YAML matrix of case by model; rubric score plus required tools | OpenRouter only; the judge sees only the final answer |
| alpic-ai/mcp-eval | MIT, 21 | YAML expected tool calls with profiles for Claude, ChatGPT, Le Chat | Public HTTP servers only; no judge |
| Arcade evals | MIT, ~1k (parent repo), active | Python suites with weighted argument critics, multi-run statistics, capture mode | Scores tool selection and arguments without executing the tool |
| mcp-jest | MIT, 18 | Snapshot tests, auto-discovered deterministic tests, compliance score | No LLM features |
| mcp-recorder | MIT, 9 | Record and replay cassettes, pytest plugin | Stalled since March 2026 |
| MCP Observatory | MIT core, 140, active | Schema drift, record, replay and verify, security scan, health score, SARIF | No benchmarking; paid tiers |
| r-huijts/mcp-server-tester | MIT, 10, stale | Claude generates and runs N tests per tool | No review step; work in progress |
| mcpevals.ai | MIT, 1 | Paste a URL; auto-generated arguments; compatibility and response times | A smoke test, not a suite |
| Specmatic MCP Auto Test | Licensing not verified | Generates positive and negative cases deterministically from schemas | Streamable HTTP only |
General evaluation platforms:
| Tool | MCP relevance | Limits |
|---|---|---|
| Promptfoo (MIT, ~24k stars, now part of OpenAI) | MCP provider for direct tool calls with assertions; red-team plugin for MCP attack classes | No functional test generation for MCP |
| DeepEval (Apache-2.0, ~17.5k stars) | Three LLM-scored MCP metrics | Scores an MCP-using app from recorded interactions; does not connect to a server |
| Braintrust, LangSmith, Arize Phoenix, Inspect AI | Tracing or generic scorers | No MCP server test features confirmed |
Framework helpers exist too: FastMCP, the TypeScript SDK and the Python SDK document in-memory client testing, and the C# SDK offers an in-memory transport sample. Benchmarks such as MCP-Universe, MCP-Bench, MCPMark and LiveMCPBench rank language models across servers; they do not test a developer’s own server. Two pieces of research are worth knowing: “Tool descriptions are smelly” found at least one defect in 97.1% of 856 tool descriptions, and Salesforce’s MCPEval verifies generated tasks by executing them, which no developer-facing tool does yet.
Audit, lint and protocol conformance
The official conformance suite (MIT; about 120 stars; very active, pre-1.0): npx @modelcontextprotocol/conformance server --url <url> tests a server, and a client mode tests clients. It carries per-revision requirement sets for 2025-11-25 and 2026-07-28, validates the wire schema of every message, and has an expected-failures baseline file for CI, a GitHub Action and SDK tier scoring. It is authoritative and tied to SDK governance, so it will track the specification automatically. It is aimed at SDK maintainers: raw pass or fail, no grade, no fix advice, server testing documented for HTTP only, and nothing about tool design, usability by an agent, security or host rules.
YawLabs mcp-compliance (MIT; 2 stars; very active) runs 88 tests for 2025-11-25 and 103 for 2026-07-28 across transport, lifecycle, tools, resources, prompts, errors, schema and security, over HTTP and stdio, with an A to F grade, JSON, SARIF, a badge and a strict CI exit code. It is the most server-author-friendly conformance tool: graded, stdio-capable, both eras. Its test catalogue is unofficial and from one small vendor, adoption is negligible, and it has no design or host checks.
Glama’s Tool Definition Quality Score (TDQS; the CLI is Apache-2.0, v0.1.0, September 2026; about 30 stars) makes one LLM call per tool and scores six dimensions from 1 to 5: purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness and completeness, plus a server-level coherence score for disambiguation, naming consistency, tool count and completeness. tdqs lint is deterministic and needs no key; tdqs score uses any OpenAI-compatible endpoint or the hosted service, with a CI gate through --fail-under. It runs continuously across Glama’s registry (15,036 servers scored as of June 2026, per the repository), so maintainers are already graded by it, and its rubric and prompts are published and research-grounded. It scores the definition only, never behaviour; it has no protocol, security or host checks; it gives justifications rather than rewritten text; and the scores depend on the rubric-plus-model pair.
Deterministic design linters:
| Tool | Licence, stars, status | What it checks | Limits |
|---|---|---|---|
| mcp-surface-lint | MIT, 0, active | 19 rules: description hygiene, tool and token budgets, naming, annotations, loose schemas, overlap clusters, CRUD mirrors, missing list limits; category scores | Static only; no fixes; no adoption |
| mcpconform | MIT, 1 | 56 cited rules for tools, registry manifest and client configs, plus data-driven profiles for Anthropic, OpenAI and Gemini API constraints; SARIF | Profiles cover LLM API limits, not host or directory rules |
| mcp-server-lint | MIT, 1, hyperactive | Static source analysis (Python, TypeScript, Go): descriptions, typing, error handling, annotation-versus-code mismatch, unsafe calls, secrets | Never runs the server; brittle parsers |
| destilabs/mcp-doctor | MIT, 16, stale since Nov 2025 | Rules based on Anthropic’s “Writing tools for agents”; calls tools and measures response tokens and pagination | Heuristic; no score or CI gate |
| MCProbe | MIT engine, 6 | 12 schema rules plus a fuzz that calls tools with bad inputs and classifies the error handling; A to F grade | Small; the hosted full report is paid ($9.90) |
| Roee-Tsur/mcp-spec-check | 16 | Readiness probe for the 2026-07-28 specification; its README claims 1 of 7,850 remote registry servers passed all three required checks | Remote HTTP only; the claim is not verified |
| zuchka/mcp-doctor, mcp-lint, MCP Tool Auditor, agent-friend | 0 to 4 stars each | The same dozen rules: missing or short descriptions, vague names, no pagination, tool count, token cost | Hobby projects; mechanical fixes at best |
The pattern: ten or more tools implement the same basic rules, and none has traction. No tool was found that uses an LLM to write concrete improvements to names, descriptions, schemas and error messages and then proves the improvement by re-scoring, and none reviews error messages or response payloads with an LLM.
Host and platform guideline validation
Each host that runs MCP servers publishes its own rules: the Claude connector directory wants title and a readOnlyHint or destructiveHint on every tool and tool names of 64 characters or fewer; ChatGPT apps reject tools without readOnlyHint, destructiveHint and openWorldHint; Codex has a 10-second startup timeout; VS Code allows at most 128 tools per chat request, Windsurf 100 in total; Kiro and Gemini CLI cap and truncate names at 64 and 63 characters. The rules conflict, roughly half can be checked statically, a quarter need live calls, and most of what reviewers actually reject needs LLM judgment.
The tools that validate against them: Alpic Beacon (May 2026; closed; part of Alpic’s platform) is a pre-submission audit for the ChatGPT store and the Claude connectors directory covering protocol plumbing, tool quality and annotations, widget rendering, CSP and LLM-driven edge-case inputs, through a dashboard or alpic audit; it is built from real submission experience and covers live behaviour, but it is two stores only, not open source, and its exact cost and metering are not verified. Manufact’s publishing checks offer a checklist on Free and end-to-end checks in simulated ChatGPT and Claude with a submission pack from $25 a month, cloud only and focused on the two app stores. MCPJam’s host compatibility and directory readiness, described above, give free deterministic grading and many client profiles, but the compatibility is a self-described prototype, the readiness runs are hosted and the rule list is not published. The vendors’ own aids are OpenAI’s chatgpt-app-submission skill, instructions a coding agent follows rather than a deterministic validator; Anthropic’s automatic scan of submissions, which is not available to developers, and the manual checklist in its mcp-server-dev plugin; mcpb validate, which checks a desktop-extension manifest against its schema only; and sunpeak (MIT, about 217 stars), which simulates the ChatGPT and Claude runtimes for MCP Apps while developers write the annotation assertions by hand.
Not found: an open, rule-pack-driven validator that covers Claude, ChatGPT and the IDE and CLI hosts (Cursor, VS Code, Codex, Gemini CLI, Kiro, Windsurf) together.
Security scanners
This is the most crowded segment and the least promising place to compete.
| Tool | Licence, stars | What it does | Notes |
|---|---|---|---|
| Snyk agent-scan (formerly Invariant mcp-scan) | Apache-2.0 CLI, ~3.1k | Discovers agent configs, scans tool descriptions for injection, poisoning, shadowing, toxic flows | De facto standard; needs a Snyk token and sends data to Snyk’s API |
| Cisco mcp-scanner | Apache-2.0, ~1.1k | YARA rules, LLM judge with your key, source scanning, dependency audit; a “readiness analyzer” with 20 heuristic rules and a 0 to 100 score | Broadest engine set; has an offline mode |
| Tencent AI-Infra-Guard | ~6.4k | Platform with agent-driven MCP scanning across 14 categories | Large, general AI security platform |
| mcp-shield | ~550 | Tool poisoning and exfiltration checks | Abandoned since April 2025 |
| Proximity, AgentSeal, Ramparts, APIsec mcp-audit | 100 to 400 each | Pattern plus LLM scanning, toxic-flow analysis, config inventory | Overlapping rule sets |
| Enkrypt AI, Akto, Backslash, Pillar | Commercial | Hosted scans, gateways, runtime defence | Pricing mostly undisclosed |
Registries also publish scores: Glama (TDQS plus sandbox monitoring), AgentSeal (a 0 to 100 trust score) and MCP Trust Checker. The OWASP MCP Top 10 is in beta and is the common mapping target.
Performance and load testing
Most search results for “k6 MCP”, “JMeter MCP” or “Gatling MCP” are MCP servers that let an AI agent drive a load tool. They do not load test MCP servers. The tools below do.
MCP Drill (MIT; 7 stars; v0.1.3) is a Go control plane with distributed workers and a React UI: a run wizard, a live dashboard, throughput, P50, P95 and P99 latency, error rate, a per-tool breakdown, connection stability and optional server CPU and memory. It has stages (preflight, baseline, ramp, soak, spike), session modes (reuse, per_request, pool, churn), a weighted operation mix, side-by-side run comparison, a bundled mock server, and it supports the 2026-07-28 protocol and legacy versions. It is the only open tool with a UI, per-tool percentiles and run comparison, and it already speaks the stateless protocol. It is Streamable HTTP only, has no OAuth flows, keeps runs in memory so there are no durable baselines, and has a tiny community.
mcp-stress (MIT; 0 stars; npm 0.2.1, February 2026) is a Deno and TypeScript CLI over stdio, Streamable HTTP and legacy SSE, with profiles (tool flood, find-ceiling, ping flood), load shapes (ramp, step, spike, sawtooth), a live browser dashboard, an HTML report, and CI assertions such as p99 < 500ms with baseline deltas. One-line install, stdio support and regression assertions; single process, no OAuth, and essentially unused.
Other load options:
| Tool | Licence, maturity | What it does | Limits |
|---|---|---|---|
| reaatech/mcp-load-test | MIT, 0 stars | CLI and library; per-tool P50 to P99, breaking-point detection, A to F grades, compare; claims OAuth client credentials | No UI; untested by third parties |
| grafana/xk6-mcp | AGPL-3.0, 21 stars | k6 extension for tools, resources, prompts | Experimental, “not officially supported”; metrics tagged by method only |
| infobip/xk6-infobip-mcp | MIT, 5 stars, active | k6 extension with per-tool tags; stateful and stateless servers | callTool only; no stdio |
| BlazeMeter jmeter-mcp-plugin | Apache-2.0, v0.1.0 | JMeter sampler for MCP operations | Shares one MCP client across threads |
| Tricentis NeoLoad 2026.2 | Commercial | Native MCP connect, list and call actions; one client per virtual user | Enterprise pricing; tools only |
| Locust recipe on Azure Load Testing | Sample code | Hand-rolled Streamable HTTP client | A recipe, not a tool |
MCPJam and lastmile’s mcp-eval report latency per run, but neither does concurrency or throughput. MCPSpec and testmcpy have small benchmark features with no adoption.
Load testing MCP is harder than load testing plain HTTP, for five reasons:
- Two protocol eras: stateful sessions before 2026-07-28, stateless after.
- One MCP call is several HTTP requests, and responses may be JSON or an SSE stream.
- Tool failures hide behind HTTP 200 as
isErrorresults. - stdio is one client per process, so concurrency means measuring spawn cost and memory.
- OAuth tokens expire mid-run, tool calls have side effects, and they hit downstream rate limits.
Observability
| Product | What it does | Model |
|---|---|---|
| AgentCat (formerly MCPcat) | SDK wrapper with session replay and agent intent capture | MIT SDK; hosted from free to $160 a month and up |
| Sentry MCP monitoring | Spans for tools, resources, prompts and transports | Paid SaaS |
| Datadog, New Relic, Grafana Cloud | MCP client or server tracing and dashboards | Paid |
| Shinzo, Moesif, Agnost AI | OpenTelemetry-based analytics for MCP servers | Mixed |
| Speakeasy Gram, Alpic, Lunar MCPX | Hosting or gateway platforms with tool-call logs and a playground | Open core or paid |
| OpenTelemetry MCP conventions | Standard span and metric names; the official Python SDK now traces automatically | Open standard |
None of these offers pre-production testing or load testing. OpenTelemetry is the neutral integration point for any deep-tracing feature.
The comparison
Counting the thirty-one tools that offer each capability outright shows where the market is: almost everything runs in CI and locally, a third of the tools inspect or store tests, and host rules are covered by one.
Column meanings: Inspect is manual calls to tools, resources and prompts; Chat an LLM playground against the server; Saved tests stored deterministic tests or assertions; AI gen an LLM drafting test cases; Judge an LLM scoring test runs; Conform protocol and specification compliance; Design a tool-definition and best-practice audit; Host checks against specific AI platform rules; Sec security scanning; Load concurrency, throughput or stress testing; CI headless runs with gates; Local every feature works without the vendor’s cloud. Yes means documented, Part partial or limited, a dash not offered, and a question mark not verified.
| Tool | Model | Inspect | Chat | Saved tests | AI gen | Judge | Conform | Design | Host | Sec | Load | CI | Local |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Official MCP Inspector | Open source | Yes | – | Planned | – | – | Planned | Part | – | – | – | Part | Yes |
| MCPJam | Open core | Yes | Yes | Yes | Yes | Yes | Yes | Part | Part | – | – | Yes | Part |
| mcp-use Inspector + Manufact | MIT + paid cloud | Yes | Yes | Cloud | ? | Cloud | – | Cloud | Cloud | – | – | Cloud | Part |
| Postman | Paid | Yes | – | Part | – | – | – | – | – | – | ? | ? | Yes |
| Apidog | Paid | Yes | – | Part | – | – | – | – | – | – | – | – | Yes |
| Insomnia | Open core | Yes | – | – | – | – | – | – | – | – | – | – | Yes |
| Glama Inspector | Free hosted | Yes | – | – | – | – | – | – | – | – | – | – | No |
| MCP Playground Online | Closed hosted | Yes | Yes | – | Part | – | – | Part | – | Part | – | – | No |
| testmcpy | Open source | Part | Yes | Yes | Yes | ? | Yes | Part | – | Part | Part | Yes | Yes |
| MCPSpec | Open source | Yes | – | Yes | – | – | – | Part | – | Part | Part | Yes | Yes |
| lastmile mcp-eval | Open source | – | – | Yes | Yes | Yes | – | – | – | – | – | Yes | Yes |
| gleanwork mcp-server-tester | Open source | – | – | Yes | Part | Yes | – | – | – | – | – | Yes | Yes |
| steviec mcp-server-tester | Open source | – | – | Yes | – | Yes | Part | – | – | – | – | Part | Yes |
| mclenhard/mcp-evals | Open source | – | – | Part | – | Yes | – | – | – | – | – | Yes | Yes |
| Arcade evals | Open source | – | – | Yes | – | ? | – | – | – | – | – | Yes | Yes |
| Promptfoo | Open source | – | – | Yes | Part | Yes | – | – | – | Yes | – | Yes | Yes |
| MCP Observatory | Open core | – | – | Yes | – | – | Part | – | – | Yes | – | Yes | Yes |
| Official conformance suite | Open source | – | – | – | – | – | Yes | – | – | – | – | Yes | Yes |
| YawLabs mcp-compliance | Open source | – | – | – | – | – | Yes | – | – | Part | – | Yes | Yes |
| Glama TDQS | Open spec + CLI | – | – | – | – | – | – | Yes | – | – | – | Yes | Yes |
| mcp-surface-lint | Open source | – | – | – | – | – | – | Yes | – | – | – | Yes | Yes |
| mcpconform | Open source | – | – | – | – | – | – | Part | Part | – | – | Yes | Yes |
| mcp-server-lint | Open source | – | – | – | – | – | – | Yes | – | Part | – | Yes | Yes |
| MCProbe | Open core | – | – | – | – | – | – | Yes | – | – | – | – | Part |
| Alpic Beacon | Closed | – | – | – | – | – | Part | Part | Yes | – | – | Part | No |
| Snyk agent-scan | Open CLI + cloud | – | – | – | – | – | – | – | – | Yes | – | Yes | No |
| Cisco mcp-scanner | Open source | – | – | – | – | – | – | Part | – | Yes | – | Yes | Yes |
| MCP Drill | Open source | – | – | – | – | – | – | – | – | – | Yes | Part | Yes |
| mcp-stress | Open source | – | – | – | – | – | Part | – | – | – | Yes | Yes | Yes |
| xk6-mcp (Grafana, Infobip) | Open source | – | – | – | – | – | – | – | – | – | Yes | Yes | Yes |
| Tricentis NeoLoad | Paid | – | – | – | – | – | – | – | – | – | Yes | Yes | Yes |
Notes on the table: MCPJam’s AI generation and judge run on its hosted backend, and its host checks are a static prototype plus hosted directory-readiness runs. Glama TDQS uses an LLM to judge tool definitions, which is counted under Design, not Judge. mcpconform’s host column refers to LLM API provider limits, not host or directory rules. Beacon’s host checks cover the ChatGPT store and the Claude directory only.
What is missing
Capability by capability, the best existing coverage and what nobody offers:
| Capability | Best existing coverage | What is still missing | How open the gap is |
|---|---|---|---|
| Interactive client | Official Inspector, MCPJam, Postman, Glama | Very little | Low |
| Deterministic functional tests | MCPJam, MCPSpec, gleanwork; the official Inspector by early 2027 | A local UI whose tests live as files in the repository | Low to medium |
| AI-generated test cases | MCPJam (hosted), testmcpy, lastmile mcp-eval | Local generation on the user’s model; cases verified by execution before review; an explicit accept or reject queue | High |
| LLM judge and scoring | MCPJam (hosted), several CLIs | A judge that runs locally on the user’s key or gateway, with a visible rubric and consistency checks | Medium to high |
| Best-practice audit | Glama TDQS, a dozen small linters | A behavioural audit (errors, payload size) and LLM-written fixes proven by re-scoring | High |
| Host guideline validation | Alpic Beacon, Manufact, MCPJam | An open, versioned rule pack; the IDE and CLI hosts; live checks | High |
| Performance testing | MCP Drill, mcp-stress, the k6 extensions | Load built from saved functional scenarios; OAuth; stdio cold start; durable baselines | Medium to high |
| One combined report | None | A single readiness report across all layers with a CI gate | High |
| .NET-native tooling | None | A dotnet tool, Aspire integration | Medium, for a smaller audience |
In more detail:
- Local-first AI evaluation. No established tool runs test generation and judging entirely on the user’s own model key or gateway. MCPJam’s judge and generation depend on its backend and send traces to third parties; the local alternatives (testmcpy, MCPSpec) have single-digit stars.
- Quality of generated tests. The most common complaint about AI-generated MCP tests is that they are generic. No developer tool executes generated cases against the server and feeds real data back before showing them to a human.
- An audit that goes beyond the definition. TDQS and the linters read
tools/listonly. Nobody combines definition lint with live probes: does every tool succeed with valid input, are errors actionable, are responses within host size limits, do annotations match real behaviour. - Fixes, not just findings. Output today is scores and lists. No tool generates concrete rewrites and then shows that tool-selection accuracy or the score improved.
- An audit linked to evals. Research shows poor descriptions reduce tool-selection success, but no tool ties a lint finding to a measured failure in a test run.
- An open multi-host rule pack. Claude and ChatGPT store submission is served by a closed tool (Beacon), a paid cloud (Manufact) and MCPJam’s hosted runs. Nothing open encodes the rules of Cursor, VS Code, Codex, Gemini CLI, Kiro and Windsurf, and nobody publishes the rules as versioned data with sources.
- Load testing connected to function. Existing load tools are standalone and tiny. None reuses the scenarios you already debugged, handles real OAuth flows, measures stdio spawn cost, or keeps durable baselines.
- One report. No tool gives a single conformance, design, host-readiness, behaviour and performance report.
- Dual-era protocol support. Many small tools stop at 2025-06-18 or 2025-11-25. A tool that handles the 2026-07-28 stateless protocol and legacy servers from day one starts ahead of most of the long tail.
- .NET. Every tool reviewed is TypeScript, Python, Go, Rust or Java, while the C# SDK is Tier 1.
What is not missing
Manual inspection and OAuth debugging; LLM chat playgrounds; YAML evals with expected tool calls, an LLM judge and a GitHub Action; basic deterministic lint rules; security scanning; protocol conformance. Each of these is served several times over, and the leaders ship fast. Two things stand out for anyone building here: both leading inspectors have had critical remote code execution flaws, because they launch local processes from a web UI; and in this category scores, badges and public leaderboards are what spread (Glama’s TDQS, the registry scans).
Method and caveats
- Repository READMEs, vendor docs, pricing pages and the npm, PyPI and NuGet registries were read on 1 October 2026. No tool was installed or run, so this reviews what each tool documents, not how well it works in practice.
- Star counts are rounded snapshots and differed slightly between sources for a few repositories (the conformance suite and TDQS among them).
- The separate licence covering MCPJam’s eval server code could not be retrieved; only the carve-out in the main LICENSE was confirmed.
- Not verified: Beacon’s cost and metering; whether Postman’s collection runner and CLI support MCP requests; Cursor’s 40-tool limit in current docs; release years for MCP Drill. The Codex and Windsurf limits come from a single read of each vendor’s docs.
- Reddit was not accessible, so user complaints come from GitHub issues, Hacker News and dev.to.