Testing MCP servers in 2026: 31 tools compared and the gaps between them

A Model Context Protocol (MCP) server is used by a language model, not by a person, so testing it means more than calling each tool by hand. Does it follow the specification? Can an agent pick the right tool from its descriptions? Will it pass review for the Claude connector directory or the ChatGPT app store, and work inside Cursor, VS Code and the command-line agents? How does it behave under load? On 1 October 2026 I reviewed about eighty tools that inspect, test, evaluate, audit, secure or load test an MCP server, open source and paid, and compared thirty-one of them capability by capability. This is what the landscape looks like, and where it is thin.

In brief

  1. Every capability exists somewhere, but no open, local tool combines them. Inspection, evals, audits, host checks and load testing are spread over about a dozen products.
  2. MCPJam is the closest thing to a complete workbench. It ships saved eval suites, AI-generated test cases that the user keeps or discards, an LLM judge, client emulation, directory-readiness checks and CI gates. Its AI features depend on its hosted backend and credits.
  3. The official MCP Inspector is no longer a slow target. Version 2 has been generally available since July 2026, ships about weekly, and has saved collections, assertions, CI flows and a conformance runner on its roadmap for November 2026 to January 2027. It has no LLM or eval work planned.
  4. Three areas are saturated: manual inspection, security scanning, and the basic “YAML evals plus LLM judge plus GitHub Action”. Protocol conformance is owned by the official suite.
  5. The gaps are real: a fully local, bring-your-own-model evaluation loop; an open, versioned rule pack for many hosts backed by live behavioural checks; fix suggestions proven by re-scoring; load testing tied to functional scenarios; one combined report; and anything .NET-native.
  6. The protocol changed fundamentally two months before this review. Specification 2026-07-28 is a stateless redesign, and many small tools are stuck on older versions.

How the ground moved in 2026

The protocol

The current MCP specification is 2026-07-28. Per its changelog it removes the initialize handshake, protocol-level sessions (Mcp-Session-Id) and ping, and makes a new server/discover call mandatory. Elicitation, sampling and roots become multi-round-trip results instead of server-initiated requests. Dynamic Client Registration is deprecated in favour of Client ID Metadata Documents, and Roots, Sampling, Logging and HTTP+SSE are in a deprecation window. Tasks moved to an extension.

The consequence for anyone testing a server: a tester must speak both the legacy era (2025-11-25 and earlier) and the modern one. Hosts lag the specification, so the two will coexist for a long time. The specification has had five revisions in about twenty months, the latest of them breaking.

Official tooling

Official assetState todayPlanned
MCP Inspector v2Web, CLI and terminal UI; both protocol eras; CI-friendly CLIRecord and replay, diff, latency view (October to November 2026); saved calls, assertions, CI, conformance runner, OTLP export (November 2026 to January 2027); plugins, multi-server (January to February 2027), per its roadmap
Conformance suiteTests any server URL against per-revision requirement sets; GitHub Action; baselinesNew specification proposals must ship conformance scenarios (SEP-2484)
MCP RegistryPreview; validates namespace and manifest onlyDelegates scanning and curation to others

Nothing on the official roadmap covers evals, LLM-as-judge, tool-design quality, host-guideline validation or load testing.

Adoption, money and churn

The MCP blog reported “close to half-a-billion downloads a month” across Tier 1 SDKs in July 2026, and Glama’s registry listed about 94,600 servers on the day of this review. The money followed: Alpic raised a $6M pre-seed in September 2025, Manufact (the company behind mcp-use) a $6.3M seed in February 2026, and MCPJam is an Open Core Ventures company. Anthropic acquired Stainless in May 2026, Arcade acquired Smithery in August 2026, and Promptfoo’s README says it is now part of OpenAI.

Dated events behind the 2026 landscapeMONEYSEPT 2025Alpic raises a $6M pre-seed roundFEB 2026Manufact, the company behind mcp-use, raises a $6.3M seed roundMAY 2026Anthropic acquires StainlessAUG 2026Arcade acquires SmitheryPROTOCOL AND OFFICIAL TOOLSJULY 2026The MCP blog reports close to half a billion SDK downloads a monthThe C# SDK reaches version 2.028 JULY 2026Specification 2026-07-28, a stateless redesignMCP Inspector v2 generally available30 SEPT 2026MCP Inspector 2.9.0 on npmOCT TO NOV 2026Inspector roadmap, record and replay, diff, latency viewNOV 2026 TO JAN 2027Inspector roadmap, saved calls, assertions, CI, conformance runner, OTLP exportJAN TO FEB 2027Inspector roadmap, plugins and multi-server

Churn is the other signal. Many small tools went quiet within three to six months of launch (lastmile’s mcp-eval, mclenhard/mcp-evals, steviec’s tester, MCPSpec, mcp-shield, destilabs/mcp-doctor), and a few vanished (Bellwether, Janix-ai’s validator).

.NET is a gap of its own. The C# SDK is Tier 1 and reached version 2.0 in July 2026, yet no .NET-native inspector, test or load tool was found; Aspire users wrap the Node-based official Inspector.

Six jobs, and who does each

A tool in this space does one or more of six jobs: inspect a server by hand, run functional tests and evals against it, audit its design and protocol conformance, check it against the rules of the hosts that will run it, scan it for security problems, and load test it. Observability sits beside them on the production side.

Does the server work, and will the hosts accept itIs it safe, does it hold under load, and how does it run in productionInspectOfficial MCP InspectorMCPJam Inspectormcp-use InspectorPostman, Apidog, InsomniaHosted playgroundsCommand-line clientsFunctional tests and evalsMCPJam evalstestmcpyMCPSpeclastmile mcp-evalgleanwork mcp-server-testerPromptfoo, Arcade evalsAudit and conformanceOfficial conformance suiteYawLabs mcp-complianceGlama TDQSDeterministic lintersHost guideline validationAlpic BeaconManufact publishing checksMCPJam host compatibilityVendor skills and checklistsSecurity scanningSnyk agent-scanCisco mcp-scannerTencent AI-Infra-GuardMCP ObservatoryLoad testingMCP Drillmcp-stressk6 and JMeter extensionsTricentis NeoLoadObservabilityAgentCatSentry, Datadog, New RelicOpenTelemetry conventions

Inspectors and interactive clients

The official MCP Inspector is open source (MIT, moving to Apache-2.0), has about 10.7k stars, and reached version 2 on 28 July 2026, with 2.9.0 on npm on 30 September. It offers a web UI, a CLI and a terminal UI on one shared core, a Docker image, and needs Node 22.19 or later. It speaks stdio, SSE and Streamable HTTP in both protocol eras, and covers tools, resources, prompts, elicitation, sampling with a stubbed response, roots, logging, completions, pagination, Tasks, Skills and MCP Apps. It debugs OAuth, including Client ID Metadata Documents and enterprise-managed authorization, has protocol, network and console monitors, and its CLI gives CI a JSON output, exit codes and a --strict schema portability lint. It is the reference client that tracks the specification first, it is free, and it ships fast. It has no LLM, chat, evals, AI test generation or judge, and none planned; the roadmap says “the Inspector is not a host”. Saved tests, assertions and conformance are planned but not shipped; there are no host rules, no load testing and no reports. It is Node-only, and version 1 had a critical remote code execution flaw (CVE-2025-49596).

MCPJam Inspector (docs) is open core: Apache-2.0 with the eval server code under a separate licence, about 2.2k stars, releases almost daily. It does manual debugging of tools, resources, prompts and elicitation with JSON-RPC logs and a guided OAuth debugger; an LLM playground with up to three models side by side and many providers including Ollama and custom endpoints; a ChatGPT Apps and MCP Apps widget emulator; eval suites with expected tool calls, argument matching, negative cases, deterministic assertions and token budgets; AI test generation, where a Generate button drafts positive and negative cases from the tool catalog and each draft is saved or discarded (docs); an LLM judge with a threshold (default 0.7) and up to ten rubric checks, plus a multi-model matrix with latency, tokens and cost; baselines, quality gates, schedules, pull request comments and JUnit, HTML and JSON reports; static host compatibility verdicts for Claude, ChatGPT, Cursor, Copilot and Codex Desktop, and directory-readiness runs against the Anthropic and OpenAI directory requirements; and protocol, OAuth, MCP Apps and Tasks conformance in the CLI and SDK. Pricing: Free with 200 credits a day, Pro at $24 a month on annual billing, Team at $199 a month, Enterprise custom; evals cost 3 credits per prompt (a credit is $0.01), and runs on your own model key are not billed. It has the broadest feature set in the category, a funded team and a real CI story, and it already ships the generate, review, store and regress loop. But the judge, rubric checks and generation need the hosted backend: the local CLI runs deterministic assertions with your own key but not the judge, and hosted judge traffic goes through OpenRouter. Host compatibility is labelled “prototype · static checks” and omits transport and auth blockers. There is no load or stress testing, only per-run latency, and the hosted app has no stdio. Early versions had an unauthenticated RCE (CVE-2026-23744).

mcp-use Inspector and the Manufact cloud: the inspector is MIT, the cloud is paid, and the framework repository as a whole has about 10.6k stars. It runs hosted, through npx, in Docker, or auto-mounted in the mcp-use dev server, and can be embedded as a React client. Streamable HTTP only; multi-server; tools, resources, prompts and elicitation; bring-your-own-key chat with OpenAI, Anthropic, Google and Ollama; saved tool calls with replay; MCP-UI and OpenAI Apps SDK widget debugging. The cloud (pricing) adds test suites and a publishing checklist on Free, end-to-end checks for ChatGPT and Claude and a submission pack from Hobby ($25 a month), Startup at $250, Enterprise from $1,000. The UI is polished and the community large, but there is no stdio, the open-source inspector has no assertions, AI generation, judge or CI mode (that value sits in the paid cloud), it is tied to the mcp-use framework funnel, and telemetry is on by default.

Postman (proprietary; Free, Solo $9, Team $19 per user a month, Enterprise custom) has an MCP request type over stdio and Streamable HTTP with tools, resources and prompts, an elicitation form, mocked sampling, an MCP Apps preview, OAuth 2.1 with a debugger, and saving to collections shared in workspaces. It is familiar to every API team and strong on collaboration and auth. It has no MCP evals, AI test generation or judge, whether MCP requests work with test scripts, the collection runner or the CLI is not documented, and it is closed source with an account required.

Apidog (proprietary freemium) covers stdio and Streamable HTTP, tools, prompts and resources, many auth types, environment variables, a raw JSON-RPC view and saving to a project. Rich auth options in a familiar workflow; no sampling or elicitation documented, no MCP test scenarios, assertions or CI.

Insomnia (Kong; Apache-2.0 core; MCP client since version 12 in November 2025) covers HTTP transport, tools, prompts and resources, automatic OAuth discovery, roots, events, notifications, elicitation, and sampling through its own LLM settings. Open-source core with Git-based project storage; stdio not mentioned; no MCP test scripts or evals.

Hosted playgrounds:

ToolWhat it doesLimits
Glama MCP InspectorFree browser inspector: tools, resources, prompts, elicitation, sampling, roots, tasks, OAuth 2.1; shareable state in the URLRemote servers only (stdio through a tunnel); no chat, no tests
MCP Playground OnlineFree inspector with JSON-RPC logs; “Agent Studio” across up to four models; an “AI Readiness Review” that generates user questions, rates tool descriptions and proposes rewrites; security scannerClosed source; remote servers only; no saved suites or CI
Cloudflare AI PlaygroundFree chat client for remote servers with OAuth and a debug logNo manual tool forms, saving, tests or CI
MCPize, Smithery playgroundsBasic inspector and chatNot independently verified

Command-line clients:

ToolLicence, starsWhat it isGap
f/mcptoolsMIT, ~1.6klist and call, shell, mock server, proxyStale since May 2025; no auth
apify/mcpcApache-2.0, ~945Persistent sessions, OAuth 2.1, JSON modeNo LLM, no tests
IBM/mcp-cliApache-2.0, ~2.0kLLM chat host with many providersA host, not a tester
philschmid/mcp-cliMIT, ~1.2kSingle binary for shell chainingNo tests
wong2/mcp-cliGPL-3.0, ~440Interactive inspector with OAuthNo tests
reloaderooMIT, ~125CLI inspection plus hot-reload proxySlow-moving
mcp-probeMIT, ~135Rust terminal debugger with a validation suiteEarly (v0.1.0)

Also in this space: FastMCP’s fastmcp dev inspector simply launches the official Inspector; VS Code and Cursor offer MCP logs and debug attach but are hosts, not test tools; Hoppscotch and Bruno have no MCP client. The name “MCP Workbench” is already used by Orkes and at least two other projects.

Functional tests and AI evaluation

MCPJam’s evals, described above, are the benchmark for this category.

testmcpy (Preset; Apache-2.0; 4 stars; active, with PyPI releases through September 2026) is a Python CLI with an optional React UI, SQLite storage, a GitHub Action and tests in YAML. It has more than forty evaluators (tool called, parameters, sequence, call count, answer contents, time and token limits), AI test generation from the CLI and the UI with a generation history, eleven model providers including Ollama, model comparison and a leaderboard, baselines, mutation and metamorphic testing, flaky detection, a load_test step, a 0 to 100 “LLM usability” score, a SARIF scan, a wrapper for the official conformance suite and JUnit output. It is the closest open-source match to a full workbench, and it runs locally. It has almost no adoption, the repository still describes itself as built for testing Superset, generated tests are written as YAML with no accept or reject queue, and LLM-judge and stdio support could not be confirmed.

MCPSpec (MIT; 7 stars; last commit March 2026, which looks stalled) is a TypeScript CLI with a React dashboard and YAML test collections: ten assertion types, variable extraction for chained calls, record and replay with diff, mock-server generation, an eight-rule security audit, a 0 to 100 “MCP Score”, a latency benchmark with P95 and P99, baselines and several reporters. Broad deterministic coverage in a single tool; nothing LLM-driven, by design; one burst of development and no traction.

lastmile-ai/mcp-eval (Apache-2.0; 25 stars; dormant since September 2025) is a Python library and CLI with tests as decorators, pytest or datasets; assertions on content, tool calls, sequence, performance and path efficiency; an LLM judge with a rubric and a minimum score; OpenTelemetry-based latency, token and cost metrics; mcp-eval generate, which writes test files from the tool list; HTML, JSON and Markdown reports; and a GitHub Action. Clean code-first design with a judge and reports, but no UI, no review step for generated tests and no run history. Users report that generated tests are generic and do not exercise the server (#36, #40), and it is unmaintained.

gleanwork/mcp-server-tester (MIT; 19 stars; active, backed by Glean) is a Playwright fixture with matchers for schema, snapshot, response size, tool calls and a judge; an “LLM host mode” with iterations and pass rates, including Claude Code and Codex as hosts; an HTML reporter with a pass-rate trend; and a non-LLM wizard that suggests expectations to accept. Solid engineering that compares servers and hosts side by side. It requires Playwright and TypeScript knowledge, ignores resources, prompts and notifications, and delegates AI authoring to coding-agent skills.

Smaller eval tools:

ToolLicence, stars, statusWhat it doesMain limits
steviec/mcp-server-testerMIT, 35, dormant since Sept 2025YAML direct-call tests and evals with required tools and an LLM judgeAnthropic only; no reports or UI
mclenhard/mcp-evalsMIT, 132, unmaintained since mid-2025LLM grades answers 1 to 5 on five criteria; GitHub ActionNo tool-call assertions; no model matrix
mcp-use/eval-actionNo licence file, 3YAML matrix of case by model; rubric score plus required toolsOpenRouter only; the judge sees only the final answer
alpic-ai/mcp-evalMIT, 21YAML expected tool calls with profiles for Claude, ChatGPT, Le ChatPublic HTTP servers only; no judge
Arcade evalsMIT, ~1k (parent repo), activePython suites with weighted argument critics, multi-run statistics, capture modeScores tool selection and arguments without executing the tool
mcp-jestMIT, 18Snapshot tests, auto-discovered deterministic tests, compliance scoreNo LLM features
mcp-recorderMIT, 9Record and replay cassettes, pytest pluginStalled since March 2026
MCP ObservatoryMIT core, 140, activeSchema drift, record, replay and verify, security scan, health score, SARIFNo benchmarking; paid tiers
r-huijts/mcp-server-testerMIT, 10, staleClaude generates and runs N tests per toolNo review step; work in progress
mcpevals.aiMIT, 1Paste a URL; auto-generated arguments; compatibility and response timesA smoke test, not a suite
Specmatic MCP Auto TestLicensing not verifiedGenerates positive and negative cases deterministically from schemasStreamable HTTP only

General evaluation platforms:

ToolMCP relevanceLimits
Promptfoo (MIT, ~24k stars, now part of OpenAI)MCP provider for direct tool calls with assertions; red-team plugin for MCP attack classesNo functional test generation for MCP
DeepEval (Apache-2.0, ~17.5k stars)Three LLM-scored MCP metricsScores an MCP-using app from recorded interactions; does not connect to a server
Braintrust, LangSmith, Arize Phoenix, Inspect AITracing or generic scorersNo MCP server test features confirmed

Framework helpers exist too: FastMCP, the TypeScript SDK and the Python SDK document in-memory client testing, and the C# SDK offers an in-memory transport sample. Benchmarks such as MCP-Universe, MCP-Bench, MCPMark and LiveMCPBench rank language models across servers; they do not test a developer’s own server. Two pieces of research are worth knowing: “Tool descriptions are smelly” found at least one defect in 97.1% of 856 tool descriptions, and Salesforce’s MCPEval verifies generated tasks by executing them, which no developer-facing tool does yet.

Audit, lint and protocol conformance

The official conformance suite (MIT; about 120 stars; very active, pre-1.0): npx @modelcontextprotocol/conformance server --url <url> tests a server, and a client mode tests clients. It carries per-revision requirement sets for 2025-11-25 and 2026-07-28, validates the wire schema of every message, and has an expected-failures baseline file for CI, a GitHub Action and SDK tier scoring. It is authoritative and tied to SDK governance, so it will track the specification automatically. It is aimed at SDK maintainers: raw pass or fail, no grade, no fix advice, server testing documented for HTTP only, and nothing about tool design, usability by an agent, security or host rules.

YawLabs mcp-compliance (MIT; 2 stars; very active) runs 88 tests for 2025-11-25 and 103 for 2026-07-28 across transport, lifecycle, tools, resources, prompts, errors, schema and security, over HTTP and stdio, with an A to F grade, JSON, SARIF, a badge and a strict CI exit code. It is the most server-author-friendly conformance tool: graded, stdio-capable, both eras. Its test catalogue is unofficial and from one small vendor, adoption is negligible, and it has no design or host checks.

Glama’s Tool Definition Quality Score (TDQS; the CLI is Apache-2.0, v0.1.0, September 2026; about 30 stars) makes one LLM call per tool and scores six dimensions from 1 to 5: purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness and completeness, plus a server-level coherence score for disambiguation, naming consistency, tool count and completeness. tdqs lint is deterministic and needs no key; tdqs score uses any OpenAI-compatible endpoint or the hosted service, with a CI gate through --fail-under. It runs continuously across Glama’s registry (15,036 servers scored as of June 2026, per the repository), so maintainers are already graded by it, and its rubric and prompts are published and research-grounded. It scores the definition only, never behaviour; it has no protocol, security or host checks; it gives justifications rather than rewritten text; and the scores depend on the rubric-plus-model pair.

Deterministic design linters:

ToolLicence, stars, statusWhat it checksLimits
mcp-surface-lintMIT, 0, active19 rules: description hygiene, tool and token budgets, naming, annotations, loose schemas, overlap clusters, CRUD mirrors, missing list limits; category scoresStatic only; no fixes; no adoption
mcpconformMIT, 156 cited rules for tools, registry manifest and client configs, plus data-driven profiles for Anthropic, OpenAI and Gemini API constraints; SARIFProfiles cover LLM API limits, not host or directory rules
mcp-server-lintMIT, 1, hyperactiveStatic source analysis (Python, TypeScript, Go): descriptions, typing, error handling, annotation-versus-code mismatch, unsafe calls, secretsNever runs the server; brittle parsers
destilabs/mcp-doctorMIT, 16, stale since Nov 2025Rules based on Anthropic’s “Writing tools for agents”; calls tools and measures response tokens and paginationHeuristic; no score or CI gate
MCProbeMIT engine, 612 schema rules plus a fuzz that calls tools with bad inputs and classifies the error handling; A to F gradeSmall; the hosted full report is paid ($9.90)
Roee-Tsur/mcp-spec-check16Readiness probe for the 2026-07-28 specification; its README claims 1 of 7,850 remote registry servers passed all three required checksRemote HTTP only; the claim is not verified
zuchka/mcp-doctor, mcp-lint, MCP Tool Auditor, agent-friend0 to 4 stars eachThe same dozen rules: missing or short descriptions, vague names, no pagination, tool count, token costHobby projects; mechanical fixes at best

The pattern: ten or more tools implement the same basic rules, and none has traction. No tool was found that uses an LLM to write concrete improvements to names, descriptions, schemas and error messages and then proves the improvement by re-scoring, and none reviews error messages or response payloads with an LLM.

Host and platform guideline validation

Each host that runs MCP servers publishes its own rules: the Claude connector directory wants title and a readOnlyHint or destructiveHint on every tool and tool names of 64 characters or fewer; ChatGPT apps reject tools without readOnlyHint, destructiveHint and openWorldHint; Codex has a 10-second startup timeout; VS Code allows at most 128 tools per chat request, Windsurf 100 in total; Kiro and Gemini CLI cap and truncate names at 64 and 63 characters. The rules conflict, roughly half can be checked statically, a quarter need live calls, and most of what reviewers actually reject needs LLM judgment.

The tools that validate against them: Alpic Beacon (May 2026; closed; part of Alpic’s platform) is a pre-submission audit for the ChatGPT store and the Claude connectors directory covering protocol plumbing, tool quality and annotations, widget rendering, CSP and LLM-driven edge-case inputs, through a dashboard or alpic audit; it is built from real submission experience and covers live behaviour, but it is two stores only, not open source, and its exact cost and metering are not verified. Manufact’s publishing checks offer a checklist on Free and end-to-end checks in simulated ChatGPT and Claude with a submission pack from $25 a month, cloud only and focused on the two app stores. MCPJam’s host compatibility and directory readiness, described above, give free deterministic grading and many client profiles, but the compatibility is a self-described prototype, the readiness runs are hosted and the rule list is not published. The vendors’ own aids are OpenAI’s chatgpt-app-submission skill, instructions a coding agent follows rather than a deterministic validator; Anthropic’s automatic scan of submissions, which is not available to developers, and the manual checklist in its mcp-server-dev plugin; mcpb validate, which checks a desktop-extension manifest against its schema only; and sunpeak (MIT, about 217 stars), which simulates the ChatGPT and Claude runtimes for MCP Apps while developers write the annotation assertions by hand.

Not found: an open, rule-pack-driven validator that covers Claude, ChatGPT and the IDE and CLI hosts (Cursor, VS Code, Codex, Gemini CLI, Kiro, Windsurf) together.

Security scanners

This is the most crowded segment and the least promising place to compete.

ToolLicence, starsWhat it doesNotes
Snyk agent-scan (formerly Invariant mcp-scan)Apache-2.0 CLI, ~3.1kDiscovers agent configs, scans tool descriptions for injection, poisoning, shadowing, toxic flowsDe facto standard; needs a Snyk token and sends data to Snyk’s API
Cisco mcp-scannerApache-2.0, ~1.1kYARA rules, LLM judge with your key, source scanning, dependency audit; a “readiness analyzer” with 20 heuristic rules and a 0 to 100 scoreBroadest engine set; has an offline mode
Tencent AI-Infra-Guard~6.4kPlatform with agent-driven MCP scanning across 14 categoriesLarge, general AI security platform
mcp-shield~550Tool poisoning and exfiltration checksAbandoned since April 2025
Proximity, AgentSeal, Ramparts, APIsec mcp-audit100 to 400 eachPattern plus LLM scanning, toxic-flow analysis, config inventoryOverlapping rule sets
Enkrypt AI, Akto, Backslash, PillarCommercialHosted scans, gateways, runtime defencePricing mostly undisclosed

Registries also publish scores: Glama (TDQS plus sandbox monitoring), AgentSeal (a 0 to 100 trust score) and MCP Trust Checker. The OWASP MCP Top 10 is in beta and is the common mapping target.

Performance and load testing

Most search results for “k6 MCP”, “JMeter MCP” or “Gatling MCP” are MCP servers that let an AI agent drive a load tool. They do not load test MCP servers. The tools below do.

MCP Drill (MIT; 7 stars; v0.1.3) is a Go control plane with distributed workers and a React UI: a run wizard, a live dashboard, throughput, P50, P95 and P99 latency, error rate, a per-tool breakdown, connection stability and optional server CPU and memory. It has stages (preflight, baseline, ramp, soak, spike), session modes (reuse, per_request, pool, churn), a weighted operation mix, side-by-side run comparison, a bundled mock server, and it supports the 2026-07-28 protocol and legacy versions. It is the only open tool with a UI, per-tool percentiles and run comparison, and it already speaks the stateless protocol. It is Streamable HTTP only, has no OAuth flows, keeps runs in memory so there are no durable baselines, and has a tiny community.

mcp-stress (MIT; 0 stars; npm 0.2.1, February 2026) is a Deno and TypeScript CLI over stdio, Streamable HTTP and legacy SSE, with profiles (tool flood, find-ceiling, ping flood), load shapes (ramp, step, spike, sawtooth), a live browser dashboard, an HTML report, and CI assertions such as p99 < 500ms with baseline deltas. One-line install, stdio support and regression assertions; single process, no OAuth, and essentially unused.

Other load options:

ToolLicence, maturityWhat it doesLimits
reaatech/mcp-load-testMIT, 0 starsCLI and library; per-tool P50 to P99, breaking-point detection, A to F grades, compare; claims OAuth client credentialsNo UI; untested by third parties
grafana/xk6-mcpAGPL-3.0, 21 starsk6 extension for tools, resources, promptsExperimental, “not officially supported”; metrics tagged by method only
infobip/xk6-infobip-mcpMIT, 5 stars, activek6 extension with per-tool tags; stateful and stateless serverscallTool only; no stdio
BlazeMeter jmeter-mcp-pluginApache-2.0, v0.1.0JMeter sampler for MCP operationsShares one MCP client across threads
Tricentis NeoLoad 2026.2CommercialNative MCP connect, list and call actions; one client per virtual userEnterprise pricing; tools only
Locust recipe on Azure Load TestingSample codeHand-rolled Streamable HTTP clientA recipe, not a tool

MCPJam and lastmile’s mcp-eval report latency per run, but neither does concurrency or throughput. MCPSpec and testmcpy have small benchmark features with no adoption.

Load testing MCP is harder than load testing plain HTTP, for five reasons:

  • Two protocol eras: stateful sessions before 2026-07-28, stateless after.
  • One MCP call is several HTTP requests, and responses may be JSON or an SSE stream.
  • Tool failures hide behind HTTP 200 as isError results.
  • stdio is one client per process, so concurrency means measuring spawn cost and memory.
  • OAuth tokens expire mid-run, tool calls have side effects, and they hit downstream rate limits.

Observability

ProductWhat it doesModel
AgentCat (formerly MCPcat)SDK wrapper with session replay and agent intent captureMIT SDK; hosted from free to $160 a month and up
Sentry MCP monitoringSpans for tools, resources, prompts and transportsPaid SaaS
Datadog, New Relic, Grafana CloudMCP client or server tracing and dashboardsPaid
Shinzo, Moesif, Agnost AIOpenTelemetry-based analytics for MCP serversMixed
Speakeasy Gram, Alpic, Lunar MCPXHosting or gateway platforms with tool-call logs and a playgroundOpen core or paid
OpenTelemetry MCP conventionsStandard span and metric names; the official Python SDK now traces automaticallyOpen standard

None of these offers pre-production testing or load testing. OpenTelemetry is the neutral integration point for any deep-tracing feature.

The comparison

Counting the thirty-one tools that offer each capability outright shows where the market is: almost everything runs in CI and locally, a third of the tools inspect or store tests, and host rules are covered by one.

Tools that offer each capability outrightTOOLS, OF 31, COUNTED AS YES IN THE TABLE BELOWInspect9Chat4Saved tests9AI test generation3LLM judge6Conformance4Design audit4Host rules1Security4Load4CI20Local24

Column meanings: Inspect is manual calls to tools, resources and prompts; Chat an LLM playground against the server; Saved tests stored deterministic tests or assertions; AI gen an LLM drafting test cases; Judge an LLM scoring test runs; Conform protocol and specification compliance; Design a tool-definition and best-practice audit; Host checks against specific AI platform rules; Sec security scanning; Load concurrency, throughput or stress testing; CI headless runs with gates; Local every feature works without the vendor’s cloud. Yes means documented, Part partial or limited, a dash not offered, and a question mark not verified.

ToolModelInspectChatSaved testsAI genJudgeConformDesignHostSecLoadCILocal
Official MCP InspectorOpen sourceYes–Planned––PlannedPart–––PartYes
MCPJamOpen coreYesYesYesYesYesYesPartPart––YesPart
mcp-use Inspector + ManufactMIT + paid cloudYesYesCloud?Cloud–CloudCloud––CloudPart
PostmanPaidYes–Part––––––??Yes
ApidogPaidYes–Part––––––––Yes
InsomniaOpen coreYes––––––––––Yes
Glama InspectorFree hostedYes––––––––––No
MCP Playground OnlineClosed hostedYesYes–Part––Part–Part––No
testmcpyOpen sourcePartYesYesYes?YesPart–PartPartYesYes
MCPSpecOpen sourceYes–Yes–––Part–PartPartYesYes
lastmile mcp-evalOpen source––YesYesYes–––––YesYes
gleanwork mcp-server-testerOpen source––YesPartYes–––––YesYes
steviec mcp-server-testerOpen source––Yes–YesPart––––PartYes
mclenhard/mcp-evalsOpen source––Part–Yes–––––YesYes
Arcade evalsOpen source––Yes–?–––––YesYes
PromptfooOpen source––YesPartYes–––Yes–YesYes
MCP ObservatoryOpen core––Yes––Part––Yes–YesYes
Official conformance suiteOpen source–––––Yes––––YesYes
YawLabs mcp-complianceOpen source–––––Yes––Part–YesYes
Glama TDQSOpen spec + CLI––––––Yes–––YesYes
mcp-surface-lintOpen source––––––Yes–––YesYes
mcpconformOpen source––––––PartPart––YesYes
mcp-server-lintOpen source––––––Yes–Part–YesYes
MCProbeOpen core––––––Yes––––Part
Alpic BeaconClosed–––––PartPartYes––PartNo
Snyk agent-scanOpen CLI + cloud––––––––Yes–YesNo
Cisco mcp-scannerOpen source––––––Part–Yes–YesYes
MCP DrillOpen source–––––––––YesPartYes
mcp-stressOpen source–––––Part–––YesYesYes
xk6-mcp (Grafana, Infobip)Open source–––––––––YesYesYes
Tricentis NeoLoadPaid–––––––––YesYesYes

Notes on the table: MCPJam’s AI generation and judge run on its hosted backend, and its host checks are a static prototype plus hosted directory-readiness runs. Glama TDQS uses an LLM to judge tool definitions, which is counted under Design, not Judge. mcpconform’s host column refers to LLM API provider limits, not host or directory rules. Beacon’s host checks cover the ChatGPT store and the Claude directory only.

What is missing

Capability by capability, the best existing coverage and what nobody offers:

CapabilityBest existing coverageWhat is still missingHow open the gap is
Interactive clientOfficial Inspector, MCPJam, Postman, GlamaVery littleLow
Deterministic functional testsMCPJam, MCPSpec, gleanwork; the official Inspector by early 2027A local UI whose tests live as files in the repositoryLow to medium
AI-generated test casesMCPJam (hosted), testmcpy, lastmile mcp-evalLocal generation on the user’s model; cases verified by execution before review; an explicit accept or reject queueHigh
LLM judge and scoringMCPJam (hosted), several CLIsA judge that runs locally on the user’s key or gateway, with a visible rubric and consistency checksMedium to high
Best-practice auditGlama TDQS, a dozen small lintersA behavioural audit (errors, payload size) and LLM-written fixes proven by re-scoringHigh
Host guideline validationAlpic Beacon, Manufact, MCPJamAn open, versioned rule pack; the IDE and CLI hosts; live checksHigh
Performance testingMCP Drill, mcp-stress, the k6 extensionsLoad built from saved functional scenarios; OAuth; stdio cold start; durable baselinesMedium to high
One combined reportNoneA single readiness report across all layers with a CI gateHigh
.NET-native toolingNoneA dotnet tool, Aspire integrationMedium, for a smaller audience

In more detail:

  1. Local-first AI evaluation. No established tool runs test generation and judging entirely on the user’s own model key or gateway. MCPJam’s judge and generation depend on its backend and send traces to third parties; the local alternatives (testmcpy, MCPSpec) have single-digit stars.
  2. Quality of generated tests. The most common complaint about AI-generated MCP tests is that they are generic. No developer tool executes generated cases against the server and feeds real data back before showing them to a human.
  3. An audit that goes beyond the definition. TDQS and the linters read tools/list only. Nobody combines definition lint with live probes: does every tool succeed with valid input, are errors actionable, are responses within host size limits, do annotations match real behaviour.
  4. Fixes, not just findings. Output today is scores and lists. No tool generates concrete rewrites and then shows that tool-selection accuracy or the score improved.
  5. An audit linked to evals. Research shows poor descriptions reduce tool-selection success, but no tool ties a lint finding to a measured failure in a test run.
  6. An open multi-host rule pack. Claude and ChatGPT store submission is served by a closed tool (Beacon), a paid cloud (Manufact) and MCPJam’s hosted runs. Nothing open encodes the rules of Cursor, VS Code, Codex, Gemini CLI, Kiro and Windsurf, and nobody publishes the rules as versioned data with sources.
  7. Load testing connected to function. Existing load tools are standalone and tiny. None reuses the scenarios you already debugged, handles real OAuth flows, measures stdio spawn cost, or keeps durable baselines.
  8. One report. No tool gives a single conformance, design, host-readiness, behaviour and performance report.
  9. Dual-era protocol support. Many small tools stop at 2025-06-18 or 2025-11-25. A tool that handles the 2026-07-28 stateless protocol and legacy servers from day one starts ahead of most of the long tail.
  10. .NET. Every tool reviewed is TypeScript, Python, Go, Rust or Java, while the C# SDK is Tier 1.

What is not missing

Manual inspection and OAuth debugging; LLM chat playgrounds; YAML evals with expected tool calls, an LLM judge and a GitHub Action; basic deterministic lint rules; security scanning; protocol conformance. Each of these is served several times over, and the leaders ship fast. Two things stand out for anyone building here: both leading inspectors have had critical remote code execution flaws, because they launch local processes from a web UI; and in this category scores, badges and public leaderboards are what spread (Glama’s TDQS, the registry scans).

Method and caveats

  • Repository READMEs, vendor docs, pricing pages and the npm, PyPI and NuGet registries were read on 1 October 2026. No tool was installed or run, so this reviews what each tool documents, not how well it works in practice.
  • Star counts are rounded snapshots and differed slightly between sources for a few repositories (the conformance suite and TDQS among them).
  • The separate licence covering MCPJam’s eval server code could not be retrieved; only the carve-out in the main LICENSE was confirmed.
  • Not verified: Beacon’s cost and metering; whether Postman’s collection runner and CLI support MCP requests; Cursor’s 40-tool limit in current docs; release years for MCP Drill. The Codex and Windsurf limits come from a single read of each vendor’s docs.
  • Reddit was not accessible, so user complaints come from GitHub issues, Hacker News and dev.to.

I’m Amir Pournasserian. I build AI and platform systems for a living, maintain FluentCMS and YeSvelte, and write here about what I find along the way.