One MCP server, eight rulebooks: what the spec, Claude, ChatGPT, Codex, VS Code, Windsurf, Kiro and Gemini CLI require

A Model Context Protocol (MCP) server that follows the specification can still be rejected by the Claude connector directory, turned away from the ChatGPT app store for a missing annotation, or have its tool names silently cut short by Gemini CLI. Each host that runs MCP servers publishes its own rules, and they do not agree with each other. While reviewing the tools that test MCP servers, I collected the rules as they stood on 1 October 2026. This is the rulebook, where it conflicts with itself, how much of it a machine can check, and what tries to check it today.

In brief

  1. Eight rulebooks, one server. The specification, the Claude connector directory, the Claude apps and Claude Code, ChatGPT apps, Codex, VS Code with Copilot, Windsurf, Kiro and Gemini CLI each add constraints of their own.
  2. They conflict. A tool name may be 128 characters long under the specification, 64 for the Claude directory and Kiro, and anything over 63 is truncated by Gemini CLI. The specification allows hyphens and dots in names; Kiro rejects them. Annotations are optional in the specification and mandatory for Claude and ChatGPT.
  3. About half of the rules can be checked statically, a quarter need live calls, and most of what reviewers actually reject needs judgment: whether a description matches behaviour, whether annotations are honest, whether a description carries an injection pattern.
  4. The rules move, so a validator has to carry them as versioned data with a source and a date, not as code.
  5. What validates against them today is closed, paid or partial: Alpic Beacon, Manufact’s publishing checks and MCPJam’s hosted readiness runs. Nothing open covers the IDE and command-line hosts.

The rules, host by host

HostConcrete, checkable rules (examples)Source
MCP specificationTool names 1 to 128 characters from A-Z a-z 0-9 _ - .; a valid inputSchema; structuredContent must match outputSchema; input errors returned as isError resultsSpecification, tools
Claude connector directoryTool names 64 characters or fewer; every tool has title plus readOnlyHint or destructiveHint; read and write split into separate tools; no prompt-injection patterns; every tool succeeds with valid input; no generic errors; reasonably sized responsesReview criteria
Claude apps and Claude CodeAbout 150,000 characters per tool result in claude.ai and Claude Desktop; 25,000 tokens by default in Claude CodeConnector docs, Claude Code docs
ChatGPT appsreadOnlyHint, destructiveHint and openWorldHint set on every tool (“a common cause of rejection”); non-promotional descriptions; minimal inputs; search and fetch tool shapes for deep researchSubmission guidelines
CodexA startup timeout of 10 seconds and a tool timeout of 60 seconds by defaultCodex MCP docs
VS Code and CopilotAt most 128 tools enabled per chat requestVS Code docs
Windsurf100 tools in totalWindsurf docs
KiroNames at most 64 characters including the server prefix, matching ^[a-zA-Z][a-zA-Z0-9_]*$Kiro docs
Gemini CLINames over 63 characters are truncated; some schema keywords are strippedGemini CLI docs

Where the rulebooks disagree

Name length. The specification allows up to 128 characters. The Claude directory and Kiro stop at 64, and Kiro counts the server prefix. Gemini CLI does not reject a longer name; it truncates it at 63, which is worse, because the server never finds out.

Longest tool name each rulebook acceptsCHARACTERS128MCPspecification64Claudedirectory64Kiro, prefixincluded63GeminiCLI, thentruncated

Characters. The specification permits letters, digits, underscore, hyphen and dot. Kiro’s pattern, ^[a-zA-Z][a-zA-Z0-9_]*$, permits none of the hyphens and dots and requires a letter first. A name that is valid everywhere else fails there.

Annotations. readOnlyHint, destructiveHint and openWorldHint are optional in the specification. The Claude directory requires title plus a read-only or destructive hint on every tool, and wants read and write operations split into separate tools. ChatGPT requires all three hints on every tool and calls their absence a common cause of rejection.

Budgets. Hosts cap different things. VS Code counts tools per chat request (128), Windsurf tools in total (100), Codex seconds (10 to start, 60 per call), and Claude the size of a result (about 150,000 characters in the apps, 25,000 tokens by default in Claude Code). A server built for one of these can exceed another without changing a line.

What a machine can check

Roughly half of the rules are checkable statically from tools/list: name length and characters, the presence of annotations and titles, a valid input schema, the tool count. About a quarter need live test calls: that every tool succeeds with valid input, that errors come back as isError results rather than generic failures, that responses stay within a host’s size limits, that a server starts within Codex’s ten seconds.

Most of what reviewers actually reject needs judgment rather than a rule: whether a description matches what the tool does, whether a readOnlyHint is honest, whether a description is promotional, whether anything in it reads as a prompt injection. That is work for a language model with a rubric, and its verdicts need a human behind them.

The host documentation moves often, and the rules above were each read once on one day. A validator that hard-codes them will rot. The rules have to be data: versioned, each with its source and the date it was read, so that a change in one host’s documentation is one record, not a release.

Descriptions are where servers fail

The rule that matters most is the one no regular expression can check: the description. “Tool descriptions are smelly” found at least one defect in 97.1% of the 856 tool descriptions it examined, and poor descriptions reduce the chance that an agent selects the right tool.

The most developed yardstick for a description is Glama’s Tool Definition Quality Score (TDQS). One model call per tool scores six dimensions from 1 to 5 (purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness and completeness) and a server-level score covers disambiguation between tools, naming consistency, tool count and completeness. The rubric and prompts are published, tdqs lint runs deterministically without a key, tdqs score runs against any OpenAI-compatible endpoint with a --fail-under gate for CI, and Glama runs it continuously across its registry (15,036 servers scored as of June 2026, per the repository), so maintainers are already being graded by it whether they know or not. It scores the definition only, never behaviour, and it gives justifications rather than rewritten text.

Anthropic’s own guidance, Writing effective tools for agents, is the basis of at least one linter’s rules, and the ChatGPT guidelines ask for descriptions that are not promotional and inputs that are minimal.

Two protocol eras

The current specification, 2026-07-28, is a stateless redesign. Per its changelog it removes the initialize handshake, protocol-level sessions and ping, makes a new server/discover call mandatory, turns elicitation, sampling and roots into multi-round-trip results, deprecates Dynamic Client Registration in favour of Client ID Metadata Documents, and puts Roots, Sampling, Logging and HTTP+SSE in a deprecation window. Hosts lag the specification, so a server has to expect the legacy era (2025-11-25 and earlier) and the modern one side by side for a long time, and so does anything that validates it.

What validates against the rules today

Alpic Beacon (May 2026; closed; part of Alpic’s platform) is a pre-submission audit for the ChatGPT store and the Claude connectors directory. It checks protocol plumbing, tool quality and annotations, widget rendering, CSP and LLM-driven edge-case inputs, through a dashboard or the alpic audit CLI. It is built from real submission experience and covers live behaviour, not only metadata. It covers two stores only, it is not open source, and its exact cost and metering are not verified.

Manufact’s publishing checks offer a publishing checklist on the free tier and, from $25 a month, end-to-end checks in simulated ChatGPT and Claude plus a submission pack. Cloud only, and focused on the two app stores.

MCPJam’s host compatibility gives static Works, Degraded, Blocked or Unknown verdicts for Claude, ChatGPT, Cursor, Copilot and Codex Desktop, and its hosted directory-readiness runs grade a server against the Anthropic and OpenAI directory requirements, with optional LLM observations that cost credits. The deterministic grading is free and the client profiles are many, but the compatibility feature describes itself as a prototype of static checks and omits transport and auth blockers, the readiness runs need the hosted backend, and the rule list is not published.

The vendors’ own aids are instructions, not validators. OpenAI publishes a chatgpt-app-submission skill that a coding agent follows to check hints and generate test cases. Anthropic scans submissions automatically, but the scanner is not available to developers; its mcp-server-dev plugin carries a manual checklist. mcpb validate checks a desktop-extension manifest against its schema only. sunpeak (MIT, about 217 stars) simulates the ChatGPT and Claude runtimes for MCP Apps, and developers write the annotation assertions by hand.

Not found: an open, rule-pack-driven validator that covers Claude, ChatGPT and the IDE and command-line hosts (Cursor, VS Code, Codex, Gemini CLI, Kiro, Windsurf) together, and nothing that publishes the rules as versioned data with sources.

Method and caveats

  • The rules were read from each vendor’s documentation on 1 October 2026. The Codex and Windsurf limits come from a single read of each vendor’s docs and should be re-checked before anyone encodes them as rules.
  • Cursor’s 40-tool limit could not be confirmed in current docs and is left out of the table.
  • No tool was installed or run. The descriptions of Beacon, Manufact and MCPJam reflect what each documents.

I’m Amir Pournasserian. I build AI and platform systems for a living, maintain FluentCMS and YeSvelte, and write here about what I find along the way.