Caching hints and pagination in MCP

Caching hints and pagination in MCP

Once the Model Context Protocol (MCP) dropped sessions, a small question got bigger: how often should a client ask your server for its tool list? With no long-lived connection in which to learn about a server once, a client would otherwise fetch the same lists over and over. The 2026-07-28 revision answers with explicit caching hints. The part that surprised me sits right next to them: the order your tools come back in can quietly cost your users money.

This is part 6 of my series on what an MCP server does under the 2026-07-28 specification. Part 5 covered _meta, resultType and explicit handles; this part covers the caching hints on a server’s results, a stable order and pagination. The facts are as I read them in October 2026.

In brief

  1. The results of six methods carry caching hints. Complete results of server/discover and the five list and read methods MUST include ttlMs and cacheScope; an input_required result carries none.
  2. A time to live (TTL) caps staleness, and a notification ends a cached copy early. ttlMs tells a client how long it may reuse a result; a list_changed notification tells a subscribed client the moment something changes.
  3. My advice: mark a list private when it depends on who’s asking. If you filter tools by the caller’s authorization scopes, a public list could reach the wrong user through a shared cache.
  4. Keep the order the same on every call. Servers SHOULD return tools in a deterministic order, which keeps client caches reliable and improves prompt-cache hit rates.
  5. Pagination uses opaque cursors. The server returns nextCursor, the client sends it back as params.cursor, and a stable order keeps the page boundaries stable.

What the hints say

The specification’s caching page names six methods whose complete results, those with resultType: "complete", MUST include ttlMs and cacheScope:

  • server/discover, the method part 3 covers
  • tools/list
  • prompts/list
  • resources/list
  • resources/templates/list
  • resources/read

An interim result with resultType: "input_required" is not cacheable and carries no hints. That matters for resources/read: the Resources page says a server MAY answer it with an input_required result. Extensions follow the same convention; skills/list in the Skills extension is one example.

FieldTypeMeaning
ttlMsinteger, in millisecondsA freshness hint: how long the client may reuse this result without asking again
cacheScope"public" or "private"Whether shared intermediaries (proxies, gateways, content delivery networks) may cache the response

The hints add to the list_changed notifications rather than replace them. A TTL puts a ceiling on staleness for a client that isn’t subscribed, and a notification invalidates the cache at once for a client that is. A client subscribes with subscriptions/listen, which gets its own part later in this series.

ServerClient cacheReuse the cached list for up to 5 minutesA feature flag enables a new tooltools/listtools + ttlMs 300000 + cacheScope privatesubscriptions/listen (toolsListChanged true)notifications/subscriptions/acknowledgednotifications/tools/list_changedtools/list (cache invalidated early)updated tools + ttlMs 300000

Choosing the scope and the TTL

The spec says which results carry the hints; the values are yours. This is how I choose them.

cacheScope

SituationScope
The list is identical for every callerpublic
The list depends on the caller’s scopes or identityprivate
A resource’s content is user-specific, such as a mailbox or personal filesprivate
Public documentation, static user interface (UI) templates, skill manifests that are the same for everyonepublic

The spec allows a list to vary by authorization, so my advice is this: if your server filters tools by the caller’s authorization scopes, mark its list results private. Otherwise a shared gateway cache could serve one user’s tool list to another. That leak isn’t catastrophic on its own, but it reveals which capabilities exist, and it confuses clients.

ttlMs

The spec sets no ranges, so this table is my rule of thumb. The spec’s own examples sit inside it: they use 5 minutes for lists and 1 hour for server/discover.

DataMy typical TTL
server/discover1 hour or more
Tool, prompt and template lists for a stable deployment5 to 60 minutes, shorter if you rely on feature flags
Static resources, such as docs and UI templatesHours
Live resources, such as status and metricsSeconds, or 0 with subscriptions

In Python

With the official Python software development kit (SDK), mcp 2.3.0, you set the hints once per method with cache_hints on MCPServer. The keys are the six methods above, and each value is a CacheHint.

import anyio
from mcp import Client
from mcp.server import CacheHint, MCPServer

mcp = MCPServer(
    "catalog",
    cache_hints={
        "server/discover": CacheHint(ttl_ms=3_600_000, scope="public"),
        "tools/list": CacheHint(ttl_ms=300_000, scope="private"),
    },
)


@mcp.tool()
def search_catalog(query: str) -> str:
    """Search the product catalog."""
    return f"No results for {query!r}"


async def main():
    async with Client(mcp) as client:
        tools = await client.list_tools()
        prompts = await client.list_prompts()  # no hint set for prompts/list
        print("tools/list  ", tools.ttl_ms, tools.cache_scope)
        print("prompts/list", prompts.ttl_ms, prompts.cache_scope)


if __name__ == "__main__":
    anyio.run(main)

Run in memory, it prints:

tools/list   300000 private
prompts/list 0 private

The last line is what the SDK does when you set nothing: ttlMs 0, which means stale at once, and cacheScope private (the same default part 3 showed for server/discover). The result stays valid on the wire and is never shared by accident. The SDK also leaves an input_required result without hints, as the spec asks.

On the client side, the SDK’s Client follows the hints by default:

  • It keeps an in-memory cache per client and reuses a result until its ttlMs runs out, capped at 24 hours.
  • It drops the result when a list-changed notification arrives.
  • Passing cache_mode="refresh" or "bypass" to a call sends it to the server anyway.

A stable order, and pagination

Servers SHOULD return tools from tools/list in a deterministic order: the same order across requests whenever the set of tools hasn’t changed. The spec gives two reasons. Clients can cache the list reliably, and prompt-cache hit rates improve when tool definitions sit in the same position of the model’s context on every turn.

The spec asks this of tools/list. My advice goes further: sort by a stable key (the name, or an explicit display order) before you serialize, and do the same for prompts, resources and skills/list. Never build the list from a dictionary or a reflection call whose order can change between processes. An unstable order silently costs your users money through prompt-cache misses. In the Python SDK, MCPServer lists tools in the order you register them, so with it the advice becomes: register your tools in a fixed order in your code.

The list methods (tools/list, prompts/list, resources/list, resources/templates/list, and extension lists such as skills/list) page with cursors:

  • The server returns nextCursor when more results exist.
  • The client passes it back as params.cursor.
  • Cursors are opaque, so a client must not parse or construct them.
ServerClientresources/listresources [1..100] + nextCursor "c2"resources/list (cursor "c2")resources [101..200] + nextCursor "c3"resources/list (cursor "c3")resources [201..230] (no nextCursor)

Pagination and a stable order go together: page boundaries are only stable if the order underneath them is. My advice: if your catalog changes often, encode a snapshot version in the cursor, so a page fetched mid-change doesn’t skip or repeat items.

MCPServer returns each list in one page, with no nextCursor. To page a large catalog in Python, write the list handler on the SDK’s low-level Server: it receives the cursor in params.cursor and returns next_cursor with each page.

Method and caveats

  • Built from my guide, written against the 2026-07-28 specification, with the facts as I read them in October 2026.
  • The Python sample was run against mcp 2.3.0 on 8 October 2026 with the SDK’s in-memory client, and the output above is what it printed. The other SDK notes (the client cache, the tool order, single-page lists, low-level paging) come from the installed package’s source and a second run.
  • Where the specification is more precise than my notes (six methods carry the hints, server/discover among them, and only complete results do), this article follows the specification.
  • The TTL table is my advice, not the spec’s; the spec’s examples quoted beside it were checked on 8 October 2026.
SeriesWhat an MCP server actually does in 2026Part 6 of 35
  1. What an MCP server does in 2026: much more than a list of tools
  2. MCP goes stateless: what changed in the 2026-07-28 specification
  3. server/discover: the one method every MCP server must implement
  4. Server instructions: the paragraph every model reads first
  5. Stateless MCP requests: _meta, resultType and explicit handles
  6. Caching hints and pagination in MCP
  7. MCP transports in 2026: stdio and Streamable HTTP

I’m Amir Pournasserian. I build AI and platform systems for a living, maintain FluentCMS and YeSvelte, and write here about what I find along the way.