Caching hints and pagination in MCP

Once the Model Context Protocol (MCP) dropped sessions, a small question got bigger: how often should a client ask your server for its tool list? With no long-lived connection in which to learn about a server once, a client would otherwise fetch the same lists over and over. The 2026-07-28 revision answers with explicit caching hints. The part that surprised me sits right next to them: the order your tools come back in can quietly cost your users money.
This is part 6 of my series on what an MCP server does under the 2026-07-28 specification. Part 5 covered _meta, resultType and explicit handles; this part covers the caching hints on a server’s results, a stable order and pagination. The facts are as I read them in October 2026.
In brief
- The results of six methods carry caching hints. Complete results of
server/discoverand the five list and read methods MUST includettlMsandcacheScope; aninput_requiredresult carries none. - A time to live (TTL) caps staleness, and a notification ends a cached copy early.
ttlMstells a client how long it may reuse a result; alist_changednotification tells a subscribed client the moment something changes. - My advice: mark a list
privatewhen it depends on who’s asking. If you filter tools by the caller’s authorization scopes, apubliclist could reach the wrong user through a shared cache. - Keep the order the same on every call. Servers SHOULD return tools in a deterministic order, which keeps client caches reliable and improves prompt-cache hit rates.
- Pagination uses opaque cursors. The server returns
nextCursor, the client sends it back asparams.cursor, and a stable order keeps the page boundaries stable.
What the hints say
The specification’s caching page names six methods whose complete results, those with resultType: "complete", MUST include ttlMs and cacheScope:
server/discover, the method part 3 coverstools/listprompts/listresources/listresources/templates/listresources/read
An interim result with resultType: "input_required" is not cacheable and carries no hints. That matters for resources/read: the Resources page says a server MAY answer it with an input_required result. Extensions follow the same convention; skills/list in the Skills extension is one example.
| Field | Type | Meaning |
|---|---|---|
ttlMs | integer, in milliseconds | A freshness hint: how long the client may reuse this result without asking again |
cacheScope | "public" or "private" | Whether shared intermediaries (proxies, gateways, content delivery networks) may cache the response |
The hints add to the list_changed notifications rather than replace them. A TTL puts a ceiling on staleness for a client that isn’t subscribed, and a notification invalidates the cache at once for a client that is. A client subscribes with subscriptions/listen, which gets its own part later in this series.
Choosing the scope and the TTL
The spec says which results carry the hints; the values are yours. This is how I choose them.
cacheScope
| Situation | Scope |
|---|---|
| The list is identical for every caller | public |
| The list depends on the caller’s scopes or identity | private |
| A resource’s content is user-specific, such as a mailbox or personal files | private |
| Public documentation, static user interface (UI) templates, skill manifests that are the same for everyone | public |
The spec allows a list to vary by authorization, so my advice is this: if your server filters tools by the caller’s authorization scopes, mark its list results private. Otherwise a shared gateway cache could serve one user’s tool list to another. That leak isn’t catastrophic on its own, but it reveals which capabilities exist, and it confuses clients.
ttlMs
The spec sets no ranges, so this table is my rule of thumb. The spec’s own examples sit inside it: they use 5 minutes for lists and 1 hour for server/discover.
| Data | My typical TTL |
|---|---|
server/discover | 1 hour or more |
| Tool, prompt and template lists for a stable deployment | 5 to 60 minutes, shorter if you rely on feature flags |
| Static resources, such as docs and UI templates | Hours |
| Live resources, such as status and metrics | Seconds, or 0 with subscriptions |
In Python
With the official Python software development kit (SDK), mcp 2.3.0, you set the hints once per method with cache_hints on MCPServer. The keys are the six methods above, and each value is a CacheHint.
import anyio
from mcp import Client
from mcp.server import CacheHint, MCPServer
mcp = MCPServer(
"catalog",
cache_hints={
"server/discover": CacheHint(ttl_ms=3_600_000, scope="public"),
"tools/list": CacheHint(ttl_ms=300_000, scope="private"),
},
)
@mcp.tool()
def search_catalog(query: str) -> str:
"""Search the product catalog."""
return f"No results for {query!r}"
async def main():
async with Client(mcp) as client:
tools = await client.list_tools()
prompts = await client.list_prompts() # no hint set for prompts/list
print("tools/list ", tools.ttl_ms, tools.cache_scope)
print("prompts/list", prompts.ttl_ms, prompts.cache_scope)
if __name__ == "__main__":
anyio.run(main)
Run in memory, it prints:
tools/list 300000 private
prompts/list 0 private
The last line is what the SDK does when you set nothing: ttlMs 0, which means stale at once, and cacheScope private (the same default part 3 showed for server/discover). The result stays valid on the wire and is never shared by accident. The SDK also leaves an input_required result without hints, as the spec asks.
On the client side, the SDK’s Client follows the hints by default:
- It keeps an in-memory cache per client and reuses a result until its
ttlMsruns out, capped at 24 hours. - It drops the result when a list-changed notification arrives.
- Passing
cache_mode="refresh"or"bypass"to a call sends it to the server anyway.
A stable order, and pagination
Servers SHOULD return tools from tools/list in a deterministic order: the same order across requests whenever the set of tools hasn’t changed. The spec gives two reasons. Clients can cache the list reliably, and prompt-cache hit rates improve when tool definitions sit in the same position of the model’s context on every turn.
The spec asks this of tools/list. My advice goes further: sort by a stable key (the name, or an explicit display order) before you serialize, and do the same for prompts, resources and skills/list. Never build the list from a dictionary or a reflection call whose order can change between processes. An unstable order silently costs your users money through prompt-cache misses. In the Python SDK, MCPServer lists tools in the order you register them, so with it the advice becomes: register your tools in a fixed order in your code.
The list methods (tools/list, prompts/list, resources/list, resources/templates/list, and extension lists such as skills/list) page with cursors:
- The server returns
nextCursorwhen more results exist. - The client passes it back as
params.cursor. - Cursors are opaque, so a client must not parse or construct them.
Pagination and a stable order go together: page boundaries are only stable if the order underneath them is. My advice: if your catalog changes often, encode a snapshot version in the cursor, so a page fetched mid-change doesn’t skip or repeat items.
MCPServer returns each list in one page, with no nextCursor. To page a large catalog in Python, write the list handler on the SDK’s low-level Server: it receives the cursor in params.cursor and returns next_cursor with each page.
Method and caveats
- Built from my guide, written against the 2026-07-28 specification, with the facts as I read them in October 2026.
- The Python sample was run against
mcp2.3.0 on 8 October 2026 with the SDK’s in-memory client, and the output above is what it printed. The other SDK notes (the client cache, the tool order, single-page lists, low-level paging) come from the installed package’s source and a second run. - Where the specification is more precise than my notes (six methods carry the hints,
server/discoveramong them, and only complete results do), this article follows the specification. - The TTL table is my advice, not the spec’s; the spec’s examples quoted beside it were checked on 8 October 2026.
SeriesWhat an MCP server actually does in 2026Part 6 of 35
- What an MCP server does in 2026: much more than a list of tools
- MCP goes stateless: what changed in the 2026-07-28 specification
- server/discover: the one method every MCP server must implement
- Server instructions: the paragraph every model reads first
- Stateless MCP requests: _meta, resultType and explicit handles
- Caching hints and pagination in MCP
- MCP transports in 2026: stdio and Streamable HTTP