CornuCornu
R&D
August 4, 2026

MCP went stateless, our MCP already was

The 2026-07-28 MCP revision makes servers stateless. Ours already was. Here's what's easy, what's hard, and what we learned scoping the migration.

There's a new MCP, and it's the biggest change since auth

On 2026-07-28, the Model Context Protocol shipped a new spec revision, and it's the largest change to the protocol since authorization was added. If you run an MCP server, here are the things to actually be aware of.

MCP is now stateless at its core. The initialize / notifications/initialized handshake is gone. The Mcp-Session-Id header is gone. Every request now carries its own protocol version and client capabilities in _meta, which means any server instance can serve any request behind a plain round-robin load balancer, with no session affinity and no sticky routing.

A few more that will touch most servers:

  • server/discover is a new RPC that servers must implement to advertise supported versions, capabilities, and identity.
  • Server-initiated requests are gone. Sampling, elicitation, and roots used to let the server call back to the client mid-request. That's replaced by the Multi-Round-Trip Request (MRTR) pattern: the server returns an input_required result, and the client retries with the extra input inline.
  • List results are now cacheable. tools/list, resources/list, etc. carry ttlMs and cacheScope, and no longer vary per connection.
  • Dynamic Client Registration (RFC 7591) is deprecated in favor of Client ID Metadata Documents (CIMD), alongside a batch of OAuth hardening (iss validation per RFC 9207, issuer-keyed credentials).
  • Roots, Sampling, and Logging are deprecated, and the legacy HTTP+SSE transport is finally reclassified as deprecated too.

About "EOL" of older versions: it's softer than a cutoff date, and that's deliberate. The 2026-07-28 release also introduced a formal feature lifecycle policy with a minimum twelve-month deprecation window. Deprecated features keep working for at least a year; the official v1 SDKs keep getting bug and security fixes for at least six months past v2. So nothing you built yesterday breaks tomorrow. The clock is real, but it's a clock, not a wall. The right posture is "plan the upgrade," not "scramble."

Our server design, and why this change is half-easy, half-hard

Cornu connects a company's information to AI tools over MCP. Our MCP endpoint is a Fastify route running the SDK's Streamable HTTP transport. Two design decisions from its original build turn out to matter a lot now.

We built it stateless from day one. To be precise about what "stateless" means here: it's the MCP transport, not the product. Cornu's documents, permissions, knowledge areas, and audit records live in durable storage exactly as before. What we never introduced was MCP session state. Our /mcp handler issues no session ID and builds a fresh MCP server per request:

// One MCP server + one transport, per request. No session store.
const server = buildMcpServer();
const transport = new StreamableHTTPServerTransport({
  sessionIdGenerator: undefined, // no session id issued or stored
  enableJsonResponse: true,      // POST -> JSON response, not SSE
});
await server.connect(transport);

We did this for boring scaling reasons: no session store to run, any task can serve any request, cheap horizontal scale. The 2026-07-28 spec now mandates exactly this shape. Our biggest architectural bet aligned with the protocol's headline change before the protocol made it.

Everything is per-request, with context flowing out-of-band. We attach the authenticated principal and a transaction-scoped database context to each request via async-local storage, so tool handlers never prop-drill it:

await runWithMcpToolCtx(ctx, async () => {
  await transport.handleRequest(request.raw, reply.raw, request.body);
});

And we validate Origin before any tool work, per the spec's DNS-rebinding warning:

if (!isAllowedOrigin(request.headers.origin)) {
  return reply.code(403).send({ error: 'origin_not_allowed' });
}

Stateless doesn't mean trustless. Removing the MCP session means every request has to stand on its own, so we authenticate each one and re-establish tenant and user context before any authorization or data access. If request #1 lands on one instance and request #2 on another, the second instance derives authorization from the request itself, never from something the first instance remembered. The application state that matters (documents, permissions, knowledge areas, audit records) lives in durable storage and is never derived from MCP transport state. An identifier the model hands us, a document_id or a knowledge_area_id, is a lookup key, never proof that the caller is allowed to see the thing it names.

What makes the migration easy

  1. We're already stateless. The headline change is a no-op for our architecture.
  2. The runtime prerequisites are already paid. The v2 SDK needs Node 20+ and zod 4. We're on both. The single most expensive part of most SDK major bumps (a zod 3→4 sweep across the codebase) is behind us.
  3. A dual-era shim keeps old clients working. The v2 Fastify adapter serves 2026-07-28 and 2025-era clients from one endpoint, so we don't strand the clients negotiating older versions today.

What makes it hard

  1. It's a package split, not a version bump. This is the trap. npm update gets you nothing: the @modelcontextprotocol/sdk latest tag is still on the old spec. The new spec lives in v2, which replaced the monolith with @modelcontextprotocol/server + @modelcontextprotocol/client and framework adapters (@modelcontextprotocol/fastify, /node, /hono). Migrating means swapping dependencies and rewriting the transport wiring, not bumping a number.

  2. The riskiest code is the glue we wrote around the old SDK. To make per-request servers work under HTTP keep-alive, we hand-rolled the transport lifecycle, including a guard that guarantees the database transaction is committed/rolled back and the connection returned to the pool, even when the SDK writes raw bytes and Fastify's normal response hooks don't fire:

} finally {
  // If we don't guarantee release here, the pg pool exhausts after a
  // handful of requests and every subsequent /mcp call hangs until the
  // load balancer times out. Ownership is taken atomically so a
  // double-release can never happen.
  const dbClient = request._dbClient;
  request._dbClient = undefined;
  if (dbClient) {
    try { await dbClient.query(ok ? 'COMMIT' : 'ROLLBACK'); }
    finally { releaseOnce(dbClient); }
  }
  await Promise.allSettled([transport.close(), server.close()]);
}

Porting this, not the happy path, is where the migration risk concentrates.

  1. Our auth path leans on a now-deprecated primitive. We built Dynamic Client Registration as a hard requirement (it's how coding-agent clients self-register). CIMD is its own project with its own OAuth-hardening checklist.

  2. The _meta move breaks our observability quietly. We log the negotiated protocol version and capabilities off the initialize params. Those fields moved to _meta. Under a new-spec client, that log line silently goes null. No error, just blind.

How we're approaching the significant changes

API names below reflect the v2 SDK's migration guide and are illustrative until the change lands.

Transport wiring: hand-rolled lifecycle → createMcpHandler

The v2 model is a factory that builds a fresh server per request, which is conceptually what we already do by hand. The rewrite collapses our connect/hijack/close dance into the adapter.

Before:

const server = buildMcpServer();
const transport = new StreamableHTTPServerTransport({
  sessionIdGenerator: undefined,
  enableJsonResponse: true,
});
await server.connect(transport);
reply.hijack();
try {
  await runWithMcpToolCtx(ctx, () =>
    transport.handleRequest(request.raw, reply.raw, request.body),
  );
} finally {
  // ...manual COMMIT/ROLLBACK + pool release + transport.close()
}

After (planned):

import { createMcpHandler } from '@modelcontextprotocol/fastify';

// The adapter owns per-request server construction and teardown.
// Our factory stays; our context wrapper stays; the lifecycle glue goes.
const mcpHandler = createMcpHandler(() => buildMcpServer(), {
  // dual-era: also serve 2025-era stateless clients
  onRequest: (req) => runWithMcpToolCtx(buildCtx(req), () => {}),
});

The non-negotiable during this rewrite: the DB commit/release guarantee has to survive the port. If the adapter's teardown doesn't give us a hook we trust, we keep our own finally around it. The pool-exhaustion failure mode is not something we'll rediscover in production.

Observability: read the version from _meta, with a fallback

Small change, high value: it keeps our dashboards honest across the transition:

Before:

protocol_version: rpcBody?.params?.protocolVersion ?? null,
capabilities:     rpcBody?.params?.capabilities ?? null,

After:

const meta = rpcBody?.params?._meta ?? {};
protocol_version:
  meta['io.modelcontextprotocol/protocolVersion'] ??
  rpcBody?.params?.protocolVersion ?? null,   // fall back to pre-2026-07-28 clients
capabilities:
  meta['io.modelcontextprotocol/clientCapabilities'] ??
  rpcBody?.params?.capabilities ?? null,

Server-initiated calls: elicitation → MRTR

We'd been eyeing elicitation for a file-upload flow. That primitive is now the MRTR input_required pattern. The nice part: handlers written the new way are auto-shimmed onto old wires, so we write once:

// Instead of a server-to-client elicitation request mid-handler:
return inputRequired({
  inputRequests: [/* the extra info we need from the client */],
});
// The client retries the original call with inputResponses inline.

Auth: DCR → CIMD

This one we're scoping as its own phase, not folding into the SDK swap. It touches the OAuth/registration path directly: adopt Client ID Metadata Documents, validate iss on authorization responses, key persisted client credentials by issuer, set application_type on registration. DCR stays as a backwards-compat fallback during the deprecation window.

Lessons from scoping the change

Verify what you actually built against. Going in, we "knew" two things that were both wrong: that we'd built against the Nov 2025 revision (we were on 2025-06-18), and that upgrading meant a clean SDK major bump (it's a package split). We pin the SDK and echo whatever version clients negotiate, so the "version we're on" wasn't a constant in our code at all. Ten minutes reading our own package.json and endpoint corrected a week's worth of wrong planning.

Architecture bets compound in your favor, if you bet on properties, not versions. We didn't go stateless because we predicted this spec. We went stateless because sessions are operational weight we didn't want. Betting on the property (any instance serves any request) rather than a protocol feature is what made a mandatory protocol change land as a no-op. That's the generalizable lesson: design for the invariant you want, and standards tend to move toward you.

The migration risk is in your glue, not the protocol. The spec diff is large but SDK-absorbed. Our genuine risk surface is the ~30 lines of lifecycle code we wrote to work around the old SDK's quirks: the pool-release guard, the keep-alive teardown. Audit the workarounds you wrote around a dependency before you audit the dependency's changelog.

Deprecation is a clock, not a wall, so instrument the transition. With a twelve-month window and version-echoing, nothing breaks on day one. The real near-term risk isn't an outage; it's your observability going blind the moment real clients start negotiating the new version. The cheapest high-value change in the whole migration is the four-line _meta fallback that keeps your telemetry truthful while everything else waits its turn.