CornuCornu
R&D
August 17, 2026

We shipped the MCP v2 migration. Here's how it actually went.

A follow-up to our MCP scoping post. The architecture bet held, the risk was exactly where we predicted, and one surprise no changelog could have warned us about. What varied from the design, why, and how Claude connects now.

A few weeks ago we wrote up the plan for migrating Cornu's MCP server from the 2025-era SDK to v2 and the new 2026-07-28 spec. That post was all design: what we thought was easy, what we thought was hard, and where we thought the risk lived. This is the follow-up from the other side. We shipped it over a handful of days in August.

The short version: our biggest architectural bet held exactly as we hoped, the migration risk landed precisely where we predicted (the glue we wrote around the old SDK, not the protocol), and we got surprised once in a way that no amount of reading the changelog could have caught. That surprise is the most useful part of this post, so it gets its own section.

What we actually shipped

Two phases mattered for the protocol:

  1. MCP SDK v2 core migration. Swap the package dependencies, rewrite the transport wiring onto the v2 adapter, keep every current client working through a dual-era compatibility path.
  2. Auth hardening (DCR toward CIMD). Adopt Client ID Metadata Documents as a forward registration path, add issuer validation per RFC 9207, key stored client credentials by issuer.

The old @modelcontextprotocol/sdk monolith is gone from the tree entirely. Every current client (Claude, ChatGPT, MCP Inspector) still connects and calls every tool. Now the interesting part: what the plan got right, and what it got wrong.

Where the plan held

The statelessness bet was a genuine no-op. We built the MCP transport stateless from day one, for boring scaling reasons, and the 2026-07-28 spec now mandates that shape. Nothing in our request model had to change to satisfy the headline change of the entire spec revision. This was the whole thesis of the last post, and it survived contact with reality.

The database commit guarantee survived the port verbatim. In the design post we called this the non-negotiable: the hand-rolled finally block that guarantees the transaction is committed or rolled back and the connection returned to the pool, even when the SDK writes raw bytes and the framework's normal response hooks may not fire. We were adamant it had to survive the rewrite untouched. It did. The only line we removed was the per-request teardown that the v2 adapter now owns internally:

} finally {
  // Race control: take ownership of request._dbClient atomically by
  // setting it to undefined BEFORE the await/release, so a double-release
  // can never happen even if the framework hook also fires.
  const dbClient = request._dbClient;
  request._dbClient = undefined;
  if (dbClient) {
    try {
      await dbClient.query(handleOk ? 'COMMIT' : 'ROLLBACK');
    } finally {
      releaseOnce(dbClient); // idempotent belt-and-suspenders
    }
  }
  // No more per-request transport.close()/server.close():
  // createMcpHandler owns that teardown now.
}

The observability fix was exactly the four-line change we scoped. The negotiated protocol version and client capabilities moved into the _meta envelope, so we read them from there and fall back to the old location for clients still on the 2025 wire:

const meta = rpcBody?.params?._meta;
const protocolVersion =
  (meta?.[PROTOCOL_VERSION_META_KEY] as string | undefined) ??
  rpcBody?.params?.protocolVersion ?? null;
const capabilities =
  meta?.[CLIENT_CAPABILITIES_META_KEY] ??
  rpcBody?.params?.capabilities ?? null;

We called this the cheapest high-value change in the whole migration, the one that keeps telemetry honest while everything else waits its turn. It shipped, and our dashboards stayed truthful across the transition.

Where it varied from the design, and why

The package split we described was wrong on paper

The design post said v2 splits the monolith into @modelcontextprotocol/server plus a framework adapter, @modelcontextprotocol/fastify, and that we would wire our route with createMcpHandler from that adapter. When we went to actually install it, the adapter turned out to be inert for our purpose. It does not export createMcpHandler, and it has no path for mounting into an existing framework app like ours. We had read the migration guide and inferred the shape; we had not yet run npm install and read the real package exports.

What actually wires v2 into a route you already own is two packages, not one: createMcpHandler lives in @modelcontextprotocol/server, and the piece that bridges it onto raw Node request and response objects is toNodeHandler() from a third package, @modelcontextprotocol/node. We dropped the framework adapter, never installed it, and built the handler once at module scope:

const mcpHandler = createMcpHandler(() => buildMcpServer(), {
  responseMode: 'json',
});

const nodeHandler = toNodeHandler({
  fetch: (request, options) => resolveMcpResponse(request, mcpHandler.fetch, options),
});

The lesson here is small but it recurs below: a .d.ts file and a migration guide tell you the intended API. They do not tell you which package actually ships it.

The surprise: legacy clients get SSE back, not JSON

This is the one worth the price of admission. Our v1 transport forced a plain JSON response for every client. The v2 adapter's compatibility fallback for 2025-era clients does not carry that setting through: it always answers those clients with a text/event-stream response, a single server-sent-event frame, regardless of the responseMode: 'json' we set. Nothing in the type definitions says this. We found it the only way it can be found, by running a live request through the new handler in a spike and watching seven of eight tests fail with SyntaxError: Unexpected token 'e', "event: mes"... as the JSON parser choked on an SSE frame.

So we wrote an unplanned shim: for legacy requests, unwrap that single SSE frame back into the JSON our older clients expect. Fine. Except the first version of that shim re-introduced the exact failure mode this whole migration was supposed to be immune to.

How the fix for the surprise re-created the original risk

The last post's central claim was that our risk lived in our glue, not the protocol. Here is that claim proving itself the hard way. The naive shim buffered every text/event-stream response with response.text() to unwrap it. That is harmless for the terminal, single-frame legacy path. But a modern client's long-lived subscription stream is also text/event-stream, and buffering it means waiting for a stream that never ends. That blocks the finally block, which means the transaction never commits and the connection never returns to the pool. A handful of concurrent subscriptions and the pool exhausts, every subsequent call hangs, and the load balancer times out at thirty seconds. That is precisely the pool-exhaustion outage the original hand-rolled guard existed to prevent.

Code review caught it before production. The fix was to unwrap only when the request is actually a legacy one, and let modern streams pass straight through:

export async function resolveMcpResponse(request, fetchImpl, options) {
  const legacy = await isLegacyRequest(request);
  const response = await fetchImpl(request, options);
  return legacy ? toJsonResponseIfLegacySse(response) : response;
}

We added a regression test that races a simulated long-lived stream against a timeout, so this specific mistake cannot come back. On our dev deployment, a real Claude subscription later held its connection open for roughly four minutes and completed cleanly, with the pool never breaking a sweat.

How AI assistants communicate now

The compatibility path is the default behavior of the v2 adapter, so Cornu carries no per-client special-casing for the transport itself. Client-era detection happens inside the adapter: a 2025-era client with no _meta envelope is served a single SSE frame, which we unwrap back to JSON to preserve our older wire contract, while a modern 2026-07-28 client is served JSON directly.

On auth, we got a pleasant surprise. We had scoped the DCR to CIMD move treating Client ID Metadata Documents as forward-looking, a draft spec nothing was calling yet, so we built it real but minimal: an SSRF-guarded fetch-validate-provision resolver wired into the authorize and token proxies, with no caching, refresh, or rotation, expecting it to sit advertised but unused while DCR kept doing the actual work. Then the live test surprised us in the good direction. Claude connected through the CIMD path, presenting a client id that is a URL pointing at its metadata document, which our resolver fetched and validated behind the SSRF guard before minting a client from it. ChatGPT connected via DCR. Both registration paths are live in practice, Claude on CIMD and ChatGPT on DCR, and whichever path a client takes, Cornu normalizes it to public PKCE.

Discovery is where the two major clients diverge most. Cornu publishes OAuth authorization-server metadata because our identity provider does not. Claude reads the root discovery path and, notably, ignores the advertised authorization endpoint and posts straight to the authorize route on the resource host. ChatGPT, by contrast, uses path-suffixed discovery, which quietly 404'd until we added aliases for it. Two major clients, two different readings of the same specs.

What we learned doing it

A migration guide describes the API. Only an install describes the packages. We planned against the intended shape and got the package split wrong. Ten minutes with the real node_modules would have corrected it earlier. The same root cause produced the SSE surprise: the type definitions were silent on a behavior that only a live request revealed. Read the docs to plan; run the code to know.

The fix for a surprise can quietly re-create the risk you started with. Our SSE-unwrap shim was a reasonable response to an unexpected behavior, and its first version reintroduced the precise pool-exhaustion outage the migration was meant to make impossible. When you add glue under pressure, hold it to the same standard as the glue you are replacing. The failure modes do not care that the code is new.

And the one from last time, now confirmed rather than predicted: bet on properties, not versions. We did not go stateless because we foresaw this spec. We went stateless because sessions were operational weight we did not want. Betting on the invariant (any instance serves any request) is what let a mandatory protocol change land as a no-op. Design for the property you want, and the standards tend to move toward you.