
In Part 1, we built an MCP Gateway that aggregates tools from multiple .NET microservices behind a single endpoint. It handles discovery, forwarding, auth propagation, and tenant selection — all stored per-session in Redis.
Then we connected ChatGPT as a client, and our session model broke.
The Problem: A New Session Every Tool Call
The MCP spec is clear about sessions. The server returns an Mcp-Session-Id header during initialize, and the client includes it in all subsequent requests. This is how the server knows which session a request belongs to — and it's how we key tenant context (which merchant and channel the user selected) in Redis.
ChatGPT’s MCP connector doesn’t do this. Instead, it creates a fresh session for every tool call. Each invocation triggers a new initialize JSON-RPC message, gets a new Mcp-Session-Id, and throws away the old one. The result:
1. User connects and selects a merchant (stored in Redis under userId:session-A)
2. User asks "show me late orders"
3. ChatGPT creates session-B, calls the tool with Mcp-Session-Id: session-B
4. Gateway looks up userId:session-B in Redis — nothing there
5. Tool returns: "No merchant selected. Call list_merchants first."
The merchant selection is gone. Every tool call starts from scratch.
This isn’t a one-off glitch. Multiple developers have reported it on the OpenAI community forums:
- Connector tool calls generating fresh MCP session each invocation — reported November 2025, noting that sessions previously worked and then regressed
- ChatGPT MCP connector refreshes token on every tool call and doesn’t persist sessions — the connector unconditionally creates a new session even when the existing one is valid
- Clarification around handling of MCP session IDs with ChatGPT
Some MCP servers have worked around this by going fully stateless (stateless_http=True in FastMCP). We can't — tenant selection is inherently stateful for admins.
The Workaround: Dual-Key Caching
The fix is simple: write the context to two Redis keys instead of one. When reading, try the session-specific key first, then fall back to a user-only key.
Writing Context
When a user selects a merchant, we write to both keys:
public async Task SetContextAsync(Guid merchantId, int channelId, string? merchantName = null)
{
var (sessionKey, userOnlyKey) = GetContextKeys();
if (sessionKey == null)
throw new InvalidOperationException("No authenticated user or MCP session.");
var context = new McpUserContext
{
MerchantId = merchantId,
ChannelId = channelId,
MerchantName = merchantName
};
// Write to session-specific key (userId:sessionId)
await cache.CacheSetAsync(context, sessionKey, "mcp:context", ContextTtl);
// Also write to user-only fallback key (userId)
if (userOnlyKey != null && userOnlyKey != sessionKey)
{
await cache.CacheSetAsync(context, userOnlyKey, "mcp:context", ContextTtl);
}
}
Reading Context
When a tool call comes in, try the session key. If it misses (because ChatGPT rotated the session), fall back to the user-only key:
public async Task<McpUserContext?> GetContextAsync()
{
var (sessionKey, userOnlyKey) = GetContextKeys();
if (sessionKey == null)
return null;
// Try session-specific key first
var result = await TryGetFromCacheAsync(sessionKey);
if (result != null)
return result;
// Fall back to user-only key
if (userOnlyKey != null && userOnlyKey != sessionKey)
{
result = await TryGetFromCacheAsync(userOnlyKey);
if (result != null)
{
logger.LogInformation(
"MCP context found via user-only fallback (session key {SessionKey} missed). " +
"Client likely creates new sessions per tool call (e.g. OpenAI ChatGPT).",
sessionKey);
}
}
return result;
}
Key Generation
The key logic returns a tuple — session-specific and user-only:
private (string? sessionKey, string? userOnlyKey) GetContextKeys()
{
var userId = /* extract from claims */;
var sessionId = httpContextAccessor.HttpContext?
.Request.Headers["Mcp-Session-Id"].ToString();
if (string.IsNullOrEmpty(sessionId))
return (userId, userId); // No session header — both keys are the same
return ($"{userId}:{sessionId}", userId);
}
Well-behaved clients like Claude Desktop still get session isolation — their userId:sessionId key hits on the first try. The fallback only activates when the session key misses, which is exactly when ChatGPT has rotated the session out from under us.
Every Tool Call Means Re-Authentication
The session rotation is just the visible symptom. Under the hood, ChatGPT’s connector does a full reset on every tool call:
1. Refresh the OAuth token — even if the current one is valid
2. Send a new initialize request — full MCP handshake, no Mcp-Session-Id from a previous session
3. Re-discover tools — ListTools runs again
4. Execute the actual tool call
That’s four HTTP round-trips per tool invocation where a well-behaved client would need one.
This makes auth server performance critical. If your OAuth token endpoint takes 500ms, your JWKS endpoint takes 200ms, and your MCP initialize + ListTools takes another 300ms, you're adding a full second of overhead to every single tool call — before any business logic runs.
For our gateway, this meant:
- JWKS caching matters. The gateway validates JWTs on every request. The .NET AddJwtBearer middleware caches the JWKS keys by default, but the refresh interval matters when you're getting hit on next tool call.
- Token issuance must be fast. Our OAuth server (MinimalOAuthServer) issues RSA-signed tokens. RSA signing is slower than HMAC, but it only happens once per token — the bigger concern is the round-trip latency to the token endpoint itself.
- Tool discovery repeats every call. The gateway connects to downstream services during startup and caches the tool list. But the MCP initialize + ListTools exchange between ChatGPT and the gateway happens fresh each time. The gateway serves this from memory, so it's fast — but it's still wasted work.
The practical impact: a tool call that takes 200ms with Claude Desktop might take 1.2 seconds with ChatGPT. Not broken, but noticeably sluggish.
No Parallel Tool Calls
There’s one more limitation that compounds the session problem: ChatGPT’s MCP connector does not support parallel tool calls.
With Claude Desktop, if the AI decides it needs to fetch late orders and check stock levels simultaneously, it can issue both tool calls in parallel. The results come back together, and the AI synthesizes them in a single response.
ChatGPT executes tools strictly one at a time. “Show me late orders and low stock products” becomes:
[call os_get_late_orders] → wait → result
[new session + re-auth]
[call ps_get_low_stock_products] → wait → result
Each call pays the full session-creation and re-auth tax. Two tool calls that would take 200ms in parallel with Claude Desktop take 2.4 seconds sequentially with ChatGPT.
This affects how you should design tools when targeting multiple clients:
- Prefer fewer, richer tools over many small ones. A get_merchant_dashboard that returns orders, stock, and revenue in one call is more practical than three separate tools when every call has a 1-second overhead.
- Batch where possible. If you have tools that are frequently called together, consider a combined tool that does both. The trade-off is a less granular tool list, but the user experience improvement is significant.
- Don’t assume parallel execution. Design your tool descriptions so the AI can get useful results from a single call rather than needing to orchestrate multiple calls.
This is a tool design concern, not a server concern. Your gateway and downstream services don’t need to change — but the tools you expose should account for sequential-only clients.
The Trade-off
The dual-key workaround has a cost: it weakens session isolation for affected users.
With the original design, a user could have Claude Desktop connected with merchant A selected and a second chat UI connected with merchant B. Each client had its own Mcp-Session-Id, so each had its own Redis key, and they didn't interfere.
With the fallback key, if that same user connects from ChatGPT, the user-only key (userId) will pick up whichever merchant was selected last — regardless of which client set it. Two ChatGPT sessions for the same user will also share context, since both fall through to the same user-only key.
In practice, this is fine for us. Support agents typically work in one client at a time, and the alternative — ChatGPT being completely unusable because context vanishes between every call — is worse. But it’s worth knowing the trade-off exists.
What OpenAI Should Fix
To be clear: this is a client bug, not a spec ambiguity. The MCP specification is explicit — clients must include the Mcp-Session-Id header in all requests after receiving it from the server. Creating a new session per tool call violates the protocol.
The community has reported this multiple times. The OpenAI forums show the issue appearing, seemingly getting fixed, and then regressing — suggesting it may be a side effect of how ChatGPT’s connector infrastructure handles request isolation internally.
Until it’s fixed, server developers have two choices:
- Go stateless. If you can avoid tying anything to the session, do it. Stateless servers are immune to this bug.
- Dual-key fallback. If you need session state (tenant selection, wizard flows, accumulated context), write it to both a session key and a user key, and fall back to the user key when the session key misses.
We went with option 2.
What’s Next
Part 1 covered building the gateway. This update covered the first real-world compatibility issue we hit.
In Part 2, we’ll look at the other side — building MCP tools in downstream services, integrating with existing service layers, and using the forwarded claims for authorization.