A local MCP server is a process one developer launches over stdio. A remote MCP server is a product boundary. Multiple clients can discover it, send authenticated requests over the internet, and ask it to touch real systems. The protocol implementation is only one part of the work; the rest is identity, authorization, tenant isolation, timeouts, idempotency, deployment, and operations.
This guide builds that boundary with TypeScript, Next.js Route Handlers, Vercel, mcp-handler 2.x, and the MCP TypeScript SDK v2. The result uses the 2026-07-28 protocol natively, supports older Streamable HTTP clients through the documented stateless fallback, and avoids the SSE-plus-Redis architecture still found in older tutorials.
An MCP server is an authorization-aware API designed for model-selected operations. Treat the model as an untrusted caller with useful context—not as the security boundary.
What changed in MCP in 2026
The 2026-07-28 specification moved the HTTP protocol to a stateless core. Modern clients can discover capabilities and send self-contained requests without a server-managed MCP session. That maps naturally onto horizontally scaled functions because any invocation can handle the next request.
The same revision hardened authorization, introduced clearer per-request metadata, made list responses cacheable, and formalized compatibility and deprecation behavior. Dynamic Client Registration is deprecated in favor of Client ID Metadata Documents, while the old HTTP-plus-SSE transport is on its removal path.
Older examples Current production baseline
@modelcontextprotocol/sdk 1.x @modelcontextprotocol/server 2.x
server.tool(...) server.registerTool(...)
raw Zod field shape z.object({...}) Standard Schema
/sse and /message routes one Streamable HTTP endpoint
Redis-backed MCP sessions stateless request handling
Dynamic Client Registration CIMD-capable authorization server
module-level connection state fresh server context per requestmcp-handler 2.x requires the v2 server package, Zod 4.2 or later, and Node.js 20 or later. Mixing v1 imports with v2 examples produces confusing type and runtime failures.
Choose Streamable HTTP for a remote server
Use stdio when an MCP host launches a local child process. Use Streamable HTTP when independent clients connect to a hosted endpoint. Streamable HTTP supports ordinary request and response behavior plus streaming when a method needs it. A simple tool server should begin stateless and add long-lived behavior only for a concrete requirement.
Stateless does not mean your product has no state. It means the protocol handler does not keep user identity, authorization decisions, or conversation state in one function instance. Durable application state still belongs in a database, object store, queue, or cache with explicit ownership and expiry.
MCP handler -> protocol parsing, tool discovery, validated calls
Identity provider -> login, consent, token issuance, client metadata
Resource server -> token verification, scopes, tenant authorization
Application DB -> durable business data and audit events
Queue or workflow -> work that outlives one HTTP request
Vercel Function -> stateless execution and response streaming
Observability -> redacted metrics, traces, and security eventsInstall the current TypeScript stack
Add mcp-handler, the v2 MCP server package, and Zod. The adapter turns a server definition into a Web-standard Request-to-Response handler that Next.js can mount directly. Pin major versions so a protocol migration does not arrive as an accidental dependency update.
pnpm add mcp-handler@^2 @modelcontextprotocol/server@^2 zod@^4
pnpm add jose
# Required runtime baseline
node --version # v20 or newerThe jose dependency is for JWT verification in the authorization example. If your identity provider publishes an official verifier, use it when it validates issuer, audience, signature, expiry, and scopes correctly. Do not decode JWT payloads without verifying their signature.
Design tools as narrow contracts
A useful tool name is action-oriented, a description tells the model exactly when the operation applies, and the schema constrains every argument. Avoid a generic execute_sql, call_api, or run_command tool. Those interfaces turn one prompt-injection mistake into broad infrastructure access.
Prefer Avoid
get_project(projectId) query_database(sql)
list_incidents(status, limit) call_internal_api(url, body)
create_ticket(title, severity) execute_action(name, args)
archive_report(reportId, reason) run_shell(command)
Every tool should define:
- one business capability
- bounded validated inputs
- the minimum required scope
- object-level authorization
- predictable output and error shapes
- timeout and retry behavior
- an audit event for consequential writesTool annotations such as readOnlyHint or destructiveHint help clients present safer UI, but they are metadata, not enforcement. Your handler must still authenticate the caller, authorize the target object, validate inputs, and require an explicit product policy for destructive operations.
Create the MCP route in Next.js
Start with a read-only project lookup. The server factory is invoked for the serving unit, and the user identity comes from verified request context rather than from a tool argument. The database query includes both the project ID and the caller’s tenant ID so an authenticated user cannot read another tenant by guessing an identifier.
import { createMcpHandler } from 'mcp-handler'
import { z } from 'zod'
import { db } from '@/lib/db'
export const runtime = 'nodejs'
export const maxDuration = 60
const handler = createMcpHandler((server) => {
server.registerTool(
'get_project',
{
title: 'Get project',
description:
'Return one project the authenticated user can access. Use the exact project ID.',
inputSchema: z.object({
projectId: z.string().uuid(),
}),
annotations: {
readOnlyHint: true,
idempotentHint: true,
openWorldHint: false,
},
},
async ({ projectId }, ctx) => {
const tenantId = ctx.http?.authInfo?.extra?.tenantId
if (typeof tenantId !== 'string') {
return {
isError: true,
content: [{ type: 'text', text: 'Not authorized.' }],
}
}
const project = await db.project.findFirst({
where: { id: projectId, tenantId },
select: { id: true, name: true, status: true, updatedAt: true },
})
if (!project) {
return {
isError: true,
content: [{ type: 'text', text: 'Project was not found.' }],
}
}
return {
content: [{ type: 'text', text: JSON.stringify(project) }],
}
},
)
})
export { handler as GET, handler as POST }Do not accept tenantId, userId, accountId, or role from tool arguments when those values express authority. Derive them from verified identity and map them through your own database. A model can supply business input; it cannot declare who it is allowed to act as.
Keep request state out of module globals
Vercel may reuse one function instance for concurrent requests and may discard it between requests. Module-level objects can be useful for immutable configuration, database pools, or cached JWKS keys, but they must not hold the current user, tenant, access token, tool result, or authorization decision.
// Safe: concurrency-aware clients and immutable configuration.
const JWKS = createRemoteJWKSet(new URL(process.env.AUTH_JWKS_URL!))
const RESOURCE_AUDIENCE = process.env.MCP_RESOURCE_AUDIENCE!
// Unsafe: values leak across users when an instance is reused.
let currentUserId: string | undefined
let currentAccessToken: string | undefined
let lastToolResult: unknown
// Keep identity and request data in ctx/http scope instead.If a tool needs resumable work, store a job with a tenant-bound identifier and return that identifier. A later tool can read status after performing the same authorization check. Do not keep an in-memory promise map and assume the same function instance will receive the next call.
Understand the OAuth boundary
Your MCP deployment is normally an OAuth resource server, not the authorization server. It exposes protected tools, publishes metadata describing which issuer can authorize access, challenges unauthenticated requests, and verifies access tokens. The identity platform handles login, user consent, client metadata, and token issuance.
MCP client -> POST /api/mcp without a valid token
MCP server -> 401 + WWW-Authenticate with resource metadata URL
Client -> GET /.well-known/oauth-protected-resource
Metadata -> trusted authorization-server issuer
Client and authorization server -> OAuth flow using CIMD when supported
Client -> retries /api/mcp with Bearer access token
MCP server -> verifies signature, issuer, audience, expiry, and scopes
Tool -> authorizes tenant and target object before doing workA valid token proves an identity and delegated scopes. Every tool still needs to check whether that identity may act on the requested record, workspace, repository, or customer account.
Verify access tokens and attach request identity
mcp-handler provides withMcpAuth to turn a verifier into the correct HTTP challenge behavior. The verifier below checks a JWT against the provider’s JWKS, expected issuer, and MCP resource audience. It extracts scopes and maps the external subject to a local tenant before returning AuthInfo.
import type { AuthInfo } from '@modelcontextprotocol/server'
import { createRemoteJWKSet, jwtVerify } from 'jose'
import { db } from '@/lib/db'
const issuer = process.env.AUTH_ISSUER!
const audience = process.env.MCP_RESOURCE_AUDIENCE!
const JWKS = createRemoteJWKSet(new URL(process.env.AUTH_JWKS_URL!))
export async function verifyMcpToken(
_request: Request,
bearerToken?: string,
): Promise<AuthInfo | undefined> {
if (!bearerToken) return undefined
try {
const { payload } = await jwtVerify(bearerToken, JWKS, {
issuer,
audience,
})
if (!payload.sub) return undefined
const membership = await db.membership.findFirst({
where: { externalSubject: payload.sub, active: true },
select: { userId: true, tenantId: true },
})
if (!membership) return undefined
const scopes = typeof payload.scope === 'string'
? payload.scope.split(' ').filter(Boolean)
: []
return {
token: bearerToken,
clientId: String(payload.client_id ?? payload.azp ?? payload.sub),
scopes,
extra: membership,
}
} catch {
return undefined
}
}Validate audience even when the signature is correct. A token minted for another API should not be accepted by the MCP resource. Validate issuer to prevent authorization-server mix-up, and keep clock skew tolerance small and deliberate.
Wrap the MCP handler with required scopes
Apply authentication before exporting the route. withMcpAuth returns 401 for missing or invalid credentials and 403 when required scopes are absent. Successful AuthInfo becomes available inside each tool through ctx.http.authInfo.
import { createMcpHandler, withMcpAuth } from 'mcp-handler'
import { verifyMcpToken } from '@/lib/mcp-auth'
const handler = createMcpHandler((server) => {
// Register tools here.
})
const authHandler = withMcpAuth(handler, verifyMcpToken, {
required: true,
requiredScopes: ['projects:read'],
resourceMetadataPath: '/.well-known/oauth-protected-resource',
})
export { authHandler as GET, authHandler as POST }A route-wide scope is appropriate when every tool has the same baseline. If tools have different sensitivity, enforce the narrow scope again inside the handler or split capabilities across endpoints. Do not grant a destructive write scope simply because one read tool shares the server.
Publish protected resource metadata
OAuth clients need a standards-based way to discover the authorization server for your resource. Serve RFC 9728 Protected Resource Metadata at the well-known path and list only issuers you trust. That issuer must publish its own authorization-server metadata and support the client-registration approach your clients need.
import {
metadataCorsOptionsRequestHandler,
protectedResourceHandler,
} from 'mcp-handler'
const handler = protectedResourceHandler({
authServerUrls: [process.env.AUTH_ISSUER!],
})
const optionsHandler = metadataCorsOptionsRequestHandler()
export { handler as GET, optionsHandler as OPTIONS }Set the production origin from trusted deployment configuration. Proxy-derived host values can generate incorrect metadata or allow host-header manipulation when an application sits behind another gateway. Test the exact public URL clients will use, including custom domains and preview deployments.
Plan for CIMD instead of building new DCR infrastructure
In the 2026 specification, Client ID Metadata Documents are the preferred direction. An OAuth client identifies itself with an HTTPS URL that hosts its metadata, and the authorization server advertises support through its RFC 8414 metadata. The MCP resource server does not register the client itself.
If you operate the authorization server, implement and validate CIMD there. If you use an identity provider, confirm which MCP clients and registration methods it supports. Dynamic Client Registration remains a compatibility path during the deprecation window, but it should not be the foundation of a new custom system.
Add write tools with idempotency and confirmation
Write tools need stronger contracts than read tools. Require a caller-generated operation ID, persist it with a unique constraint, and return the stored result on replay. This protects against client retries, network ambiguity, and models calling the same operation twice.
server.registerTool(
'create_ticket',
{
title: 'Create support ticket',
description: 'Create one support ticket after the user approves the action.',
inputSchema: z.object({
operationId: z.string().uuid(),
title: z.string().trim().min(4).max(120),
severity: z.enum(['low', 'medium', 'high']),
}),
annotations: {
readOnlyHint: false,
destructiveHint: false,
idempotentHint: true,
openWorldHint: false,
},
},
async (input, ctx) => {
const tenantId = requireTenant(ctx)
const ticket = await createTicketOnce({
tenantId,
operationId: input.operationId,
title: input.title,
severity: input.severity,
})
return {
content: [{ type: 'text', text: JSON.stringify(ticket) }],
}
},
)The client or host should obtain explicit user confirmation before consequential actions, but the server must not assume every host renders that UI. Apply server-side policy for destructive operations, limit the blast radius, and create an audit record containing the subject, tenant, tool, target, decision, and outcome.
Bound every downstream call
One MCP request can fan out to databases, SaaS APIs, model providers, or internal services. Put a timeout on every hop, limit response sizes, and use allowlists for outbound destinations. A user-controlled URL fetcher is an SSRF primitive even when the input arrived through a typed tool schema.
const ALLOWED_HOSTS = new Set(['api.example.com'])
export async function fetchWithBudget(path: string) {
const url = new URL(path, 'https://api.example.com')
if (!ALLOWED_HOSTS.has(url.hostname)) {
throw new Error('Destination is not allowed')
}
const response = await fetch(url, {
signal: AbortSignal.timeout(8_000),
headers: { Accept: 'application/json' },
})
if (!response.ok) throw new Error('Upstream request failed')
const length = Number(response.headers.get('content-length') ?? 0)
if (length > 1_000_000) throw new Error('Upstream response is too large')
return response.json()
}Treat tool output as untrusted data. External documents, tickets, and web pages can contain instructions intended to manipulate the model. Return only the fields required for the task, label external content clearly, and never let retrieved text expand the tool’s authority.
Configure Vercel for the workload you actually have
A normal database-backed tool usually needs seconds, not minutes. Set maxDuration to a realistic ceiling and give every downstream call a shorter budget. Increasing function duration is not a substitute for moving durable work into a queue or workflow system.
Fast read or write tool
-> complete inside the MCP request
-> strict downstream timeouts
-> maxDuration around the measured tail latency
Long document or research job
-> validate and enqueue
-> return a tenant-bound job ID
-> separate status/result tools
-> persist checkpoints outside function memory
Streaming result
-> Streamable HTTP response
-> handle client cancellation
-> still respect the Vercel duration ceilingFluid compute improves concurrency and makes I/O-heavy functions more efficient, but instances remain disposable. Choose a region close to the database, reuse concurrency-safe clients, and expect reconnects or retries when an instance reaches its lifetime or a connection closes.
Separate preview and production authorization
Vercel preview URLs are valuable for testing, but OAuth metadata, redirect URIs, resource audiences, and cookie domains are origin-sensitive. Use separate OAuth clients or issuers for preview and production. Never let a preview deployment accept production audience tokens simply because both builds share code.
Development
resource: http://localhost:3000/api/mcp
issuer: local or dedicated development tenant
Preview
resource: stable preview/testing domain
issuer: non-production tenant and client metadata
Production
resource: https://mcp.example.com/api/mcp
issuer: production authorization server
Never share:
production signing secrets, resource audience, or write credentialsUse a stable custom domain for the production resource. If every deployment generates a new resource URL, client metadata, token audiences, and allowlists become difficult to reason about. Promote code between environments; do not promote identity configuration by copying secrets.
Connect a client to the deployed endpoint
After deployment, the complete Next.js route is the MCP server URL. Modern clients connect to that HTTPS URL with Streamable HTTP. Stdio-only clients can use a trusted remote bridge, but the bridge becomes part of the credential path and should be reviewed accordingly.
{
"mcpServers": {
"fieldnotes-projects": {
"url": "https://mcp.example.com/api/mcp"
}
}
}Use the MCP Inspector and at least two real target clients. Test discovery, OAuth challenge handling, tool schemas, successful calls, denied scopes, expired tokens, malformed inputs, and cancellation. Compatibility claims should come from actual clients, not only from a curl request that returns 200.
Test authorization as an adversary would
Most serious MCP bugs live around the tool rather than inside the protocol parser. Build tests that swap tenant identifiers, replay write operations, use tokens with the wrong audience, and insert hostile instructions into upstream content.
[ ] Missing token receives 401 and a valid metadata challenge
[ ] Invalid signature, issuer, audience, and expiry are rejected
[ ] Missing required scope receives 403
[ ] User A cannot access User B's object by ID
[ ] Client-supplied tenant IDs never determine authority
[ ] Replayed operation IDs return one stored result
[ ] Concurrent writes create one business operation
[ ] Tool inputs reject extra, oversized, and malformed values
[ ] Downstream timeouts finish before the function deadline
[ ] Outbound URLs cannot reach private or unapproved hosts
[ ] External prompt-injection text cannot grant a new capability
[ ] Access tokens and sensitive arguments never enter logs
[ ] Preview tokens fail against production and vice versa
[ ] Modern and supported legacy clients discover the same tools
[ ] Client cancellation stops expensive downstream workMake observability useful without recording credentials
Record protocol era, tool name, request ID, hashed subject identifier, tenant identifier, authorization decision, duration, downstream dependency, result class, and token usage where relevant. Do not record access tokens, authorization headers, secrets, raw customer documents, or unrestricted tool arguments.
Traffic
requests by protocol era, client, and tool
Reliability
p50/p95/p99 duration, timeout rate, cancellation rate, dependency errors
Security
401s, 403s, tenant-boundary denials, SSRF denials, rate-limit events
Product
tool discovery-to-call rate, validation failures, confirmed writes
Cost
function duration, active CPU, database time, external API usage
Never log
Bearer tokens, authorization headers, signing secrets, full user contentUse a correlation ID across the MCP request, audit record, queue message, and downstream calls. Keep security audit events immutable and separate from high-volume debug logs. Retention should reflect the sensitivity of tool inputs and the support obligations of the product.
Migration checklist for older MCP servers
[ ] Replace @modelcontextprotocol/sdk with the v2 packages in use
[ ] Upgrade Zod to 4.2 or later
[ ] Replace server.tool with server.registerTool
[ ] Wrap input schemas with z.object(...)
[ ] Replace /sse and /message with one Streamable HTTP route
[ ] Remove Redis used only for MCP transport sessions
[ ] Create request-scoped server and auth context
[ ] Read auth data from ctx.http.authInfo
[ ] Publish RFC 9728 protected resource metadata
[ ] Confirm the authorization server's CIMD support
[ ] Validate issuer and resource audience, not only signature
[ ] Retest every client against the production URL
[ ] Keep the legacy fallback only for clients you intentionally support
[ ] Delete compatibility code after its documented retirement dateProduction checklist
[ ] mcp-handler 2.x and MCP TypeScript SDK v2 are version-pinned
[ ] One Streamable HTTP endpoint serves the intended protocol eras
[ ] Every tool has a narrow capability and strict schema
[ ] No generic SQL, shell, URL fetch, or arbitrary API tools exist
[ ] OAuth resource metadata uses the public production origin
[ ] Tokens validate signature, issuer, audience, expiry, and scopes
[ ] Tenant and user identity come only from verified auth context
[ ] Every object lookup includes tenant authorization
[ ] Write tools require replay-safe operation IDs
[ ] Consequential operations produce immutable audit events
[ ] Outbound hosts, response sizes, and durations are bounded
[ ] Long-running work moves to a durable queue or workflow
[ ] Function duration and region match measured dependencies
[ ] Preview and production use separate identity configuration
[ ] Logs and traces redact credentials and sensitive content
[ ] Rate limits apply by subject, tenant, client, and expensive tool
[ ] Inspector plus real-client compatibility tests pass
[ ] Abuse response and token-revocation procedures have ownersThe implementation principle
The simplest remote MCP server is one stateless HTTPS endpoint backed by narrow tools. The safest one derives identity from a verified token, derives tenant access from server-side data, and treats every tool call as an API authorization decision. Vercel can scale that endpoint; it cannot decide what the caller should be allowed to do.
Build the transport with mcp-handler and the current TypeScript SDK, but spend most of the design effort on the contracts around it: explicit scopes, object authorization, idempotent writes, bounded dependencies, durable long-running work, and redacted evidence. That is what turns an MCP demo into infrastructure another team can trust.
A tool schema controls what a model may ask for. Verified identity plus server-side authorization controls what the system may actually do. Production MCP needs both.
