DEV FIELDNOTES
AI infrastructure field guide 012Updated September 28, 2026

Build a Remote MCP Server with TypeScript and Vercel: OAuth, Streamable HTTP, and Production Deployment

Build and deploy a production remote MCP server with TypeScript, Next.js, Vercel, Streamable HTTP, OAuth, tenant isolation, and safe tools.

A secure remote MCP server connecting multiple AI clients through an authorization gateway to isolated tools, APIs, and databases

A local MCP server is a process one developer launches over stdio. A remote MCP server is a product boundary. Multiple clients can discover it, send authenticated requests over the internet, and ask it to touch real systems. The protocol implementation is only one part of the work; the rest is identity, authorization, tenant isolation, timeouts, idempotency, deployment, and operations.

This guide builds that boundary with TypeScript, Next.js Route Handlers, Vercel, mcp-handler 2.x, and the MCP TypeScript SDK v2. The result uses the 2026-07-28 protocol natively, supports older Streamable HTTP clients through the documented stateless fallback, and avoids the SSE-plus-Redis architecture still found in older tutorials.

The production mental model

An MCP server is an authorization-aware API designed for model-selected operations. Treat the model as an untrusted caller with useful context—not as the security boundary.

What changed in MCP in 2026

The 2026-07-28 specification moved the HTTP protocol to a stateless core. Modern clients can discover capabilities and send self-contained requests without a server-managed MCP session. That maps naturally onto horizontally scaled functions because any invocation can handle the next request.

The same revision hardened authorization, introduced clearer per-request metadata, made list responses cacheable, and formalized compatibility and deprecation behavior. Dynamic Client Registration is deprecated in favor of Client ID Metadata Documents, while the old HTTP-plus-SSE transport is on its removal path.

Old tutorials versus the current stack
Older examples                     Current production baseline
@modelcontextprotocol/sdk 1.x        @modelcontextprotocol/server 2.x
server.tool(...)                    server.registerTool(...)
raw Zod field shape                 z.object({...}) Standard Schema
/sse and /message routes            one Streamable HTTP endpoint
Redis-backed MCP sessions           stateless request handling
Dynamic Client Registration         CIMD-capable authorization server
module-level connection state       fresh server context per request
Version-match the tutorial before copying code

mcp-handler 2.x requires the v2 server package, Zod 4.2 or later, and Node.js 20 or later. Mixing v1 imports with v2 examples produces confusing type and runtime failures.

Choose Streamable HTTP for a remote server

Use stdio when an MCP host launches a local child process. Use Streamable HTTP when independent clients connect to a hosted endpoint. Streamable HTTP supports ordinary request and response behavior plus streaming when a method needs it. A simple tool server should begin stateless and add long-lived behavior only for a concrete requirement.

Stateless does not mean your product has no state. It means the protocol handler does not keep user identity, authorization decisions, or conversation state in one function instance. Durable application state still belongs in a database, object store, queue, or cache with explicit ownership and expiry.

Responsibility map
MCP handler       -> protocol parsing, tool discovery, validated calls
Identity provider  -> login, consent, token issuance, client metadata
Resource server    -> token verification, scopes, tenant authorization
Application DB     -> durable business data and audit events
Queue or workflow  -> work that outlives one HTTP request
Vercel Function    -> stateless execution and response streaming
Observability      -> redacted metrics, traces, and security events

Install the current TypeScript stack

Add mcp-handler, the v2 MCP server package, and Zod. The adapter turns a server definition into a Web-standard Request-to-Response handler that Next.js can mount directly. Pin major versions so a protocol migration does not arrive as an accidental dependency update.

Terminal
pnpm add mcp-handler@^2 @modelcontextprotocol/server@^2 zod@^4
pnpm add jose

# Required runtime baseline
node --version  # v20 or newer

The jose dependency is for JWT verification in the authorization example. If your identity provider publishes an official verifier, use it when it validates issuer, audience, signature, expiry, and scopes correctly. Do not decode JWT payloads without verifying their signature.

Design tools as narrow contracts

A useful tool name is action-oriented, a description tells the model exactly when the operation applies, and the schema constrains every argument. Avoid a generic execute_sql, call_api, or run_command tool. Those interfaces turn one prompt-injection mistake into broad infrastructure access.

Tool contract rules
Prefer                              Avoid
get_project(projectId)              query_database(sql)
list_incidents(status, limit)       call_internal_api(url, body)
create_ticket(title, severity)      execute_action(name, args)
archive_report(reportId, reason)    run_shell(command)

Every tool should define:
- one business capability
- bounded validated inputs
- the minimum required scope
- object-level authorization
- predictable output and error shapes
- timeout and retry behavior
- an audit event for consequential writes

Tool annotations such as readOnlyHint or destructiveHint help clients present safer UI, but they are metadata, not enforcement. Your handler must still authenticate the caller, authorize the target object, validate inputs, and require an explicit product policy for destructive operations.

Create the MCP route in Next.js

Start with a read-only project lookup. The server factory is invoked for the serving unit, and the user identity comes from verified request context rather than from a tool argument. The database query includes both the project ID and the caller’s tenant ID so an authenticated user cannot read another tenant by guessing an identifier.

app/api/mcp/route.ts
import { createMcpHandler } from 'mcp-handler'
import { z } from 'zod'
import { db } from '@/lib/db'

export const runtime = 'nodejs'
export const maxDuration = 60

const handler = createMcpHandler((server) => {
  server.registerTool(
    'get_project',
    {
      title: 'Get project',
      description:
        'Return one project the authenticated user can access. Use the exact project ID.',
      inputSchema: z.object({
        projectId: z.string().uuid(),
      }),
      annotations: {
        readOnlyHint: true,
        idempotentHint: true,
        openWorldHint: false,
      },
    },
    async ({ projectId }, ctx) => {
      const tenantId = ctx.http?.authInfo?.extra?.tenantId
      if (typeof tenantId !== 'string') {
        return {
          isError: true,
          content: [{ type: 'text', text: 'Not authorized.' }],
        }
      }

      const project = await db.project.findFirst({
        where: { id: projectId, tenantId },
        select: { id: true, name: true, status: true, updatedAt: true },
      })

      if (!project) {
        return {
          isError: true,
          content: [{ type: 'text', text: 'Project was not found.' }],
        }
      }

      return {
        content: [{ type: 'text', text: JSON.stringify(project) }],
      }
    },
  )
})

export { handler as GET, handler as POST }

Do not accept tenantId, userId, accountId, or role from tool arguments when those values express authority. Derive them from verified identity and map them through your own database. A model can supply business input; it cannot declare who it is allowed to act as.

Keep request state out of module globals

Vercel may reuse one function instance for concurrent requests and may discard it between requests. Module-level objects can be useful for immutable configuration, database pools, or cached JWKS keys, but they must not hold the current user, tenant, access token, tool result, or authorization decision.

Safe versus unsafe module state
// Safe: concurrency-aware clients and immutable configuration.
const JWKS = createRemoteJWKSet(new URL(process.env.AUTH_JWKS_URL!))
const RESOURCE_AUDIENCE = process.env.MCP_RESOURCE_AUDIENCE!

// Unsafe: values leak across users when an instance is reused.
let currentUserId: string | undefined
let currentAccessToken: string | undefined
let lastToolResult: unknown

// Keep identity and request data in ctx/http scope instead.

If a tool needs resumable work, store a job with a tenant-bound identifier and return that identifier. A later tool can read status after performing the same authorization check. Do not keep an in-memory promise map and assume the same function instance will receive the next call.

Understand the OAuth boundary

Your MCP deployment is normally an OAuth resource server, not the authorization server. It exposes protected tools, publishes metadata describing which issuer can authorize access, challenges unauthenticated requests, and verifies access tokens. The identity platform handles login, user consent, client metadata, and token issuance.

Authorization flow
MCP client -> POST /api/mcp without a valid token
MCP server -> 401 + WWW-Authenticate with resource metadata URL
Client -> GET /.well-known/oauth-protected-resource
Metadata -> trusted authorization-server issuer
Client and authorization server -> OAuth flow using CIMD when supported
Client -> retries /api/mcp with Bearer access token
MCP server -> verifies signature, issuer, audience, expiry, and scopes
Tool -> authorizes tenant and target object before doing work
Authentication is not object authorization

A valid token proves an identity and delegated scopes. Every tool still needs to check whether that identity may act on the requested record, workspace, repository, or customer account.

Verify access tokens and attach request identity

mcp-handler provides withMcpAuth to turn a verifier into the correct HTTP challenge behavior. The verifier below checks a JWT against the provider’s JWKS, expected issuer, and MCP resource audience. It extracts scopes and maps the external subject to a local tenant before returning AuthInfo.

lib/mcp-auth.ts
import type { AuthInfo } from '@modelcontextprotocol/server'
import { createRemoteJWKSet, jwtVerify } from 'jose'
import { db } from '@/lib/db'

const issuer = process.env.AUTH_ISSUER!
const audience = process.env.MCP_RESOURCE_AUDIENCE!
const JWKS = createRemoteJWKSet(new URL(process.env.AUTH_JWKS_URL!))

export async function verifyMcpToken(
  _request: Request,
  bearerToken?: string,
): Promise<AuthInfo | undefined> {
  if (!bearerToken) return undefined

  try {
    const { payload } = await jwtVerify(bearerToken, JWKS, {
      issuer,
      audience,
    })

    if (!payload.sub) return undefined

    const membership = await db.membership.findFirst({
      where: { externalSubject: payload.sub, active: true },
      select: { userId: true, tenantId: true },
    })
    if (!membership) return undefined

    const scopes = typeof payload.scope === 'string'
      ? payload.scope.split(' ').filter(Boolean)
      : []

    return {
      token: bearerToken,
      clientId: String(payload.client_id ?? payload.azp ?? payload.sub),
      scopes,
      extra: membership,
    }
  } catch {
    return undefined
  }
}

Validate audience even when the signature is correct. A token minted for another API should not be accepted by the MCP resource. Validate issuer to prevent authorization-server mix-up, and keep clock skew tolerance small and deliberate.

Wrap the MCP handler with required scopes

Apply authentication before exporting the route. withMcpAuth returns 401 for missing or invalid credentials and 403 when required scopes are absent. Successful AuthInfo becomes available inside each tool through ctx.http.authInfo.

app/api/mcp/route.ts — authorization wrapper
import { createMcpHandler, withMcpAuth } from 'mcp-handler'
import { verifyMcpToken } from '@/lib/mcp-auth'

const handler = createMcpHandler((server) => {
  // Register tools here.
})

const authHandler = withMcpAuth(handler, verifyMcpToken, {
  required: true,
  requiredScopes: ['projects:read'],
  resourceMetadataPath: '/.well-known/oauth-protected-resource',
})

export { authHandler as GET, authHandler as POST }

A route-wide scope is appropriate when every tool has the same baseline. If tools have different sensitivity, enforce the narrow scope again inside the handler or split capabilities across endpoints. Do not grant a destructive write scope simply because one read tool shares the server.

Publish protected resource metadata

OAuth clients need a standards-based way to discover the authorization server for your resource. Serve RFC 9728 Protected Resource Metadata at the well-known path and list only issuers you trust. That issuer must publish its own authorization-server metadata and support the client-registration approach your clients need.

app/.well-known/oauth-protected-resource/route.ts
import {
  metadataCorsOptionsRequestHandler,
  protectedResourceHandler,
} from 'mcp-handler'

const handler = protectedResourceHandler({
  authServerUrls: [process.env.AUTH_ISSUER!],
})

const optionsHandler = metadataCorsOptionsRequestHandler()

export { handler as GET, optionsHandler as OPTIONS }

Set the production origin from trusted deployment configuration. Proxy-derived host values can generate incorrect metadata or allow host-header manipulation when an application sits behind another gateway. Test the exact public URL clients will use, including custom domains and preview deployments.

Plan for CIMD instead of building new DCR infrastructure

In the 2026 specification, Client ID Metadata Documents are the preferred direction. An OAuth client identifies itself with an HTTPS URL that hosts its metadata, and the authorization server advertises support through its RFC 8414 metadata. The MCP resource server does not register the client itself.

If you operate the authorization server, implement and validate CIMD there. If you use an identity provider, confirm which MCP clients and registration methods it supports. Dynamic Client Registration remains a compatibility path during the deprecation window, but it should not be the foundation of a new custom system.

Add write tools with idempotency and confirmation

Write tools need stronger contracts than read tools. Require a caller-generated operation ID, persist it with a unique constraint, and return the stored result on replay. This protects against client retries, network ambiguity, and models calling the same operation twice.

A replay-safe write tool
server.registerTool(
  'create_ticket',
  {
    title: 'Create support ticket',
    description: 'Create one support ticket after the user approves the action.',
    inputSchema: z.object({
      operationId: z.string().uuid(),
      title: z.string().trim().min(4).max(120),
      severity: z.enum(['low', 'medium', 'high']),
    }),
    annotations: {
      readOnlyHint: false,
      destructiveHint: false,
      idempotentHint: true,
      openWorldHint: false,
    },
  },
  async (input, ctx) => {
    const tenantId = requireTenant(ctx)

    const ticket = await createTicketOnce({
      tenantId,
      operationId: input.operationId,
      title: input.title,
      severity: input.severity,
    })

    return {
      content: [{ type: 'text', text: JSON.stringify(ticket) }],
    }
  },
)

The client or host should obtain explicit user confirmation before consequential actions, but the server must not assume every host renders that UI. Apply server-side policy for destructive operations, limit the blast radius, and create an audit record containing the subject, tenant, tool, target, decision, and outcome.

Bound every downstream call

One MCP request can fan out to databases, SaaS APIs, model providers, or internal services. Put a timeout on every hop, limit response sizes, and use allowlists for outbound destinations. A user-controlled URL fetcher is an SSRF primitive even when the input arrived through a typed tool schema.

lib/fetch-with-budget.ts
const ALLOWED_HOSTS = new Set(['api.example.com'])

export async function fetchWithBudget(path: string) {
  const url = new URL(path, 'https://api.example.com')
  if (!ALLOWED_HOSTS.has(url.hostname)) {
    throw new Error('Destination is not allowed')
  }

  const response = await fetch(url, {
    signal: AbortSignal.timeout(8_000),
    headers: { Accept: 'application/json' },
  })

  if (!response.ok) throw new Error('Upstream request failed')

  const length = Number(response.headers.get('content-length') ?? 0)
  if (length > 1_000_000) throw new Error('Upstream response is too large')

  return response.json()
}

Treat tool output as untrusted data. External documents, tickets, and web pages can contain instructions intended to manipulate the model. Return only the fields required for the task, label external content clearly, and never let retrieved text expand the tool’s authority.

Configure Vercel for the workload you actually have

A normal database-backed tool usually needs seconds, not minutes. Set maxDuration to a realistic ceiling and give every downstream call a shorter budget. Increasing function duration is not a substitute for moving durable work into a queue or workflow system.

Vercel execution choices
Fast read or write tool
  -> complete inside the MCP request
  -> strict downstream timeouts
  -> maxDuration around the measured tail latency

Long document or research job
  -> validate and enqueue
  -> return a tenant-bound job ID
  -> separate status/result tools
  -> persist checkpoints outside function memory

Streaming result
  -> Streamable HTTP response
  -> handle client cancellation
  -> still respect the Vercel duration ceiling

Fluid compute improves concurrency and makes I/O-heavy functions more efficient, but instances remain disposable. Choose a region close to the database, reuse concurrency-safe clients, and expect reconnects or retries when an instance reaches its lifetime or a connection closes.

Separate preview and production authorization

Vercel preview URLs are valuable for testing, but OAuth metadata, redirect URIs, resource audiences, and cookie domains are origin-sensitive. Use separate OAuth clients or issuers for preview and production. Never let a preview deployment accept production audience tokens simply because both builds share code.

Environment boundary
Development
  resource: http://localhost:3000/api/mcp
  issuer: local or dedicated development tenant

Preview
  resource: stable preview/testing domain
  issuer: non-production tenant and client metadata

Production
  resource: https://mcp.example.com/api/mcp
  issuer: production authorization server

Never share:
  production signing secrets, resource audience, or write credentials

Use a stable custom domain for the production resource. If every deployment generates a new resource URL, client metadata, token audiences, and allowlists become difficult to reason about. Promote code between environments; do not promote identity configuration by copying secrets.

Connect a client to the deployed endpoint

After deployment, the complete Next.js route is the MCP server URL. Modern clients connect to that HTTPS URL with Streamable HTTP. Stdio-only clients can use a trusted remote bridge, but the bridge becomes part of the credential path and should be reviewed accordingly.

Client configuration
{
  "mcpServers": {
    "fieldnotes-projects": {
      "url": "https://mcp.example.com/api/mcp"
    }
  }
}

Use the MCP Inspector and at least two real target clients. Test discovery, OAuth challenge handling, tool schemas, successful calls, denied scopes, expired tokens, malformed inputs, and cancellation. Compatibility claims should come from actual clients, not only from a curl request that returns 200.

Test authorization as an adversary would

Most serious MCP bugs live around the tool rather than inside the protocol parser. Build tests that swap tenant identifiers, replay write operations, use tokens with the wrong audience, and insert hostile instructions into upstream content.

Production test matrix
[ ] Missing token receives 401 and a valid metadata challenge
[ ] Invalid signature, issuer, audience, and expiry are rejected
[ ] Missing required scope receives 403
[ ] User A cannot access User B's object by ID
[ ] Client-supplied tenant IDs never determine authority
[ ] Replayed operation IDs return one stored result
[ ] Concurrent writes create one business operation
[ ] Tool inputs reject extra, oversized, and malformed values
[ ] Downstream timeouts finish before the function deadline
[ ] Outbound URLs cannot reach private or unapproved hosts
[ ] External prompt-injection text cannot grant a new capability
[ ] Access tokens and sensitive arguments never enter logs
[ ] Preview tokens fail against production and vice versa
[ ] Modern and supported legacy clients discover the same tools
[ ] Client cancellation stops expensive downstream work

Make observability useful without recording credentials

Record protocol era, tool name, request ID, hashed subject identifier, tenant identifier, authorization decision, duration, downstream dependency, result class, and token usage where relevant. Do not record access tokens, authorization headers, secrets, raw customer documents, or unrestricted tool arguments.

Recommended operational signals
Traffic
  requests by protocol era, client, and tool

Reliability
  p50/p95/p99 duration, timeout rate, cancellation rate, dependency errors

Security
  401s, 403s, tenant-boundary denials, SSRF denials, rate-limit events

Product
  tool discovery-to-call rate, validation failures, confirmed writes

Cost
  function duration, active CPU, database time, external API usage

Never log
  Bearer tokens, authorization headers, signing secrets, full user content

Use a correlation ID across the MCP request, audit record, queue message, and downstream calls. Keep security audit events immutable and separate from high-volume debug logs. Retention should reflect the sensitivity of tool inputs and the support obligations of the product.

Migration checklist for older MCP servers

Moving from SDK v1 and old Vercel adapters
[ ] Replace @modelcontextprotocol/sdk with the v2 packages in use
[ ] Upgrade Zod to 4.2 or later
[ ] Replace server.tool with server.registerTool
[ ] Wrap input schemas with z.object(...)
[ ] Replace /sse and /message with one Streamable HTTP route
[ ] Remove Redis used only for MCP transport sessions
[ ] Create request-scoped server and auth context
[ ] Read auth data from ctx.http.authInfo
[ ] Publish RFC 9728 protected resource metadata
[ ] Confirm the authorization server's CIMD support
[ ] Validate issuer and resource audience, not only signature
[ ] Retest every client against the production URL
[ ] Keep the legacy fallback only for clients you intentionally support
[ ] Delete compatibility code after its documented retirement date

Production checklist

Before sharing the server URL
[ ] mcp-handler 2.x and MCP TypeScript SDK v2 are version-pinned
[ ] One Streamable HTTP endpoint serves the intended protocol eras
[ ] Every tool has a narrow capability and strict schema
[ ] No generic SQL, shell, URL fetch, or arbitrary API tools exist
[ ] OAuth resource metadata uses the public production origin
[ ] Tokens validate signature, issuer, audience, expiry, and scopes
[ ] Tenant and user identity come only from verified auth context
[ ] Every object lookup includes tenant authorization
[ ] Write tools require replay-safe operation IDs
[ ] Consequential operations produce immutable audit events
[ ] Outbound hosts, response sizes, and durations are bounded
[ ] Long-running work moves to a durable queue or workflow
[ ] Function duration and region match measured dependencies
[ ] Preview and production use separate identity configuration
[ ] Logs and traces redact credentials and sensitive content
[ ] Rate limits apply by subject, tenant, client, and expensive tool
[ ] Inspector plus real-client compatibility tests pass
[ ] Abuse response and token-revocation procedures have owners

The implementation principle

The simplest remote MCP server is one stateless HTTPS endpoint backed by narrow tools. The safest one derives identity from a verified token, derives tenant access from server-side data, and treats every tool call as an API authorization decision. Vercel can scale that endpoint; it cannot decide what the caller should be allowed to do.

Build the transport with mcp-handler and the current TypeScript SDK, but spend most of the design effort on the contracts around it: explicit scopes, object authorization, idempotent writes, bounded dependencies, durable long-running work, and redacted evidence. That is what turns an MCP demo into infrastructure another team can trust.

The rule worth keeping

A tool schema controls what a model may ask for. Verified identity plus server-side authorization controls what the system may actually do. Production MCP needs both.

Model Context Protocol: 2026-07-28 specificationMCP TypeScript SDK v2MCP TypeScript SDK: Protocol versionsVercel mcp-handlermcp-handler: AuthorizationVercel: Model Context ProtocolVercel: Fluid computeVercel: Streaming FunctionsNext.js: Route HandlersRFC 9728: OAuth 2.0 Protected Resource Metadata