Build a Streaming Claude Chat in Next.js
Adding an AI assistant to a Next.js app is mostly plumbing: keep the API key on the server, stream tokens to the browser as they arrive, and handle the cases where generation stops early. This tutorial builds that plumbing with the official Anthropic TypeScript SDK and the App Router — one Route Handler, one Client Component, no extra framework.
Everything below targets Next.js 16 and React 19, the same versions this site runs on.
1. Install and configure
npm install @anthropic-ai/sdk zodCreate an API key in the Claude Console, then add it to .env.local alongside the model you want to use:
# Server-only. Never prefix with NEXT_PUBLIC_.
ANTHROPIC_API_KEY=sk-ant-...
# Copy the current ID from Anthropic's models overview page.
ANTHROPIC_MODEL=Which model should you pick?
Claude comes in three tiers. Opus is the most capable and the right default when answer quality matters most. Sonnet balances capability with speed and cost, and suits most product features. Haikuis the fastest and cheapest, ideal for high-volume, simple tasks such as classification or short replies. Model IDs change as new versions ship, so copy the current one from Anthropic's models overview page rather than from any blog post — including this one.
Want this built for you? Use code CLAUDE10 for 10% off your first Claude API integration project. Start a project
2. The streaming Route Handler
This is the only file that talks to Claude. It validates the request, opens a stream with client.messages.stream(), and forwards each text delta to the browser as plain text.
import Anthropic from "@anthropic-ai/sdk";
import { z } from "zod";
// Reads ANTHROPIC_API_KEY from the environment.
const client = new Anthropic();
const Body = z.object({
messages: z
.array(
z.object({
role: z.enum(["user", "assistant"]),
content: z.string().min(1).max(4000),
})
)
.min(1)
.max(20), // cap history so one visitor can't send a novel
});
const SYSTEM =
"You are the assistant on a web developer's portfolio. " +
"Answer questions about web development clearly and briefly.";
export async function POST(request: Request) {
const parsed = Body.safeParse(await request.json().catch(() => null));
if (!parsed.success) {
return Response.json({ error: "Invalid request" }, { status: 400 });
}
const stream = client.messages.stream({
model: process.env.ANTHROPIC_MODEL!,
max_tokens: 16000,
system: SYSTEM,
messages: parsed.data.messages,
});
const encoder = new TextEncoder();
const body = new ReadableStream<Uint8Array>({
async start(controller) {
try {
for await (const event of stream) {
if (
event.type === "content_block_delta" &&
event.delta.type === "text_delta"
) {
controller.enqueue(encoder.encode(event.delta.text));
}
}
// Always check why generation stopped before trusting the output.
const final = await stream.finalMessage();
if (final.stop_reason === "refusal") {
controller.enqueue(
encoder.encode("\n\nSorry — I can't help with that one.")
);
}
controller.close();
} catch (error) {
if (error instanceof Anthropic.RateLimitError) {
controller.enqueue(encoder.encode("Busy right now — try again shortly."));
controller.close();
} else {
controller.error(error);
}
}
},
// The visitor closed the tab: stop paying for tokens nobody will read.
cancel() {
stream.abort();
},
});
return new Response(body, {
headers: {
"Content-Type": "text/plain; charset=utf-8",
"Cache-Control": "no-store",
},
});
}Three details worth pausing on:
- Only
text_deltaevents are forwarded. The stream also carries message metadata and, on models that think before answering, thinking blocks. Filtering by delta type keeps the client simple. finalMessage()after the loop gives you the assembled response — includingstop_reasonand tokenusagefor logging — without re-implementing accumulation yourself.cancel()aborts the upstream request. When a visitor navigates away mid-answer, the browser cancels the response stream; aborting stops generation instead of finishing a reply nobody will read.
3. The chat component
The client posts the conversation, then reads the response body chunk-by-chunk with a TextDecoderStream, appending each chunk to the last message. The API is stateless, so the full history is sent every turn.
"use client";
import { useState, type FormEvent } from "react";
type Message = { role: "user" | "assistant"; content: string };
export default function Chat() {
const [messages, setMessages] = useState<Message[]>([]);
const [input, setInput] = useState("");
const [busy, setBusy] = useState(false);
async function send(e: FormEvent) {
e.preventDefault();
const text = input.trim();
if (!text || busy) return;
const history: Message[] = [...messages, { role: "user", content: text }];
setMessages([...history, { role: "assistant", content: "" }]);
setInput("");
setBusy(true);
try {
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: history }),
});
if (!res.ok || !res.body) throw new Error(`HTTP ${res.status}`);
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
for (;;) {
const { value, done } = await reader.read();
if (done) break;
// Append each chunk to the last (assistant) message.
setMessages((prev) => {
const last = prev[prev.length - 1];
return [...prev.slice(0, -1), { ...last, content: last.content + value }];
});
}
} catch {
setMessages((prev) => [
...prev.slice(0, -1),
{ role: "assistant", content: "Something went wrong. Please retry." },
]);
} finally {
setBusy(false);
}
}
return (
<div>
<ul aria-live="polite">
{messages.map((m, i) => (
<li key={i} data-role={m.role}>
{m.content}
</li>
))}
</ul>
<form onSubmit={send}>
<label htmlFor="chat-input" className="sr-only">Message</label>
<input
id="chat-input"
value={input}
onChange={(e) => setInput(e.target.value)}
placeholder="Ask anything…"
/>
<button disabled={busy}>{busy ? "Thinking…" : "Send"}</button>
</form>
</div>
);
}The aria-live="polite" region means screen readers announce the reply as it streams in — an accessibility detail that is easy to forget and costs one attribute.
4. Cut costs with prompt caching
If your system prompt is long — product documentation, a style guide, an FAQ — every request re-sends it. Prompt caching lets Claude reuse that stable prefix across requests at a fraction of the input cost. Keep the system prompt byte-for-byte identical between calls (no timestamps or per-user IDs in it) and add one field:
const stream = client.messages.stream({
model: process.env.ANTHROPIC_MODEL!,
max_tokens: 16000,
// Caches the stable prefix (your long system prompt) between requests.
cache_control: { type: "ephemeral" },
system: LONG_SYSTEM_PROMPT,
messages: parsed.data.messages,
});Verify it is working by logging final.usage.cache_read_input_tokens: it should be non-zero from the second request onwards. Very short prompts fall below the minimum cacheable size and are silently not cached.
Production checklist
- 1Keep the key on the serverThe API key lives in a server-only env var and is only touched inside the Route Handler. Anything prefixed NEXT_PUBLIC_ ships to every visitor's browser.
- 2Validate and cap the inputThe Zod schema above limits message length and history depth. Without it, a single request can carry as much text as the model accepts — and you pay for every token.
- 3Rate-limit the routePer-IP limits in proxy.ts or at your edge provider. A public chat endpoint without one is an open invitation to run up your bill.
- 4Check stop_reasonend_turn is the happy path. max_tokens means the reply was cut off; refusal means Claude declined. Handle both explicitly instead of showing a half-sentence.
- 5Catch typed errorsThe SDK throws classes such as Anthropic.RateLimitError and Anthropic.AuthenticationError. Branch on the class, not on the message text. The SDK already retries transient failures twice by default.
- 6Never hard-code the modelRead it from configuration, so moving to a newer model is a one-line env change and a redeploy rather than a code change.
Where to go from here
The same route grows naturally into tool use — letting Claude call your own functions to look up orders, search docs or book meetings — and into structured outputs when you need JSON back instead of prose. The Claude API documentation covers both with TypeScript examples.
And if you would rather have Claude write this kind of plumbing for you, read how I use Claude Code as a full-stack developer.
Common questions
- How do I stream Claude responses in Next.js?
- Call client.messages.stream() from the Anthropic TypeScript SDK inside an App Router Route Handler, forward each text_delta event into a ReadableStream, and return it as the Response body. The browser reads it chunk-by-chunk with a TextDecoderStream.
- Where should the Anthropic API key live?
- In a server-only environment variable such as ANTHROPIC_API_KEY, used only inside Route Handlers or Server Actions. Never prefix it with NEXT_PUBLIC_, which would ship it to every visitor's browser.
- Which Claude model should I use for a chat feature?
- Opus is the most capable, Sonnet balances capability with speed and cost, and Haiku is the fastest and cheapest. Keep the model ID in configuration so you can switch tiers or upgrade versions without a code change.
- How can I reduce Claude API costs?
- Use prompt caching for long, stable system prompts, cap conversation length and message size, rate-limit public endpoints, and pick the smallest model tier that meets your quality bar.