Type something to search...
How to Build an AI Chatbot with Next.js and the Vercel AI SDK?

How to Build an AI Chatbot with Next.js and the Vercel AI SDK?

Adding a chat assistant to a website used to mean wiring up a WebSocket server, parsing a provider's streaming format by hand, and keeping message state in sync between the browser and the backend. Each model provider had its own SDK and its own response shape, so switching models meant rewriting the integration. Most of that work has nothing to do with what makes your chatbot useful.

The Vercel AI SDK removes that plumbing. It gives you one server API for calling language models from any supported provider, a streaming protocol that works over a normal HTTP response, and a React hook that manages the conversation on the client. Combined with Next.js route handlers, you can have a streaming chatbot running in about 50 lines of code, then extend it with system prompts, tools, persistence, and abuse protection.

This article covers building a production-ready AI chatbot with Next.js 16 and the AI SDK: project setup, the chat route handler, the useChat client component, rendering message parts, adding a system prompt and tools the model can call, saving conversations, handling errors and limits, and deploying with working streaming.

How the Pieces Fit Together

The AI SDK is split into packages that map to where code runs:

PackageRuns whereWhat it does
aiServerCore functions such as streamText, message conversion, tools
@ai-sdk/reactClientThe useChat hook for chat state, streaming, and status
@ai-sdk/anthropicServerProvider adapter for Claude models (other providers available)
zodBothSchemas for tool inputs

A request flows like this:

  1. The user types a message, and useChat adds it to the local message list.
  2. The hook POSTs the full conversation to your route handler at /api/chat.
  3. The route handler converts the UI messages into model messages and calls streamText with your chosen model.
  4. The model's output streams back as a UI message stream over the HTTP response.
  5. useChat applies each chunk to the assistant message as it arrives, re-rendering the UI token by token.

The API key never leaves the server. The browser talks only to your own route handler, which is the same reason you would put any third-party secret behind a server endpoint.

Setting Up the Project

Start with a new App Router project, or use an existing one:

# Terminal
npx create-next-app@latest ai-chat --typescript --app --tailwind
cd ai-chat
npm install ai@6 @ai-sdk/react@3 @ai-sdk/anthropic@3 zod

The versions are pinned on purpose. The code in this guide targets AI SDK 6 (the ai@6 package with @ai-sdk/react@3 and the version 3 providers), which is the most widely deployed release. AI SDK 7 renames several APIs used here, for example stepCountIs becomes isStepCount and onFinish becomes onEnd, and it moves stream responses to standalone helpers. If you install without version numbers you get version 7, so either pin as shown or follow the official AI SDK 7 migration guide when you adapt the examples.

Add your provider API key to .env.local. The Anthropic provider reads ANTHROPIC_API_KEY automatically:

# .env.local
ANTHROPIC_API_KEY=sk-ant-xxxxxxxxxxxxxxxxxxxxx

Do not prefix it with NEXT_PUBLIC_. That prefix inlines a value into the client bundle, where anyone can read it and spend your credits. For a refresher on how Next.js handles these values, see how to use environment variables in Next.js.

The examples use Anthropic's Claude, but the AI SDK supports many providers through the same interface. To use another one, install its provider package, such as @ai-sdk/openai or @ai-sdk/google, and change one line in the route handler.

Creating the Chat Route Handler

The route handler receives the conversation, calls the model, and returns a stream. Create it at app/api/chat/route.ts, which is the default endpoint useChat posts to:

// app/api/chat/route.ts
import { anthropic } from "@ai-sdk/anthropic";
import { convertToModelMessages, streamText, type UIMessage } from "ai";

// Allow streaming responses up to 30 seconds on platforms that enforce limits
export const maxDuration = 30;

export async function POST(request: Request) {
  const { messages }: { messages: UIMessage[] } = await request.json();

  const result = streamText({
    model: anthropic("claude-sonnet-5-5"),
    messages: await convertToModelMessages(messages),
  });

  return result.toUIMessageStreamResponse();
}

What each part does:

  • UIMessage[] is the message format the client sends. Each message has an id, a role, and an array of parts, which can hold text, tool calls, files, and other content.
  • convertToModelMessages turns UI messages into the format language models expect, dropping UI-only data. It is asynchronous in AI SDK 6, so await it.
  • streamText starts the model call and returns immediately with a result object. It does not wait for the full response.
  • toUIMessageStreamResponse() wraps the stream in a standard Response using the protocol useChat understands.

maxDuration is a route segment config that tells serverless platforms how long the function may run. On a self-hosted Node.js server it has no effect, because requests are not time-limited by the platform. For background on functions that run per request, see how to use serverless functions with Next.js.

Building the Chat Interface with useChat

The client component uses useChat to manage messages and stream state. It must be a Client Component because it uses state and event handlers:

// app/chat.tsx
"use client";

import { useChat } from "@ai-sdk/react";
import { useState } from "react";

export function Chat() {
  const { messages, sendMessage, status, stop, error, regenerate } = useChat();
  const [input, setInput] = useState("");

  const isBusy = status === "submitted" || status === "streaming";

  function handleSubmit(event: React.FormEvent<HTMLFormElement>) {
    event.preventDefault();
    const text = input.trim();
    if (!text || isBusy) return;
    sendMessage({ text });
    setInput("");
  }

  return (
    <div className="mx-auto flex h-dvh max-w-2xl flex-col p-4">
      <div className="flex-1 space-y-4 overflow-y-auto">
        {messages.map((message) => (
          <div
            key={message.id}
            className={message.role === "user" ? "text-right" : "text-left"}
          >
            <div
              className={`inline-block max-w-[85%] whitespace-pre-wrap rounded-lg px-4 py-2 ${
                message.role === "user" ? "bg-blue-600 text-white" : "bg-gray-100 text-gray-900"
              }`}
            >
              {message.parts.map((part, index) =>
                part.type === "text" ? <span key={index}>{part.text}</span> : null,
              )}
            </div>
          </div>
        ))}

        {status === "submitted" && <p className="text-sm text-gray-500">Thinking...</p>}

        {error && (
          <div className="rounded bg-red-50 p-3 text-sm text-red-700">
            Something went wrong.{" "}
            <button type="button" onClick={() => regenerate()} className="underline">
              Try again
            </button>
          </div>
        )}
      </div>

      <form onSubmit={handleSubmit} className="mt-4 flex gap-2">
        <input
          value={input}
          onChange={(event) => setInput(event.target.value)}
          placeholder="Ask me anything..."
          className="flex-1 rounded border px-3 py-2"
          disabled={status === "error"}
        />
        {isBusy ? (
          <button type="button" onClick={() => stop()} className="rounded border px-4 py-2">
            Stop
          </button>
        ) : (
          <button type="submit" className="rounded bg-black px-4 py-2 text-white">
            Send
          </button>
        )}
      </form>
    </div>
  );
}

Then render it from a Server Component page:

// app/page.tsx
import { Chat } from "./chat";

export default function Home() {
  return <Chat />;
}

Run npm run dev, open the page, and send a message. The reply appears token by token as the model generates it.

Understanding status and message parts

useChat exposes a status value that drives most UI decisions:

StatusMeaningTypical UI
submittedRequest sent, waiting for the first chunkShow a thinking indicator
streamingChunks are arrivingShow a stop button
readyThe response finished, ready for the next messageEnable the input
errorThe request failedShow an error and a retry

Messages are rendered from parts rather than a single content string. A short answer has one text part. Once you add tools, a single assistant message can contain text, then a tool call, then a tool result, then more text, in the order they happened. Rendering parts in order keeps the conversation readable as it grows more capable.

The input state is managed by your own component, which gives you full control over the form, including multiline inputs, attachments, or suggested prompts that call sendMessage directly.

Rendering Markdown

Models often answer with Markdown lists, headings, and code blocks. Plain text rendering shows the raw syntax. Install react-markdown and render text parts through it:

// app/message-text.tsx
"use client";

import ReactMarkdown from "react-markdown";

export function MessageText({ text }: { text: string }) {
  return (
    <div className="prose prose-sm max-w-none">
      <ReactMarkdown>{text}</ReactMarkdown>
    </div>
  );
}

Replace the span in the chat component with MessageText. react-markdown does not render raw HTML by default, which is the safe choice for model output.

Adding a System Prompt

A general-purpose model needs instructions to become your assistant. Pass a system prompt to streamText. Keep it on the server, where users cannot read or change it:

// app/api/chat/route.ts
import { anthropic } from "@ai-sdk/anthropic";
import { convertToModelMessages, streamText, type UIMessage } from "ai";

export const maxDuration = 30;

const SYSTEM_PROMPT = `You are the support assistant for Web Solution Master, a web development studio.
Answer questions about Next.js, WordPress, hosting, and DNS clearly and briefly.
If a question is unrelated to web development, politely say you can only help with web topics.
When you are not sure, say so instead of guessing.`;

export async function POST(request: Request) {
  const { messages }: { messages: UIMessage[] } = await request.json();

  const result = streamText({
    model: anthropic("claude-sonnet-5-5"),
    system: SYSTEM_PROMPT,
    messages: await convertToModelMessages(messages),
    maxOutputTokens: 1024,
    temperature: 0.3,
  });

  return result.toUIMessageStreamResponse();
}

maxOutputTokens caps the length of each reply, which bounds cost and latency. A lower temperature makes answers more consistent, which suits support bots better than creative writing.

Letting the Model Call Tools

Tools turn a chatbot from something that only talks into something that can look things up and act. You describe a function with a name, a description, and an input schema. The model decides when to call it, the SDK runs your execute function on the server, and the result goes back to the model so it can write the final answer.

// app/api/chat/route.ts
import { anthropic } from "@ai-sdk/anthropic";
import {
  convertToModelMessages,
  stepCountIs,
  streamText,
  tool,
  type UIMessage,
} from "ai";
import { z } from "zod";
import { searchPosts } from "@/lib/search";

export const maxDuration = 30;

export async function POST(request: Request) {
  const { messages }: { messages: UIMessage[] } = await request.json();

  const result = streamText({
    model: anthropic("claude-sonnet-5-5"),
    system:
      "You are a helpful assistant for a web development blog. Use the searchPosts tool to find relevant articles and link to them in your answers.",
    messages: await convertToModelMessages(messages),
    tools: {
      searchPosts: tool({
        description: "Search the blog for articles that match a query",
        inputSchema: z.object({
          query: z.string().describe("Keywords to search for"),
        }),
        execute: async ({ query }) => {
          const posts = await searchPosts(query, 5);
          return posts.map((post) => ({ title: post.title, url: post.url }));
        },
      }),
    },
    stopWhen: stepCountIs(5),
  });

  return result.toUIMessageStreamResponse();
}

stopWhen: stepCountIs(5) lets the model take several steps, for example search, read the results, and then answer, while capping the loop so a confused model cannot call tools forever. Without a multi-step stop condition, the response ends right after the tool call and the user never sees an answer.

searchPosts is your own function. It could query a database, a search index, or a vector store. Because execute runs in the route handler, it can use secrets and server-only modules safely.

On the client, tool calls appear as parts with a type of tool- followed by the tool name. Show progress while the tool runs:

// app/chat.tsx (inside the parts map)
{message.parts.map((part, index) => {
  switch (part.type) {
    case "text":
      return <MessageText key={index} text={part.text} />;
    case "tool-searchPosts":
      return part.state === "output-available" ? null : (
        <p key={index} className="text-sm italic text-gray-500">
          Searching the blog...
        </p>
      );
    default:
      return null;
  }
})}

Tool parts move through states such as input-streaming, input-available, output-available, and output-error, so you can show a loading line, the result, or an error message for each call.

Saving Conversations

useChat keeps messages in React state, so a page refresh loses the conversation. To persist chats, give each conversation an ID and save messages when a response finishes. The toUIMessageStreamResponse method accepts an onFinish callback that receives the complete message list:

// app/api/chat/route.ts (persistence variant)
import { anthropic } from "@ai-sdk/anthropic";
import { convertToModelMessages, streamText, type UIMessage } from "ai";
import { saveChat } from "@/lib/chat-store";

export async function POST(request: Request) {
  const { id, messages }: { id: string; messages: UIMessage[] } = await request.json();

  const result = streamText({
    model: anthropic("claude-sonnet-5-5"),
    messages: await convertToModelMessages(messages),
  });

  return result.toUIMessageStreamResponse({
    originalMessages: messages,
    onFinish: async ({ messages: allMessages }) => {
      await saveChat({ id, messages: allMessages });
    },
  });
}

On the client, pass the conversation ID and the stored messages to the hook:

// app/chat/[id]/chat-view.tsx
"use client";

import { useChat } from "@ai-sdk/react";
import type { UIMessage } from "ai";

export function ChatView({ id, initialMessages }: { id: string; initialMessages: UIMessage[] }) {
  const { messages, sendMessage, status } = useChat({
    id,
    messages: initialMessages,
  });

  // ...render as before
  return null;
}

The page that renders ChatView is a Server Component that loads the saved messages from your database by id. Store messages as JSON in a single column per chat, or normalize them into a table, whichever fits your data layer. Verify on the server that the signed-in user owns the chat ID before loading or saving, so one user cannot read another user's conversation by guessing an ID.

Protecting Your Endpoint and Your Budget

A public chat endpoint is an open door to your API bill. Before launch:

  • Require authentication for anything beyond a demo. Check the session at the start of the route handler and return 401 for anonymous requests.
  • Rate limit per user or IP. A simple counter in Redis or your database, such as 20 messages per hour, stops scripted abuse.
  • Cap the conversation length. Long histories are sent on every request and cost more each time. Trim to the most recent messages before calling convertToModelMessages.
  • Limit output with maxOutputTokens, and set a spending limit in your provider's console as a final safeguard.
  • Validate the request body. Reject unexpected shapes or oversized payloads before they reach the model.
// app/api/chat/route.ts (guard at the top of POST)
const MAX_MESSAGES = 30;

if (!Array.isArray(messages) || messages.length === 0) {
  return new Response("Invalid request", { status: 400 });
}

const recentMessages = messages.slice(-MAX_MESSAGES);

Pass recentMessages to convertToModelMessages instead of the full list.

Deploying with Working Streaming

On Vercel, streaming works without configuration. When self-hosting behind a reverse proxy such as Nginx, buffering is the most common reason a chatbot appears to wait and then dump the whole answer at once. Disable proxy buffering for the chat route, or set the X-Accel-Buffering: no header from Next.js. The full setup is covered in how to self-host a Next.js app on a VPS with Nginx and PM2.

Also set the provider API key in your hosting environment, not in a committed file, and confirm the route handler runs on the Node.js runtime, which is the default.

Common Problems and Fixes

  • The reply arrives all at once instead of streaming. A proxy or CDN is buffering the response. Disable buffering in Nginx, and check that nothing in between compresses or collects the full body before forwarding it.
  • 401 or "API key is missing" errors. The key is not set in the environment where the server runs, or the server was not restarted after editing .env.local. Restart npm run dev and check your host's environment settings.
  • The model calls a tool and then stops without answering. No multi-step stop condition was set. Add stopWhen: stepCountIs(5) or a similar limit.
  • Messages render as empty bubbles. The component reads a content field that does not exist. Render message.parts and handle the text part type.
  • "Cannot use useChat in a Server Component." The chat file is missing the "use client" directive. Add it to the top of the component that calls the hook.
  • Timeouts on long answers in serverless hosting. Raise maxDuration within your plan's limits, or lower maxOutputTokens.

AI Chatbot with Next.js FAQ

No. The AI SDK is an open-source TypeScript library that works in any Node.js environment, including self-hosted servers, Docker containers, and other cloud platforms. Vercel maintains it, but nothing in the SDK requires deploying to Vercel.

Yes. Install the provider package for the model you want and change the model passed to streamText. The route handler, the streaming protocol, and the useChat client code stay the same, which is one of the main reasons to use the SDK.

Use a route handler for chat. useChat posts to an HTTP endpoint and reads a streaming response from it, which is what route handlers are designed for. Server Actions are better suited to form mutations that return a single result.

Give it a tool that searches your content, or retrieve relevant passages before calling the model and include them in the system prompt. This approach, often called retrieval-augmented generation, keeps answers grounded in your documentation instead of the model's general knowledge.

Yes, as long as it is read from a server environment variable without the NEXT_PUBLIC prefix. Route handlers run only on the server, so the key never reaches the browser. Never pass the key to a Client Component or call the provider directly from the browser.

Require sign-in, rate limit requests per user, trim long conversation histories, cap output tokens on every call, and set a monthly spending limit in your provider's console. Logging token usage per user also helps you spot abuse early.

Conclusion

The Vercel AI SDK lets Next.js do what it is good at: a route handler keeps secrets and model calls on the server, and a Client Component handles the interactive conversation. streamText and toUIMessageStreamResponse stream the model's output over a standard HTTP response, and useChat turns that stream into live, token-by-token messages with built-in status, stop, and retry handling.

From that base, a system prompt gives the assistant a focused role, tools let it search and act on your own data, onFinish persistence keeps conversations across visits, and authentication, rate limits, and token caps keep the endpoint safe to expose. Because the provider sits behind a single model parameter, you can test different models without touching the rest of the app.

Here are some useful references for going deeper on building AI chatbots with Next.js:

  1. AI SDK Docs: Next.js App Router getting started — the official setup for route handlers and useChat.
  2. AI SDK Docs: Chatbot — useChat options, status values, message parts, and error handling.
  3. AI SDK Docs: Tool calling — defining tools, input schemas, and multi-step calls.
  4. Anthropic Docs: Claude models overview — available Claude models and their capabilities.
  5. Next.js Docs: Route Handlers — how route handlers receive requests and stream responses.
Tags :
Share :

Related Posts

Building Powerful Desktop Applications with Next.js

Building Powerful Desktop Applications with Next.js

Next.js is a popular React framework known for its capabilities in building server-side rendered (S

Continue Reading
Can use custom server logic with Next.js?

Can use custom server logic with Next.js?

Next.js, a popular React framework for building web applications, has gained widespread adoption for its simplicity, performance, and developer-frien

Continue Reading
TypeScript with Next.js?

TypeScript with Next.js?

Next.js has emerged as a popular React framework for building robust web applications, offering developers a powerful set of features to enhance thei

Continue Reading