AI Assistant Platform
a multi-agent assistant: a router picks the right specialist for every message, and the answer streams back token by token.

- role
- sole author
- stack
- TypeScript · Node.js 20 · Next.js 14 · AWS Lambda · DynamoDB · Claude · MCP
The problem
One general-purpose prompt does everything a little badly. A coding question needs a sandbox, a market question needs fresh data and a disclaimer, and small talk needs neither. The goal was an assistant that sends each message to an agent built for it, shows its work while it happens, and stays cheap to run.
How it works
- The composer posts the message to
/chatover Server-Sent Events. Auth, rate limits, validation and a topic blocklist all run before the stream opens, so failures come back as ordinary HTTP errors. - A router agent on Claude Haiku classifies the message with a forced tool call whose options are generated from the agent registry. If it errors or takes longer than 2 seconds, a keyword fallback takes over, so routing never fails.
- The chosen specialist on Claude Sonnet answers: a Generic agent with web search, a Coding agent with a sandboxed code interpreter, or a Financial agent with web search and a disclaimer. Every agent can also call tools from MCP servers the user connects.
- Tokens, tool calls and results stream to the UI as typed events, and a live pipeline panel shows each stage as it runs.
- After the answer, a second Haiku call pulls out anything worth remembering and shows a "Memory updated" chip.

Decisions worth calling out
- One codebase, two transports. Handlers don't know whether they run behind Express locally or Lambda in production, so local behaviour is deployed behaviour. Chat runs on a Lambda Function URL in streaming mode, because API Gateway can't stream.
- SSE over WebSockets. The stream only flows one way, and WebSockets would need a connection table and more infrastructure.
- Errors by phase. Before the stream starts, errors are plain HTTP. After it starts, they arrive as in-band events, and partial answers are saved flagged as truncated.
- Custom JWT auth. Short-lived access tokens, rotating refresh tokens stored hashed, and reuse detection that revokes the whole token family. Login timing doesn't reveal which emails exist.
- DynamoDB single table. Users, conversations, messages, memories, settings and MCP servers share one table, and every read is a query, never a scan.
- Adding an agent is one file. Register it and the router prompt picks it up automatically.
Features
Streaming with stop (the partial answer is kept), edit and regenerate with branching, searchable history, Markdown with LaTeX and syntax highlighting, cross-session memory the user controls, custom instructions, an artifacts panel for code, image and file attachments, dictation, user-connected MCP servers, feedback with comments, incognito chats and rebindable shortcuts.
By the numbers
- about $0.0009 to route a message and $0.015 for a typical answer
- router decision in about 300 ms, with a 2 s timeout
- the code sandbox has no network, no
requireand a 5 s limit - 137 tests across backend and frontend