Let's Build a Streaming Web Chat (Express Proxy + Plain JS)
In this tutorial you will build a small chat page in about 100 lines. A tiny Express server holds your secret key and relays requests to the Uncensored Chatbot API, while a plain JavaScript front end reads the streamed reply word by word. No framework, no build step, and your key never touches the browser.
Updated
Key points
- Never call the API from browser code; a server route keeps the key private and lets you validate input.
- The proxy only needs to pass the server-sent events through; the page parses data lines until [DONE].
- Use textContent while streaming and sanitize before rendering markdown.
- Cap history length and max_tokens on the server so one user cannot spend your whole balance.
What we are building, and why a proxy
Picture a single page with a text box and a scrolling log. You type a line, the page posts the conversation to /api/chat on your own server, and your server forwards it upstream with the real key attached. The reply streams back, and each fragment appears the instant it arrives. We will give the bot a little personality, a lighthouse keeper called Wren, so the demo feels alive.
You need Node 18 or newer (for built-in fetch) and an API key. If you do not have one yet, grab the free trial from the key page: new accounts get $0.50 of credit for 7 days, no payment details needed. The base URL is https://api.uncensoredchatbotapi.com/v1 and the model id is uncensored.
Why you need the proxy
It is tempting to paste the key into front-end code and call the API directly. Please don't. Anything shipped to a browser can be read by anyone who opens the developer tools, and a leaked key means a drained balance. A proxy fixes this and gives you three extra benefits.
- Secrecy. The key lives in an environment variable on the server.
- Control. You decide the system prompt, the history length and
max_tokens, so users cannot override them. - A place for rules. Per-user limits, age gates and logging policy all belong in this layer. The safety guide builds on this very route.
Step 1: set up the project
Create a folder, initialize it as an ES-module project and install Express. Everything else is built in.
mkdir lantern-chat && cd lantern-chat
npm init -y
npm pkg set type=module
npm install express
mkdir publicYou will end up with two things: server.js at the root and a public folder holding the page and its script.
Step 2: write the Express proxy
The server serves static files and exposes one POST route. Read it in three parts. First it sanitizes the incoming history: only user and assistant roles pass, each message is trimmed to 4,000 characters, and only the last 30 turns are kept. Second, it prepends your own system message, which the browser can never change. Third, it calls the upstream endpoint with stream: true and pipes the bytes straight back.
import express from "express";
const app = express();
app.use(express.json({ limit: "256kb" }));
app.use(express.static("public"));
const UPSTREAM = "https://api.uncensoredchatbotapi.com/v1/chat/completions";
const SYSTEM = "You are Wren, a dry-witted night-shift lighthouse keeper. Stay in character.";
app.post("/api/chat", async (req, res) => {
const history = Array.isArray(req.body.messages) ? req.body.messages : [];
// Keep only well-formed turns and cap the history we forward.
const turns = history
.filter((m) => ["user", "assistant"].includes(m.role) && typeof m.content === "string")
.slice(-30)
.map((m) => ({ role: m.role, content: m.content.slice(0, 4000) }));
if (turns.length === 0) return res.status(400).json({ error: "no messages" });
const upstream = await fetch(UPSTREAM, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "uncensored",
messages: [{ role: "system", content: SYSTEM }, ...turns],
stream: true,
max_tokens: 500,
temperature: 0.8,
}),
});
if (!upstream.ok) {
const info = await upstream.json().catch(() => ({}));
return res.status(upstream.status).json(info);
}
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
for await (const chunk of upstream.body) res.write(chunk);
res.end();
});
app.listen(3000, () => console.log("http://localhost:3000"));Two choices are worth explaining. Passing the raw bytes through means you do not need to understand the event format on the server at all. And returning the upstream status for failures lets the page react sensibly; for example, a 402 means your balance is empty and a 429 means someone is going too fast.
Step 3: the page and the stream reader
Now the front end. Save this as public/index.html; it is deliberately bare so you can restyle it later.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Lantern Chat</title>
<style>
body { font: 16px system-ui; max-width: 640px; margin: 2rem auto; padding: 0 1rem; }
#log { min-height: 320px; border: 1px solid #ccc; padding: 1rem; white-space: pre-wrap; }
.me { color: #234; font-weight: 600; }
form { display: flex; gap: .5rem; margin-top: 1rem; }
input { flex: 1; padding: .6rem; }
</style>
</head>
<body>
<h1>Lantern Chat</h1>
<div id="log"></div>
<form id="form">
<input id="text" autocomplete="off" placeholder="Say something...">
<button>Send</button>
</form>
<script src="chat.js"></script>
</body>
</html>Next, public/chat.js. The interesting part is the read loop. Network chunks do not respect line boundaries, so we keep a buffer, split on newlines, and hold back the last partial line until more data arrives. Each complete line that starts with data: is JSON, except the final [DONE] marker.
const log = document.getElementById("log");
const form = document.getElementById("form");
const input = document.getElementById("text");
const history = [];
function addLine(cls, text) {
const div = document.createElement("div");
div.className = cls;
div.textContent = text; // textContent, never innerHTML, for untrusted text
log.appendChild(div);
return div;
}
form.addEventListener("submit", async (e) => {
e.preventDefault();
const text = input.value.trim();
if (!text) return;
input.value = "";
history.push({ role: "user", content: text });
addLine("me", "You: " + text);
const bubble = addLine("bot", "");
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: history }),
});
if (!res.ok) {
bubble.textContent = "(the chat is unavailable right now)";
history.pop();
return;
}
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "", reply = "";
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop(); // keep a partial line for the next read
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const data = line.slice(6).trim();
if (data === "[DONE]") continue;
const json = JSON.parse(data);
const piece = json.choices?.[0]?.delta?.content;
if (piece) {
reply += piece;
bubble.textContent = reply;
}
}
}
history.push({ role: "assistant", content: reply });
});Notice the history array. The API keeps no memory between calls, so the page resends the whole conversation each time, and the server trims it. Start the app and open it in a browser:
export API_KEY="paste-your-key-here"
node server.js
A note on markdown rendering
Chat models love asterisks, lists and the occasional code fence. Our demo shows raw text, which is safe. When you want pretty output, render markdown with a library, but follow two rules. Run the HTML through a sanitizer before inserting it with innerHTML, since a model, or a user who tricks it, can emit tags and attributes you did not intend. And render incrementally with care: re-parsing the whole reply on every fragment is fine for short messages, but half-finished markdown can flicker, so some builders show plain text while streaming and swap to formatted output when the stream ends.
Roleplay formatting is a separate quirk. Many characters wrap actions in asterisks, such as *adjusts the lamp*. Decide whether your app styles those as italics, and tell the character in the system prompt which convention to follow. Our persona design guide shows prompt wording for that.
Step 4: smoke-test the route
Before blaming the browser, poke the proxy directly. If the terminal works, the server is fine and any remaining bug lives in the page.
curl -N http://localhost:3000/api/chat \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Wren, is the fog coming in?"}]}'The -N flag disables curl's own buffering, so you should see data: lines appear gradually, ending with data: [DONE]. If you get a JSON error instead, read its status: 401 means the key in your environment is wrong, 402 means the balance is empty, and a 404 means the upstream URL has a typo.
| Symptom | Likely cause | Fix |
|---|---|---|
| Page shows the unavailable message | Server returned a non-200 status | Run the curl test and read the status |
| Text appears only at the end | A proxy layer buffers the response | Disable buffering for the route |
| Garbled characters | Decoder used without stream: true | Keep the option set on TextDecoder.decode |
| Reply cuts off mid-sentence | max_tokens reached | Raise the cap on the server |
With those four fixes you can diagnose nearly every first-run problem in under a minute. Keep the curl command in your notes; it is also a handy health check after you change the proxy later.
Small touches that make it feel finished
A chat that streams is already pleasant, but a few details separate a demo from a product people return to.
- Disable the send button while a reply is arriving. Double submits create interleaved histories that confuse both the user and the model.
- Add a stop button. Create an
AbortController, pass its signal tofetch, and callabort()on click. Save whatever text arrived so far as the assistant turn. - Persist the log. Keep the history in
sessionStorageso a refresh does not wipe the conversation, and offer a clear-chat button that empties it. - Auto-scroll sensibly. Follow the bottom only when the user is already near it; otherwise let them read older lines in peace.
- Show a gentle failure. Replace the empty bubble with a retry link that resends the last user message.
Each of these is a dozen lines of plain JavaScript, so resist reaching for a framework until your interface genuinely needs components, routing or shared state.
Hardening and next steps
You now have a working chat. Before real users arrive, add a few safeguards in server.js. Limit requests per IP or session, because a single key is allowed 300 requests per minute in total, and one eager user could use the lot. Handle upstream errors explicitly: a 503 with upstream_busy deserves a friendly retry button, and a 403 content_blocked deserves a clear message rather than a blank bubble. Keep max_tokens modest; 500 is plenty for chat, while the allowed maximum is 16,000 per request.
Then think about cost. As an illustration, assume each turn sends 1,200 prompt tokens and receives 250 completion tokens. That is about $0.0003 for input and $0.00025 for output, roughly $0.00055 a turn. Those token counts are assumptions, so measure your own with the usage chunk. The docs cover that field, and the guide to hosted uncensored LLMs explains what to expect from this kind of service.
Questions and answers
Can I call the API straight from the browser?
Technically yes, but you would expose your key to every visitor. Use a server route like the one in this tutorial so the key stays in an environment variable.
Does the server need to parse the stream?
No. It can forward the bytes unchanged, and the page reads the data lines. Parse on the server only if you want to log text or filter output.
Why does my reply arrive all at once?
Something in between is buffering. Check that the request sets stream to true and that any reverse proxy or compression layer is not holding the response.
How do I give the bot memory?
Resend the conversation in messages on every call, trimming the oldest turns when it grows. The 100,000-token context is shared with the reply.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.