How two-way chat actually talks to the server
Introduction
This post is written for two different readers. The first half is readable without a systems-programming background; the second half is for implementers (Go).
A lot of people wonder “why does a message on Slack or LINE show up instantly?” without ever getting to the mechanism underneath. This post starts there, walking through the difference between HTTP and WebSocket with diagrams. Partway through it shifts into the concrete story of the anonymous chat app I mentioned in the previous post — written in Go — so keep that shift in mind as you read on.
An ordinary web page runs on HTTP request/response, but a push from the server, like chat needs, is hard to pull off with HTTP alone. This post covers how HTTP and WebSocket actually differ, using the anonymous chat app I built as the case study for how I split the two.
First, the premises of the service I’m using as the example — change these and the design choices change with them.
| People per room | Two or three (matched per topic) |
| Concurrent connections, at most | 1,000 (hard-capped in the app layer) |
| Servers | Two small instances (one Go process per instance) |
| Language / library | Go / gorilla/websocket |
| Supporting services | Redis (Pub/Sub, presence, matching queue), PostgreSQL |
Here is what this post covers:
- HTTP’s communication model: request/response
- WebSocket’s communication model: two-way, always-on
- HTTP vs WebSocket, side by side
- Implementation: how I split it for anonymous chat
- The decision I use to pick between them
- Wrapping up
- References
HTTP’s communication model: request/response
HTTP is fundamentally “the client sends a request, the server sends back a response.” It’s like exchanging letters — nothing comes back unless you send something first.
How to read the diagram: every arrow starts at the client and points to the server. Any arrow going the other way is always a “reply” to a request.
- Communication is always client-initiated. The server can’t send data on its own
- Stateless. Each request stands alone; the server doesn’t remember the previous one
- Every request carries headers (auth info, cookies, and so on — a few hundred bytes to a few KB)
Faking real-time over HTTP (and where it breaks down)
| Approach | How it works | The problem |
|---|---|---|
| Polling | Repeat GET /messages on a timer | A request goes out even with nothing new. Delay equals the interval |
| Long polling | Hold the response open until something new arrives | Connections pile up. Timeout handling gets complicated |
Building “arrives the instant the other person sends it” out of either one is rough — on efficiency, on latency, and on implementation complexity all at once.
WebSocket’s communication model: two-way, always-on
WebSocket is a protocol that, once a connection is established, keeps it open so either side can send a message at any time (RFC 6455). Not letters — a phone call that never hangs up.
The handshake starts as an ordinary HTTP request
GET /ws/rooms/0193f0a1-... HTTP/1.1
Host: example.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13
The client asks to switch protocols with Upgrade: websocket, and if the server agrees, it replies with 101 Switching Protocols.
HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
How to read the diagram: after the handshake, the S->>C arrow fires without any request from the client. That’s the whole point of WebSocket.
After the handshake, traffic becomes binary “frames” with as little as 2 bytes of header. Compare that to HTTP sending a few hundred bytes of headers on every single exchange — the difference adds up fast with frequent messages.
HTTP vs WebSocket, side by side
| HTTP | WebSocket | |
|---|---|---|
| Direction | Client → server | Two-way |
| Connection | Opened and closed per request | Stays open once established |
| Server-initiated push | Not possible (faked via polling etc.) | Native |
| Header overhead | A few hundred bytes to a few KB, every time | Once, up front. 2–14 bytes after that |
| State | Stateless | Stateful (holds the connection) |
| Implementation / ops difficulty | Low | High (reconnection, scaling, etc.) |
| Good fit for | CRUD, page rendering | Chat, notifications, live updates |
Implementation: how I split it for anonymous chat
From here it’s the actual thing I built. Each section covers a piece of the design.
Where HTTP and WebSocket actually split
I didn’t make everything WebSocket. Here’s how the endpoints actually broke down.
| Function | Endpoint | Protocol |
|---|---|---|
| Issue / extend an anonymous session | POST /api/session | HTTP |
| Join the matching queue / check status / leave | /api/matching | HTTP |
| List topics | GET /api/topics | HTTP |
| Conversation inside a room | GET /ws/rooms/{roomID} | WebSocket |
There’s exactly one WebSocket. Waiting to be matched is handled by HTTP polling, since a few seconds of delay is fine when you’re just waiting for a match to happen — opening a WebSocket here would mean holding a connection for everyone in the queue, before they’ve even started talking.
One more difference from typical chat apps: there’s no message-history API. Messages are saved to the database, but that’s for handling reports and disclosure requests — there’s no history screen shown to participants. So the common “HTTP for history, WebSocket for new messages” pattern never happened; WebSocket ended up being the only channel for anything conversation-related.
What the handshake checks
WebSocket auth has to happen during the HTTP handshake, since there’s no way to send headers after the connection is open. In the implementation, it happens in this order.
func (h *Handler) handleRoomSocket(w http.ResponseWriter, r *http.Request) {
// 1. Origin
if !h.checkOrigin(r) {
http.Error(w, "forbidden", http.StatusForbidden)
return
}
// 2. Concurrent connection count
release, err := h.acquire(r)
if err != nil {
writeCapacityError(w, err)
return
}
defer release()
// 3. Session
token, err := h.sessions.Authenticate(r)
...
// 4. Is this person actually in this conversation?
adm, err := h.svc.Admit(r.Context(), r.PathValue("roomID"), token)
...
// 5. Only now: 101 Switching Protocols
ws, err := h.upgrader.Upgrade(w, r, nil)
h.svc.Serve(r.Context(), ws, adm)
}
Each step in this order exists for a reason.
- Origin gets checked first: because CORS doesn’t apply to WebSocket (more on this below).
- Connection count comes before authentication: a connection rejected for being over capacity can be turned away with a single Redis call.
- Membership check (
Admit): won’t connect someone to a conversation ID they aren’t part of. A nonexistent conversation and someone else’s conversation both come back as the same 403 — telling them apart would leak whether a given conversation exists.
Why Origin gets checked first
Because WebSocket doesn’t get HTTP’s CORS (same-origin policy) protection.
fetch() calling another site’s API gets blocked by CORS, but new WebSocket(...) has no such mechanism. A connection can be opened from a page on any site at all.
// Code sitting on an attacker's page. CORS doesn't stop this
const ws = new WebSocket("wss://our-chat.example/ws/rooms/0193f0a1-...");
ws.onmessage = (e) => { /* read the conversation contents */ };
ws.send("..."); // and impersonate the victim
There are two directions of damage.
- Reading (
ws.onmessage→ forwarded viafetch): the room’s conversation flows to the attacker’s server. Since it’s anonymous chat, no identity leaks, but the content of the conversation itself does. The other person in the room has no way to know they’re being watched. - Writing (
ws.send): the attacker can speak through the victim’s connection. Other participants see it as something the victim said — usable for trolling, or for getting the victim reported by putting policy-violating messages in their mouth.
The important part: the attacker never needs to steal a cookie. The cookie’s value is unreadable to them (it’s HttpOnly, so JavaScript can’t read it either). The connection still goes through, because of a browser behavior: a request to a domain automatically carries that domain’s cookies. The attacker’s page just has to say “connect to our-chat.example” — the victim’s own browser attaches the credentials without being asked. From the server’s side, this is indistinguishable from the legitimate user connecting normally. The attacker never obtained anyone else’s cookie — the victim’s own cookie is being used inside the victim’s own browser. Same shape as CSRF, just the WebSocket version — hence the name, Cross-Site WebSocket Hijacking.
Easy to get wrong: the attacker is not logging in as the victim from their own browser. They never have the cookie’s value, so they can’t reproduce the connection on their own machine. The connection is always made from the victim’s browser; the attacker’s script just forwards whatever conversation it receives to the attacker’s own server. Which means the attack only works while the victim has the attacker’s page open — close that tab and the connection ends too. (Someone stealing the token itself and impersonating the victim outright is a different risk, XSS being the usual route — HttpOnly cookies are the defense against that one.)
This is also where it differs from fetch(). CORS works by letting the request go out but blocking JavaScript from reading the response. WebSocket has no equivalent, so whatever comes back can just be read directly.
In practice, there are three layers of defense
That said, this attack doesn’t succeed on its own against this implementation.
| # | Defense | How it helps |
|---|---|---|
| 1 | Session cookie is SameSite=Lax | A connection opened from another site’s page doesn’t carry the cookie |
| 2 | Conversation IDs are unguessable UUIDs | An attacker can’t learn the URL to connect to |
| 3 | Origin check | Cuts off any connection not coming from my own frontend, with a 403 |
Layer 3 is there anyway because there’s a configuration where layer 1 disappears. If the frontend is served from a different origin, the cookie has to be SameSite=None — and at that point, the first defense is simply gone. The Origin check doesn’t depend on cookie settings, so it survives that case too.
Origin is a header the browser attaches, and page-level JavaScript can’t rewrite it. If it doesn’t match my own frontend’s origin, the connection gets a 403 before 101 Switching Protocols is ever returned.
There’s a reason for the ordering, too. Session validation, which comes right after, extends how long that session stays valid. Put the Origin check after it, and merely having the attacker’s page open keeps renewing someone else’s session. Hence “Origin first, then auth.”
When a limit is hit, the response is split: 503 if the whole service is full, 429 if it’s just that one IP connecting too much. That way the client can tell “it’s busy” apart from “it’s specifically me.” Both come with Retry-After: 10.
One connection, held by three goroutines
Once a WebSocket connection is established, three loops start running in parallel for that single connection.
There’s a reason writing is centralized into writeLoop. gorilla/websocket only allows one Writer per connection at a time. If two or more goroutines call WriteMessage on the same connection at once, a frame that’s midway through being sent gets interrupted by another frame and corrupts. So everything to send goes through a channel to writeLoop, and that’s the only place that actually writes to the socket.
writeLoop waits on three things at once.
select {
case msg := <-c.out: // an event arrived from the room
...
case <-ping.C: // 30-second ping tick
...
case <-c.done: // the connection ended
return
}
Delivery is per-room, over Redis Pub/Sub
The typical sample code broadcasts to every connected client, but here delivery only ever goes to the room for one conversation (2–3 people).
The catch is that each process only holds the connections it accepted itself. With two servers (meaning two Go processes running), the other person in the same conversation might well be connected to the other process. Process 1’s memory has no idea what connections process 2 is holding.
That’s what Redis is for.
What Redis is
An extremely fast data store that keeps everything in memory. Unlike PostgreSQL, which writes to disk, its contents basically disappear if the process crashes or restarts — but in exchange, reads and writes are dramatically faster, and multiple processes can touch the same data.
I use it for three things — places to put data that’s fine to lose, but needs to be shared across processes.
| Use | Redis feature |
|---|---|
| Delivering messages to a room | Pub/Sub |
| Who’s currently in this room | Sorted Set (score = last-seen time) |
| Matching wait queue | Sorted Set (score = time they started waiting) |
Anything that would be a problem to lose — the conversation record, the topic list — lives in PostgreSQL instead.
What Pub/Sub is
One of Redis’s features — short for Publish / Subscribe.
- There’s a named channel that acts as a pathway
- Publish: throw data onto a channel
- Subscribe: register to listen on a channel, and whatever gets published there flows to you
The key property is that the publisher doesn’t need to know who’s listening. Process 1 just publishes to “this conversation’s channel,” and every process subscribed to that channel — including itself — gets the data. No need to look up which process the other person happens to be connected to.
Channel names are built from the conversation ID, in the shape chat:room:{conversationID}. Different conversations get different channels, so messages never bleed across conversations.
How to read the diagram: participant 1 and participant 2 are in the same conversation, so even though they’re connected to different processes, they’re both subscribed to the same channel, chat:room:xxx. So when participant 1 speaks, process 1 publishes to Redis, it flows to process 2, and participant 2 receives it. The other conversation shown bottom-right is subscribed to chat:room:yyy, so it never sees this message.
The publish also comes back to the publishing process itself, so delivery to the sender goes through the exact same path. Not special-casing the sender keeps the code down to one path instead of two.
The structure inside a process
// Keeps one room per conversation ID; each room holds one Redis subscription
func (h *hub) attach(ctx context.Context, conversationID string, c *conn) error {
r, ok := h.rooms[conversationID]
if !ok {
sub, err := h.store.Subscribe(context.WithoutCancel(ctx), conversationID)
if err != nil {
return err
}
r = &room{conversationID: conversationID, sub: sub, conns: ...}
h.rooms[conversationID] = r
go r.run() // keeps reading the subscribed channel, fans out to its connections
}
r.conns[c] = struct{}{}
return nil
}
Each process keeps exactly one subscription per conversation. Even if three people in the same conversation land on the same process, that’s still one subscription, and it closes once that conversation has zero connections left on this process.
context.WithoutCancel is there because a subscription needs to outlive “the HTTP request that created it.” Pass the request’s own context straight through, and the subscription would die before the connection does, even while it’s still alive.
Presence tracking also lives in Redis, since I need to count “who’s in this room right now” across processes. It goes into a per-conversation Sorted Set (score = last-seen time), and entries past their TTL get dropped when read. That has a nice side effect: when a process dies, whatever presence data it left behind fades away on its own.
Slow connections get dropped, not waited for
Picture this in a three-person room.
- A: talking normally
- B: just went into a subway tunnel, signal is weak, not keeping up with incoming data
- C: talking normally
When A sends a message, the process tries to deliver it to both B and C. If B’s connection is backed up and the design is “wait until B can receive it,” then C doesn’t get it either. C’s screen — who has nothing to do with B’s dead zone — freezes until B’s signal comes back. The conversation stops working for all three people.
So I chose not to wait. Once 32 messages have queued up waiting to reach B, B’s connection gets closed.
func (c *conn) enqueue(msg outgoing) {
select {
case c.out <- msg:
case <-c.done:
default:
// buffer (32 items) is full = this connection isn't keeping up
slog.Warn("closing a connection that is not keeping up", ...)
c.close()
}
}
From the users’ side, this is what happens.
| What happens | |
|---|---|
| A and C | Nothing. The conversation continues as normal |
| B | Connection drops. Reconnecting returns them to the conversation, but they miss whatever was said while disconnected |
There’s also the option of “drop just the one message and keep the connection open,” which I didn’t take. B’s screen would keep showing a conversation with holes in it, and B wouldn’t even know. Closing the connection means the client side can actually detect the disconnect.
If B reconnects within 30 seconds, the conversation itself doesn’t end (a conversation is designed to end once it’s down to one participant for 30 seconds straight).
The decision I use to pick between them
- Page rendering, form submission, CRUD, file transfer → HTTP
- Chat, notifications, live updates, collaborative editing → WebSocket
- Only need one-way push from the server → Server-Sent Events is also worth a look (lighter than WebSocket, and since it rides on top of HTTP, load balancer configuration is simpler)
Even in this service, the only thing that actually became WebSocket was the one conversation channel. “A real-time service” and “a service built entirely on WebSocket” turned out to be two different things — that’s the thing that stuck with me most from actually building it.
Wrapping up
- HTTP is one-way, client-initiated communication; the server can’t push
- WebSocket upgrades via an HTTP handshake, then stays two-way, always-on, and framed lightly
- The practical approach is making only the parts that need real-time into WebSocket, and leaving the rest on HTTP. Here that meant waiting, sessions, and topics on HTTP, and only the conversation itself on WebSocket
- In exchange, there’s more to decide: handshake-time auth, delivery across multiple processes, how to handle slow connections
- Operational concerns — reconnection grace periods, ping/pong, load balancer settings — are each big enough on their own that I’ve left them out of this post
“WebSocket is a phone call that stays connected; HTTP is redialing every time” is a decent mental shortcut for deciding between them.