Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering the reactor pattern,poll()multiplexing, non-blocking socket flags, state-driven connection readiness, and graceful connection lifecycle management.
Opening & Scope:
"In Radis, rather than spawning a thread per connection which creates massive memory and context-switching overhead, I implemented a single-threaded reactor pattern using Linux poll() and non-blocking sockets to manage concurrent client traffic."
Step 1: Enforcing Non-Blocking Mode with fcntl:
"First, for every accepted connection and the listening socket itself, I explicitly configure O_NONBLOCK via fd_set_nb() (server.cpp:48-55). Before and after calling fcntl, I clear and verify errno to catch descriptor failures immediately. In non-blocking mode, if a read() or write() syscall cannot immediately transfer bytes, it returns -1 with errno == EAGAIN or EWOULDBLOCK instead of blocking the entire process."
Step 2: Connection State Machine (struct Conn):
"Second, I model each client connection with struct Conn (server.cpp:78-98). Each connection tracks its active file descriptor, dynamic byte buffers for incoming and outgoing data, an intrusive timer list node, and explicit application intent flags: want_read, want_write, and want_close. This separates network readiness from application state."
Step 3: Dynamic pollfd Assembly & Syscall Interruption:
"Third, in each iteration of the event loop inside main() (server.cpp:846-869), the server constructs a std::vector<struct pollfd>. The listening socket is always placed at index 0 listening for POLLIN. For active connections, we conditionally subscribe to POLLIN if want_read is set, POLLOUT if want_write is set, and unconditionally listen for POLLERR. If poll() is interrupted by an operating system signal (errno == EINTR), the loop safely retries without crashing."
Step 4: Non-Blocking Read/Write Execution & Opportunistic Writes:
"Fourth, when `poll()` detects activity, handle_read() (server.cpp:703-741) ingests up to 64 KB chunks into the `incoming` buffer, feeding requests into the parser. If the request generates a response, the connection flips `want_read` to false and `want_write` to true. Rather than waiting for the next `poll()` cycle, the server immediately attempts an opportunistic synchronous write via handle_write() (server.cpp:680-701). In typical network conditions, small responses are written immediately, eliminating an entire event loop round-trip."
Step 5: Resource Teardown and Zombie Prevention:
"Finally, when a client closes the socket (read() returning 0), encounters a network error, or issues an invalid protocol payload, want_close is set. The server calls conn_destroy() (server.cpp:140-147), which closes the file descriptor, nulls out the lookup table in fd2conn, detaches the connection from the idle linked list, and deallocates the connection memory—preventing descriptor leaks and dangling pointers."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Non-Blocking Sockets | Configured O_NONBLOCK via fcntl |
Read/write syscalls return EAGAIN immediately if buffers are full/empty. |
Prevents a slow or malicious client from blocking the single server thread. | server.cpp:48-55 |
| 2. State Representation | Created struct Conn state container |
Encapsulates buffers, timer nodes, and want_read/want_write intent flags. |
Decouples network I/O events from high-level application protocol parsing. | server.cpp:78-98 |
| 3. Event Multiplexing | Assembled pollfd list for poll() |
Populates event masks dynamically based on Conn state flags; checks EINTR. |
Minimizes kernel wakeups by polling only for events the application actually needs. | server.cpp:846-869 |
| 4. Opportunistic I/O | Attempted immediate handle_write after read |
If client response is generated, invokes write() without re-polling. |
Drastically cuts median round-trip latency by avoiding an unnecessary poll() cycle. |
server.cpp:680-741 |
| 5. Connection Teardown | Centralized conn_destroy lifecycle |
Closes FD, clears fd2conn index, detaches idle node, and frees heap memory. |
Eliminates file descriptor leaks, dangling pointers, and memory growth. | server.cpp:140-147 |
poll() instead of spawning a thread per client or using a thread pool?Answer: In high-concurrency in-memory databases, memory throughput and CPU cache coherence are the primary performance bottlenecks, not CPU cores. Spawning a thread per client incurs prohibitive memory overhead (stack allocations of 2–8 MB per thread) and degrades throughput through OS preemption and thread context switches. Furthermore, multi-threading requires locking or latching across the entire hash table and trees, degrading latency under contention. A single-threaded reactor guarantees that all table lookups, insertions, and progressive rehashing steps execute locklessly and deterministically in single-digit microseconds.
Answer: In non-blocking TCP sockets, send buffers are almost always empty and ready to accept bytes. If an event loop subscribes to
POLLOUTon all sockets unconditionally,poll()returns immediately on every iteration, driving CPU utilization to 100% in a busy-wait spin. We solved this with explicit state intent flags (want_read,want_write): a connection only registersPOLLOUTwhenoutgoingdata is actually buffered. For reads, when a slow or fragmented packet arrives with fewer than 4 bytes,handle_readkeeps the partial slice inincomingand yields topoll(), avoiding thread blocking.
EMFILE), and verify under load?Answer: When
read()returns -1, the server checkserrno:EAGAINorEWOULDBLOCKindicates transient socket buffer exhaustion and yields safely, whereas any other error (such asECONNRESET) orread() == 0signifies that the remote peer sent a TCP FIN packet and closed the connection. Whenaccept()fails withEMFILEorENFILE(process descriptor table full), we log the failure and allow the event loop to continue without crashing. To verify stability, our Python stress harness (test_cmds.py) spawned concurrent connections executing thousands of pipelined queries, verifying zero FD leaks vialsof.