Topic 02: Binary Framing, TCP Pipelining & Type Serialization

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering length-prefixed protocol framing, stream buffer preservation for TCP pipelining, defensive input bounds, and type-tagged binary serialization.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Length-Prefixed Framing 4-byte prefix on all messages Verifies buffer has $\ge 4$ bytes and $4 + len \le size$; enforces 32 MB cap. Eliminates delimiter parsing ambiguities and guards against memory exhaustion. server.cpp:643-678
2. Argument Deserialization Parsed [nstr][len, data]... Sequential reading via pointer bounds checks (read_u32, read_str). Validates structural integrity and blocks malformed payload attacks. server.cpp:149-165
3. TCP Pipelining Looped try_one_request + buf_consume Slices parsed commands from vector front; retains remaining bytes in-place. Enables high throughput by executing batched commands in a single network trip. server.cpp:74-76, 730
4. Two-Pass Framing Reserved 4-byte header offset Serializes body first, then computes and backpatches final byte length. Avoids pre-calculating dynamic response sizes or performing redundant memory copies. server.cpp:623-641
5. Type-Tagged Serialization Prefixed payloads with 1-byte tags Dispatches typed encoders (out_nil, out_str, out_int, out_dbl, out_arr). Provides unambiguous data typing over raw byte sockets without JSON/text overhead. server.cpp:218-268

❓ Anticipated Interview Questions & Crisp Answers

Q1: Why build a custom length-prefixed binary protocol rather than standard HTTP/JSON or Redis text RESP?

Answer: Text protocols like HTTP/JSON or Redis RESP require byte-by-byte scanning to find line breaks (\r\n), delimiter escaping, and expensive string-to-number conversions (atoi/atof). Binary length-prefixed framing allows $O(1)$ size validation, zero-copy pointer arithmetic, and direct memory transfers via memcpy. Furthermore, binary serialization handles arbitrary binary blobs (images, serialized protobufs, raw keys) without escape characters or base64 bloat.

Q2: What bugs or packet corruption occurred during TCP pipelining, and how did buf_consume solve them?

Answer: When pipelined clients batch commands, a single read() syscall can return 64 KB containing three complete commands followed by a partial 10-byte fragment of a fourth command. If the server reset or wiped the buffer after executing command 1, commands 2, 3, and the partial fragment would be lost, causing immediate protocol desynchronization. If the server moved bytes naively with overlapping buffers, memory corruption could occur. We implemented buf_consume(incoming, 4 + len) using memmove() to slide the unparsed remainder to the buffer front, executing in a while (try_one_request(conn)) loop until no complete frames remain.

Q3: How did you prevent buffer overflow and integer truncation attacks when decoding length headers?

Answer: Untrusted clients can send negative integers, 0xFFFFFFFF (4 GiB), or values exceeding available RAM. If parsed as a signed integer, a large length prefix could cause integer overflow during buffer math (4 + len). We enforced strict unsigned 32-bit arithmetic (read_u32), capped payloads with k_max_msg = 32 * 1024 * 1024 (32 MB), and capped argument counts with k_max_args = 200 * 1000. Any violation immediately sets want_close = true and tears down the connection without memory allocation.