Module 11 · Lesson 1

WiFi (802.11): Architecture and Generations

20 min read
Article

Ten modules of this course built a toolkit. Module 11 spends it. Nothing in this lesson is a new mechanism — every piece of WiFi is something you already know under a different name: OFDM from M10-L2, MIMO from M10-L3, QAM from M5-L4, spectral efficiency from M6-L3, interference from M9-L3, and the TDD-like single channel M10-L4 already put in its table. What is new is the assembly, and one mechanism WiFi owns almost alone among the systems in this course: nobody schedules it. There is no base station handing out slots. Every radio decides for itself when to talk, and the entire design follows from that single decision.

A Standard With No Referee

802.11 was published in 1997 by the IEEE, and its founding constraint was regulatory rather than technical: it was designed for unlicensed spectrum, which anyone may transmit in without a licence and therefore without coordination. Compare that with the cellular systems of M10-L1, where an operator owns the band, owns every base station in it, and can therefore tell each handset exactly which slot and which subcarriers to use. A WiFi access point owns nothing. Its neighbour’s access point, three walls away, is on the same channel and has never heard of it, and neither has the microwave oven.

So the protocol has to be polite instead of scheduled. Politeness in radio has a name — carrier sense multiple access with collision avoidance, CSMA/CA — and the rest of this lesson is mostly a study of what it costs. That cost is the honest headline: WiFi’s advertised numbers are physical-layer rates, and the fraction of them a user sees depends on how many other radios are talking. Hold on to the distinction between PHY rate (what one frame’s bits travel at while it is on the air) and throughput (payload bits delivered per second of wall-clock time). They differ by a factor of two to ten.

Infrastructure, Ad-Hoc, and Four Acronyms

A group of 802.11 stations that talk to each other is a basic service set, a BSS, and there are two ways to form one. In an infrastructure BSS every frame goes through an access point, which also bridges to a wired network; two laptops in the same room send to each other via the AP, using the air twice. In an independent BSS — ad-hoc, or IBSS — stations talk directly with no AP at all. Nearly everything deployed is infrastructure, and the reason is not performance but the mechanisms of the next section: an AP is a single place to synchronise timing, buffer for sleeping devices, and hold the security association.

Association is the handshake that joins a station to a BSS: the AP beacons about ten times a second, the station either listens for beacons (passive scanning) or sends probe requests (active scanning), then authenticates, then associates, then runs the key exchange. Roaming is doing that again with a different AP in the same ESS, and here is the fact that matters — in plain 802.11 the station decides when to roam, entirely on its own, and the network cannot compel it. That is why a laptop clings to a distant AP at 6 Mbit/s while a better one sits overhead, and it is a direct consequence of the same no-referee design. M11-L2 will show cellular handover doing the opposite: the network decides, and the phone obeys.

CSMA/CA: Listen, Wait, Hope

Wired Ethernet solved the same sharing problem with CSMA/CD — collision detection. A transmitting station keeps listening, notices the voltage anomaly when someone else starts, aborts within a few bit times, and retries. WiFi cannot do this, and the reason is arithmetic you have already done twice in this course.

Why a radio cannot hear while it shouts

Detecting a collision means hearing a remote signal while transmitting. Take an access point transmitting at the common 20 dBm and receiving a signal near its sensitivity floor of −82 dBm for a mid-rate mode. Its own transmission arrives at its own receiver essentially at full power, so the receiver would have to see a wanted signal 102 dB below what its own power amplifier is putting out, at the same instant, on the same channel:

Why 802.11 Avoids Collisions Instead of Detecting Them
P_{tx} - P_{sens} = 20 - (-82) = 102\ \text{dB}
This is exactly M10-L4’s in-band full duplex problem wearing a different hat. That lesson priced a handset case at 123 dB (23 dBm transmit against a −100 dBm wanted signal); the WiFi AP case above is 102 dB, no easier in kind. Collision detection is same-channel, same-instant transmit-and-receive, so a radio that could do it could also do in-band full duplex — which M10-L4 established is a laboratory result. Hence the CA: 802.11 cannot detect a collision, so it must try to avoid one, and it learns a collision happened only by the absence of an acknowledgement.

That last sentence is the expensive one. On Ethernet a collision is discovered in microseconds and costs a fragment of a frame. On WiFi it is discovered after the whole frame has been sent and the ACK has failed to arrive — so a collision wastes the entire frame plus a timeout. Everything odd about the DCF timing below is an attempt to make that rare.

The DCF, step by step

The distributed coordination function is the basic access rule, and “distributed” is the whole point: no station is in charge, and every station runs the identical algorithm. Using 802.11a/n numbers at 5 GHz, where the slot time is 9 µs and SIFS is 16 µs:

  1. Sense. Listen. If the channel is busy, wait for it to go idle
  2. DIFS. Once idle, wait a distributed interframe space = SIFS + 2 slots = 16 + 2×9 = 34 µs. Nothing may be sent before this expires, which reserves the shorter intervals for higher-priority traffic
  3. Back off. Pick a random integer uniformly from 0 to CW, where the contention window starts at CWmin = 15, and count that many idle slots down. Two stations that both waited only DIFS would collide every time; the random backoff is what breaks the tie, and its average is 15/2 = 7.5 slots = 67.5 µs
  4. Transmit the frame when the counter reaches zero. Freeze the counter, do not restart it, if someone else starts first — so a station that has already waited a long time keeps its place in the queue
  5. SIFS, then ACK. The receiver waits a short interframe space of 16 µs — deliberately shorter than DIFS, so the acknowledgement always wins the channel against any new transmission — and returns a 14-byte ACK
  6. No ACK means retry. Double the contention window (15, 31, 63, … up to 1023) and go back to step 1. This is binary exponential backoff: the busier it gets, the more politely everyone waits

Two consequences are worth naming. First, the ACK is per frame and mandatory, which is why WiFi survives a channel that loses frames constantly. Second, because CW doubles on failure but resets to CWmin on success, DCF adapts to load without anyone measuring it.

The hidden node, and RTS/CTS

Carrier sense has a hole in it, and it is geometric. Suppose stations A and C are each within range of access point B but not of each other — opposite ends of a corridor, or either side of a wall. A senses the channel, hears nothing, and transmits. C senses the channel, also hears nothing, because A is hidden from it, and transmits too. Both frames arrive at B and destroy each other. Carrier sense worked perfectly at both stations and the collision happened anyway: sensing tells you about the air at your own antenna, not at the receiver’s.

The fix is to reserve the channel through the one radio both stations can hear. A station sends a short RTS (request to send, 20 bytes) naming the duration it needs; the AP answers with a CTS (clear to send, 14 bytes), which everyone in range of the AP hears, including the hidden station. Each station maintains a countdown called the network allocation vector from the duration field, and treats the channel as busy for that long even though it can hear nothing at all — virtual carrier sense. The trade is honest and it is why RTS/CTS is off by default for small frames: two extra frames and two extra SIFS cost roughly 40 to 60 µs of airtime per transmission, which is worth paying to protect a 1500-byte frame and absurd to pay to protect a 100-byte one. Note also the mirror-image problem, the exposed node: a station that defers because it hears a transmission whose receiver is nowhere near it wastes airtime on a collision that could not have happened. Both are the near-far geometry of M9-L3 reappearing as a protocol problem rather than a power problem.

What one frame actually costs

Now put numbers on it, because the overhead is fixed in time while the frame shrinks as the PHY rate grows — and that is the single most important thing to understand about fast WiFi. One successful exchange with a single station and no collisions is:

One DCF Cycle
T_{cycle} = \text{DIFS} + \overline{T_{bo}} + T_{data} + \text{SIFS} + T_{ACK}
Every term but Tdata is a constant number of microseconds, set by the slot time and the interframe spaces — and even Tdata contains a constant, the PHY preamble, which is 20 µs of training symbols sent at a fixed low rate no matter how fast the payload that follows is.
Throughput and Efficiency
S = \frac{L_{payload}}{T_{cycle}} \qquad \eta = \frac{S}{R_{PHY}}
Now repeat the calculation at 600 Mbit/s, holding every other term fixed: Tdata = 20 + 12 288/600 = 40.5 µs, Tcycle = 34 + 67.5 + 40.5 + 16 + 24.7 = 182.7 µs, and throughput = 12 000/182.7 µs = 65.7 Mbit/s — only 11.0% of the PHY rate. Eleven times the PHY rate bought 2.1 times the throughput. Making the radio faster makes the protocol less efficient, because the fixed overhead does not shrink.

This is why every generation from 802.11n onward makes frame aggregation mandatory in practice: put many payloads inside one transmission and pay the overhead once. Thirty-two 1500-byte frames in one aggregated MPDU at 600 Mbit/s take 20 + 32×12 320/600 = 677 µs of air, so Tcycle = 34 + 67.5 + 677 + 16 + 30.7 = 825.2 µs for 32 × 12 000 = 384 000 payload bits — 465 Mbit/s, or 77.6% of the PHY rate. That is the honest reason a modern link can approach the commonly quoted 60–70% of PHY rate: not because the overhead is small, but because it is amortised over dozens of frames. Without aggregation the same link would deliver 11%.

Be careful which number you quote. “WiFi delivers about 60–70% of its PHY rate” is a fair rule of thumb for an aggregated modern link with one active station. It is not a property of CSMA/CA, and the construction above shows both failure modes: a single unaggregated 1500-byte frame at 54 Mbit/s gets 57% and the same frame at 600 Mbit/s gets 11.0%. Add contending stations and it falls further. Any efficiency figure for WiFi is meaningless without stating the frame size, the aggregation and the number of stations.

The Generations, Rebuilt From Their Numerology

The marketing names — WiFi 4, 5, 6, 7 — were invented in 2018 by the Wi-Fi Alliance because 802.11ax means nothing to a shopper. The engineering names are the amendment letters, and every peak rate any of them advertises is constructible from four numbers you already have: how many data subcarriers the channel holds, how many bits each carries, the code rate, and how many spatial streams run in parallel.

Every WiFi Peak Rate, From First Principles
R_{PHY} = \frac{N_{sd}\,\log_2 M \; r \; N_{ss}}{T_{sym}}
Nsd is the number of data subcarriers (M10-L2 counted them: 48 of 64 in a 20 MHz 802.11a channel, 234 of 256 in WiFi 6), log₂M is bits per QAM symbol (M5-L4: 6 for 64-QAM, 8 for 256-QAM, 10 for 1024-QAM, 12 for 4096-QAM), r is the code rate, Nss is the number of spatial streams (M10-L3), and Tsym is the OFDM symbol period including its cyclic prefix. Nothing else. Four verifications follow, and all four land on the published figure exactly.

One row of the table below cannot be built this way, and the exception is instructive. 802.11b’s 11 Mbit/s is not OFDM at all — it is direct-sequence spread spectrum with complementary code keying, a single wide carrier rather than 64 narrow ones, so it has no subcarrier count and no symbol period to divide by. Every rate from 802.11a onward is OFDM, and every one of them yields to the formula above.

AmendmentMarketing nameYearBandsMax channelTop modulationMax streamsPHY peak
802.11—19972.4 GHz22 MHzDSSS, DBPSK/DQPSK12 Mbit/s
802.11b—19992.4 GHz22 MHzCCK (not OFDM)111 Mbit/s
802.11a—19995 GHz20 MHz64-QAM, r = 3/4154 Mbit/s
802.11g—20032.4 GHz20 MHz64-QAM, r = 3/4154 Mbit/s
802.11nWiFi 420092.4 + 5 GHz40 MHz64-QAM, r = 5/64600 Mbit/s
802.11acWiFi 52013 / 20165 GHz only160 MHz256-QAM, r = 5/686933 Mbit/s
802.11axWiFi 6 / 6E20212.4 + 5 (+ 6) GHz160 MHz1024-QAM, r = 5/689608 Mbit/s
802.11beWiFi 720242.4 + 5 + 6 GHz320 MHz4096-QAM, r = 5/61646 118 Mbit/s

Read the table as three levers pulled over and over, because that is nearly all it is: wider channels (20 → 40 → 80 → 160 → 320 MHz, a factor of 16), denser constellations (6 → 8 → 10 → 12 bits, a factor of 2), and more spatial streams (1 → 4 → 8 → 16, a factor of 16). Multiply: 16 × 2 × 16 = 512, so those three alone predict 54 × 512 ≈ 27 600 Mbit/s. The published ratio is 46 118/54 = 854, so something else contributed 854/512 = 1.67, and it is two things. The code rate rose from 3/4 to 5/6, which is (5/6)/(3/4) = 1.111. And the symbol itself got more efficient: 802.11a delivers 48 data subcarriers per 4.0 µs = 12.0 per µs in 20 MHz, while WiFi 7 delivers 3920 per 13.6 µs = 288.2 per µs in 320 MHz, which normalised to 20 MHz is 288.2/16 = 18.0 per µs — a factor of 18.0/12.0 = 1.501, and that is precisely M10-L2’s longer-symbol argument turning up as money. Check: 1.111 × 1.501 = 1.668, and 512 × 1.668 = 854 ✓.

Checking the rates against M6-L3

M6-L3 lists WiFi 6 at ~8 b/s/Hz per stream with 1024-QAM and a rate-5/6 code. Construct it and the two figures disagree slightly, in a way worth understanding. One WiFi 6 stream in a 20 MHz channel is 234 × 10 × (5/6) / 13.6 µs = 1950 bits / 13.6 µs = 143.4 Mbit/s, which over the nominal 20 MHz is 143.4/20 = 7.17 b/s/Hz. The difference is entirely OFDM overhead: 10 × 5/6 = 8.33 b/s/Hz is what the modulation and code deliver on a subcarrier, and multiplying by M10-L2’s useful fraction for WiFi 6, 0.914 × (12.8/13.6) = 0.860, gives 8.33 × 0.860 = 7.17 ✓. So M6-L3’s ~8 is the modulation-times-code figure rounded, and 7.17 is the delivered figure. Both are defensible; they are not the same quantity, and the honest way to quote WiFi 6 is 7.2 b/s/Hz per stream delivered.

Three Clean Channels in 83.5 Megahertz

Everything above assumed a channel was available. In the 2.4 GHz band that assumption is the problem, and the arithmetic is short. The band runs 2400 to 2483.5 MHz, so it is 2483.5 − 2400 = 83.5 MHz wide. Divide by 20 and you get 83.5/20 = 4.175, which looks like four channels — and that is the number people repeat. It is wrong, for two reasons that compound.

First, the channel centres sit on a 5 MHz grid: channel 1 is 2412 MHz, channel 2 is 2417, up to channel 11 at 2462 in North America and channel 13 at 2472 in Europe. Adjacent channel numbers therefore overlap almost completely — channels 1 and 2 are 5 MHz apart and each occupies about 22 MHz. Second, non-overlap requires the centres to be at least the occupied bandwidth apart, and on a 5 MHz grid the smallest legal step of at least 22 MHz is 25 MHz, or five channel numbers:

How Many Non-Overlapping Channels Fit
N_{clean} = \left\lfloor \frac{83.5 - 22}{25} \right\rfloor + 1 = 3
Place the first centre 11 MHz above the band edge, then step 25 MHz. The result is channels 1, 6 and 11 — 2412, 2437 and 2462 MHz — occupying 2401 to 2473 MHz. A fourth would need centre 2487, which occupies up to 2498 MHz and is 14.5 MHz outside the band, so it does not exist. Three, not four, and every access point in range of yours is on one of them.

Compare the other bands and the congestion explains itself. 5 GHz offers roughly 500 MHz across three sub-bands, which is about 25 non-overlapping 20 MHz channels — but much of it is subject to dynamic frequency selection, a regulatory duty to detect radar and vacate the channel, which some devices avoid entirely. The 6 GHz band opened for WiFi 6E runs 5925 to 7125 MHz: 1200 MHz, which is 1200/20 = 59 channels of 20 MHz (59 × 20 = 1180 ≤ 1200 ✓), or 1200/160 = 7 channels of 160 MHz, or just three of WiFi 7’s 320 MHz (3 × 320 = 960 ≤ 1200 ✓). Note what that last figure means: WiFi 7’s widest channel is so wide that even the enormous 6 GHz band holds only three of them, so 320 MHz is a feature for a quiet house and not for an apartment block. And 2.4 GHz gets worse still, because M11-L4 will show Bluetooth, BLE and Zigbee living in the same 83.5 MHz.

WiFi 6: Five Features, One Enemy

WiFi 6 is the first generation whose headline is not the peak rate. Its stated design goal was performance in dense deployments — a lecture hall, a stadium, an apartment block — and every one of its features attacks the cost the airtime arithmetic above exposed. Read them as answers rather than as a feature list:

The longer symbol (M10-L2’s 78.125 kHz spacing and 12.8 µs useful period) is the enabler underneath three of those five, and it also raised the OFDM efficiency from 60% to 86% and cut the guard-interval overhead from 20% to 5.9%. It is not listed as a feature on any box, and it is the largest single engineering improvement in the generation.

WiFi 7: Wider, Denser, and in Two Places at Once

802.11be, ratified in 2024, pulls the same levers once more and adds one genuinely new idea:

Where This Goes

Hold on to one sentence, because the next two lessons are built against it: WiFi is contended and cellular is scheduled. A WiFi station asks the air for permission and sometimes loses; a cellular device is told when to transmit and never collides. That single difference explains almost every other contrast between them — why WiFi is cheap and unlicensed and degrades unpredictably in a crowd, and why cellular is expensive and licensed and holds up. M11-L2 and M11-L3 walk the cellular generations, where the scheduler in the base station does the job DCF does with dice. M11-L4 returns to this band with Bluetooth, BLE and Zigbee, all of them sharing the same three clean channels of 2.4 GHz, and asks how they coexist.

Key Takeaways

Previous: Duplexing (FDD vs. TDD) Overview Next: Cellular — From 2G to 4G LTE