Module 9 ยท Lesson 2

Bit Error Rate (BER)

15 min read
Article

The previous lesson built the noise floor and turned it into a signal-to-noise ratio. Module 5 turned that ratio into an error probability: Pb = Q(√(2Eb/N₀)) for BPSK and QPSK, Pb = Q(√(Eb/N₀)) for OOK and BFSK, and a 3 dB gap between the two families (M5-L2). This lesson does not re-derive any of that. It asks three questions the formulas leave open: what exactly is the Q in them, what shape does the curve have, and — the one that decides real designs — what value of BER is actually good enough. The last answer is the surprising one: in LTE and 5G, nobody is trying to make the raw BER small.

What BER Actually Means

The bit error rate is the expected number of bits the receiver gets wrong, divided by the number of bits sent. Nothing more. It is dimensionless, it lives between 0 and 0.5 for any useful link, and it is written either as a decimal (0.000001) or a power of ten (10−6), the second being universal in engineering because the interesting range spans nine decades.

Definition of Bit Error Rate
\text{BER} = \frac{\mathbb{E}[\text{bits received in error}]}{\text{bits transmitted}}
The expectation matters. BER is a probability, not a count — a property of the channel and the receiver, which a measurement only estimates. The measured ratio of errors to bits is called the bit error ratio by careful standards documents, and it converges on the underlying rate only as the test gets long.

That distinction sounds pedantic until you try to measure a small one. Errors at a low BER are rare, independent events, so the count in a fixed-length test is Poisson-distributed, and the fractional uncertainty of a Poisson count of N errors is 1/√N. To get to roughly 10% confidence you need N ≈ 100 errors, because 1/√100 = 0.1. Fewer than about 10 errors and your estimate is worth almost nothing — a single unlucky burst moves it by a factor of two.

Worked example: how long to measure 10−9

Suppose the true BER is 10−9 and you want those 100 errors. You need 100 / 10−9 = 1011 bits — a hundred billion of them. At a line rate of 1 Gbps that is 1011/109 = 100 seconds, which is tolerable. Now change one number at a time:

The last line is the practical point. Very low BERs are not measured; they are extrapolated. You measure the curve where errors are plentiful — 10−4 to 10−6 — then deliberately degrade the link, fit the measured points to the theoretical shape, and read off where the curve would cross 10−12. That is what the rest of this lesson is for: the theoretical shape is what makes the extrapolation legitimate.

BER Is Not the Rate Your User Feels

Bits do not travel alone. They travel in frames, blocks and packets, and a packet with one bad bit is usually a packet thrown away entirely, because the cyclic redundancy check at the end of it fails. So the number that determines what a user experiences is not the bit error rate but the packet error rate (PER), also met as frame error rate (FER) and, in LTE and 5G, as block error rate (BLER). If bit errors are independent, the arithmetic linking them is elementary:

Packet Error Rate from Bit Error Rate
\text{PER} = 1 - (1-p)^{L} \approx 1 - e^{-Lp}
A packet survives only if every one of its L bits survives, hence (1−p)L. The exponential form is the small-argument approximation, accurate whenever Lp is not large, and it makes the scaling obvious: PER is set by the product Lp, so a longer packet is a more fragile packet at the same BER.

Take a full-size Ethernet frame, 1500 bytes, so L = 1500 × 8 = 12,000 bits, on a link with a perfectly respectable BER of p = 10−6. Then Lp = 12,000 × 10−6 = 0.012, and PER = 1 − e−0.012 = 1 − 0.988072 = 0.011928, i.e. 1.19 × 10−2 — about 1.2% of packets lost. The exact binomial form gives 1 − (1 − 10−6)12000 = 0.0119283, agreeing to five figures.

Sit with that number, because it is the most useful thing on this page. “One error in a million bits” sounds like an excellent link. It is a link losing more than one packet in a hundred — enough to cripple a TCP transfer, since classic TCP congestion control reads loss as congestion and backs off. The learner’s instinct that BER 10−6 is fine is wrong by a factor of ten thousand in the quantity that matters, and the factor is exactly L. Run the same frame at p = 10−9 and Lp = 1.2 × 10−5, so the PER is 1.2 × 10−5: about one packet in 83,000, which is a healthy link.

The Q-Function

Every formula in M5-L2 was a Q of something, and it is time to say what Q is. Thermal noise is Gaussian (M9-L1), and a decision error happens when the noise sample is large enough, and in the right direction, to push the received point across the decision boundary. So the error probability is the area in the tail of a Gaussian, and Q is the name of that area for the standardised case — zero mean, unit variance:

The Gaussian Tail Probability
Q(x) = P(Z > x) = \frac{1}{\sqrt{2\pi}}\int_{x}^{\infty} e^{-t^2/2}\, dt
Read Q(x) as “the probability that a standard normal variable exceeds x”, and read x itself as a distance measured in standard deviations. That is exactly how M5-L3 used it: the argument was dmin/2σ, a voltage gap divided by a noise sigma. Q has no closed form, which is why it is tabulated and approximated rather than evaluated.

Four properties carry almost all the practical weight, and they are worth checking rather than memorising. Q(0) = 0.5, because half of a symmetric distribution lies above its mean — and this is the sanity check that a BER can never exceed 0.5 for a sensible receiver. Q(x) + Q(−x) = 1, from the same symmetry, which is what lets you write a two-sided error as twice a one-sided tail. Q is strictly decreasing, so it has an inverse Q−1, and we will need it in a moment. And Q is a rescaled complementary error function, which is how you compute it in any language that has erfc:

Q in Terms of erfc, and a Usable Approximation
Q(x) = \tfrac{1}{2}\,\mathrm{erfc}\!\left(\tfrac{x}{\sqrt{2}}\right) \qquad Q(x) \approx \frac{e^{-x^2/2}}{x\sqrt{2\pi}}, \;\; x \gg 1
The erfc identity is exact; the second expression is the standard large-argument approximation. It is always high, and it tightens as x grows: at x = 3 it gives 1.477 × 10−3 against a true 1.350 × 10−3 (+9.4%), at x = 4 it gives 3.346 × 10−5 against 3.167 × 10−5 (+5.6%), and at x = 6 it is 1.013 × 10−9 against 9.866 × 10−10 (+2.6%). For back-of-envelope work in the interesting region that is plenty.

Here is the table, computed with the erfc identity. Read it once and the waterfall shape stops being mysterious — look at what one extra unit of x does:

xQ(x)Order of magnitudeRatio to previous row
05.000 × 10−1a coin flip—
11.587 × 10−110−1÷ 3.2
22.275 × 10−210−2÷ 7.0
31.350 × 10−310−3÷ 16.9
43.167 × 10−510−5÷ 42.6
52.867 × 10−710−7÷ 110
69.866 × 10−1010−9÷ 291
71.280 × 10−1210−12÷ 771

Two entries are worth committing to memory because they are the anchors of every BER discussion. Q(4.753) = 1.00 × 10−6 and Q(5.998) = 1.00 × 10−9. Notice how little the argument had to grow: three more decades of reliability cost 5.998 − 4.753 = 1.245 in x, about 26% more distance from the decision boundary. And look at the last column — the divisor grows every row, which is the numerical fingerprint of the e−x²/2 factor. The tail does not decay exponentially; it decays like a Gaussian, which is far faster.

The Waterfall, and Why the Cliff Exists

Plot BER on a logarithmic vertical axis against Eb/N₀ in decibels on a linear horizontal axis, and every coherent modulation scheme draws the same picture: a lazy shoulder at low SNR where the link is useless, then a plunge so steep it is universally called the waterfall. The name is descriptive, and its steepness is measurable. Using the BPSK/QPSK law and the required-Eb/N₀ arithmetic from the next section:

So “about a decade of BER per decibel” is a fair rule of thumb around 10−6, and the curve keeps getting steeper below that. Now read those numbers backwards, which is the direction the world actually moves in: losing 1 dB of margin near 10−6 multiplies your error rate by ten, and losing 3 dB multiplies it by roughly a thousand.

That is the whole explanation of the cliff effect that M5-L1 introduced as a bare observation — digital links working perfectly and then failing abruptly, digital television freezing rather than fading into snow. There is no separate cliff mechanism. The cliff is the waterfall, seen from the outside: a fading channel whose received power wanders a few decibels (M8-L3) drags the BER across several decades, and because the error-correcting code either succeeds or fails, the user sees a service that is either perfect or absent. Analog degrades smoothly because a 3 dB SNR loss makes the audio 3 dB noisier. Digital does not, because a 3 dB loss makes it a thousand times wronger.

The waterfall shape is what licenses extrapolation. Because the theoretical curve is fixed and steep, a handful of measured points at 10−4 and 10−5 pin the whole curve down: if your measurements lie parallel to theory a couple of decibels to the right, the implementation loses those decibels everywhere, and the 10−12 point follows. If the measured points are not parallel — if the curve flattens into a floor — something other than Gaussian noise is limiting the link, and no amount of extra power will fix it. That flattening is the signature of interference or a hardware imperfection, and it is M9-L3’s subject.

Required Eb/N₀ for a Target BER

In design you rarely want the BER at a given SNR. You want the reverse: given a target BER, what Eb/N₀ must the link deliver? For BPSK and QPSK, invert Pb = Q(√(2Eb/N₀)) in two steps — apply Q−1 to both sides, then square and halve:

Inverting the Coherent BER Law
\frac{E_b}{N_0} = \tfrac{1}{2}\left[Q^{-1}(P_b)\right]^2 \qquad \frac{E_b}{N_0} = \left[Q^{-1}(P_b)\right]^2
The left form is BPSK and QPSK; the right is OOK and BFSK, which lack the factor of 2. The two required ratios therefore differ by exactly a factor of 2 at every target BER, which is 10 log₁₀(2) = 3.01 dB — the M5-L2 gap, now visible as a constant horizontal offset between two curves rather than as a claim about one operating point.

Work the 10−6 case in full. Q−1(10−6) = 4.7534, from the table above. Square it: 4.7534² = 22.595. Halve it: Eb/N₀ = 11.2975 as a linear ratio. In decibels, 10 log₁₀(11.2975) = 10.53 dB. And that is a genuine consistency check on the course, not a coincidence: M6-L4’s table lists 10.5 dB as the uncoded QPSK Eb/N₀ requirement at BER 10−6, derived there from the square-QAM family. Two different routes, the same number. Here is the family:

Target BERQ−1(Pb)Eb/N₀ linear (BPSK/QPSK)dB (BPSK/QPSK)dB (OOK/BFSK)
10−22.32632.7064.32 dB7.33 dB
10−33.09024.7756.79 dB9.80 dB
10−43.71906.9168.40 dB11.41 dB
10−64.753411.29810.53 dB13.54 dB
10−95.997817.98712.55 dB15.56 dB
10−127.034524.74213.93 dB16.94 dB

Check the last column against the fourth in any row and the difference is 3.01 dB, every time. Then check the compression down the column: going from 10−3 to 10−12 — nine decades of reliability — costs 13.93 − 6.79 = 7.14 dB, less than a factor of six in power. Reliability is astonishingly cheap in a clean Gaussian channel, and that fact is easy to forget when the same 7 dB is impossible to find in a link budget (M9-L4). Higher-order modulation shifts these numbers up without changing their character: M6-L4 lists 14.4 dB for 16-QAM, 18.8 dB for 64-QAM and 23.5 dB for 256-QAM at the same 10−6, because packing more bits per symbol crowds the constellation and shrinks dmin (M5-L3).

Coding Gain: Buying dB with Bandwidth

Everything above is uncoded. Add forward error correction (M6-L2, M6-L3) and the curve moves bodily to the left: the same BER is reached at lower Eb/N₀. The horizontal distance between the coded and uncoded curves, measured at a stated target BER, is the coding gain, and it is one of the few genuinely free-feeling wins in radio — except that it is not free, and the price is exactly the one M6-L4’s triangle predicted.

There is a ceiling on this, and it is worth computing so the numbers above do not look arbitrary. A rate-r code cannot beat its own capacity bound, Eb/N₀ ≥ (2r − 1)/r. For r = 1/2 that is (1.41421 − 1)/0.5 = 0.8284, which is 10 log₁₀(0.8284) = −0.82 dB. Since uncoded BPSK needs 12.55 dB at 10−9, the largest coding gain any rate-1/2 code could ever show at that BER is 12.55 − (−0.82) = 13.37 dB. Real LDPC lands one to two decibels short of the bound, which is precisely where the 10–11 dB figure comes from. And as r → 0 the bound approaches ln 2 = −1.59 dB, the absolute floor from M6-L1 — below which no code of any rate can operate.

How Good Is Good Enough?

Now the engineering question. There is no universal acceptable BER, because the tolerance belongs to the application, not the radio. What a human ear forgives, a bank ledger does not.

ApplicationTypical delivered targetWhy that level
Digital voice (speech codec)≈ 10−3An occasional corrupted sample is masked by the codec or heard as a faint click; intelligibility survives. Latency matters far more than perfection, so retransmission is the wrong tool.
Video streaming≈ 10−6 after FECA lost packet corrupts a whole macroblock and propagates until the next key frame, so the visible artefact outlasts the error. Buffering allows some retransmission, but not much.
File transfer, web, financeEffectively 10−12 or betterOne wrong bit can be one wrong digit. In practice reliability comes from CRC plus retransmit-until-correct at a higher layer, not from the radio alone.
Deep-space telemetry10−6 or lower on almost no powerRetransmission costs hours of round-trip time, so the answer is very low code rates, huge antennas, and every decibel of coding gain that exists.

Read the middle column carefully and a pattern appears: the numbers are delivered targets, measured at the top of the stack after coding and retransmission have done their work. Almost nowhere is that number the raw BER on the air interface, and this is where naive intuition fails hardest.

Modern systems aim for a high raw BER on purpose

LTE and 5G do not try to keep the raw error rate low. They deliberately run the physical layer hot — a modulation and coding scheme chosen so aggressively that the first transmission of a block fails about 10% of the time (M6-L4). A 10% block error rate would be an alarming number if it were the final answer. It is not: HARQ retransmits the failed block with extra parity, the receiver combines the attempts, and the residual error rate after one or two retries falls to 10−3 or below at the transport layer.

The logic is throughput, and it is the same logic as M6-L4’s adaptive modulation. A link tuned so that it never errs is a link that was transmitting at a lower rate than the channel could support, and every one of those unused bits is gone forever. A link tuned to fail 10% of the time and repair the failures is running much closer to capacity, and pays for it in a little delay. So raw BER is an intermediate quantity, not a goal. It is the input to a code and a retransmission scheme, and the right question is never “is my BER low?” but “does my BER, fed through this code and this HARQ scheme, deliver the reliability and delay this application needs?”

Which Eb/N₀ and which BER. Two ambiguities cause most confusion in datasheets. First, pre- and post-decoding BER are different numbers that both get called BER; a receiver reporting 10−2 at its demodulator output and 10−9 after the decoder is behaving normally. Second, Eb is energy per information bit, not per channel bit, which is what makes coding gain a fair comparison at all — and Eb/N₀ relates to SNR through SNR = (Eb/N₀)(Rb/B) from M9-L1. State which bit, which side of the decoder, and which bandwidth, or the number means nothing.

Key Takeaways

Previous: Thermal Noise and SNR Overview Next: Sources of Interference