Skip to main content
Back to blog

Packet Loss Is a Full Queue Before It Is a Broken Line

Technical

Packet Loss Is a Full Queue Before It Is a Broken Line article illustration

A packet loss figure reads like an accusation. Some share of what you sent did not come back, and the natural conclusion is that something along the way is broken.

Sometimes it is. But a full queue produces the same number by design, and so does a test that stopped waiting too early, and neither of those needs repairing. Knowing which one you are holding is the entire diagnosis.

Why a device drops a packet on purpose

Network equipment holds packets in a queue while it waits for the outgoing link to become free. That queue has a size, and when it is full the device has to do something with the next packet that arrives.

What it does is drop it. RFC 7567, the IETF’s best current practice on queue management, published as BCP 197 in July 2015 and replacing RFC 2309, names the traditional arrangement tail drop and defines it by what happens at the moment the queue fills: “the packet that arrived most recently (i.e., the one on the tail of the queue) is dropped when the queue is full.”

That is not a malfunction. RFC 7567 describes what the drop does: “In the current Internet, dropped packets provide a critical mechanism indicating congestion notification to hosts.” Absent explicit marking, discarding a packet is how a device in the middle tells a sender at the edge to ease off.

The document also states what the queue is there for. “The goal of buffering in the network is to absorb data bursts and to transmit them during the (hopefully) ensuing bursts of silence.” A queue is shock absorption. When the shock outlasts the absorber, something has to give, and what gives is a packet.

Worth noting what RFC 7567 lists as the drawback here, because it is not the dropping. It is that tail drop “signals congestion (via a packet drop) only when the queue has become full”, so the queue sits full for long stretches and everything behind it waits. That failure mode shows up as delay rather than loss, and it belongs to what counts as a good ping, which covers bufferbloat and the jitter it produces.

The sender is listening for exactly this

A drop is only useful if something acts on it, and something does.

RFC 5681, the September 2009 specification of TCP’s congestion control algorithms, says where its own signal comes from: “Also, note that the algorithms specified in this document work in terms of using loss as the signal of congestion.” A transfer that meets a dropped packet slows down, retransmits and builds back up. The data still arrives. What it cost you was time.

This inverts the usual reading of a loss figure. Loss on a saturated link is the control system engaging. Loss on an idle link, where nothing is competing and no queue should be anywhere near full, is the reading that deserves attention, because the ordinary explanation does not apply to it.

It also explains why a transfer can be slow while no single number looks broken, which is one reason a transfer’s result and a path’s capacity are different quantities, set out in latency, bandwidth and throughput.

A packet is lost when you stop waiting for it

Loss is not an observable event. Nothing arrives to tell you a packet died, so every measurement of loss is really a measurement of patience.

The IETF pins this down. RFC 7680, Internet Standard 82, published in January 2016 and making RFC 2680 obsolete, defines one-way packet loss using a parameter it calls Tmax, “a loss threshold waiting time”. The method turns on that threshold: “If the packet fails to arrive within a reasonable period of time, Tmax, the one-way packet loss is taken to be one.” The document adds that what counts as reasonable is itself a parameter of the metric.

A parameter, not a fact about the network. It spells out the consequence of choosing badly: “if the loss threshold is set too low, then many packets may be counted as lost.” It also requires the threshold to be reported alongside the figure: “The threshold, Tmax, between a large finite delay and loss (or other methodology to distinguish between finite delay and loss) MUST be reported.”

So a loss percentage on its own is an incomplete number. A packet that took four seconds and a packet that never existed look identical in the total, and the waiting time is what separates them.

The instrument gets counted too

RFC 7680 lists three sources of error in a loss measurement. The third is the measuring equipment itself: “The last source of error, resource limits, cause the packet to be dropped by the measurement instrument and counted as lost when in fact the network delivered the packet in reasonable time.”

The reply arrived. The thing measuring was too busy to notice. That goes down as loss.

The tool also changes what you are measuring. Command-line ping and traceroute send ICMP, and a router treats producing an ICMP reply as a chore rather than as traffic, which is why a hop can answer badly while forwarding perfectly. Reading a traceroute covers that in full, including what a missing reply does and does not prove. A browser cannot send ICMP at all, so a browser-based test measures round trips of ordinary web requests instead. Neither is your real traffic. Both are proxies for it, and they fail in different directions.

A Wi-Fi link that loses a frame can hide the fact from you, because it retransmits underneath IP before the packet is ever declared missing.

RFC 3819, the IETF’s advice to subnetwork designers, published as BCP 89 in July 2004, names the trade that buys: “Many subnetwork designers have opportunities to reduce the probability of packet loss, e.g., with FEC, ARQ, and interleaving, at the cost of increased delay.” It is specific about which technique costs what. “While ARQ increases delay variance, FEC does not.”

ARQ is the retransmission one, and it is what Wi-Fi does. So a wireless link in trouble can hand you a clean loss figure while its delay figures come apart, which is the opposite of the signature you went looking for. The document asks for all three to be held down together: “Subnetwork designers should therefore minimize all three parameters (delay, delay variance, and packet loss) as much as possible.”

Read the delay variation alongside the loss, not after it.

Finding a real one

Once the ordinary explanations are exhausted, localising a genuine fault is mostly a matter of changing one thing at a time.

Start on a cable. Our Ping Test page puts wireless first in its own list of things to rule out, and a wired run takes the whole retransmitting link layer out of the picture. If the cable is clean and the wireless is not, you have finished.

Then compare idle against loaded. Run our Speed Test in one tab to fill the line and the Ping Test in another while it works. Loss that appears only when the line is full is the queue behaving as RFC 7567 describes. Loss that is there on a quiet line is not, and that is the one worth chasing. The matching test for delay under load is in what counts as a good ping.

Take more than one sample. RFC 7680 builds its percentage out of a sequence of individual observations rather than a single one, and for good reason: one packet failing to arrive is an event, while a rate needs a set. Our Ping Test sends twenty round trips, which is a set rather than a verdict.

Change the destination. Loss belongs to a path, not to your connection, and a second destination costs nothing to try. If one target loses packets and another does not, the fault sits somewhere along the first path and possibly nowhere near you.

All four steps do the same job: exhausting the innocent explanations before accepting the guilty one. A queue that fills under load is the network working, a measurement is only as good as the waiting time behind it, and a wireless link reports its own trouble in a different column. When all of that has been ruled out and the loss is still there, on a cable, on an idle line, to more than one destination, you have something worth reporting. Until then you have a figure, and a figure is not yet a fault.