Latency numbers
Jeff Dean’s list with the measurements brought forward, the sources named, and storage, cryptography, GPUs and model inference added. Every row points at where its number came from.
- nanoseconds silicon
- microseconds devices and local network
- milliseconds distance and mechanics
- seconds people notice
Updated 1 September 2026 / 24 sources / scale is logarithmic, 0.1 ns to 100 s
What these numbers are
Medians on an idle machine, rounded for memorisation rather than precision. They are for deciding whether a design is off by an order of magnitude, not for predicting production latency.
What they are not
They contain no queueing, no contention and no tail. A saturated system behaves nothing like this table. Treat every figure as the best case the hardware allows.
Cycles travel, nanoseconds do not
Cache latency has stayed near-constant in cycles for a decade while clocks moved, so cycle counts age well and nanosecond figures age badly. Both are given where it matters.
Measure before you trust
Storage and network figures vary by more than 2x across vendors and instance types. Where a number matters, benchmark your own hardware with fio or a load generator and replace the row.
Processor and memory
Measured on current x86 server and desktop parts. Cycle counts are stable across generations; nanoseconds are not, because clocks move.
Storage
Device-level numbers at low queue depth. Filesystem, RAID and virtualisation layers each add their own overhead on top.
Network
Round trip times unless stated otherwise. Propagation is physics and cannot be optimised away; everything else can.
Software operations
Per mebibyte, single core, on data already resident in memory.
Language model inference
Added to the canonical list in May 2026. These move faster than anything else on this page, so treat them as a snapshot.
Human perception
The reason any of the rest matters. These thresholds have not changed since Nielsen measured them in 1993.
What follows from the table
1 ms of round trip buys about 50 km
Fibre costs 4.9 us per km each way, so a millisecond covers roughly 100 km of glass. Real routes run 1.5 to 2x the great-circle distance, so halve it. No protocol removes this.
Durability costs about 150x
The same 8 KiB write is 2 us buffered and 300 us with fsync. If a design needs more write throughput, the question is almost always how many fsyncs it issues, not how fast the drive is.
Random access costs about 100x on flash
Sequential 8 KiB from NVMe is 1 us; random is 100 us. On a spinning disk the same ratio is about 350x: 250 MiB/s sequential against 0.7 MiB/s random.
Queueing beats latency above 70% utilisation
For an M/M/1 queue, response time is service time divided by (1 - utilisation). At 50% load you pay 2x, at 80% you pay 5x, at 95% you pay 20x. Every number on this page is an unloaded number.
p99 is not p50
Anything that crosses a network or a device boundary typically shows a p99 five to ten times its median. Capacity plans built on averages fail at the tail.
Check the exponent, then the coefficient
Back-of-envelope work only needs the order of magnitude right. Keep the units attached as a checksum and round aggressively.
Sources
Numbers were taken from these and not from memory. Where two sources disagreed, both are cited and the range is shown in the table. Entries marked further reading back no single row.
- 1Latency Numbers Every Programmer Should Know
Jeff Dean and Peter Norvig, maintained by Jonas Boner. The canonical list. Rows for language model inference were added in May 2026.
- 2Teach Yourself Programming in Ten Years, answers table further reading
Peter Norvig. The original source of the numbers.
- 3napkin-math
Simon Eskildsen. Benchmarked, referenced and actively maintained. The single-host rows were re-measured on 8 March 2026 on a GCP c4-standard-48-lssd instance (Intel Xeon 6985P-C, 24 physical cores, 180 GB RAM, Ubuntu 22.04.5). Most throughput figures on this page come from here.
- 4Interactive latency page further reading
Colin Scott. The same numbers extrapolated across years, useful for seeing which of them actually move.
- 57-cpu.com
Per-microarchitecture cache and memory latency measured in cycles and nanoseconds. The reference for cache hierarchy figures.
- 6Zen 5 cache latencies
HWCooling. L1d 48 KiB at 4 cycles, L2 1 MiB at 14 cycles, L3 about 46 cycles.
- 7Ryzen 9950X, Zen 5 on desktop
Chips and Cheese. Sampled memory latency under real workloads.
- 8Running gaming workloads through Zen 5
Chips and Cheese. Comparison of Zen 5 and Lion Cove cache latency.
- 9Intel Memory Latency Checker results
Ice Lake server: L1 1.5 ns, L2 4.1 ns, L3 22.5 ns.
- 10Zen 5 variants, clock for clock
Chips and Cheese. Pointer chase through a 1 GB array measured above 128 ns.
- 11Dissecting the NVIDIA Hopper architecture
arXiv 2501.12084. Global memory latency in cycles for A100, H800 and RTX 4090.
- 12
- 13NVMe queue depth explained
Optane queue-depth-1 read latency under 10 us; NAND NVMe under 50 us.
- 14Micron 7500 NVMe SSD product brief
Typical read 70 us, typical write 15 us, p99 read 80 us, measured with fio at 4 KiB and queue depth 1.
- 15Seagate Nytro XP1920SE specification
Average 4 KiB queue-depth-1 read latency 75 us, write 12 us.
- 16NVMe latency, typical numbers
simplyblock. NVMe 20 to 70 us, SATA SSD 100 to 200 us, HDD 5 to 10 ms.
- 17Calculating optical fibre latency
m2optics. 4.9 us per km at a refractive index of 1.47.
- 18Dissecting latency in the internet's fibre infrastructure
arXiv 1811.10737. Uses 204,000 km/s and compares fibre types.
- 19Hashing vs encryption, SHA-256 and AES-256 compared
SHA-NI at 1.8 cycles per byte; 2.8 to 3.2 GB/s measured on a Ryzen 9 7950X.
- 20OpenSSL Cookbook, performance chapter
Ivan Ristic. aes-128-gcm near 2 GB/s per core at 8 KiB blocks, and far slower at 16 B.
- 21wolfCrypt benchmarks
Ed25519 sign 53 us, verify 184 us on a 2.5 GHz Core i7.
- 22Response Times: The 3 Important Limits
Jakob Nielsen, 1993. 0.1 s, 1 s and 10 s.
- 23Introducing RAIL
Smashing Magazine. Adds the 16 ms animation frame to Nielsen's thresholds and splits the 100 ms budget into 50 ms of work plus rendering.
- 24cloudping.co further reading
Live inter-region round trip times between AWS regions. Use this instead of a static table when the exact pair matters.
Storage, network and software throughput figures are dominated by Simon Eskildsen's napkin-math, re-measured on 8 March 2026. Cache and memory figures come from 7-cpu.com, Chips and Cheese and Intel MLC runs. Perception thresholds are Nielsen's, via the RAIL model. Re-check anything in the inference section first; it has the shortest half-life on this page.