cpus.me

Latency numbers

Jeff Dean’s list with the measurements brought forward, the sources named, and storage, cryptography, GPUs and model inference added. Every row points at where its number came from.

1 ns1 us1 ms1 sL1 hitMain memoryNVMe random readHDD seekSame-zone round tripIntercontinental round tripFrontier LLM first token60 Hz frameAttention lost
  • nanoseconds silicon
  • microseconds devices and local network
  • milliseconds distance and mechanics
  • seconds people notice

Updated 1 September 2026 / 24 sources / scale is logarithmic, 0.1 ns to 100 s

What these numbers are

Medians on an idle machine, rounded for memorisation rather than precision. They are for deciding whether a design is off by an order of magnitude, not for predicting production latency.

What they are not

They contain no queueing, no contention and no tail. A saturated system behaves nothing like this table. Treat every figure as the best case the hardware allows.

Cycles travel, nanoseconds do not

Cache latency has stayed near-constant in cycles for a decade while clocks moved, so cycle counts age well and nanosecond figures age badly. Both are given where it matters.

Measure before you trust

Storage and network figures vary by more than 2x across vendors and instance types. Where a number matters, benchmark your own hardware with fio or a load generator and replace the row.

Processor and memory

Measured on current x86 server and desktop parts. Cycle counts are stable across generations; nanoseconds are not, because clocks move.

Operation
Latency
Notes
Src
One CPU cycle at 5 GHz
0.2 ns
The floor. Everything below is measured in multiples of this.
L1 data cache hit
1 ns
4 to 5 cycles. Zen 5 holds 4 cycles at 48 KiB; Intel Lion Cove matches it.
L2 cache hit
2 to 4 ns
Zen 5 about 14 cycles at 1 MiB, Lion Cove 16 to 17 cycles.
Branch mispredict
3 to 5 ns
10 to 20 cycles of pipeline refill. Same order as an L2 hit.
Non-cryptographic hash, 64 B
10 ns
5 GiB/s sustained.
L3 or last-level cache hit
10 to 25 ns
40 to 55 cycles. Grows with cache size; Ice Lake server measures 22.5 ns.
Random memory read, many in flight
20 ns
3 GiB/s. Out-of-order execution hides most of the real DRAM latency.
Uncontended mutex lock and unlock
25 ns
Contended, expect hundreds of ns to several us.
Main memory, pointer chase
80 to 130 ns
The honest DDR5 load-to-use number. Measured above 128 ns on Ryzen AI 9 HX 370.
Cryptographic hash, 64 B
100 ns
1 GiB/s.
System call
300 ns
Higher on kernels with speculative-execution mitigations enabled.
GPU global memory read (HBM)
400 ns
About 600 core cycles: A100 566, H800 656, RTX 4090 571.
Sequential memory read, 1 MiB, all threads
5 us
200 GiB/s aggregate.
Context switch
1 to 10 us
The switch itself is cheap; the cache and TLB pollution afterwards is not.
Sequential memory read, 1 MiB, one thread
50 us
20 GiB/s. A single core cannot saturate modern memory.

Storage

Device-level numbers at low queue depth. Filesystem, RAID and virtualisation layers each add their own overhead on top.

Operation
Latency
Notes
Src
NVMe sequential read, 8 KiB
1 us
8 GiB/s, so 1 MiB lands in about 100 us.
NVMe sequential write, 8 KiB, no fsync
2 us
3 GiB/s. This is a write to the page cache, not to durable media.
Optane random read, queue depth 1
under 10 us
Discontinued, but still the reference floor for persistent media.
NVMe random read, 4 KiB, queue depth 1
20 to 75 us
Micron 7500 MAX: 70 us typical, 80 us at p99. Seagate XP1920SE: 75 us.
NVMe random read, 8 KiB
100 us
70 MiB/s effective. Random access costs roughly 100x sequential.
SATA SSD random read, 4 KiB
100 to 200 us
The protocol, not the flash, is the difference from NVMe.
NVMe sequential write, 8 KiB, with fsync
300 us
30 MiB/s. Durability costs about 150x over the buffered write.
HDD sequential read, 1 MiB
2 ms
250 MiB/s once the head is already in place.
HDD seek and rotation
10 ms
Essentially unchanged for twenty years. It is a mechanical limit.
HDD random read, 8 KiB
10 ms
0.7 MiB/s effective. Reading 1 GiB at random takes about half an hour.
Blob storage GET returning 304
30 ms
S3 or GCS conditional request. The cheapest thing you can ask object storage.
Blob storage GET, 128 KiB, one connection
80 ms
About 100 MiB/s single stream; 2 to 5 GiB/s with concurrent range reads.
Blob storage LIST
100 ms
One 1000-key page.
Blob storage PUT, 128 KiB, one connection
200 ms
Writes cost roughly 12x reads per operation, and 2.5x the latency.

Network

Round trip times unless stated otherwise. Propagation is physics and cannot be optimised away; everything else can.

Operation
Latency
Notes
Src
Fibre propagation, 1 km, one way
4.9 us
Refractive index about 1.468, so light moves at roughly 204,000 km/s.
Put 1 KiB on the wire at 1 Gbit/s
8.2 us
Serialisation delay, before Ethernet, IP and TCP framing. Not propagation.
Proxy hop (Envoy, nginx, HAProxy)
50 us
Budget one of these per service boundary.
TCP echo, 32 KiB
50 us
500 MiB/s through the socket layer.
Round trip, same zone or VPC
250 us
Premium cloud networking, 25 GiB/s available alongside it.
Round trip, same region
250 us to 1 ms
Crossing availability zones is the expensive part.
Round trip, same datacentre
500 us
The 2012 figure. Still a safe upper bound to design against.
Redis, Memcached or MySQL point query
500 us
Dominated by the round trip, not the lookup.
US Central to US East
25 ms
US Central to US West
40 ms
US East to US West
60 ms
EU West to US East
80 ms
EU West to US Central
100 ms
California to Netherlands and back
150 ms
The original long-haul landmark from the 2012 list.
EU West to Singapore
160 ms
US West to Singapore
180 ms
Inter-region throughput per stream is only about 25 MiB/s.

Software operations

Per mebibyte, single core, on data already resident in memory.

Operation
Latency
Notes
Src
Ed25519 sign
50 us
wolfCrypt on a 2.5 GHz i7. Tuned libraries do better.
Ed25519 verify
180 us
Verification costs roughly 3.5x signing.
Non-cryptographic hash, 1 MiB
200 us
5 GiB/s.
SHA-256, 1 MiB, one core
350 us
2.8 to 3.2 GB/s with SHA-NI on a Ryzen 9 7950X.
AES-GCM, 1 MiB, one core
500 us
AES-128 reaches about 2 GB/s with AES-NI on 8 KiB blocks; AES-256 is slower, and 16 B buffers are 10x worse.
Decompress 1 MiB
1 ms
1 GiB/s.
Serialise 1 MiB, flat wire format
1 ms
1 GiB/s. Protobuf-class encoding, or simdjson-class parsing.
Compress 1 MiB
2 ms
500 MiB/s. Each extra 1x of compression ratio costs about 10x in speed.
Sort 1 MiB of 64-bit integers
2 ms
500 MiB/s.
Serialise 1 MiB, general purpose
10 ms
100 MiB/s. This is where JSON lives.

Language model inference

Added to the canonical list in May 2026. These move faster than anything else on this page, so treat them as a snapshot.

Operation
Latency
Notes
Src
Local model, generate one token
15 ms
Small model on a consumer GPU.
Frontier model, generate one token
20 ms
Hosted output, so about 50 tokens per second.
Local model, time to first token
75 ms
Small model, short prompt.
Local model on CPU, generate one token
100 ms
No GPU. Roughly 7x slower than the same model on one.
Specialised inference hardware, time to first token
250 ms
Frontier model, time to first token
1 s
Short prompt, no prompt cache.
Frontier model, short response
3 s
About 100 output tokens.
Frontier model, prefill 100K tokens
10 s
No prompt cache. Caching is the single biggest lever here.
Frontier model, reasoning response
30 s
One call, with thinking.

Human perception

The reason any of the rest matters. These thresholds have not changed since Nielsen measured them in 1993.

Operation
Latency
Notes
Src
One frame at 60 Hz
16.7 ms
Budget about 10 ms of your own work; the browser needs the rest.
Feels instantaneous
100 ms
Handle the input within 50 ms to render something by 100 ms.
Keeps the flow of thought
1 s
Past this, attention drifts off the task.
Limit of attention
10 s
Past this, people leave or switch to something else.

What follows from the table

1 ms of round trip buys about 50 km

Fibre costs 4.9 us per km each way, so a millisecond covers roughly 100 km of glass. Real routes run 1.5 to 2x the great-circle distance, so halve it. No protocol removes this.

Durability costs about 150x

The same 8 KiB write is 2 us buffered and 300 us with fsync. If a design needs more write throughput, the question is almost always how many fsyncs it issues, not how fast the drive is.

Random access costs about 100x on flash

Sequential 8 KiB from NVMe is 1 us; random is 100 us. On a spinning disk the same ratio is about 350x: 250 MiB/s sequential against 0.7 MiB/s random.

Queueing beats latency above 70% utilisation

For an M/M/1 queue, response time is service time divided by (1 - utilisation). At 50% load you pay 2x, at 80% you pay 5x, at 95% you pay 20x. Every number on this page is an unloaded number.

p99 is not p50

Anything that crosses a network or a device boundary typically shows a p99 five to ten times its median. Capacity plans built on averages fail at the tail.

Check the exponent, then the coefficient

Back-of-envelope work only needs the order of magnitude right. Keep the units attached as a checksum and round aggressively.

Sources

Numbers were taken from these and not from memory. Where two sources disagreed, both are cited and the range is shown in the table. Entries marked further reading back no single row.

  1. 1
    Latency Numbers Every Programmer Should Know

    Jeff Dean and Peter Norvig, maintained by Jonas Boner. The canonical list. Rows for language model inference were added in May 2026.

  2. 2
    Teach Yourself Programming in Ten Years, answers table further reading

    Peter Norvig. The original source of the numbers.

  3. 3
    napkin-math

    Simon Eskildsen. Benchmarked, referenced and actively maintained. The single-host rows were re-measured on 8 March 2026 on a GCP c4-standard-48-lssd instance (Intel Xeon 6985P-C, 24 physical cores, 180 GB RAM, Ubuntu 22.04.5). Most throughput figures on this page come from here.

  4. 4
    Interactive latency page further reading

    Colin Scott. The same numbers extrapolated across years, useful for seeing which of them actually move.

  5. 5
    7-cpu.com

    Per-microarchitecture cache and memory latency measured in cycles and nanoseconds. The reference for cache hierarchy figures.

  6. 6
    Zen 5 cache latencies

    HWCooling. L1d 48 KiB at 4 cycles, L2 1 MiB at 14 cycles, L3 about 46 cycles.

  7. 7
    Ryzen 9950X, Zen 5 on desktop

    Chips and Cheese. Sampled memory latency under real workloads.

  8. 8
    Running gaming workloads through Zen 5

    Chips and Cheese. Comparison of Zen 5 and Lion Cove cache latency.

  9. 9
    Intel Memory Latency Checker results

    Ice Lake server: L1 1.5 ns, L2 4.1 ns, L3 22.5 ns.

  10. 10
    Zen 5 variants, clock for clock

    Chips and Cheese. Pointer chase through a 1 GB array measured above 128 ns.

  11. 11
    Dissecting the NVIDIA Hopper architecture

    arXiv 2501.12084. Global memory latency in cycles for A100, H800 and RTX 4090.

  12. 12
  13. 13
    NVMe queue depth explained

    Optane queue-depth-1 read latency under 10 us; NAND NVMe under 50 us.

  14. 14
    Micron 7500 NVMe SSD product brief

    Typical read 70 us, typical write 15 us, p99 read 80 us, measured with fio at 4 KiB and queue depth 1.

  15. 15
    Seagate Nytro XP1920SE specification

    Average 4 KiB queue-depth-1 read latency 75 us, write 12 us.

  16. 16
    NVMe latency, typical numbers

    simplyblock. NVMe 20 to 70 us, SATA SSD 100 to 200 us, HDD 5 to 10 ms.

  17. 17
    Calculating optical fibre latency

    m2optics. 4.9 us per km at a refractive index of 1.47.

  18. 18
    Dissecting latency in the internet's fibre infrastructure

    arXiv 1811.10737. Uses 204,000 km/s and compares fibre types.

  19. 19
    Hashing vs encryption, SHA-256 and AES-256 compared

    SHA-NI at 1.8 cycles per byte; 2.8 to 3.2 GB/s measured on a Ryzen 9 7950X.

  20. 20
    OpenSSL Cookbook, performance chapter

    Ivan Ristic. aes-128-gcm near 2 GB/s per core at 8 KiB blocks, and far slower at 16 B.

  21. 21
    wolfCrypt benchmarks

    Ed25519 sign 53 us, verify 184 us on a 2.5 GHz Core i7.

  22. 22
    Response Times: The 3 Important Limits

    Jakob Nielsen, 1993. 0.1 s, 1 s and 10 s.

  23. 23
    Introducing RAIL

    Smashing Magazine. Adds the 16 ms animation frame to Nielsen's thresholds and splits the 100 ms budget into 50 ms of work plus rendering.

  24. 24
    cloudping.co further reading

    Live inter-region round trip times between AWS regions. Use this instead of a static table when the exact pair matters.

Storage, network and software throughput figures are dominated by Simon Eskildsen's napkin-math, re-measured on 8 March 2026. Cache and memory figures come from 7-cpu.com, Chips and Cheese and Intel MLC runs. Perception thresholds are Nielsen's, via the RAIL model. Re-check anything in the inference section first; it has the shortest half-life on this page.