Host-to-GPU bandwidth on an 8x RTX 4090 server without peer-to-peer
On a server with eight RTX 4090 cards, the GPUs cannot copy to each other directly. Every byte that one GPU sends to another, whether a KV cache moving from a prefill GPU to a decode GPU or a tensor-parallel all-reduce, goes down to host memory and back up. The bandwidth between host memory and the GPUs is therefore the budget for any multi-GPU serving on this machine, so I measured it before designing anything around it.
The short answer: each CPU socket carries about 25 GB/s from host to GPU, and that number does not grow when more GPUs on the same socket copy at once.
The machine
- 8 x RTX 4090, 24 GB each, NVIDIA driver 550.144.03.
- Two Intel Xeon Gold 5320 sockets. GPUs 0-3 belong to socket 0, GPUs 4-7 to socket 1.
- No peer-to-peer:
nvidia-smi topo -p2p rreportsCNS(“chipset not supported”) for every one of the 56 ordered GPU pairs, and there is no NVLink.
The PCIe layout matters more than anything else here. lspci -tv and sysfs show that on each
socket the four GPUs sit behind one PCIe switch, and the switch reaches the CPU through a single
PCIe 4.0 x16 link. That link carries at most 31.5 GB/s in each direction before protocol
overhead (16 GT/s x 16 lanes x 128/130).
How I measured
- One process per GPU. Each process pins its CPU threads to the GPU’s socket and allocates two
256 MiB pinned host buffers, one to copy from and one to copy into, on the GPU’s NUMA node. Page placement is read back from
/proc/self/numa_mapson every run, and every buffer page was on the intended node. - All processes wait on one barrier, then copy back to back for 3 seconds, either host to device only, or both directions at once on two CUDA streams per GPU.
- Each process records its own window start and end on the host’s monotonic clock and counts the bytes it finished. An aggregate is the sum of each process’s bytes divided by its own window, and it is reported only if the window shared by all processes covers at least 99% of every process’s window. Without that check the processes may not have copied at the same time, and the sum would not be a bandwidth anyone can get (more on that below).
- Five rounds. The table shows medians; the spread between rounds is under 0.5% in every cell.
Results
| GPUs copying at once | Host to GPU only | Both directions at once (to GPU + to host) |
|---|---|---|
| 1 | 25.2 GB/s | 16.5 + 19.7 GB/s |
| 2 on one socket | 25.3 GB/s in total | 17.9 + 18.6 GB/s in total |
| 4 on one socket | 25.8 GB/s in total | 18.3 + 18.3 GB/s in total |
| 2, one per socket | 50.5 GB/s in total | 33.0 + 39.3 GB/s in total |
| 4, two per socket | 50.7 GB/s in total | 35.8 + 37.3 GB/s in total |
| all 8 | 51.6 GB/s in total | 36.7 + 36.6 GB/s in total |
What the numbers say:
- One socket carries about 25 GB/s from host to GPU, whether one, two or four of its GPUs copy. A second GPU on the same socket does not add bandwidth: each of the two gets half, and each of four gets a quarter. This is consistent with the shared x16 uplink being the limit, but the measurement does not rule out other resources the four GPUs of a socket share.
- The two sockets do not share that limit. Any set that spans both sockets gets about 50 GB/s, twice the single-socket number.
- Copying both ways at once does not double the bandwidth. One GPU alone moves 25 GB/s one way, but only 16.5 + 19.7 = 36 GB/s when it copies both ways. A whole socket moves the same 36 GB/s in total.
- Buffer placement did not matter for one GPU. In an earlier run on GPU 5, with some other GPUs busy, the buffer on the GPU’s own NUMA node, the CPU threads and the buffer both on the other socket, and only the CPU threads on the other socket all gave 25.3 GB/s host to device and 26.3 GB/s device to host.
What this means for moving a KV cache
Without peer-to-peer, handing a KV cache from one GPU to another takes two copies: the first GPU copies it down to host memory, the second copies it up.
A back-of-envelope estimate, not a measurement: if the two copies run one after the other at the single-GPU rates above (26.3 GB/s down, 25.3 GB/s up), the transfer runs at about 1 / (1/26.3 + 1/25.3) = 12.9 GB/s. A 1 GiB KV cache then takes about 83 ms. Overlapping the two copies in chunks can do better, but point 3 is a warning against assuming each direction keeps its full one-way rate while the other is busy. What a real serving engine achieves is the next thing I am measuring.
A measurement mistake worth sharing
An earlier version of this benchmark reported 37 GB/s for two GPUs on one socket, and about 38-41 GB/s for four. Both numbers should have been a warning on their own: they are more than one x16 uplink can carry in one direction. The script had added up each process’s median copy rate, but the processes had no common start, and each repetition did a copy to the GPU followed by a copy back. The sum mixed directions and periods when a process copied alone.
The method above replaced it: all processes start on a barrier, each direction is counted separately, and no aggregate is reported unless the copy windows overlap. With it, the per-socket total is about 25 GB/s in all five rounds.