
Researchers measure the true speed costs of letting graphics chips talk directly to networks
A new study separates hardware network limits from software overheads across modern computing platforms.
5 Oct 2026
Modern artificial intelligence uses Mixture-of-Experts models. These models rely on quick, small messages sent across clusters. Many systems now let graphics processing units talk straight to network cards using remote direct memory access, which lets a chip read or write another machine's memory without involving the operating system. Software libraries rely on this path, but the exact hardware costs have stayed poorly documented outside of raw program code.
Four researchers tested this direct boundary. They are Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay, and Didem Unat. Their paper reviews how graphics cards place queues, build network work requests, ring notification doorbells, and handle completed tasks.
The authors built two small test tools called mini-gda and mini-proxy. Mini-gda handles direct submission from graphics hardware, while mini-proxy routes requests through a central processor. They compared these basic tools with several common software libraries. The tests ran on Nvidia H100, H200, B200, and GB200 systems.
The basic direct path from the graphics chip can issue an action in 0.7 microseconds. It finishes the transfer in 4.0 microseconds. Issue speed changes depending on the clock speed of the graphics processor's streaming multiprocessors.
Existing communication libraries add extra delays. They add up to 4.6 microseconds to the issue time. This overhead comes from managing queues, setting memory order rules, and tracking when actions finish.
A tuned setup using a central processor can match or beat the idle direct graphics path. However, this method requires a dedicated processor core. The working state of that core dictates its overall speed and delay.
Mixing traffic creates steep slowdowns on both setups. Sharing a single queue with large data transfers increases latency by ten to one thousand times.
Reaching the InfiniBand hardware limit of 260 million messages per second requires grouping doorbells together and using multiple queues at once. These steps consume hardware resources. Communication code can lower block residency on the graphics processor even when the code is not active.
Large networks face other bottlenecks. When all nodes talk to all other nodes across roughly 3,000 active links, the network card message rate drops by 59 percent. Because of these factors, the chosen submission method alone cannot predict total networking performance.