TCP, the data transport protocol underpinning the web and most cloud computing, is poorly suited for emerging AI workloads, according to John Ousterhout, a retired Stanford University computer science professor. In a recent presentation at the AI Engineer World's Fair, Ousterhout argued that a new protocol called Homa offers a solution. The protocol began as a PhD dissertation published in 2019 by Behnam Montazeri, now a Google staff engineer, and Ousterhout has since adopted promoting it as his "life's mission."

Homa delivers the 99th percentile of latency for shorter messages at 92 microseconds, which is 13 times faster than TCP's 1.2 milliseconds, based on packets moving through a 100 Gbps network at 80% utilization. Even for the longest messages, Homa outperforms TCP by a factor of two, Ousterhout said. The protocol is message-based rather than stream-based, with the receiver managing congestion control by explicitly scheduling when packets should be sent. It prioritizes shorter messages over longer ones using a shortest-remaining-processing-time algorithm.

Ousterhout is currently drafting an IETF standardization document for the protocol and working to integrate Homa into the Linux kernel. In March, the protocol was backported to Red Hat Enterprise Linux versions 8 and 9.5. "TCP, for all the amazing things it has done, is not a good match for datacenters," Ousterhout told The Register. He's also helping large companies investigate Homa's applicability, including work with one major financial services firm on a prototype. Installing Homa requires compiling it from its GitHub source and adding the module to Linux kernels on clients and servers, with no reboot needed, and it can run alongside TCP while making remaining TCP applications faster.

The protocol addresses a growing problem for AI workloads, where even millisecond delays cause expensive GPUs to sit idle. TCP's data model relies on byte streams, a continuous flow with no differentiation between messages, leaving receivers with limited visibility into how much traffic is incoming. Senders must guess how much to slow their output based on acknowledgment timing. For latency-sensitive AI tasks—including weight gradients, model weights, KV cache entries, and checkpoints—this becomes problematic as large data transfers increasingly share bandwidth with short bursts from agents and control tasks like metadata coordination and cache lookups. The AI community isn't alone in its frustration: the high-performance database community uses DPDK to bypass TCP, storage area networks turned to NVMe over Fabrics, Google created QUIC for HTTP/3, and high-frequency trading and gaming sectors have pursued RDMA fabrics and Amazon's Scalable Reliable Datagram.

Ousterhout is pushing ahead despite skepticism, including a 2023 position paper by network architect Ivan Pepelnjak that questioned Homa's performance characterizations of TCP and criticized it as a solution seeking a problem. For now, TCP remains the dominant protocol, but frontier labs developing large language models require top-tier network performance, making Homa's latency advantages increasingly relevant. Whether Homa was simply ahead of its time or has finally found its problem to solve remains an open question. Enterprises weighing protocol transitions will need to balance the operational complexity of adopting new infrastructure against the performance demands of workloads that weren't imagined when TCP was designed decades ago.