Major AI companies are moving beyond traditional GPUs and adopting specialized silicon from vendors like Cerebras and Marvell to power their inference workloads, according to a report from The Register this week. The shift marks the end of GPUs' "exclusive reign on the rack" in data centers, as companies pursue alternatives that offer faster token generation, lower operating costs, and better power efficiency. Meanwhile, Microsoft has finally restored a long-requested Windows 11 feature after a five-year absence, allowing users to move the taskbar for the first time since the operating system launched.

Cerebras, which builds dinner-plate-sized chips packed with 44 gigabytes of ultra-fast SRAM, has launched its next-generation CS-4 system featuring three chips in a modular rack-scale design. The new architecture doubles clock speeds and performance compared to the previous generation while tripling the compute density per rack, though each chip still consumes 15 kilowatts of power compared to 1.2 kilowatts for modern Nvidia GPUs. Google has partnered with Marvell to design custom Tensor Processing Units alongside its existing relationship with Broadcom, giving the cloud giant multiple IP providers to compete on performance and pricing. Waymo disclosed it's developing a custom machine learning accelerator for autonomous vehicles to replace FPGAs, aiming to reduce latency between detecting road obstacles and taking evasive action while improving compute density and ease of development.

The report notes that Cerebras has positioned itself as the alternative to Nvidia's $20 billion Groq acquisition for companies seeking high-speed, low-cost token generation. Major partners including AWS and AMD are expected to deploy Cerebras systems to serve coding assistants and other applications requiring hundreds or thousands of tokens per second. According to the report, "fast tokens are also cheap tokens" because faster generation drives down the operating cost of hardware and enables providers to charge less while maintaining profits. The industry has shifted toward distributing workloads across two types of chips rather than relying on a single GPU or custom accelerator to handle everything.

This hardware diversification reflects growing competition and the need for specialized solutions as AI workloads mature beyond training into inference at scale. Cloud providers are increasingly working with multiple IP houses to avoid monopoly pricing, similar to concerns that emerged after Broadcom's VMware acquisition when licensing fees rose sharply. Cerebras founder Andrew Feldman has been vocal in his criticism of Nvidia, previously labeling the company an "AI arms dealer" for continuing GPU sales into China. The move toward custom silicon by companies like Google and Waymo suggests that off-the-shelf solutions, while abundant from vendors including Intel Mobileye, Qualcomm, Nvidia, and AMD, can't always match the latency, power efficiency, or cost profile that specific applications demand.

On the Windows front, Microsoft has added taskbar positioning to the Release Preview channel, meaning the feature should reach the public within weeks or months after being absent since Windows 95. Users can now place the taskbar left, right, top, or bottom and adjust its size, along with Start menu customization options. The company is also testing a redesigned right-click context menu in the experimental channel that allows Windows 10 muscle memory to work again, and has updated Task Manager to show which processes are using the neural processing unit so users can identify inefficient workloads burning excess power on the CPU and GPU instead. The report characterizes these changes as Microsoft responding to user demands after years of focusing on features like Copilot and advertisements that customers didn't request, with Windows boss Pavan Davuluri pledging earlier this year to rethink how the company approaches the taskbar, Start menu, and AI integration. The delayed restoration of basic customization options suggests enterprises evaluating chip partners for inference may face similar strategic questions about vendor lock-in and the long-term cost of platform decisions made for convenience rather than control.