Nvidia’s next-generation Kyber rack-scale AI architecture, intended for its 2027 Rubin Ultra chips, has reportedly been delayed until 2028 due to significant manufacturing difficulties with a key circuit board. This setback challenges Nvidia’s aggressive annual release schedule and could provide rivals like AMD and Google with a crucial technical opening in the high-end AI market. The reported delay also comes after a backup plan to combine existing racks was rejected by cloud providers due to operational complexities.
NVIDIA's highly anticipated Kyber rack-scale architecture, engineered to house its cutting-edge 2027 Rubin Ultra chips, has reportedly been pushed back by over a year to 2028. This news, according to research firm SemiAnalysis, marks the latest in a series of reported setbacks that are casting a shadow over the AI behemoth's ambitious product roadmap.
Kyber is designed as a colossal server cabinet capable of integrating 144 of Nvidia's most powerful chips into a single, cohesive unit. This configuration allows them to function as one giant supercomputer, delivering the immense processing power required by AI companies to train and execute their most sophisticated models.
The innovative design positions graphics processing units (GPUs) in compute trays that are mounted vertically rather than horizontally. This approach significantly boosts density and reduces data latency. The system was originally slated to debut alongside Vera Rubin Ultra, Nvidia's next-generation rack-scale system, in 2027.
However, the primary hurdle stems from difficulties in manufacturing a crucial circuit board that forms the core of the system, SemiAnalysis revealed in a recent post. "Kyber NVL144 rack architecture has been delayed to 2028 as the PCB midplane remains challenging from a manufacturability standpoint," the firm stated, referring to a specialized, multi-layer printed circuit board essential for connecting electronic modules within the system. Furthermore, the NVL576 – a larger system designed to link eight racks via optical connections – is also likely to face delays or be limited to small production volumes.
Nvidia has not yet responded to requests for comment regarding these reported delays.
This latest postponement exacerbates growing pressures across Nvidia's diverse product lines, intensifying concerns that the company's aggressive annual release cadence is encountering the inherent limits of manufacturing capabilities. Adding to the complications, a contingency plan to bolt together two of Nvidia's current-generation racks to achieve similar power was also abandoned. This backup solution faced strong rejection from cloud customers who deemed its design awkward and its operation overly costly. SemiAnalysis noted, "It has since been cancelled due to heavy pushback from CSPs [cloud service providers] and hyperscalers over its odd design and heavy operational burden."
Consequently, Nvidia is left without "no proven solution to expand the scale-up world size for Rubin Ultra," SemiAnalysis warned. This void could present a rare technical advantage for rivals such as Advanced Micro Devices and Google, whose in-house chip developments are already securing significant business from leading AI laboratories, particularly at the high end of the market.
Despite these challenges, Nvidia's current-generation Rubin systems are in full production and are expected to begin shipping this fall to a roster of eight prominent cloud partners, including Amazon Web Services, Microsoft Azure, and Google Cloud. SemiAnalysis also projects a robust performance for Nvidia, forecasting its data-center compute revenue to exceed Wall Street consensus by 20% in the second half of fiscal 2027. Following the news, shares of Nvidia saw minor fluctuations in premarket trading, registering a slight dip of less than 0.1% to $194.79.
