GPT-5.6 Sol Debuts Tomorrow: Inference Speed Hits 750 Tokens/s as Cerebras Pours Billions into European Expansion

OpenAI will officially launch its flagship model GPT-5.6 Sol on July 10, achieving inference speeds of up to 750 tokens per second on Cerebras wafer-scale hardware. Industry analysts speculate the model runs across 70 to 100 wafers and employs a disruptive lightweight architecture design to push physical limits. Simultaneously, AI chip company Cerebras announced a multi-billion-dollar push into Europe, planning to build 200 megawatts of data center capacity by the end of 2027, directly challenging Nvidia's market dominance. These two developments signal that OpenAI is accelerating the construction of a full-stack AI empire spanning from models to infrastructure through deep coupling of its proprietary Jalapeño chip with top-tier third-party hardware.
GPT-5.6 Sol Debuts Tomorrow: Inference Speed Hits 750 Tokens/s as Cerebras Pours Billions into European Expansion

A software-defined and hardware-reshaped AI arms race is heating up. OpenAI announced it will officially launch its latest flagship model, GPT-5.6 Sol, this Thursday, with inference speeds reaching up to 750 tokens per second on Cerebras Systems (CBRS.US) wafer-scale hardware. Meanwhile, Cerebras also announced it will invest billions of dollars in a major European expansion to build regional AI computing centers, directly challenging Nvidia's (NVDA.US) dominance.

OpenAI CEO Sam Altman confirmed the launch plan on social media. GPT-5.6 Sol will feature an all-new Ultra multi-agent mode and Max inference intensity, setting new records across core benchmarks in coding, biology, and cybersecurity. According to OpenAI's published pricing strategy, Sol's input cost is $5 per million tokens and output cost is $30 per million tokens, positioning it as the highest-end product in the GPT-5.6 series.

This launch is not merely a model iteration — it reveals that OpenAI is building a full-stack AI empire through deep coupling of proprietary chips and top-tier third-party hardware.

The Physics Breakthrough at 750 Tokens/s: Brute-Force Elegance Across a Hundred Wafers

What does "750 tokens per second" inference speed actually mean? For humans, it's equivalent to reading and outputting roughly 500 to 600 Chinese characters in a single second. Complex agent tasks that previously required minutes of waiting can now be completed almost in the blink of an eye.

Achieving this speed is not as simple as stuffing a model onto a single chip. Veteran technical expert Bleys Goodson, through rigorous analysis, pointed out that GPT-5.6 Sol very likely spans 70 to 100 Cerebras wafer-scale chips. Industry estimates suggest that to achieve healthy inference serving characteristics, OpenAI and Cerebras adopted an extraordinarily lavish deployment approach — placing each layer of the neural network individually on an entire Cerebras wafer.

Massive wafer counts alone aren't enough. A key feature of Cerebras chip architecture is its enormous on-chip SRAM, which is extremely fast but precious in capacity. If OpenAI used traditional heavy KV-cache mechanisms, this expensive SRAM bandwidth would be instantly exhausted. Therefore, outside observers speculate that OpenAI almost certainly restructured the model around specific hardware, possibly adopting a lightweight cache architecture similar to DeepSeek V4, or a hybrid SSM design combining linear-time-complexity models like Mamba with Transformers, to completely shed the historical burden of KV-caches.

Additionally, well-known developer John Lam proposed another theory: OpenAI may be using traditional GPUs to handle attention computation while leveraging massive Cerebras wafers to brute-force the feed-forward neural network calculations. This speculation is not unfounded — Cerebras previously achieved near-1,000 tokens/s running a trillion-parameter MoE model on its CS-3 system when deploying Kimi K2.6, with inter-layer all-to-all communication bandwidth reportedly over 200 times that of Nvidia's NVLink on NVL72.

Cerebras Expands into Europe Simultaneously: Spending Billions to Challenge Nvidia Head-On

While software capabilities surge forward, the physical land-grab for hardware infrastructure is also accelerating. Cerebras Systems announced on July 9 that it plans to invest billions of dollars in Europe, aiming to activate its first regional data center by the end of 2026 and achieve total installed capacity of 200 megawatts by the end of 2027.

Cerebras co-founder and CEO Andrew Feldman told AFP on the sidelines of the RAISE Summit in Paris: "These are massive, multi-billion-dollar expansions." He emphasized that European demand for computing power to run generative AI is "staggering... growing so fast we can barely keep up."

According to the plan, Cerebras will rapidly build infrastructure in France and the Nordic region, with data centers in Norway and Finland already in the pipeline. Some capacity is expected to support the computing needs of partner OpenAI. Feldman noted that by placing data centers across Europe, Cerebras can meet all unique European requirements, including data sovereignty.

This expansion plan comes as European AI infrastructure investment accelerates. European enterprises and research institutions have long sought to reduce dependence on computing capacity concentrated in the US and Asia. While Nvidia claims its technology powers over 90% of announced AI factory projects in Europe, Cerebras is attempting to break this monopoly with its wafer-scale engine.

Buoyed by the European expansion news, Cerebras shares rose as much as 6% in pre-market trading. The company completed its IPO on Nasdaq in May this year, with an offering price of $185 per share, opening at $350, raising a total of $5.5 billion — one of the largest IPOs in Wall Street history.

Full-Stack Empire Ambitions: From the Jalapeño Chip to Gigawatt-Scale Data Centers

OpenAI's ambitions extend far beyond models themselves. Previously, OpenAI officially unveiled its first-ever proprietary AI inference chip — Jalapeño. This custom ASIC, designed specifically for large model inference, went from design to tape-out in just nine months. Behind it stands an extraordinarily powerful industry alliance: OpenAI personally handled the underlying architecture design, Broadcom provided chip implementation and interconnect technology support, and Celestica was responsible for final board manufacturing and rack-level physical integration.

Jalapeño not only runs OpenAI's own models — its architecture is also compatible with large language models across the entire industry, demonstrating significant platform ambitions. The deep synergy between this chip and Cerebras hardware has allowed OpenAI to thoroughly understand the critical points of dedicated inference architecture and transform that knowledge into a controllable underlying platform.

According to OpenAI's grand blueprint, the first gigawatt-scale super data centers will begin deployment from late 2026 onward, in partnership with core collaborators including Microsoft. The entire electricity consumption of a mid-sized city will be used to power the next generation of inference racks.

In terms of model performance, GPT-5.6 Sol Ultra ranked first on the coding benchmark Terminal-Bench 2.1 with a score of 91.9%, ahead of Claude Mythos 5's 88.0%. On the biology domain benchmark GeneBench v1, Sol achieved superior results compared to predecessors while using fewer tokens; on the cybersecurity benchmark ExploitBench, Sol competed effectively using only about one-third of the output tokens of rivals.

OpenAI equipped this series with its most robust safety guardrail system to date, investing over 700,000 A100-equivalent GPU hours in automated red-teaming. According to OpenAI's preparedness framework, GPT-5.6 Sol did not cross the "critical" threshold for cybersecurity.

As GPT-5.6 Sol races forward at 750 tokens/s on Cerebras wafers, the physical constraints between software and hardware are being shattered, and a new era of real-time intelligence is accelerating toward us.

Add to Google Preferred Sources

Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.







More Related News