Moonshot AI Races Toward $30 Billion Hong Kong IPO After Kimi K3 Demand Crashes Its Servers

Moonshot AI, the Beijing-based startup behind the viral Kimi chatbot, is accelerating plans for a Hong Kong initial public offering that could value the company at more than $30 billion, just as overwhelming demand for its newly launched Kimi K3 model forces it to halt new user subscriptions.
The dual developments—a blockbuster funding round and a capacity crunch—underscore how a single model release is reshaping the global AI landscape. Kimi K3, unveiled on July 16, is a 2.8-trillion-parameter open-weight system that has stunned the industry by matching or surpassing leading US models on critical coding benchmarks. The fallout triggered a sharp selloff in global semiconductor stocks, with the Philadelphia Semiconductor Index plunging 10% in a single week and entering bear market territory after falling more than 20% from its June peak.
Moonshot has circulated a shareholder resolution seeking approval for a listing that could occur within six months, according to people familiar with the matter. The company is simultaneously finalizing a new private funding round that would push its valuation past $30 billion—a nearly sevenfold increase from the $4.3 billion valuation it held in December. Goldman Sachs (GS) and China International Capital Corp. are in discussions to play roles in the offering.
A Model That Broke the Internet—and Its Own Servers
Within 48 hours of its release, Kimi K3 demand had pushed Moonshot’s GPU clusters to their absolute limits. The company announced on July 19 that it was immediately suspending new consumer subscriptions to preserve service quality for existing members. "User requests have far exceeded projections and are approaching the carrying capacity of our existing clusters," the company stated.
To manage the strain, Moonshot is restructuring its product tiers, separating the core Kimi service—spanning web, mobile app, and the Kimi Work desktop client—from Kimi Code, a programming-focused offering that consumes significantly more compute resources. The company said it is "advancing at full speed" on infrastructure expansion but provided no timeline for when new subscriptions would reopen.
The technical specifications explain the frenzy. Kimi K3 employs a Mixture-of-Experts architecture with 896 expert modules, activating only 16 per request. It features a 1-million-token context window, native multimodal capabilities for processing images and text, and a novel KDA (Kimi Delta Attention) hybrid linear attention mechanism that slashes memory consumption. Independent evaluator Artificial Analysis scored K3's intelligence index at 57, placing it behind only Anthropic's Claude Fable 5 at 60 and OpenAI's GPT-5.6 Sol at 59—making it the first Chinese open-weight model to reach that tier. On Arena.ai's frontend coding benchmark, K3 claimed the number one spot outright.
Wall Street's "DeepSeek 2.0" Panic
The market reaction was swift and brutal. JPMorgan dubbed the event "DeepSeek 2.0," a reference to the January 2025 moment when DeepSeek's R1 model erased $589 billion from Nvidia's (NVDA) market cap in a single day. This time, Nvidia fell 2.2%, Applied Materials dropped 5.6%, and memory chip makers Micron, Samsung Electronics, and SK Hynix had already entered bear markets earlier in July. Taiwan's benchmark index slid more than 6%, while Japanese equities lost 4%. Hong Kong-listed AI rival Z.AI plummeted 30% in its steepest single-day decline since its January listing, and MiniMax Group shed 16%.
The core fear gripping traders: if a Chinese startup operating under US semiconductor export controls can produce a model that rivals the best from OpenAI and Anthropic at a fraction of the cost, then the trillion-dollar capital expenditure plans of Silicon Valley giants may be built on sand. Kimi K3's API pricing is set at $3 per million input tokens and $15 per million output tokens—roughly 60% cheaper than Claude Fable 5's $10/$50 pricing.
Yet several analysts argue the selloff is more about crowded positioning than a fundamental collapse in AI infrastructure demand. The Philadelphia Semiconductor Index had surged 68% year-to-date before the pullback. EPFR data showed US corporate insiders sold $77.6 billion in stock during the first half of the year, the second-highest level in two decades. "K3 was the catalyst, not the cause," one Wall Street strategist noted.
The Jevons Paradox and the Real Compute Equation
A critical detail often lost in the panic: Kimi K3 is not a lightweight model. Its 2.8 trillion parameters require more than 1.5 terabytes of HBM memory capacity at 4-bit floating-point precision, demanding deployment across GPU super-nodes with at least 64 cards. This is heavy infrastructure, not a cost-free miracle.
Morgan Stanley's chief China equity strategist noted that combined capital expenditure from the top five global cloud providers—Amazon (AMZN), Google (GOOGL), Meta (META), Microsoft (MSFT), and Oracle (ORCL)—is expected to exceed $800 billion in 2026 and could reach $1 trillion to $1.2 trillion in 2027. The investment thesis hinges on the Jevons Paradox: as AI becomes cheaper, consumption explodes. Kimi K3's own server crash is Exhibit A.
"K3 proves that sparse architectures and training method improvements can boost compute efficiency, which triggered panic selling in Nvidia, Broadcom, and TSMC," one AI industry analyst said. "But the flaw in that logic is that total training compute demand depends on the product of per-training cost and training frequency. Even if K3 cuts single-training costs by 40%, if US-China model competition doubles the number of training runs, total compute demand actually grows 20%."
There are also genuine limitations. Third-party testing reveals that K3 generates roughly twice as many output tokens as comparable models to complete the same evaluation suite, partially offsetting its per-token price advantage. Its output speed of approximately 62 tokens per second lags the 73 tokens-per-second median of peer models. And while Moonshot demonstrated K3 autonomously designing a functional chip in 48 hours, the design used a 45-nanometer process—several generations behind the cutting-edge 2nm and 3nm nodes where Cadence (CDNS) and Synopsys (SNPS) maintain an iron grip.
From $100 Million to $300 Million in Three Months
Moonshot's IPO ambitions are backed by explosive revenue growth. The company's annual recurring revenue hit $300 million in June, tripling from $100 million in March. API business now accounts for over 70% of total revenue, marking a decisive shift from consumer subscriptions to high-margin enterprise contracts. On the day after K3's release, the company recorded its largest single-day ARR increase in history, according to President Zhang Yutong.
Founded in early 2023 by Yang Zhilin, a 34-year-old former Tsinghua University professor who earned his PhD at Carnegie Mellon University and worked at Meta and Google Brain, Moonshot has long operated in the shadow of better-known Chinese AI labs. The company's name—"Yue Zhi An Mian" in Chinese—is a nod to Pink Floyd's "Dark Side of the Moon."
K3's release has reshuffled the competitive hierarchy. Bernstein analysts now rank Moonshot as the temporary leader among Chinese AI labs, ahead of Zhipu AI, Alibaba's Qwen, and DeepSeek—the latter having fallen behind after dedicating six months to adapting its models for domestic Chinese chips. The next catalysts, analysts say, will be Zhipu's next-generation pre-trained model and Alibaba's upcoming Apsara Cloud Conference.
A Market in Transition
The Kimi K3 moment is less about any single model's capabilities and more about what it reveals: the AI industry's center of gravity is shifting from pure compute scale toward engineering efficiency. Wall Street's three-year narrative—more GPUs equal more intelligence equal more value—is being rewritten in real time.
"Kimi didn't change AI," one commentator wrote. "Kimi changed how capital understands AI."
For Moonshot, the immediate challenge is operational: adding enough GPU capacity to meet surging demand without degrading service. For global markets, the question is whether the AI capital expenditure supercycle can survive the dawning reality that world-class models no longer require world-dominating budgets. The answer, as Kimi K3's own overwhelmed servers suggest, may be that cheaper models don't reduce demand—they unleash it.
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.