--- date: 2026-07-20 subject: "Moonshot pauses Kimi K3 subscriptions | SenseTime Galaxy Plan bundles chip vendors | Enflame shows glass CoPoS" --- **Moonshot pauses new subscriptions on Kimi K3** after demand overwhelmed its GPUs within 48 hours of the model's July 17 launch, leaving existing subscribers unaffected while pausing new subscriptions. **SenseTime and nearly 20 partners launched** the "Galaxy Plan" at WAIC 2026 to co-build five domestic 10,000-card AI clusters using chips from Cambricon, Huawei Ascend, Moore Threads and others, with co-founder Yang Fan reporting utilization gains of 85% to 152% and 1.25 times the inference cost-performance of Nvidia H-series accelerators. **Enflame and Xianfeng unveiled** what they described as China's first glass-substrate CoPoS advanced packaging sample for AI chips at WAIC 2026, without disclosed die dimensions, yield rates or volume timeline. # 1. Top Stories - **Moonshot pauses new Kimi consumer signups after K3 demand hits GPU cluster capacity** — Three days after its Kimi K3 launch (covered July 17), Moonshot AI suspended new consumer subscriptions to its Kimi chatbot on July 19, saying GPU cluster capacity had been pushed to its limit within 48 hours. Existing paid subscribers kept their access, per Moonshot. Moonshot's daily sales grew at least sixfold since K3's release, and annualized recurring revenue reached $300 million in June, up from $200 million in April, per Bloomberg. K3 topped the front-end coding arena at 1,679, above Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol, and led SWE Marathon, a benchmark for long-horizon software engineering tasks, per QbitAI. [QbitAI](https://www.qbitai.com/2026/07/455179.html) - **SenseTime and nearly 20 Chinese chip partners launch "Galaxy Plan" to co-build five 10,000-card AI clusters** — SenseTime and nearly 20 Chinese chip and infrastructure partners launched the "Galaxy Plan" at WAIC 2026 on July 18, committing to co-build five domestic 10,000-card AI compute clusters. Named partners include chip makers Cambricon, Huawei Ascend and Moore Threads and others, alongside infrastructure firms such as SiliconFlow, per SenseTime. SenseTime co-founder Yang Fan said heterogeneous mixed inference, blending chips from multiple domestic vendors in one deployment, lifted chip utilization on mainstream Chinese silicon between 85% to 152%. Yang said the setup delivered 1.25 times the inference cost-performance of Nvidia H-series accelerators, and SenseTime's platform now serves 2.42 trillion tokens per day, on track for 10 trillion by Q4 2026. [ITHome](https://www.ithome.com/0/979/115.htm) [36Kr](https://www.36kr.com/p/3900754954749572) - **Enflame and Xianfeng show China's first glass-substrate advanced packaging sample for AI chips at WAIC** — Chinese AI accelerator designer Enflame Technology and domestic packaging specialist Xianfeng Technology jointly unveiled what they described as China's first glass-substrate CoPoS advanced packaging sample for AI compute chips at WAIC 2026 on July 18. Advanced packaging connects individual chip dies with high-density interconnects to form a single accelerator module; CoPoS assembles those dies on a common panel-level substrate. The demonstration is a sample rather than a production line; neither company disclosed die dimensions, yield rates or a target-volume timeline. [ITHome](https://www.ithome.com/0/979/096.htm) # 2. Policy & Regulation - **Shanghai signs 32 AI project deals worth 40.9 billion yuan at WAIC close** — Shanghai held a signing ceremony for 32 AI project deals totaling 40.9 billion yuan ($5.7 billion) at the closing of WAIC 2026 on July 20. The headline figure came from a Securities Times readout of the closing ceremony. Neither outlet published a per-project breakdown, list of counterparties or a deployment timeline. [Yicai](https://www.yicai.com/brief/103283450.html) [state-affiliated] # 3. AI & Foundation Models - **Chinese social platform Xiaohongshu and Peking University open-source UltraEP load balancer for mixture-of-experts models** — Xiaohongshu and Peking University released UltraEP, an open-source load-balancing framework for large mixture-of-experts model training and inference, on July 20. Mixture-of-experts (MoE) architectures route each token to a small subset of specialist sub-networks; a handful of hotspot experts often throttle throughput while the rest of the cluster sits idle. UltraEP tracks routing decisions in real time and dynamically replicates hotspot experts within each microbatch and each model layer, per Xiaohongshu. Xiaohongshu reported UltraEP averages 94.3% of theoretical ideal throughput and runs 1.49 times faster than existing MoE frameworks; the code is already deployed in the company's large-scale pre-training production systems. [Jiemian](https://www.jiemian.com/article/14797341.html) # 4. Chips & Semiconductors - **Sanctioned Chinese GPU maker Moore Threads trains Peking University AI model to top Stanford's WorldScore for 37 days** — Moore Threads, a Chinese GPU maker on the U.S. Entity List, said its MTT S5000 accelerators trained Peking University's EvoPhys-World to hold the top ranking on Stanford's WorldScore benchmark for 37 consecutive days. EvoPhys-World is a five-dimensional world-generation model; WorldScore is a Stanford academic suite that scores video-generation and world-model quality. Founder Zhang Jianliang told a WAIC 2026 forum on July 18 that Moore Threads' cluster linear-scaling efficiency reached 95%, with effective training time under checkpoint recovery above 90%. On inference, three MTT S5000 nodes combined with two mainstream international GPU nodes averaged 1.87 times the tokens-per-GPU-per-second of an all-international-GPU baseline on Kimi K2.5, per Sina Finance. [Yicai](https://www.yicai.com/brief/103281286.html) [state-affiliated] [Sina Finance](https://finance.sina.com.cn/jjxw/2026-07-20/doc-iniikhxa4995579.shtml) # 5. Research and Development - **Shanghai startup Buchou Quantum unveils 1,500-qubit atomic-quantum AI compute platform at WAIC** — Shanghai-based Buchou Quantum unveiled "Liangchou No. 1," which it described as China's first atomic-quantum AI foundation platform, at WAIC 2026. Atomic quantum computing uses individual neutral atoms held in optical traps as qubits, one of several competing quantum-hardware approaches alongside superconducting circuits and trapped ions. Buchou reported specifications including a qubit count above 1,500, single-qubit fidelity of 99.9%, Rydberg two-qubit gate fidelity above 99% and atomic loss rate below 0.1%. The platform, branded Q-Cub, pairs the neutral-atom quantum processor with a GPU cluster for state simulation and AI inference plus a classical CPU for scheduling, per Buchou. [QbitAI](https://www.qbitai.com/2026/07/455136.html) - **China's industry ministry announces advances in compute and telecommunications** — The Ministry of Industry and Information Technology (MIIT) said China's intelligent compute capacity reached 2,185 exaflops at end-June 2026, up 177% year on year, at a July 20 State Council Information Office briefing. MIIT licenses telecom operators and sets industrial policy; bureau head Xie Cun delivered the readout. China ran 5.102 million 5G base stations and 32.86 million gigabit fixed-line ports as of end-June, alongside more than 80,000 5G industrial private networks. MIIT said it approved 6GHz-band experimental spectrum in May for the state-led IMT-2030 promotion group and is accelerating phase two of 6G technology trials, per CGTN. [ITHome](https://www.ithome.com/0/979/093.htm) [CGTN](https://news.cgtn.com/news/2026-05-08/China-approves-6GHz-band-trial-spectrum-for-6G-development-1MZ1Z8eUkW4/p.html) [state media]