← China Tech Daily
About Editorial Standards Corrections Privacy Policy
Subscribe →
China Tech · Daily chinatechdaily.org
JULY 20, 2026
Your morning brief on Chinese technology. Published by the AI Policy Institute
Bottom Line Up Front

Moonshot pauses new subscriptions on Kimi K3 after demand overwhelmed its GPUs within 48 hours of the model's July 17 launch, leaving existing subscribers unaffected while pausing new subscriptions. SenseTime and nearly 20 partners launched the "Galaxy Plan" at WAIC 2026 to co-build five domestic 10,000-card AI clusters using chips from Cambricon, Huawei Ascend, Moore Threads and others, with co-founder Yang Fan reporting utilization gains of 85% to 152% and 1.25 times the inference cost-performance of Nvidia H-series accelerators. Enflame and Xianfeng unveiled what they described as China's first glass-substrate CoPoS advanced packaging sample for AI chips at WAIC 2026, without disclosed die dimensions, yield rates or volume timeline.

ITOP STORIES
 
Moonshot pauses new Kimi consumer signups after K3 demand hits GPU cluster capacity
Three days after its Kimi K3 launch (covered July 17), Moonshot AI suspended new consumer subscriptions to its Kimi chatbot on July 19, saying GPU cluster capacity had been pushed to its limit within 48 hours. Existing paid subscribers kept their access, per Moonshot. Moonshot's daily sales grew at least sixfold since K3's release, and annualized recurring revenue reached $300 million in June, up from $200 million in April, per Bloomberg. K3 topped the front-end coding arena at 1,679, above Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol, and led SWE Marathon, a benchmark for long-horizon software engineering tasks, per QbitAI.
Read at QbitAI ↗
SenseTime and nearly 20 Chinese chip partners launch "Galaxy Plan" to co-build five 10,000-card AI clusters
SenseTime and nearly 20 Chinese chip and infrastructure partners launched the "Galaxy Plan" at WAIC 2026 on July 18, committing to co-build five domestic 10,000-card AI compute clusters. Named partners include chip makers Cambricon, Huawei Ascend and Moore Threads and others, alongside infrastructure firms such as SiliconFlow, per SenseTime. SenseTime co-founder Yang Fan said heterogeneous mixed inference, blending chips from multiple domestic vendors in one deployment, lifted chip utilization on mainstream Chinese silicon between 85% to 152%. Yang said the setup delivered 1.25 times the inference cost-performance of Nvidia H-series accelerators, and SenseTime's platform now serves 2.42 trillion tokens per day, on track for 10 trillion by Q4 2026.
Read at ITHome ↗·Read at 36Kr ↗
Enflame and Xianfeng show China's first glass-substrate advanced packaging sample for AI chips at WAIC
Chinese AI accelerator designer Enflame Technology and domestic packaging specialist Xianfeng Technology jointly unveiled what they described as China's first glass-substrate CoPoS advanced packaging sample for AI compute chips at WAIC 2026 on July 18. Advanced packaging connects individual chip dies with high-density interconnects to form a single accelerator module; CoPoS assembles those dies on a common panel-level substrate. The demonstration is a sample rather than a production line; neither company disclosed die dimensions, yield rates or a target-volume timeline.
Read at ITHome ↗
IIPOLICY & REGULATION
 
Shanghai signs 32 AI project deals worth 40.9 billion yuan at WAIC close
Shanghai held a signing ceremony for 32 AI project deals totaling 40.9 billion yuan ($5.7 billion) at the closing of WAIC 2026 on July 20. The headline figure came from a Securities Times readout of the closing ceremony. Neither outlet published a per-project breakdown, list of counterparties or a deployment timeline.
Read at Yicaistate media ↗
IIIAI & FOUNDATION MODELS
 
Chinese social platform Xiaohongshu and Peking University open-source UltraEP load balancer for mixture-of-experts models
Xiaohongshu and Peking University released UltraEP, an open-source load-balancing framework for large mixture-of-experts model training and inference, on July 20. Mixture-of-experts (MoE) architectures route each token to a small subset of specialist sub-networks; a handful of hotspot experts often throttle throughput while the rest of the cluster sits idle. UltraEP tracks routing decisions in real time and dynamically replicates hotspot experts within each microbatch and each model layer, per Xiaohongshu. Xiaohongshu reported UltraEP averages 94.3% of theoretical ideal throughput and runs 1.49 times faster than existing MoE frameworks; the code is already deployed in the company's large-scale pre-training production systems.
Read at Jiemian ↗
IVCHIPS & SEMICONDUCTORS
 
Sanctioned Chinese GPU maker Moore Threads trains Peking University AI model to top Stanford's WorldScore for 37 days
Moore Threads, a Chinese GPU maker on the U.S. Entity List, said its MTT S5000 accelerators trained Peking University's EvoPhys-World to hold the top ranking on Stanford's WorldScore benchmark for 37 consecutive days. EvoPhys-World is a five-dimensional world-generation model; WorldScore is a Stanford academic suite that scores video-generation and world-model quality. Founder Zhang Jianliang told a WAIC 2026 forum on July 18 that Moore Threads' cluster linear-scaling efficiency reached 95%, with effective training time under checkpoint recovery above 90%. On inference, three MTT S5000 nodes combined with two mainstream international GPU nodes averaged 1.87 times the tokens-per-GPU-per-second of an all-international-GPU baseline on Kimi K2.5, per Sina Finance.
Read at Yicaistate media ↗·Read at Sina Finance ↗
VRESEARCH AND DEVELOPMENT
 
Shanghai startup Buchou Quantum unveils 1,500-qubit atomic-quantum AI compute platform at WAIC
Shanghai-based Buchou Quantum unveiled "Liangchou No. 1," which it described as China's first atomic-quantum AI foundation platform, at WAIC 2026. Atomic quantum computing uses individual neutral atoms held in optical traps as qubits, one of several competing quantum-hardware approaches alongside superconducting circuits and trapped ions. Buchou reported specifications including a qubit count above 1,500, single-qubit fidelity of 99.9%, Rydberg two-qubit gate fidelity above 99% and atomic loss rate below 0.1%. The platform, branded Q-Cub, pairs the neutral-atom quantum processor with a GPU cluster for state simulation and AI inference plus a classical CPU for scheduling, per Buchou.
Read at QbitAI ↗
China's industry ministry announces advances in compute and telecommunications
The Ministry of Industry and Information Technology (MIIT) said China's intelligent compute capacity reached 2,185 exaflops at end-June 2026, up 177% year on year, at a July 20 State Council Information Office briefing. MIIT licenses telecom operators and sets industrial policy; bureau head Xie Cun delivered the readout. China ran 5.102 million 5G base stations and 32.86 million gigabit fixed-line ports as of end-June, alongside more than 80,000 5G industrial private networks. MIIT said it approved 6GHz-band experimental spectrum in May for the state-led IMT-2030 promotion group and is accelerating phase two of 6G technology trials, per CGTN.
Read at ITHome ↗·Read at CGTNstate media ↗

Thanks for reading. See you tomorrow.— Daniel

Comments or corrections? hello@theaipi.org

 
Editor-in-Chief:Daniel Colson
Managing Editor:Min Goodman-Cheng
Lead Software Engineer:Christopher Käck
Editorial Advisor:Carrie Adams
Research Manager:Philip Wieczorek
Operations Specialist:Joanne Chua
Research Assistants:Claude Code + Codex
 

View in browser · LinkedIn

AI Policy Institute · 700 Pennsylvania Avenue SE, Washington, DC 20003

China TechDaily
About Editorial Standards Corrections Privacy Policy
© 2026 AI POLICY INSTITUTE · WASHINGTON D.C.