The Shifting Global Compute Landscape: China's Rise in AI Hardware and Software
U.S. export controls on advanced AI chips have acted as a catalyst for China to develop a self-sufficient AI ecosystem, accelerating the production of domestic hardware and the creation of highly compute-efficient open-weight models. This shift is moving the global AI landscape from a U.S.-centric model toward one where China maintains a viable, parallel infrastructure for training and deployment.
The State of Global Compute
Demand for advanced AI chips continues to rise globally, but the long-standing dominance of NVIDIA is being challenged by China's strategic push for domestic self-sufficiency. The next generation of Chinese open-weight AI models are increasingly powered by domestic chips, shifting norms for global AI training and deployment.
This transition is driven by national security concerns in both the U.S. and China, leading to restrictions on chips and rare earth resources. As U.S. export controls tightened, the rollout of Chinese-produced chips accelerated, resulting in more models being optimized for local hardware and a surge in the adoption of compute-efficient open-weight models.
The Catalyst: U.S. Export Controls and Domestic Innovation
Export controls established by the Biden administration in 2022, intended to stall China's AI progress by limiting access to high-end GPUs, instead laid the foundation for a burgeoning domestic industry. Chinese AI labs responded to the threat of being cut off with a surge of innovation in both hardware and software.
The Rise of Domestic Hardware
China has developed several notable advanced chips to replace NVIDIA GPUs, including:
- Huawei Ascend: Initially launched in 2018, with expanded deployment throughout 2024 and 2025.
- Cambricon Technologies and Baidu Kunlun: Other key players in the domestic chip landscape.
The Growth of Open-Weight Models
The scarcity of compute led Chinese labs to prioritize architectural efficiency and open collaboration. This pragmatic "non-NVIDIA first" approach resulted in world-class open-weight models such as Qwen, DeepSeek, GLM, and Kimi. The ability to run these models locally has created a feedback loop between chipmakers and researchers, leading to hardware-optimized models, such as those optimized for the Ascend platform.
Technical Breakthroughs in Compute-Constrained Environments
Compute limitations incentivized significant architectural and infrastructural advancements that have now influenced global AI development.
Algorithmic and Architectural Efficiency
- DeepSeek: Introduced Multi-head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA) to reduce inference costs without sacrificing performance. DeepSeek also developed Group Relative Policy Optimization (GRPO), a reinforcement learning (RL) methodology that reduces compute costs compared to Proximal Policy Optimization (PPO). GRPO has been adopted by Meta researchers and praised by OpenAI's Jerry Tworek for accelerating RL research.
- Linear Attention: Researchers like Peng Bo have championed Linear Attention as a Transformer successor, seen in models like RWKV and scaled in commercial models like MiniMax M1 and Qwen-Next.
Open Infrastructure and Engineering
Chinese labs have shifted from corporate secrecy to sharing engineering secrets to accelerate progress:
- Kimi: Developed the Mooncake serving system for prefill/decoding disaggregation.
- StepFun: Enhanced this with Attention-FFN Disaggregation (AFD) in Step3.
- ByteDance: Contributed verl, an open-source library for production-grade RL training.
- Baidu: Published detailed technical reports on overcoming engineering challenges for Ernie 4.
The Emergence of a Parallel Software Ecosystem
NVIDIA's dominance was built not just on hardware, but on the CUDA software ecosystem. China is now developing backend-neutral alternatives to challenge this reliance.
From Sufficiency to Demand
The demonstration of DeepSeek's R1 model running seamlessly on Huawei's Ascend cloud sparked a market-wide race to optimize domestic models for domestic chips. This proved that domestic silicon was sufficient for frontier AI, shifting the perception of local hardware from a niche alternative to a demanded asset.
Hardware-Software Co-Design
Researchers are now co-developing models with chip vendors. For example, DeepSeek-V3.1's FP8 precision format was designed specifically for next-gen domestic chips, and DeepSeek-V3.2 utilizes TileLang-based kernels for portability across multiple hardware vendors.
Challenging the CUDA Stack
New software layers are emerging to replace the NVIDIA stack:
- FlagGems (BAAI) and TileLang: Backend-neutral alternatives to CUDA and cuDNN.
- Huawei Collective Communication Library (HCCL): A substitute for NVIDIA's NCCL.
Timeline of Compute Controls and Market Shifts
- October 2022: U.S. Bureau of Industry and Security (BIS) introduces controls on A100 and H100 GPUs.
- Late 2022–2023: NVIDIA creates compliant variants (A800, H800, RTX 4090D) with reduced bandwidth.
- Late 2023–2024: BIS upgrades frameworks to Total Processing Performance (TPP) and performance density, targeting the A800/H800.
- January 2025: The "DeepSeek moment" accelerates the adoption of Ascend, Cambricon, and Kunlun chips.
- April–May 2025: The U.S. issues licensing requirements for NVIDIA chips, charging $5.5 billion, before rescinding the AI Diffusion Rule in May.
- August 2025: The U.S. Commerce Department implements a 15% revenue-sharing arrangement for H20 licenses.
- Late 2025: Chinese regulators reportedly instruct firms to cancel NVIDIA orders to prioritize domestic accelerators.