Nvidia's lead in AI model training is under pressure as Chinese companies like Huawei, Cambricon, Moore Threads, and Biren Technology push to develop homegrown chips and software, aiming to reduce dependence on foreign hardware.
China's biggest tech firms are moving quickly to swap out Nvidia processors in their AI systems, but the gap in both hardware and software remains significant. Demand for domestic chips is rising fast, with companies such as Huawei, Cambricon Technologies, Moore Threads, and Biren Technology reporting strong sales and new installations. Still, replacing Nvidia is not straightforward, especially for advanced model training. According to Reuters, Chinese chipmakers including Huawei and Cambricon have sharply increased prices for current and next-generation processors as shortages of high-bandwidth memory put pressure on local supply chains. This price surge highlights both the competition and the strain on manufacturing in China.
Nvidia's CUDA software platform sits at the heart of this competition. CUDA is tightly integrated with Nvidia's GPUs and networking hardware, so switching to other chips means more than just swapping parts. Developers have to rewrite code, change workflows, and retrain teams-a costly and time-consuming process. Despite years of U.S. export controls, Chinese AI developers still rely on Nvidia hardware for training large models, since moving entire software stacks is often too complex and expensive. Reuters notes that OpenAI's latest ChatGPT version was trained on about 100,000 Nvidia chips, showing how deeply Nvidia is embedded in global AI research, including at places like MIT and Stanford.
Domestic alternatives
Huawei has promoted its Ascend processors and SuperNode systems as the leading domestic option. Its Compute Architecture for Neural Networks (CANN) aims to mirror much of CUDA's functionality, making it easier for developers to switch. In a recent rollout, Huawei's Ascend 950PR and 950DT chips powered inference for DeepSeek V4, with support from the model's creators. This kind of hardware-software integration is essential for any company hoping to match Nvidia's ecosystem, not just its raw speed. But prices have jumped: Reuters reports the Ascend 950DT now sells for over 250,000 yuan per chip, the 950PR has gone from about 60,000 to more than 80,000 yuan, and the Ascend 910C has risen from around 90,000 yuan at the start of the year to over 110,000 yuan. These increases reflect both supply limits and strong demand, especially as high-bandwidth memory remains scarce.
Cambricon Technologies focuses on machine-learning accelerators for servers and data centers. Its chips are built for the parallel processing needed in AI training and inference. As access to Nvidia's top chips has tightened, Cambricon and Huawei have become the main suppliers for China's AI server market. Moore Threads, founded in 2020, is developing general-purpose GPUs for AI, graphics, and scientific work, targeting organizations that once depended on foreign hardware. Biren Technology, despite manufacturing setbacks from export restrictions, expects first-half revenue to grow by over 2,100 percent as demand for domestic chips soars. Still, Biren's processors have not matched Nvidia's in performance or software support. Peer-reviewed studies in journals such as Nature stress the need for hardware and software to be designed together to achieve breakthroughs in AI model training-a challenge Chinese firms are now tackling directly.
Technical and market barriers
Replacing Nvidia hardware is only part of the problem. Nvidia's lead is reinforced by its mature developer ecosystem, detailed documentation, and established support. For Chinese chip makers, building something similar means investing in software tools, libraries, and community support. The transition is further complicated by the need to keep existing AI models and workflows running, many of which were built for Nvidia's architecture. DeepSeek is reportedly using Huawei chips for some training, but these are still not as capable as Nvidia's latest and are roughly on par with Nvidia chips from about four years ago, according to Reuters. This performance gap matters for research centers like CERN and Harvard, which need the most advanced computing for their AI-driven work.
Some numbers show the scale of change. Biren Technology says its first-half revenue could rise by as much as 2,107 percent, driven by orders from Chinese cloud providers and research groups. Moore Threads has seen quick adoption of its GPUs among organizations looking to localize their computing. Cambricon and Huawei now make up a growing share of AI server deployments in China, though exact market share figures are not public. Despite these gains, Nvidia's hardware and software are still the default for training the most advanced models, especially those needing large-scale parallelism and high memory bandwidth. China's AI chip market is now estimated at about $90 billion, reflecting both the scale of local investment and the continued reliance on foreign technology.
Strategic consequences
The push for domestic AI chips is about more than technical competition. U.S. export controls have forced Chinese companies to speed up efforts to become self-reliant, changing supply chains and investment priorities. The result is a fragmented global market, with different hardware and software ecosystems competing for users. For developers and researchers, this split brings new risks: code and models optimized for one platform may not work easily on another, and the lack of standard tools can slow progress on both sides. Recent advances in chip interoperability, such as custom AI chips connecting to Nvidia's infrastructure through NVLink Fusion, show that even new entrants still depend on Nvidia's server ecosystem rather than fully replacing it.
These trends mirror broader patterns in AI and robotics, where hardware and software lock-in can shape who leads innovation. As seen in the previous investigation of generalist robot models, the ability to adapt across platforms is increasingly important. In AI chips, though, the technical and institutional momentum behind Nvidia's ecosystem remains a major obstacle for challengers. The Max Planck Society and other research groups have called for open standards and cross-platform compatibility to speed up scientific progress in AI.
Nvidia's strength in the AI chip market is not just about faster processors. Its advantage comes from a complete ecosystem-hardware, software, networking, and developer support-which makes switching costly for organizations. Chinese companies have made real progress in hardware and early deployments, but the lack of a mature, widely used software stack still limits their reach. The current wave of domestic chip adoption in China is driven as much by necessity as by technical achievement, and the performance gap in large-scale model training remains wide. Until a domestic ecosystem can match Nvidia's in both capability and ease of use, the global AI infrastructure will stay divided by more than just hardware.
Model training is the process where machine-learning systems adjust their internal settings to fit patterns in data. For large AI models, this needs specialized hardware that can handle huge numbers of parallel calculations efficiently. Training involves running data through the model many times, updating weights to reduce error. It is computationally demanding and usually relies on GPUs or dedicated accelerators. Once trained, models can be used for inference-making predictions or generating outputs-on a range of hardware, but the training phase sets the technical requirements and often locks in the hardware and software ecosystem. For more detail on the computational demands of AI model training, see this Science journal article on large-scale neural network optimization.