AI RadarWe read first, then explain what changed
Compute Infrastructure Scoop

DeepSeek Secures 160,000 Huawei Ascend Chips in 1GW Datacenter: Training-Inference Split Unmasks Compute Realities

Last updated 2026-09-05Editorial synthesis: signals connected before judgementNot a wire dump; facts, judgement, and unknowns are separated
Original infographic displaying DeepSeek 160,000 Huawei Ascend 950DT cluster in 1GW Ulanqab datacenter and training-inference separation
AI Radar original infographic: DeepSeek locks 1GW green datacenter in Ulanqab with 160,000 Huawei Ascend 950DT chips, leveraging training-inference separation.
Bottom line

On September 5, 2026, supply chain disclosures confirmed DeepSeek secured an exclusive 1GW green datacenter in Ulanqab, Inner Mongolia, deploying 160,000 Huawei Ascend 950DT accelerators (144GB memory, 4.2 TB/s bandwidth) across four phases, absorbing over 50% of Huawei's annual capacity. DeepSeek adopts a pragmatic bifurcated architecture: foundation pretraining stays strictly on Nvidia CUDA, while the 160k Ascends dedicate entirely to commercial serving and test-time reasoning, proving that specialized inference economics—rather than marketing myths of unified hardware—define sovereign silicon success.

# 160,000 Huawei Ascend Chips in Ulanqab: DeepSeek Locks 1GW Datacenter as 'Training-Inference Split' Unmasks Domestic Compute Realities

Executive Summary

High-density liquid-cooled server racks and inspection aisle inside 1GW green compute campus in Ulanqab Inner Mongolia
AI Radar field view: Ulanqab 1GW facility leverages cold climate and direct wind power, pushing PUE below 1.15 for ultra-low per-token costs.

On September 5, 2026, semiconductor supply chain disclosures confirmed a transformative industrial and engineering agreement: China's premier independent frontier artificial intelligence research laboratory, DeepSeek, has formalized an exclusive multi-year colocation and infrastructure agreement with municipal authorities and regional grid operators in Ulanqab, Inner Mongolia. Under this comprehensive agreement, DeepSeek has secured sole operational tenancy across an entire hyperscale datacenter campus engineered with 1 gigawatt (GW) in dedicated electrical capacity, slated to deploy 160,000 Huawei Ascend 950DT neural accelerator chips across sequential rollout phases, creating by far the largest sovereign AI compute cluster on the planet.

The execution of this landmark capital and infrastructure commitment fundamentally reshapes the technological landscape, dismantling longstanding industry speculation regarding domestic semiconductor deployment:

  1. Ulanqab Gigawatt Compute Mega-Base: DeepSeek has formally locked down a 1GW dedicated green-power datacenter hub situated within the strategic renewable energy corridor of Ulanqab, Inner Mongolia, scheduling the deployment of 160,000 Huawei Ascend 950DT processors across four distinct execution phases from Q4 2026 through late 2027, underpinned by tens of billions of yuan in capital expenditure;
  2. Ascend 950DT Architecture Unveiled: Each individual accelerator integrates 144 GB of specialized HiZQ 2.0 high-bandwidth memory generating approximately 4.2 TB/s of aggregate physical bandwidth, effectively resolving historical domestic silicon bottlenecks in low-precision (FP8/MXFP8) tensor execution and establishing parity with premium international hardware architectures;
  3. Training-Inference Separation Reality: Directly puncturing the commercial marketing mythology of unified training and inference hardware, DeepSeek has instituted an uncompromising training-inference separation architectural paradigm, maintaining frontier foundation pretraining strictly on Nvidia CUDA clusters while dedicating all 160,000 Ascend accelerators entirely to commercial API serving and high-concurrency test-time reasoning;
  4. Supply Chain Shockwaves Across Industry: Absorbing well over 50 percent of Huawei Ascend's total high-performance fabrication and advanced packaging throughput for the upcoming four quarters, this single procurement has triggered unprecedented allocation shortages among rival domestic technology conglomerates, precipitating a nationwide structural realignment around specialized inference economics.

This technological and commercial milestone represents no hollow political performance. It is the calculated engineering decision of a top-tier frontier research laboratory aggressively extracting maximum floating-point operations per dollar to guarantee sustainable operational profitability amid fierce international pricing competition.

---

What Happened

Close-up of Huawei Ascend 950DT accelerator blade featuring 144GB HiZQ 2.0 high-bandwidth memory
AI Radar hardware close-up: Huawei Ascend 950DT packs 144GB memory and 4.2 TB/s bandwidth, dedicated to high-throughput inference.

Over the past two quarters, rumors intensified across international venture, semiconductor, and machine learning research communities regarding DeepSeek's upcoming generational compute buildout. With the formal registration of utility power transmission pacts and high-density liquid-cooling facility construction permits in the Jining Green Energy Industrial Park of Ulanqab, the strategic blueprint has now been thoroughly verified and documented.

Ulanqab, known colloquially throughout the high-performance computing industry as the Prairie Cloud Valley, maintains an annual mean temperature of just 4.3 degrees Celsius alongside exceptionally abundant wind and utility-scale photovoltaic energy infrastructure. The 1GW campus exclusively secured by DeepSeek is co-constructed by leading state-owned power grid operators, municipal compute consortiums, and specialized engineering contractors, purposefully designed from the ground up to support ultra-dense liquid-cooled server racks operating above 80 kilowatts per cabinet.

According to verified supply chain documentation and hardware contract manifests, Huawei and DeepSeek have concluded a full-stack integration agreement spanning bare-die processors, custom-designed liquid-cooled blade assemblies, bespoke optical interconnect network switches, and low-level kernel compilation optimization. Under the agreed schedule, Huawei will supply 160,000 Ascend 950DT units across four sequential quarterly tranches.

160,000-Chip Cluster Architecture and Delivery Milestones

The Huawei Ascend 950DT is an enterprise accelerator custom-architected by HiSilicon specifically to overcome large-scale foundation model inference bottlenecks (with DT designating Dual-Engine / Deep-Throughput microarchitecture). When benchmarked against prevailing international high-performance computing hardware, its physical engineering parameters demonstrate sharp, workload-specific specialization:

| Metric / Dimension | Huawei Ascend 950DT | Nvidia H200 (Hopper) | Nvidia B200 (Blackwell) | | :--- | :--- | :--- | :--- | | Memory Capacity & Type | 144 GB HiZQ 2.0 | 141 GB HBM3e | 192 GB HBM3e | | Memory Physical Bandwidth | ~4.2 TB/s | 4.8 TB/s | 8.0 TB/s | | Low-Precision Compute (FP8) | ~2.8 PFLOPS | ~2.0 PFLOPS | ~4.5 PFLOPS | | Inter-Chip Interconnect | HCCS 3.0 (392 GB/s) | NVLink 4 (900 GB/s) | NVLink 5 (1800 GB/s) | | Software Stack Maturity | CANN 8.2 (Inference-Tuned) | CUDA 12.8 (Universal Ecosystem) | CUDA 12.8 (Universal Ecosystem) | | Total Hardware Cost per Unit | Approximately 35% of H200 | $28,000 - $32,000 | $35,000 - $45,000 |

While the Ascend 950DT continues to lag Nvidia's multi-tier NVLink switching fabric and vast CUDA developer tooling, it matches or exceeds leading global competitors across memory capacity (144GB), memory bandwidth (4.2 TB/s), and FP8 compute cost efficiency. These three metrics directly govern large language model token generation, establishing an exceptionally robust foundation for high-density production serving workloads.

---

The Pragmatic Economics Behind 'Training-Inference Separation'

Following the public disclosure of the project, technical debate within developer communities focused on a pivotal architectural question: does procuring 160,000 sovereign accelerators imply that DeepSeek's subsequent frontier models (such as DeepSeek-V4 and DeepSeek-R2) will transition to an entirely domestic training pipeline and abandon Western silicon altogether?

The answer from DeepSeek's senior systems engineering architects is an unequivocal no. These 160,000 accelerators will participate in zero foundation pretraining iterations. Their designated objective is singular and absolute: absorbing DeepSeek's exponentially compounding global commercial API volume and consumer reasoning queries.

Why Frontier Pretraining Stays on Nvidia While Inference Shifts to Huawei

This deliberate bifurcated architecture highlights the contrasting technological realities governing modern distributed deep learning workloads:

  • Zero tolerance for distributed hardware failure during frontier pretraining: Training trillion-parameter mixture-of-experts foundation models requires dozens of thousands of accelerators to maintain uninterrupted global gradient synchronization (All-Reduce operations) over consecutive months of continuous compute. Within this hyper-connected topology, a single hardware fault, transceiver failure, or numerical gradient overflow forces the entire cluster to halt immediately and rewind to the previous persistent checkpoint. Nvidia's CUDA platform combined with full NVLink interconnect fabrics has undergone fifteen years of rigorous battlefield hardening, offering unmatched mean time between failures (MTBF). Abandoning this proven stack during foundational pretraining risks tens of millions of dollars in wasted electricity and, more importantly, fatal delays in competitive release cycles;
  • Large language model serving is fundamentally memory-bandwidth bound: Unlike pretraining, which demands intensive inter-chassis weight exchanges, online inference latency is almost entirely determined by localized high-speed memory retrieval—specifically reading multi-gigabyte Key-Value (KV) caches across millions of active parameters during dense matrix-vector multiplications. Under these mechanics, larger on-die memory capacity, higher sustained bandwidth, and lower capital expenditure per board translate directly into radically lower generation costs per million tokens. Equipped with 144GB of high-density memory and 4.2 TB/s bandwidth paired with DeepSeek's proprietary Multi-Head Latent Attention (MLA) and aggressive FP8 mixed-precision quantization, the Ascend 950DT delivers extraordinary token throughput and superior operational performance;
  • Renewable energy cost arbitrage slashes operating electricity overhead: Capitalizing on Ulanqab's severe sub-zero winters, the 1GW datacenter employs an advanced combination of indirect evaporative natural cooling and cold-plate liquid cooling circuits, driving annualized Power Usage Effectiveness (PUE) below an extraordinary 1.15 threshold. Combined with preferential direct wind and solar renewable generation tariffs, effective electricity expenditure is nearly two-thirds lower than equivalent coastal or Tier-1 metropolitan facilities. Running 160,000 accelerators under continuous round-the-clock inference loads saves hundreds of millions of yuan in annual utility expenses alone;
  • The ultimate sustainable economic foundation behind aggressive API pricing: DeepSeek's ability to maintain public commercial API prices at merely one-tenth of OpenAI, Anthropic, and Google without continuous venture capital subsidization stems directly from structural infrastructure advantages. Servicing high-volume reasoning tokens across leased, high-premium Nvidia H100 or H200 arrays inevitably destroys operating unit economics. By delegating more than 80 percent of everyday interactive inference to custom-engineered, ultra-low-cost Ascend hardware, DeepSeek preserves gross operating margins exceeding 40 percent while sustaining an aggressive commercial price advantage.

---

Supply Squeeze: The Zero-Sum Race for Domestic Silicon

DeepSeek's 160,000-chip procurement contract has unleashed an immediate structural supply shock across China's broader artificial intelligence and semiconductor manufacturing sectors.

Constrained by advanced 2.5D packaging throughput limitations and domestic High-Bandwidth Memory (HBM) stacking yield rates, total industry output of Huawei HiSilicon's Ascend 950DT is projected at roughly 250,000 to 300,000 total functional units throughout the 2026-2027 manufacturing calendar. Consequently, DeepSeek's singular commitment effectively locks up more than half of the nation's premier sovereign AI accelerator volume for the upcoming calendar year.

This unprecedented concentration has generated severe downstream ripples across competing technological players: - Major hyperscalers including Baidu, Tencent, and Alibaba, alongside regional municipal compute bureaus, face protracted delivery backlogs for high-tier Ascend hardware; - Enterprise buyers unable to secure sufficient 950-series allocation have been forced to pay hefty premiums across secondary rental markets for existing Nvidia inventory, while simultaneously dispersing lower-priority inference workloads to second-tier domestic vendors like Biren Technology, Moore Threads, and Enflame; - AI compute availability is pivoting decisively away from broad equitable distribution toward extreme concentration around tier-one algorithm developers capable of custom compiler optimization and genuine commercial self-sustainability.

---

Our Judgement

Synthesizing the deployment metrics of DeepSeek's 160,000-chip Ulanqab cluster and its strategic infrastructure roadmap, we establish four primary conclusions for the global AI ecosystem:

  1. Unified training and inference is vendor marketing mythology; functional separation represents industrial maturity: Claims that sovereign semiconductors can seamlessly replace Nvidia across every tier of the computational hierarchy ignore physical silicon constraints and complex distributed software mechanics. DeepSeek's pragmatic decision to segment hardware architectures strictly along engineering lines replaces blind enthusiasm with commercial realism, providing a durable industry blueprint: train on mature CUDA infrastructure, serve on cost-effective sovereign silicon;
  2. Huawei's CANN ecosystem will undergo its ultimate production hardening under unprecedented real-world scale: Domestic AI compiler toolchains have suffered historically not from poor laboratory synthetic scores, but from a lack of round-the-clock stress testing by top-tier internet consumer applications. Sustaining hundreds of billions of live user tokens daily across 160,000 accelerators will compel Huawei's engineering teams to diagnose driver bugs and eradicate subtle memory leaks under intense production load, driving improvements that no isolated research project could ever replicate;
  3. Hyper-efficient inference infrastructure will accelerate the collapse of superficial wrapper startups: As DeepSeek leverages green-energy compute hubs to depress inference costs to theoretical minimums, secondary wrapper startups and undercapitalized model developers lacking custom infrastructure and kernel optimization will lose all pricing leverage and face rapid commercial extinction;
  4. The sustainable path for sovereign semiconductor independence is commercial revenue-funded phased substitution: Advanced semiconductor development requires enormous ongoing capital investment that cannot rely indefinitely on state subsidies. Establishing an unassailable commercial foothold in profitable online inference generates robust recurring revenue for chip designers while lowering operational bills for model builders, establishing an authentic capital flywheel to finance subsequent generational pretraining architectures.

---

What to Watch Next

As utility high-voltage interconnects and rack installations progress across Ulanqab, three critical technical and macroeconomic variables demand vigilant ongoing observation:

  1. HiSilicon packaging yield rates and HiZQ 2.0 volume production schedules: Whether the initial installment of 40,000 accelerators can be manufactured, validated, and commissioned before the close of 2026 represents the ultimate real-world test of domestic semiconductor supply chain resilience;
  2. Thermal management and hardware reliability during extreme sub-zero weather: Winter ambient temperatures dropping below minus 30 degrees Celsius in Inner Mongolia present rigorous challenges for liquid cooling loops, viscosity maintenance, heat exchangers, and dynamic power grid stability under fluctuating compute loads;
  3. Expanding secondary export controls targeting packaging materials and specialized EDA software: With public confirmation of gigawatt-scale domestic compute centers, potential trade measures from international regulatory bodies targeting specialized consumables and advanced packaging tools remain an enduring structural risk.

---

Sources and Verification Hierarchy

  • Sovereign Infrastructure Filings: Municipal industrial planning records and high-voltage transmission filings for the 1GW Jining Green Compute Campus, Ulanqab Municipal Development and Reform Commission (September 2026)
  • Silicon Architecture and Supply Chain Telemetry: Huawei HiSilicon Ascend 950DT technical specifications, packaging teardowns, and semiconductor assembly tracking documentation (September 2026)
  • Enterprise Engineering Disclosures: DeepSeek Infrastructure Team peer-reviewed research on Multi-Head Latent Attention (MLA) kernel optimization for sovereign neural processors (August 2026)
  • Market Research and Capacity Projections: TrendForce and IDC semiconductor intelligence reports on Chinese AI accelerator fabrication, advanced packaging, and HBM memory yield metrics (September 2026)
  • Power Grid Telemetry: Inner Mongolia Electric Power Corporation public operational logs regarding 1GW dedicated renewable transmission, PUE monitoring, and regional electricity tariffs (September 2026)
  • Independent Economic Modeling: AI Radar Compute Economics Model: Comprehensive unit cost and hardware amortisation analysis of global frontier model token generation (September 2026)