-+ 0.00%
-+ 0.00%
-+ 0.00%

Guohai Securities: Supernode reconstructs intelligent computing base, domestic AI accelerates breakthrough

Zhitongcaijing·07/23/2026 09:01:05
Listen to the news

The Zhitong Finance App learned that Guohai Securities released a research report saying that the conditions for large-scale deployment of domestic supernodes are gradually maturing, the pace of introducing downstream verification may be better than expected, and the penetration rate may increase at an accelerated pace. Coupled with the high technical barriers of supernodes, which is beneficial to the increase in both the share and profitability of leading OEMs, the bank is optimistic about leading server OEMs and the accelerated expansion of the industry chain, and maintains a “recommended” rating for the computer industry.

Guohai Securities's main views are as follows:

High bandwidth, low latency, and unified memory addressing to improve the efficiency of computing power utilization

1. Large model training has dual rigid bottlenecks in communication and video storage. High-frequency all-to-all communication in the training-side MoE model overwhelms a large amount of GPU computing power; KvCache on the inference side expands with the context, causing problems such as large capacity, high frequency of small packets, sensitive latency, and high concurrency. Traditional server clusters have poor utilization of computing power, spawning the immediate need for supernode architectures.

2. Increase communication bandwidth and achieve high-speed interconnection between AI acceleration cards. Scale-Up is responsible for high bandwidth and low latency communication between chips within the cabinet (carrying strong TP/EP coupling, 100 nanosecond delay), and multiple technology routes (bus/Ether/proprietary interconnect) in parallel, and accelerate breakthroughs and development of multiple interconnection technology routes.

3. Multi-layer technology collaboration improves the computing power efficiency of the whole machine. It relies on increasing communication bandwidth to achieve inter-card interconnection speed, relying on unified memory addressing to achieve memory pooling, optimizing the network topology to reduce the number of communication hops, superposition software stack scheduling optimization, shorten communication time, and greatly reduce single-token reasoning and training costs. At the same time, supernode implementation requires comprehensive evaluation of indicators such as cooling, operation and maintenance, and long-term ROI.

Overseas supernode benchmark full-stack architecture Vera Rubin platform

1. The hardware completely covers all training, push, storage and scheduling scenarios, and the introduction of LPU pushes up the upper limit of inference performance. The complete NVL72 platform includes Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 CPO and GroQ 3 LPU 7 chips, and 5 Vera Rubin NVL72, Groq 3 LPX, Vera CPU, BlueField-4 STX storage, Spectrum-6 SPX Ethernet rack. The Rubin architecture GPU inference computing power is increased by 5 times compared to the previous generation Blackwell, and Rubin and LPU collaboration can achieve a 35 times performance improvement.

2. The STX rack makes up for the shortcomings of long contextual inference storage, and lays out CPU racks and Ethernet CPO switches. The STX rack is equipped with the CMX context memory storage platform, which is designed to support the rapid expansion of KVCache required for long contexts and agent workloads. In response to future CPU requirements for proxy AI mass sandboxes and intensive learning, high-density Vera CPU cabinets were launched; Spectrum series CPO Ethernet switches were mass-produced simultaneously, and silicon technology may be upgraded to standard Nvidia network infrastructure in the future.

Domestic supernode deployment conditions are gradually maturing, and server OEMs are expected to directly benefit

1. Multi-line innovation of domestic internet agreements, various manufacturers. Domestic inter-card interconnection technology routes are multiple and parallel. Self-developed interconnection agreements such as Huawei UB, Haiguang Information Hyswitch, Mu Xi MetaXLink, and Ali Alink are competing for iterative breakthroughs, and the domestic high-speed interconnection system is being improved at an accelerated pace. Starting in 2025, chip, server, cloud, and ICT manufacturers will centrally release self-developed supernodes, while Huawei, Ali, and Zhongke Shuguang have all launched self-developed machines.

2. Real machine deployment and OEM orders in the 10,000 Ka cluster hit the real estate industry in batches, and there is plenty of room for rapid growth. Zhongke Shuguang achieved the deployment of a 100,000 card cluster, and Huaqin technology orders gradually began to be implemented in batches. Downstream demand may exceed expectations, and the penetration rate is expected to accelerate upward. The bank expects the domestic supernode market to reach 341% CAGR in 2025-2028.

3. Raise technical barriers and be optimistic about leading machines and the entire industry chain. The supernode integrates high-speed interconnection, liquid cooling, and machine structure design. The technical threshold is significantly higher than that of traditional servers, and the share and gross margin of leading OEMs is expected to increase.

Risk warning: Downstream industry demand recovery falls short of expectations, AI models fall short of expectations, risk of raw material price fluctuations, increased market competition, risk of exchange rate fluctuations, focus on the company's performance falling short of expectations.