DeepSeek-V4 Family Exposed in Research Paper on Ascend SuperPOD Training

data center

A Stealthy Debut: The DeepSeek-V4 Family Shows Up on Hugging Face

On July 23, the AI research community got an unexpected signal: a paper titled "SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD" rose to the top of Hugging Face's daily papers ranking. With 28 upvotes and an unusual 65 co-authors, the posting—submitted by user 'crazyofapple'—quickly drew attention. The paper doesn't just introduce a novel training technique; it casually reveals the existence of an entire new model family from one of the world's most closely watched AI labs. DeepSeek, the Chinese AI company that shook the industry with its cost-efficient R1 and V3 models, now appears to have moved on to a next-generation version, dubbed V4.

What the Paper Actually Says: Full-Parameter Post-Training Takes Center Stage

The research focuses on a method called SLAI T-Rex, designed for full-parameter post-training of large language models. Unlike parameter-efficient fine-tuning (PEFT) methods like LoRA, which update only a small fraction of a model's weights, full-parameter approaches modify every parameter during post-training. This is computationally demanding but can lead to better adaptation, especially when aligning a model to new tasks or instruction formats. The paper demonstrates this technique on the DeepSeek-V4 family, which strongly suggests that such models already exist in a pre-trained form and are being adapted for downstream use. The term "post-training" implies that the base V4 model had been pre-trained previously, likely on a massive dataset, and this work focuses on alignment and fine-tuning.

The SLAI T-Rex method presumably addresses challenges in training stability or efficiency when applying full-parameter updates at scale. While the paper's abstract and full details remain behind the Hugging Face listing, the high community engagement—12 saves or comments within hours—indicates immediate impact. What stands out is that all this work was performed not on standard NVIDIA GPUs, but on Huawei's Ascend SuperPOD, a cluster of Ascend AI accelerators. This is a deliberate, public demonstration of an alternative hardware ecosystem.

data center

Why Huawei Ascend? The Geopolitics of AI Infrastructure

Since 2022, U.S. export controls have restricted the sale of advanced NVIDIA chips, like the A100 and H100, to China. In response, Chinese tech giants have accelerated investment in domestic hardware. Huawei's Ascend series, particularly the 910B and newer 910C processors, have emerged as the leading alternative. The Ascend SuperPOD is Huawei's counterpart to NVIDIA's DGX SuperPOD—a high-density, interconnection-optimized system for training large models. DeepSeek's previous models, including V3, were trained on NVIDIA H800s, which are slightly modified H100s designed to comply with earlier export restrictions. By now moving to Ascend for post-training, DeepSeek signals that Huawei's infrastructure has reached a maturity sufficient for cutting-edge LLM work.

The paper's authorship likely includes engineers from both DeepSeek and Huawei. Such a collaboration would be strategic: Huawei gets to showcase its platform with a high-profile model, and DeepSeek secures a reliable hardware supply chain free from U.S. sanctions. This is a significant geopolitical milestone. If the training results are competitive, it could accelerate the adoption of Ascend not just in China but globally, as more labs seek to diversify away from a single vendor.

Implications for the AI Industry: A Shift in Hardware Dependency

The quiet unveiling of DeepSeek-V4 on an all-Chinese hardware stack has several layers of meaning. First, it demonstrates that advanced LLMs can be built without NVIDIA—something many Western labs still consider infeasible at scale. DeepSeek had already challenged the industry with its efficient training techniques on older hardware; now it is proving that even state-of-the-art post-training can run on non-CUDA architectures. Huawei's CANN (Compute Architecture for Neural Networks) software stack is a direct competitor to CUDA, and this paper serves as a public proof point that CANN can handle full-parameter updates on models with likely hundreds of billions of parameters.

supercomputing cluster

Second, the mention of the DeepSeek-V4 family suggests that DeepSeek is already iterating beyond V3. The community has been anticipating a successor to the open-source V3 model, but the company has not made a formal announcement. This paper is effectively a controlled leak, perhaps intentional, to build anticipation and demonstrate hardware independence. For developers and enterprises, this means the next generation of DeepSeek models could be available soon, and they may run natively on Ascend hardware, reducing inference costs for Chinese users and any international buyers of Huawei's AI accelerators.

What to Watch: Availability, Benchmarks, and the NVIDIA Response

The paper's appearance raises more questions than it answers. There are no benchmarks, so the community doesn't yet know how V4 compares to V3, GPT-4o, or Claude 3.5. The "family" designation implies multiple model sizes, possibly catering to different computational budgets. The full-parameter post-training approach could lead to better instruction following or reasoning capabilities, areas where V3 already impressed. However, without seeing actual performance data, the paper is primarily an infrastructure teaser.

Nevertheless, the broader signal is unmistakable: the AI hardware moat that NVIDIA has enjoyed is under active erosion. DeepSeek's move coincides with growing interest in AMD's Instinct GPUs, Intel Gaudi, and Google's TPUs. But Huawei's Ascend is the only ecosystem backed by a major AI lab openly using it for frontier model training. A successful V4 release trained and post-trained on Ascend could fracture the CUDA monopoly, giving rise to a truly multi-vendor AI hardware landscape. That would benefit the entire industry by lowering costs and preventing supply chain bottlenecks.

The Hugging Face paper is likely just the beginning. Watch for a formal DeepSeek V4 announcement, detailed benchmarks, and the inevitable scrutiny from analysts tracking the impact of U.S. export controls. For now, the AI world has been put on notice: the next generation of open-weight models may run on chips you didn't expect.

345tool Editorial Team
345tool Editorial Team

We are a team of AI technology enthusiasts and researchers dedicated to discovering, testing, and reviewing the latest AI tools to help users find the right solutions for their needs.

我们是一支由 AI 技术爱好者和研究人员组成的团队,致力于发现、测试和评测最新的 AI 工具,帮助用户找到最适合自己的解决方案。

コメント

Loading comments...