Quick Dive: What You'll Learn
I've spent years in the AI hardware space, and few chips have sparked as much debate as the Huawei Ascend series. When the Ascend 910 first hit benchmarks, I remember the shockwaves it sent through the industry — a Chinese chip claiming 256 TOPS of AI compute? That was unheard of. But the real story is messier, more nuanced. Let me walk you through what made the Ascend famous, and where it fell short.
The Raw Computing Power That Shocked the Industry
The Ascend 910 was Huawei's flagship — a monster designed for AI training. At its peak, it delivered 256 TFLOPS of half-precision floating-point performance. That number put it ahead of NVIDIA's V100 and even rivaled the early A100 in certain metrics. I remember pulling up the specs for the first time and thinking, "Is this real?"
But raw TOPS isn't everything. The Ascend 910 achieved this through a Da Vinci architecture that used a 3D Cube matrix multiplier. It was tailor-made for tensor operations — think convolutions and matrix multiplications. In practice, that meant training ResNet-50 or BERT models could be faster than on comparable NVIDIA cards, if the software stack played nice.
The catch? Power draw. The Ascend 910 consumed around 310W under load, similar to its competitors, but thermal management in early server designs was a headache. I’ve seen data centers where they had to rework cooling because the hot spots were unpredictable.
A Purpose-Built Architecture for AI Workloads
The Da Vinci core is what set Ascend apart. Unlike NVIDIA's general-purpose CUDA cores, Huawei designed the Ascend around a dedicated AI compute unit. Each core had a 3D Cube that could perform 16x16x16 matrix operations in one clock cycle. That’s efficient for AI, but it meant the chip was less flexible for non-AI tasks.
Let me break down the key architectural elements:
- Cube Unit: Handles matrix multiply-accumulate. Peak efficiency when batch sizes are large.
- Vector Unit: For element-wise operations and activation functions.
- Scalar Unit: Handles control flow and data movement.
This specialization gave Ascend a huge advantage in throughput-per-watt for training. But it also made the chip a nightmare for anything that didn't fit perfectly into the 3D Cube paradigm — like sparse models or unconventional layer types.
The Chinese Answer to NVIDIA's Dominance
For years, the AI accelerator market was a two-horse race: NVIDIA vs. everyone else. Huawei's Ascend was the first serious challenger from China. The Chinese government pushed it as a way to reduce dependence on foreign chips, especially after the trade bans. I’ve talked to engineers at Chinese tech companies who were required to test Ascend for certain projects. That wasn't just about performance — it was geopolitical.
Compared to NVIDIA's A100, the Ascend 910 had slightly lower raw FP32 performance but matched or exceeded it in INT8 inference. The real battle, however, was in the ecosystem. NVIDIA had CUDA, cuDNN, and a decade of software maturity. Huawei had CANN (Computer Architecture Neural Network) and MindSpore — and they were rough.
I'll be honest: porting a TensorFlow model to Ascend took me three times longer than expected. The documentation was sparse, and many operators weren't supported. But when it worked, the performance was impressive.
Practical Applications and Real-World Deployments
Where did Ascend actually shine? I saw it deployed in three main areas:
- Cloud Training: China's largest cloud providers (Huawei Cloud, Tencent) used Ascend 910 clusters for internal AI training. I visited a Huawei Cloud data center in Shenzhen where they had racks of Atlas 900 training nodes — each node packed with eight Ascend 910 chips. The throughput for image classification tasks was neck-and-neck with DGX-1 systems.
- Edge Inference: The Ascend 310, a lower-power chip, ended up in smart cameras and industrial robots. I tested an Edge Computing box with an Ascend 310 that could run real-time object detection at 30 FPS on 1080p video — at only 8W. That was genuinely impressive.
- Autonomous Driving: Huawei's MDC platform integrated Ascend chips for in-vehicle AI. I rode in a prototype bus that used the MDC 510. It handled lane detection and obstacle avoidance smoothly, though the system still felt prototype-ish.
Here's a quick comparison table I put together based on public specs and my own testing:
| Model | Peak FP16 (TFLOPS) | Power (W) | Key Strength |
|---|---|---|---|
| Ascend 910 | 256 | 310 | High throughput for dense models |
| NVIDIA A100 | 312 | 400 | Software maturity & flexibility |
| Google TPU v4 | 275 | ~300 | Optimized for TensorFlow |
| Intel Habana Gaudi | 265 | ~350 | Good for inference-heavy workflows |
The table doesn't capture the full picture though. In practice, the Ascend 910's actual FLOPS utilization was often lower because of software overhead. I’ve seen benchmarks where it only hit 60-70% of theoretical peak on custom models.
Software Ecosystem: The Weakest Link?
Let’s not sugarcoat it. The Ascend’s biggest weakness was — and still is — its software stack. CANN (Compute Architecture for Neural Networks) is the underlying framework. It’s supposed to be the equivalent of CUDA, but it’s far less polished. I remember spending an entire weekend just trying to compile a simple ONNX model for Ascend. The error messages were cryptic, and community support was thin.
Huawei’s own AI framework, MindSpore, works seamlessly with Ascend, but hardly anyone uses MindSpore outside of China. For PyTorch and TensorFlow users, there’s an adapter called TorchAir and TensorFlow on Ascend, but they lag behind the upstream versions. I’ve had models that ran fine on NVIDIA but failed on Ascend due to unsupported operations.
That said, Huawei did improve things over time. They released better profilers, debug tools, and a model zoo. But the gap with CUDA is still huge. If you're a startup with limited engineering resources, the learning curve might kill the project.
How the Ascend Stack Compares to NVIDIA CUDA
For those considering a switch, here's the practical reality: if your models use standard layers (Conv, BN, ReLU, Linear), Ascend will likely work fine after some tweaking. But if you rely on custom CUDA kernels, you're out of luck. There's no direct equivalent to CUDA’s inline PTX or kernel fusion capabilities. I once tried to port a model with a custom attention mechanism — after three days of frustration, I gave up and rented an A100 cluster instead.
Inference, though, is a different story. The Ascend 310's power efficiency makes it a great choice for edge deployment. I know several companies that use it for smart retail and security cameras. The trade-off is worth it if latency and power are more important than model flexibility.
Frequently Asked Questions
Fact-checking note: This article draws from Huawei's official technical papers, benchmark reports from MLPerf, and conversations with engineers at Chinese cloud providers. The performance numbers match publicly available data as of the last revision. I've personally tested both the Ascend 310 and 910 in lab conditions.
Reader Comments