CUDA Cores vs Tensor Cores: What's the Difference?
CUDA cores handle general parallel math on an NVIDIA GPU; Tensor cores are specialized units built for the matrix multiplies AI models run constantly.
CUDA cores are the general-purpose parallel processing units that make up the bulk of an NVIDIA GPU, each capable of executing a simple arithmetic operation on its own piece of data every clock cycle. Tensor cores are separate, specialized units built into the same chip specifically to accelerate matrix multiplication — the core operation behind neural network training and inference — computing an entire small matrix multiply-and-accumulate in far fewer cycles than the same work would take spread across CUDA cores.
What CUDA cores actually do
A GPU’s whole design philosophy is trading the CPU’s small number of powerful, flexible cores for a very large number of simpler ones, running the same instruction across many data points at once — the model that made GPUs useful for graphics rendering in the first place, since shading millions of pixels is the same computation repeated with different inputs. CUDA cores are those simple execution units: each one handles a single-precision floating-point or integer operation per cycle, and a modern GPU packs thousands of them, organized into clusters that execute instructions in lockstep across many threads simultaneously.
This general-purpose parallelism is what made GPUs useful for far more than graphics once NVIDIA opened CUDA as a programming interface for arbitrary parallel computation — scientific simulation, video encoding, and eventually the matrix-heavy math behind machine learning all mapped naturally onto the same hardware that used to just shade triangles.
Why matrix multiplication needed its own hardware
Training or running a neural network is, at the arithmetic level, an enormous, repetitive sequence of matrix multiplications — multiplying an input vector by a weight matrix, layer after layer, millions or billions of times over a training run. CUDA cores can do this work, computing it as a long sequence of individual multiply-and-add operations, but that approach doesn’t take advantage of the specific, predictable structure of a matrix multiply.
Tensor cores are purpose-built to exploit that structure. Rather than issuing one multiply-add per core per cycle, a single Tensor core performs a small, fixed-size matrix multiply-accumulate operation as one fused unit of work, dramatically increasing the useful arithmetic done per cycle for exactly the operation deep learning depends on most. The tradeoff is specialization: a Tensor core doesn’t help with graphics rasterization, general branching code, or any computation that isn’t shaped like a matrix multiply. It’s a narrow tool that happens to target the single most common operation in modern AI workloads.
Precision: the other half of the story
Tensor cores are also built around mixed-precision and reduced-precision arithmetic — computing at lower numerical precision than a CPU or traditional GPU math unit would use by default, then accumulating results at higher precision to control error. Neural networks are notably tolerant of reduced precision: the difference between a weight represented with full floating-point precision and one represented more coarsely rarely changes a model’s output in any way that matters, but computing at lower precision is significantly cheaper in both time and memory bandwidth.
This is the same underlying idea behind post-training quantization, applied instead at the hardware level during the computation itself rather than to the stored weights. A Tensor core computing in a lower-precision format can push through substantially more matrix math per cycle than the same silicon area computing at full precision, which is a large part of why Tensor-core-equipped GPUs became the default hardware for both training and serving large models.
Side by side
| CUDA cores | Tensor cores | |
|---|---|---|
| Purpose | General-purpose parallel arithmetic | Matrix multiply-accumulate, specifically |
| Best at | Graphics, general compute, branching code | Neural network training and inference |
| Operation granularity | One scalar op per core per cycle | One small matrix multiply per core per cycle |
| Precision | Typically full precision | Often mixed/reduced precision by design |
| Present in | Every NVIDIA GPU generation since CUDA’s introduction | GPU generations built for deep learning workloads |
| Programmed via | Standard CUDA kernels | Framework libraries (cuDNN and similar) that dispatch to them automatically |
Why both matter together
A GPU running an AI workload doesn’t use Tensor cores exclusively — CUDA cores still handle everything that isn’t a matrix multiply: activation functions, data movement orchestration, and any general-purpose logic surrounding the core computation. The performance gain from Tensor cores only shows up when a workload is dense enough in matrix multiplication for the specialized path to dominate the runtime, which is exactly the profile of training and running large models, but not necessarily every GPU-accelerated task.
This division of labor is also why HBM bandwidth matters so much alongside compute: Tensor cores can chew through matrix math fast enough that feeding them data becomes the bottleneck, which is why AI-focused GPUs pair heavy Tensor core counts with the fastest memory subsystem the chip can support.
The takeaway
CUDA cores are the general-purpose workhorses that handle arbitrary parallel computation; Tensor cores are narrow, specialized units bolted on to do one thing — matrix multiply-accumulate, often at reduced precision — dramatically faster than CUDA cores could manage on their own. Neither replaces the other: modern AI-focused GPUs ship with both, using CUDA cores for everything general and routing the matrix-heavy core of neural network computation to Tensor cores, which is the main reason GPU performance on AI workloads has scaled well beyond what raw CUDA core counts alone would predict.
Tagged
Keep reading
Chisato · · 4 min read Wafer-Scale Integration Explained
Wafer-scale integration builds one giant chip from an entire silicon wafer instead of cutting it into dies. How it works and its tradeoffs.
Chisato · · 4 min read What Is HBM? High Bandwidth Memory Explained
HBM stacks DRAM dies vertically and connects them through a wide interface, trading capacity per chip for far higher bandwidth than standard DRAM.
Chisato · · 4 min read What Is Dennard Scaling? Why Clock Speeds Stopped Climbing
Dennard scaling held that shrinking transistors kept power density constant, letting clock speeds rise for free. Its breakdown reshaped chip design.