Articles

CUDA Cores vs Tensor Cores: What's the Difference?

CUDA cores handle general parallel math on an NVIDIA GPU; Tensor cores are specialized units built for the matrix multiplies AI models run constantly.

Chisato Chisato · · 4 min read
Close-up of an NVIDIA GPU circuit board

CUDA cores are the general-purpose parallel processing units that make up the bulk of an NVIDIA GPU, each capable of executing a simple arithmetic operation on its own piece of data every clock cycle. Tensor cores are separate, specialized units built into the same chip specifically to accelerate matrix multiplication — the core operation behind neural network training and inference — computing an entire small matrix multiply-and-accumulate in far fewer cycles than the same work would take spread across CUDA cores.

What CUDA cores actually do

A GPU’s whole design philosophy is trading the CPU’s small number of powerful, flexible cores for a very large number of simpler ones, running the same instruction across many data points at once — the model that made GPUs useful for graphics rendering in the first place, since shading millions of pixels is the same computation repeated with different inputs. CUDA cores are those simple execution units: each one handles a single-precision floating-point or integer operation per cycle, and a modern GPU packs thousands of them, organized into clusters that execute instructions in lockstep across many threads simultaneously.

This general-purpose parallelism is what made GPUs useful for far more than graphics once NVIDIA opened CUDA as a programming interface for arbitrary parallel computation — scientific simulation, video encoding, and eventually the matrix-heavy math behind machine learning all mapped naturally onto the same hardware that used to just shade triangles.

Why matrix multiplication needed its own hardware

Training or running a neural network is, at the arithmetic level, an enormous, repetitive sequence of matrix multiplications — multiplying an input vector by a weight matrix, layer after layer, millions or billions of times over a training run. CUDA cores can do this work, computing it as a long sequence of individual multiply-and-add operations, but that approach doesn’t take advantage of the specific, predictable structure of a matrix multiply.

Tensor cores are purpose-built to exploit that structure. Rather than issuing one multiply-add per core per cycle, a single Tensor core performs a small, fixed-size matrix multiply-accumulate operation as one fused unit of work, dramatically increasing the useful arithmetic done per cycle for exactly the operation deep learning depends on most. The tradeoff is specialization: a Tensor core doesn’t help with graphics rasterization, general branching code, or any computation that isn’t shaped like a matrix multiply. It’s a narrow tool that happens to target the single most common operation in modern AI workloads.

Precision: the other half of the story

Tensor cores are also built around mixed-precision and reduced-precision arithmetic — computing at lower numerical precision than a CPU or traditional GPU math unit would use by default, then accumulating results at higher precision to control error. Neural networks are notably tolerant of reduced precision: the difference between a weight represented with full floating-point precision and one represented more coarsely rarely changes a model’s output in any way that matters, but computing at lower precision is significantly cheaper in both time and memory bandwidth.

This is the same underlying idea behind post-training quantization, applied instead at the hardware level during the computation itself rather than to the stored weights. A Tensor core computing in a lower-precision format can push through substantially more matrix math per cycle than the same silicon area computing at full precision, which is a large part of why Tensor-core-equipped GPUs became the default hardware for both training and serving large models.

Side by side

CUDA coresTensor cores
PurposeGeneral-purpose parallel arithmeticMatrix multiply-accumulate, specifically
Best atGraphics, general compute, branching codeNeural network training and inference
Operation granularityOne scalar op per core per cycleOne small matrix multiply per core per cycle
PrecisionTypically full precisionOften mixed/reduced precision by design
Present inEvery NVIDIA GPU generation since CUDA’s introductionGPU generations built for deep learning workloads
Programmed viaStandard CUDA kernelsFramework libraries (cuDNN and similar) that dispatch to them automatically

Why both matter together

A GPU running an AI workload doesn’t use Tensor cores exclusively — CUDA cores still handle everything that isn’t a matrix multiply: activation functions, data movement orchestration, and any general-purpose logic surrounding the core computation. The performance gain from Tensor cores only shows up when a workload is dense enough in matrix multiplication for the specialized path to dominate the runtime, which is exactly the profile of training and running large models, but not necessarily every GPU-accelerated task.

This division of labor is also why HBM bandwidth matters so much alongside compute: Tensor cores can chew through matrix math fast enough that feeding them data becomes the bottleneck, which is why AI-focused GPUs pair heavy Tensor core counts with the fastest memory subsystem the chip can support.

The takeaway

CUDA cores are the general-purpose workhorses that handle arbitrary parallel computation; Tensor cores are narrow, specialized units bolted on to do one thing — matrix multiply-accumulate, often at reduced precision — dramatically faster than CUDA cores could manage on their own. Neither replaces the other: modern AI-focused GPUs ship with both, using CUDA cores for everything general and routing the matrix-heavy core of neural network computation to Tensor cores, which is the main reason GPU performance on AI workloads has scaled well beyond what raw CUDA core counts alone would predict.

Chisato Chisato · · 4 min read

Wafer-Scale Integration Explained

Wafer-scale integration builds one giant chip from an entire silicon wafer instead of cutting it into dies. How it works and its tradeoffs.

#Hardware #Chips #Semiconductors
Chisato Chisato · · 4 min read

What Is HBM? High Bandwidth Memory Explained

HBM stacks DRAM dies vertically and connects them through a wide interface, trading capacity per chip for far higher bandwidth than standard DRAM.

#Hardware #Semiconductors #Computer Science