What Is a TPU? Tensor Processing Units Explained
A TPU is a chip built specifically for the matrix math behind neural networks, using a systolic array instead of a general-purpose GPU pipeline.
A TPU, or Tensor Processing Unit, is a chip designed specifically to accelerate the matrix multiplications that make up the bulk of neural network training and inference. Unlike a GPU, which started as a general-purpose graphics chip later repurposed for machine learning, a TPU was built from the ground up with one job in mind: run tensor operations as fast and as efficiently as possible.
What a TPU is actually optimized for
Neural networks, at their core, spend most of their compute time on one operation repeated at enormous scale: multiplying and accumulating matrices. A transformer’s attention mechanism, a convolutional layer, a fully connected layer — nearly all of it reduces to matrix multiplication. General-purpose processors handle this, but they carry overhead for instruction decoding, branching, and flexible memory access that isn’t needed for this one repetitive task.
A TPU strips that overhead away. It’s an application-specific integrated circuit — see what an ASIC is for the broader category — purpose-built around a hardware structure called a systolic array, which is specifically shaped to perform matrix multiplication efficiently.
The systolic array: why TPUs are structured differently
A systolic array is a grid of small processing units, each one directly connected to its neighbors, arranged so that data flows through the grid in a synchronized rhythm — values are loaded once and reused across many multiply-accumulate operations as they pass from cell to cell, rather than being re-fetched from memory for every calculation.
This matters because moving data is often the real bottleneck in matrix math, not the arithmetic itself. A conventional processor has to repeatedly read operands from memory for each multiplication; a systolic array loads a value once and lets it ripple through many computations as it flows across the grid, dramatically cutting the number of memory accesses relative to the number of operations performed. That’s the core architectural bet a TPU makes: keep data moving through a specialized grid rather than shuttling it back and forth to memory.
TPU vs GPU
| TPU | GPU | |
|---|---|---|
| Origin | Purpose-built for tensor math | Originally built for graphics rendering, adapted for ML |
| Core structure | Systolic array | Thousands of parallel cores, CUDA cores and tensor cores |
| Flexibility | Narrow — optimized almost entirely for matrix ops | Broad — general parallel compute, graphics, and ML |
| Software ecosystem | Tied closely to specific ML frameworks | Broad ecosystem across ML, graphics, and general compute |
| Best fit | Large-scale, well-defined training or inference workloads | Mixed workloads, research flexibility, wide framework support |
A useful way to think about the difference from CPU vs GPU vs TPU: a CPU is general-purpose and flexible with the least parallelism, a GPU is broadly parallel and still fairly flexible, and a TPU narrows that flexibility even further in exchange for efficiency on exactly the operation neural networks need most. Modern GPUs have closed some of this gap by adding dedicated tensor cores of their own, but a TPU’s entire architecture is built around that one workload from the start rather than added on top of a graphics pipeline.
Where TPUs fit relative to other accelerators
TPUs sit alongside other AI-specific accelerators in a broader trend toward specialized silicon. An NPU, for instance, targets similar matrix-heavy workloads but is typically built for lower-power, on-device inference rather than large-scale data-center training — same underlying motivation (specialize the hardware to the math), different deployment target. A DPU addresses a different bottleneck entirely — offloading networking and storage processing rather than accelerating the model computation itself.
This specialization trend is a direct response to the limits of general-purpose compute for AI workloads: as model sizes and training costs have grown, the efficiency gap between a general processor and hardware purpose-built for the dominant operation has become large enough to justify designing, fabricating, and maintaining entirely separate chip families.
Tradeoffs of going purpose-built
The efficiency TPUs gain from specialization comes with a real cost: reduced flexibility. A chip built almost entirely around matrix multiplication is excellent at exactly that and comparatively weak at anything else — general control flow, irregular computation, or workloads that don’t map cleanly onto dense matrix operations. Software support is also narrower than for GPUs, which benefit from decades of tooling built for graphics and general compute that carried over to machine learning almost by accident.
That tradeoff mirrors the classic ASIC vs FPGA decision in chip design more broadly: an ASIC-style, fixed-function chip like a TPU is faster and more power-efficient once you’ve committed to the workload, but it can’t be reconfigured after the fact the way a more flexible or programmable design can.
The takeaway
A TPU accelerates neural network computation by dedicating its entire hardware structure — a systolic array — to matrix multiplication, cutting the memory traffic that slows general-purpose processors down on the same task. That specialization makes TPUs efficient at large, well-defined training and inference workloads, at the cost of the flexibility a GPU still offers for more varied or experimental work. Which one makes sense depends less on raw performance numbers and more on how well-defined and matrix-heavy the workload actually is.
Keep reading
Chisato · · 7 min read Nvidia Vera CPU Specs: Olympus Cores, SPEC Benchmarks
Nvidia detailed its Vera CPU — 88 custom Olympus cores, 1.2 TB/s memory, and SPEC CPU 2026 scores that edge AMD's Epyc dual-socket flagship.
Chisato · · 4 min read Wafer-Scale Integration Explained
Wafer-scale integration builds one giant chip from an entire silicon wafer instead of cutting it into dies. How it works and its tradeoffs.
The Lycoris Team · · 6 min read Apple iPhone 18 Pro Event: Foldable iPhone, New CEO
Apple's September 9 'Surprise and shine' event is set to unveil the iPhone 18 Pro and a first foldable iPhone — the first launch under new CEO John Ternus.