What Is Memory-Mapped I/O? MMIO Explained
Memory-mapped I/O maps device registers into the CPU's address space, letting hardware be read and written with ordinary load and store instructions.
Memory-mapped I/O (MMIO) is a technique that gives peripheral devices — a GPU, a network card, a disk controller — their own addresses within the CPU’s regular memory address space, so software can talk to them using ordinary load and store instructions instead of special-purpose I/O instructions. From the CPU’s point of view, writing a command to a device register looks identical to writing to RAM; the difference is invisible at the instruction level and handled entirely by the memory system underneath.
The alternative: port-mapped I/O
Before MMIO was the default, many architectures used port-mapped I/O, which gives devices a separate address space entirely, accessed with dedicated instructions (IN and OUT on x86) rather than the CPU’s normal load/store instructions. Port-mapped I/O keeps device addresses cleanly separate from memory addresses, but it requires the instruction set to carry special I/O instructions and means every piece of software that talks to hardware needs to use a different code path than the one it uses for regular memory access. x86 still supports port I/O for legacy compatibility, but modern devices overwhelmingly use MMIO instead.
How MMIO actually works
A region of the physical address space is reserved and wired, at the hardware level, not to RAM but to a device’s control and status registers. When the CPU issues a normal store instruction to an address inside that region, the memory controller routes the write to the device instead of to RAM; a load from that address returns whatever the device currently reports, rather than data that was previously written there.
// A simplified example: writing to a device register mapped at a fixed address.
volatile uint32_t *device_ctrl = (uint32_t *)0xFEA00000;
*device_ctrl = 0x1; // Looks like a memory write; actually a command to hardware.
The volatile qualifier in that example matters more here than almost anywhere else in C: without it, a compiler is free to assume that writing the same value to the same address twice in a row is redundant and eliminate the second write as an optimization — which is exactly the wrong assumption for a device register, where writing the same value twice might genuinely need to happen (to trigger an action twice) or where reading the register might have side effects the compiler has no way to know about. volatile tells the compiler: don’t optimize away or reorder accesses to this location, because reads and writes here have effects beyond just storing a value.
Why devices need this at all
A network card, for instance, exposes registers for things like “here’s the address of the next packet to transmit,” “start transmitting now,” and “how many bytes were received.” None of that data behaves like RAM — writing “start transmitting” doesn’t store a value for later retrieval, it triggers hardware to act. MMIO gives the CPU a uniform way to reach all of this: the same load/store instructions, the same addressing modes, and the same caches and pipeline machinery that handle ordinary memory access also handle device access, with the memory controller quietly routing requests to the right destination based on address.
MMIO and caching: a hazard, not a feature
Ordinarily, caching a memory address is a pure performance win — CPU caches exist precisely to avoid repeatedly going out to slower memory. For MMIO addresses, caching is actively dangerous: if a “status register” read gets served from a stale cache line instead of hitting the actual device, software can end up reading old status data forever, never seeing that an operation has completed. For this reason, MMIO address ranges are marked as uncacheable (or “device memory,” in ARM’s terminology) at the page-table level, and the CPU deliberately skips its normal caching behavior for any access that lands in that range — one of the rare cases where the memory system is told to be slower on purpose, because correctness depends on every access reaching the real device.
MMIO vs DMA: two different jobs
MMIO and direct memory access (DMA) are often mentioned together but solve different halves of the device-communication problem. MMIO is how the CPU issues commands to a device and reads its status — small, low-latency, register-sized transfers. DMA is how bulk data moves between a device and RAM without the CPU copying every byte itself — a network card or disk controller uses DMA to write incoming data directly into a memory buffer, and only uses MMIO to tell the CPU “that transfer is done” (typically via an interrupt) or to check a status register confirming it. A high-throughput device almost always uses both: MMIO for control, DMA for the actual payload.
Where address translation fits in
The addresses software uses to reach MMIO regions are virtual addresses like any other, translated to physical addresses through the same page tables and TLB used for regular memory. The operating system maps a device’s physical MMIO region into a process’s (usually the kernel’s, or a driver’s) virtual address space explicitly, the same mechanism underlying virtual memory generally — the distinguishing detail is simply that the physical address on the other end of the mapping belongs to a device, not to RAM controlled by the memory controller in the ordinary sense.
The takeaway
Memory-mapped I/O lets software talk to hardware devices using the same load and store instructions it already uses for RAM, by giving device registers real addresses in the CPU’s address space and letting the memory controller route accesses to the right place. It trades away caching — MMIO regions are deliberately marked uncacheable, since a stale cache line reading device state would be actively wrong — in exchange for a uniform, general-purpose way to reach any device without dedicated I/O instructions. Where MMIO handles commands and status, DMA handles the bulk data transfer alongside it; real device drivers almost always use both.
Keep reading
Chisato · · 4 min read Thermal Interface Materials Explained
Thermal interface material fills microscopic gaps between a chip and its heatsink so heat can actually transfer to the cooler.
Chisato · · 4 min read Clock Speed vs. IPC: What Actually Makes a CPU Fast
Clock speed measures cycles per second; IPC measures work done per cycle. Real CPU performance is the product of both, not either one alone.
Chisato · · 4 min read UMA vs NUMA: Memory Architecture Explained
UMA gives every CPU core equal-latency memory access; NUMA gives each core faster access to its local memory bank. How the two architectures differ.