Copy-on-Write Explained: How Lazy Copying Works
Copy-on-write shares data until someone modifies it, then copies only what changed. How CoW powers fork(), snapshots, containers, and immutable data.
Copy-on-write (CoW) is an optimization where multiple owners share the same copy of some data, and a private copy is made only at the moment one of them tries to modify it. Until a write happens, a “copy” costs almost nothing: it’s just another reference to the original. The expensive work is deferred until it’s actually needed, and often it never is.
The idea shows up at nearly every layer of a computer, from the operating system’s memory manager to filesystems, container images, databases, and the collection types in your programming language. Once you recognize the pattern, a lot of seemingly magical performance tricks become easy to explain.
The core idea
Imagine two variables pointing at the same large buffer. A naive copy duplicates every byte up front. Copy-on-write instead does three things:
- Share. Both owners point to the same underlying data, and the data is marked as shared (by a reference count, a read-only flag, or a page-table permission).
- Detect the write. When either owner tries to modify the data, the system notices the data is shared.
- Copy, then write. It makes a private copy for the writer, points the writer at the copy, and applies the modification there. The other owner keeps the original, untouched.
The payoff comes from the common case where copies are made but rarely modified, or modified only in small parts. If the unit of copying is small, such as a single memory page or a single tree node, a write only duplicates that unit rather than the whole structure.
Copy-on-write in operating systems: fork()
The classic example is the Unix fork() system call, which creates a child process that starts as a duplicate of its parent. Copying the parent’s entire address space would be slow and wasteful, especially since the child often calls exec() immediately to run a different program.
So the kernel doesn’t copy. It gives the child its own page tables that map to the same physical pages as the parent, and marks those pages read-only in both processes. As long as both only read, they share memory. When either one writes to a page, the CPU raises a page fault, the kernel allocates a fresh physical page, copies the 4 KB (or larger) page into it, updates that process’s mapping to point at the copy with write permission, and resumes the instruction. This is built directly on top of virtual memory, which is what lets two processes see the same virtual addresses backed by different, or shared, physical frames.
Some databases rely on this for snapshots. A process can fork a child, let the child serialize a frozen, point-in-time view of memory to disk, and keep serving writes in the parent. Pages the parent changes get copied; everything else stays shared.
Filesystems and snapshots
Copy-on-write filesystems never overwrite a block in place. When a file changes, the new data is written to a fresh block, and the metadata that points at it is updated, which itself is written to a fresh location, all the way up to the root. The old blocks remain valid until nothing references them.
That design gives you two features nearly for free:
- Cheap snapshots. A snapshot is just a saved pointer to an old root. Creating one is instant and initially consumes no extra space; it only grows as the live filesystem diverges.
- Crash consistency. Because old data is never half-overwritten, a crash mid-write leaves either the old tree or the new tree, not a corrupted mix.
The same “never modify in place” principle appears in databases. Copy-on-write B-trees write modified pages to new locations and swap the root atomically, which lets readers keep a consistent view while writers proceed. It’s a close cousin of MVCC, where multiple row versions coexist so readers and writers don’t block each other.
Containers and image layers
Container images are stacks of read-only layers. When a container starts, the runtime adds a thin writable layer on top. If the container modifies a file that lives in a lower layer, the file is first copied up into the writable layer and then changed there. Every container started from the same image shares the read-only layers on disk, which is why you can run dozens of containers from one image without dozens of copies. The details of how those layers are built and cached are covered in Docker image layers and caching.
Copy-on-write in programming languages
Many languages and libraries use CoW for value types that would be expensive to copy eagerly. Assigning a large array to a new variable may only increment a reference count; the buffer is duplicated the first time one of the variables mutates it while the count is above one. Rust exposes the pattern directly through its Cow type, which holds either a borrowed reference or an owned value and clones only when you ask for a mutable version.
A related technique is structural sharing in persistent data structures. Instead of copying a whole tree on update, you copy only the path from the root to the changed node and reuse every other subtree. Immutable collection libraries, editor buffers built on a rope, and state management in UI frameworks all lean on this. JavaScript’s non-mutating array methods give you immutability semantics, though they produce full copies rather than shared structure.
Eager copy vs copy-on-write
| Eager copy | Copy-on-write | |
|---|---|---|
| Cost at copy time | Proportional to data size | Near constant |
| Cost at first write | None | Copy of the affected unit |
| Memory when copies are unmodified | Full duplicate | Shared |
| Write latency | Predictable | Occasional spikes on first write |
| Implementation complexity | Simple | Needs sharing detection and bookkeeping |
The trade-offs
Copy-on-write isn’t free. It moves cost around, and sometimes it moves cost to a worse place.
- Write spikes. The first write to each shared unit triggers a copy. In a forked process under heavy write load, nearly every page may end up copied, and memory use can approach double the original in the worst case.
- Granularity matters. If the copy unit is large, a one-byte write copies the whole unit. This is why large memory pages can make fork-based snapshots more expensive: a tiny write duplicates a much bigger page.
- Fragmentation. CoW filesystems and B-trees write new data to new locations, so files that are frequently updated in place can scatter across the disk over time.
- Hidden costs in tight loops. A language-level CoW type that silently copies on mutation can surprise you if two references are alive inside a hot loop. The copy happens, just not where you expected it.
- Reference counting overhead. Detecting “is this shared?” usually means maintaining counts, which costs something on every copy and drop, and atomics if the data crosses threads.
When copy-on-write is the right tool
Reach for it when copies are frequent, writes are rare or localized, and you want snapshots or isolation without paying for full duplication. Avoid relying on it when nearly every copy will be fully modified anyway, since you’ll pay the bookkeeping and still copy everything. Related memory techniques, such as memory-mapped I/O, often combine with CoW mappings so that a process can read a file through shared pages and get private copies only of the pages it changes.
The takeaway
Copy-on-write makes copying cheap by sharing data until the first modification, then duplicating only the part being changed. It’s the reason fork() is fast, filesystem snapshots are instant, and containers share image layers. The cost doesn’t disappear; it shifts to the first write of each shared unit. Understand the copy granularity and your write patterns, and CoW becomes one of the most useful tricks in systems design.
Keep reading
The Lycoris Team · · 5 min read LRU vs LFU: Cache Eviction Policies Compared
LRU evicts whatever hasn't been used in the longest time; LFU evicts whatever has been used the fewest times. How each policy behaves and when to pick it.
The Lycoris Team · · 5 min read TCP Congestion Control Explained: Slow Start to BBR
TCP congestion control decides how fast a sender transmits without overloading the network. Slow start, AIMD, CUBIC, and BBR explained.
The Lycoris Team · · 4 min read The Rope Data Structure: Editing Huge Strings Efficiently
A rope is a binary tree of string chunks that makes inserting, deleting, and slicing huge strings fast, which is why text editors use it over plain arrays.