Chisato · · 4 min read Wafer-Scale Integration Explained
Wafer-scale integration builds one giant chip from an entire silicon wafer instead of cutting it into dies. How it works and its tradeoffs.
Topic
104 posts tagged “Hardware”.
Chisato · · 4 min read Wafer-scale integration builds one giant chip from an entire silicon wafer instead of cutting it into dies. How it works and its tradeoffs.
Chisato · · 4 min read HBM stacks DRAM dies vertically and connects them through a wide interface, trading capacity per chip for far higher bandwidth than standard DRAM.
Chisato · · 4 min read Thermal interface material fills microscopic gaps between a chip and its heatsink so heat can actually transfer to the cooler.
Chisato · · 4 min read Dennard scaling held that shrinking transistors kept power density constant, letting clock speeds rise for free. Its breakdown reshaped chip design.
Chisato · · 4 min read Clock speed measures cycles per second; IPC measures work done per cycle. Real CPU performance is the product of both, not either one alone.
The Lycoris Team · · 6 min read Apple's September 9 'Surprise and shine' event is set to unveil the iPhone 18 Pro and a first foldable iPhone — the first launch under new CEO John Ternus.
Chisato · · 4 min read Wi-Fi 6E opened the 6 GHz band; Wi-Fi 7 adds Multi-Link Operation and wider channels on top of it. How the two standards actually differ.
Chisato · · 4 min read UMA gives every CPU core equal-latency memory access; NUMA gives each core faster access to its local memory bank. How the two architectures differ.
Chisato · · 5 min read Overclocking runs a CPU, GPU, or memory beyond its rated clock speed for more performance, trading power, heat, and stability margin to get it.
The Lycoris Team · · 5 min read Floating-point numbers can't represent most decimals exactly in binary — that's why 0.1 + 0.2 gives 0.30000000000000004, and what IEEE 754 actually stores.
Chisato · · 5 min read OpenAI reportedly bought tens of thousands of Mac minis and Studios to train computer-use agents, while Anthropic rents Apple silicon via AWS.
Chisato · · 5 min read Clock gating and power gating both cut chip energy use by disabling idle circuit blocks, but one stops the clock while the other cuts power entirely.
Chisato · · 4 min read A side-channel attack recovers secrets from a system's physical behavior — timing, power draw, cache access — rather than breaking its algorithm directly.
Chisato · · 5 min read NVMe SSDs talk directly over PCIe for massive parallel throughput; SATA SSDs are bottlenecked by an interface built for spinning disks. Here's why it matters.
Chisato · · 4 min read Yield is the share of chips on a wafer that work. It's the single biggest driver of chip cost, and why new process nodes start out expensive.
Chisato · · 5 min read 2.5D packaging places dies side by side on an interposer; 3D packaging stacks dies vertically through silicon vias. How advanced packaging works.
Chisato · · 4 min read SerDes converts parallel data to a serial stream for transmission and back, letting chips move data over fewer, faster wires with less overhead.
Chisato · · 4 min read Chip fabrication turns a silicon wafer into a working processor through photolithography, etching, doping, and packaging — here's the full process.
Chisato · · 6 min read Apple refreshed the Mac Studio with M5 Max and M5 Ultra and the Mac mini with M6, adding up to 512GB of unified memory and big on-device AI gains. What to know.
Chisato · · 4 min read An ASIC is fixed-function silicon built for one job at massive scale; an FPGA is reconfigurable hardware you can rewire after it ships. How to choose.
Chisato · · 4 min read Silicon photonics moves data with light instead of electrical signals on standard chip manufacturing lines. How it works and why AI data centers need it.
Chisato · · 4 min read False sharing happens when threads on different cores write to unrelated variables that share one cache line, silently wrecking performance.
Chisato · · 4 min read Binning sorts identical die designs by how well each one actually performs after manufacturing, turning natural process variation into a full product lineup.
Chisato · · 4 min read Instruction pipelining overlaps a CPU's fetch, decode, and execute stages so multiple instructions are in flight at once. How it works.
Chisato · · 4 min read DRAM stores each bit as a charge in a capacitor that needs constant refreshing; SRAM stores each bit in a transistor circuit that holds its state.
Chisato · · 4 min read A TPU is a chip built specifically for the matrix math behind neural networks, using a systolic array instead of a general-purpose GPU pipeline.
Chisato · · 4 min read A hardware security module is a dedicated device that generates, stores, and uses cryptographic keys so private keys never leave secure hardware.
Chisato · · 4 min read HBM stacks DRAM vertically next to the chip for extreme bandwidth per watt; GDDR spreads fast chips around the board. How the two approaches differ.
Chisato · · 4 min read CUDA cores handle general parallel math on an NVIDIA GPU; Tensor cores are specialized units built for the matrix multiplies AI models run constantly.
Chisato · · 4 min read UCIe is an open standard for wiring chiplets from different vendors into one package, defining the physical layer, protocols, and packaging it needs.
Chisato · · 4 min read Memory-mapped I/O maps device registers into the CPU's address space, letting hardware be read and written with ordinary load and store instructions.
Chisato · · 5 min read Cache associativity determines where a memory block can be placed in a CPU cache. How direct-mapped, fully associative, and set-associative caches compare.
Chisato · · 5 min read A memory controller is the circuit that manages every read and write between the CPU and main memory, setting the ceiling on bandwidth and latency.
Chisato · · 4 min read DDR5 raises transfer rates, splits each module into two independent sub-channels, and moves voltage regulation onto the module itself. Here's what that buys you.
Chisato · · 4 min read Von Neumann shares one memory for code and data; Harvard architecture splits them. How the two computer architectures differ, and where each is used.
Kurumi · · 4 min read Lenovo's Q1 FY2027 revenue rose 43% to $26.9B and net income jumped 176% as AI PCs and servers drove a record quarter, sending shares up about 19%.
The Lycoris Team · · 4 min read Amdahl's Law says a program's speedup from parallelizing is capped by the fraction that must stay sequential — no matter how many cores you add.
Chisato · · 5 min read Memory channel configuration multiplies bandwidth between RAM and the CPU. How single-, dual-, and quad-channel setups differ, and why slot order matters.
Chisato · · 4 min read The northbridge and southbridge were the two chips that routed data between a CPU, memory, and peripherals before modern SoCs absorbed their jobs.
The Lycoris Team · · 4 min read Write amplification is when a system writes more data physically than the logical write requested, wearing out storage faster and hurting throughput.
Chisato · · 4 min read Virtual memory gives every process its own private address space, mapped to physical RAM by the OS and CPU — enabling isolation, swapping, and overcommit.
Chisato · · 4 min read NVLink and PCIe both move data to and from GPUs, but NVLink trades PCIe's universality for far higher bandwidth between GPUs specifically.
Chisato · · 5 min read Simultaneous multithreading lets one physical CPU core run two instruction streams at once, filling idle execution units to raise throughput.
Chisato · · 4 min read A TLB is a small CPU cache that stores recent virtual-to-physical address translations, avoiding a slow page-table walk on every memory access.
Kurumi · · 5 min read SanDisk's fiscal Q4 2026 revenue jumped 372% to $8.97B with 84.6% gross margin as datacenter NAND booms. Here are the numbers, guidance, and what to watch.
Chisato · · 5 min read Cache coherence keeps each CPU core's private cache consistent with the others. The MESI protocol is the classic mechanism that makes it work.
Chisato · · 5 min read Memory interleaving spreads consecutive addresses across multiple memory banks so the system can access them in parallel instead of one at a time.
Chisato · · 5 min read Endianness decides whether a multi-byte number's most or least significant byte is stored first in memory. Why it matters and how to spot it.
Chisato · · 4 min read Speculative execution lets a CPU guess ahead and run instructions before it knows they're needed, buying speed at the cost of the timing side channels behind Spectre and Meltdown.
Chisato · · 5 min read Confidential computing uses hardware-isolated enclaves to keep data encrypted even while it's being processed, not just at rest or in transit.
Chisato · · 4 min read A systolic array is a grid of processing elements that pass data to their neighbors in rhythm, built to accelerate matrix multiplication in AI chips like TPUs.
Chisato · · 4 min read SIMD lets a CPU apply one instruction to multiple data points at once. How vectorization works, why compilers auto-vectorize loops, and its limits.
Chisato · · 4 min read DMA lets peripherals move data to and from memory without the CPU copying every byte, freeing the processor to do other work during transfers.
Chisato · · 4 min read Branch prediction guesses which way an if-statement will go before the CPU knows, and out-of-order execution reorders instructions to keep pipelines full.
Chisato · · 4 min read Thermal throttling automatically reduces a chip's clock speed when it gets too hot, trading performance for safety. How it works and how to spot it.
Chisato · · 4 min read TDP is the amount of heat a cooling system must dissipate for a chip, not a hard limit on its power draw. Why TDP and actual power draw often diverge.
Chisato · · 4 min read An ASIC is a chip custom-built for one task, trading flexibility for speed and power efficiency. How ASICs differ from GPUs and FPGAs, and when to use one.
Chisato · · 4 min read China has begun limited mass production of home-grown immersion DUV lithography machines, with first units bound for SMIC, Hua Hong and CXMT. What it changes.
Chisato · · 4 min read UEFI is the firmware that initializes hardware and boots the OS on modern computers, replacing BIOS with faster boot times, larger disk support, and Secure Boot.
Kurumi · · 6 min read SK Hynix reports Q2 2026 earnings July 29 with record profit expected on HBM demand — the memory supercycle's first real test this season. What to watch.
Chisato · · 4 min read CXL is an interconnect standard that lets CPUs, GPUs, and memory devices share coherent memory over PCIe, enabling memory pooling and expansion.
Chisato · · 4 min read NUMA gives each CPU its own local memory bank, so access speed depends on which processor is asking. How NUMA nodes and remote access latency work.
Chisato · · 5 min read RAID combines multiple drives into one logical unit for redundancy, speed, or both. How RAID 0, 1, 5, 6, and 10 trade off capacity, speed, and safety.
Chisato · · 4 min read PCIe (PCI Express) is the high-speed serial bus connecting GPUs, SSDs, and network cards to a CPU. How lanes, generations, and bandwidth work.
Chisato · · 4 min read SSDs store data in flash memory chips with no moving parts; HDDs use spinning magnetic platters. How that difference plays out in speed, cost, and durability.
Chisato · · 4 min read Bandwidth measures how much data memory moves per second; latency measures how long one access takes. Why chips need both, not just one.
Chisato · · 7 min read Nvidia detailed its Vera CPU — 88 custom Olympus cores, 1.2 TB/s memory, and SPEC CPU 2026 scores that edge AMD's Epyc dual-socket flagship.
The Lycoris Team · · 4 min read A quantum computer uses qubits in superposition and entanglement to explore many possible states at once, rather than one bit value at a time.
Chisato · · 4 min read ECC memory detects and corrects single-bit errors in RAM automatically, using extra parity bits — critical for servers where silent corruption is costly.
Chisato · · 4 min read DDR and GDDR are both DRAM, but optimized for opposite goals: DDR minimizes latency for CPUs, GDDR maximizes bandwidth for GPUs. Here's how they diverge.
Chisato · · 4 min read A DPU is a specialized chip that offloads networking, storage, and security tasks from the CPU. How data processing units fit alongside CPUs and GPUs.
Chisato · · 4 min read A process node like '5nm' or '3nm' names a chipmaker's manufacturing generation, not a literal measurement anymore. Here's what the number means.
Chisato · · 6 min read Japan and Nvidia launched Noetra, a 140MW Vera Rubin AI factory with 27,500 GPUs, to build sovereign robotics foundation models under the FRONTia plan.
Kurumi · · 5 min read Intel is the first to mass-produce logic chips on ASML's High-NA EUV, dual-qualifying select 18A Panther Lake layers. Here's what shipped and why it matters.
Chisato · · 4 min read A TPM is a dedicated chip that generates and stores cryptographic keys in hardware, isolated from the operating system. Here's what it actually does.
Chisato · · 5 min read RISC and CISC are two philosophies for CPU instruction sets — simple fixed-length instructions versus fewer, complex ones. How they differ and why.
Chisato · · 4 min read A system on chip packs a CPU, GPU, memory controller, and other components onto one die — the design behind phones, laptops, and most modern chips.
Chisato · · 5 min read ARM and x86 are the two dominant CPU instruction set architectures — how RISC vs CISC design differs and why it affects power and performance.
Chisato · · 4 min read EUV lithography uses 13.5nm-wavelength light to etch the finest features on modern chips. How it works and why it's a chokepoint in chip manufacturing.
Chisato · · 4 min read SRAM is fast, expensive, six-transistor memory used for CPU caches; DRAM is slower, cheaper, one-transistor memory used for main system memory.
Chisato · · 5 min read L1, L2, and L3 CPU caches sit between the processor and main memory, trading capacity for speed at each level. How the hierarchy actually works.
Chisato · · 4 min read An FPGA is a chip whose logic circuits can be reconfigured after manufacturing, sitting between fixed-function ASICs and general-purpose CPUs in flexibility.
Chisato · · 5 min read A chiplet is a small, self-contained die that's packaged together with others to form one chip. How chiplets work and why the industry moved to them.
Chisato · · 4 min read Moore's Law is the observation that transistor density on a chip roughly doubles every couple of years. Why it drove decades of gains, and why it's slowing.
Chisato · · 4 min read CPUs excel at sequential logic, GPUs at parallel math, and TPUs at the specific matrix operations behind neural networks. Here's how they compare.
Chisato · · 5 min read Windows on Arm was famous for broken apps. In 2026, most software runs natively and Prism emulation covers the rest. What works, what doesn't, how to check.
Kurumi · · 4 min read Humanoid robots are arriving with $20,000 price tags and rental plans. What a robot worker really costs to build and run — and when it beats a human wage.
Kurumi · · 4 min read AI ambition is measured in gigawatts. What one actually costs to build and power for a year — a back-of-the-envelope teardown of tech's priciest machine.
Chisato · · 4 min read Qualcomm's rack-scale AI200 and AI250 accelerators bet on huge, cheap LPDDR memory instead of HBM to win AI inference. How the design works and who's buying.
Chisato · · 4 min read Qualcomm's second-generation laptop chip jumps to 18 cores, 5 GHz, and an 80 TOPS NPU. What changed from the Snapdragon X Elite — and whether it matters.
Kurumi · · 3 min read Hyperscalers are pouring record sums into AI data centers, chips, and power. What's driving the capex boom, who profits, and the risk if demand stalls.
Chisato · · 4 min read An NPU is a processor built for one job: running AI models fast at very low power. What TOPS numbers actually mean and why every new laptop ships with one.
Kurumi · · 2 min read Micron has whipsawed in 2026 — record highs on AI memory demand, sharp drops on rate fears, AI-capex doubts, and a Google compression breakthrough. What's moving it.
Chisato · · 6 min read Google TurboQuant compresses AI model memory ~6x with no accuracy loss or retraining, and speeds attention up to 8x. How it works and what it means for HBM.
Kurumi · · 3 min read Micron and Anthropic signed a four-pillar agreement — memory co-design, a multi-year supply deal, Claude adoption, and a Series H investment. Here's what it means.
Kurumi · · 6 min read The AI memory supercycle, explained: why HBM demand outran supply, how DRAM pricing turned, what could end the boom, and what it means for chip stocks.
The Lycoris Team · · 3 min read China unveiled a $295 billion, five-year national AI infrastructure plan — one of the largest state AI commitments ever. Here's the scale and the strategic stakes.
Chisato · · 3 min read A GPU packs thousands of small cores built for parallel arithmetic. Originally for graphics, it's now the engine behind training and running AI models.
Chisato · · 3 min read High-Bandwidth Memory stacks DRAM dies vertically beside the processor, delivering far more bandwidth than DDR5 or GDDR — and AI hardware depends on it.
Kurumi · · 2 min read Samsung, SK Hynix, and Micron are racing to mass-produce HBM4 and win NVIDIA's orders. Inside the next phase of the memory supercycle — and who's ahead.
Chisato · · 2 min read AMD's Instinct MI400 brings 432GB of HBM4 and a full-rack Helios system to challenge NVIDIA in 2026. Here's what the MI455X packs and why it matters.
Chisato · · 2 min read OpenAI and NVIDIA unveiled a landmark deal: at least 10 gigawatts of NVIDIA systems and up to $100 billion in investment, starting on the Vera Rubin platform.
Chisato · · 2 min read NVIDIA unveiled Vera Rubin — a platform of six new chips designed to work as a single AI supercomputer — while its Vera CPU enters full production. What's coming.
The Lycoris Team · · 6 min read RISC-V is a free, open instruction set architecture anyone can implement without royalties. How it works, why it matters, and where it's already winning.