At a glance
Where the flow bends
- 02Architecture
Architects start from the math inside neural networks and from how fast memory can feed it, and they decide early how to split the design across several chips in one package.
Architecture work centers on the matrix unit (often a systolic array), on-chip SRAM size, HBM bandwidth, and how to partition the design across dies. Performance models of real workloads drive these choices before any RTL exists.
Workload models set the balance of MACs, SRAM, and HBM stacks. Chiplet partitioning fixes die-to-die bandwidth, latency, and protocol (UCIe or proprietary) on day one, and those choices are expensive to reverse once the package and interposer are committed.
- 04Verification
Simulating a chip this big on ordinary computers is far too slow, so teams run the design on racks of reprogrammable hardware and test real AI software on it months before the chip exists.
RTL simulation cannot boot a software stack or run full models, so teams add hardware emulation and FPGA prototyping and verify the compiler and runtime against the RTL.
Emulation and prototyping capacity becomes a schedule resource. Multi-die systems also need package-level verification of die-to-die links, and performance correlation against the architecture model sits next to functional coverage as a signoff criterion.
- 06Design for test
With this much silicon, some parts will have defects. Designers add spare parts plus tests that find bad pieces and switch the spares in.
Large dies add redundant cores, spare interconnect lanes, and repairable SRAM. Multi-die packages also need each die tested before assembly so a bad die does not ruin the whole package.
Known-good-die test, die-to-die lane repair, memory BISR, and harvesting of partly defective dies drive DFT. A defect found after assembly scraps the interposer and the attached HBM stacks too.
- 07Floorplanning
The chip outline is as big as the factory can print, and its edges are crowded with connections to memory and to neighboring chips.
Dies sit near the reticle limit, and the floorplan is driven by shoreline: HBM PHYs and die-to-die interfaces must sit on particular edges, lined up with the interposer wiring.
Bump maps and PHY placement are co-designed with the package. Die-to-die bandwidth per millimeter of die edge bounds how much shoreline each interface consumes, so the I/O ring often sets the die outline.
- 08Power planning
These chips draw so much power that delivering it without the voltage sagging, and removing the heat, are among the hardest problems.
Very high current at low supply voltage makes IR drop and electromigration first-order concerns. Dense power grids eat routing tracks, and backside power delivery moves the grid underneath the transistors.
Dynamic IR drop from synchronized MAC activity, package PDN resonance, and thermal density dominate. Backside power frees front-side tracks for signals but changes the heat path and debug access.
- 12Signoff
Final checks cover the whole package, because several chips and memory stacks must work as one.
Signoff extends past the die: timing, power, and signal integrity across die-to-die links and the interposer, plus package thermal analysis.
Multi-die STA with interface budgets, package-level SI/PI, thermal-aware IR and EM analysis, and physical verification across interposer and hybrid-bonded interfaces.
- 13GDS & tapeout
These chips use the newest and most expensive manufacturing, so a mistake costs a great deal of money and months of time.
Leading-edge mask sets are expensive and the market window is short. Teams tape out on tight schedules and reuse chiplets across products to spread the cost.
Mask NRE and wafer cost at leading nodes push teams to put I/O or base dies on older nodes and reuse them. Tapeout also starts a package assembly flow with its own lead time and test steps.
Modern AI runs mostly on one kind of math: multiplying large grids of numbers. An AI accelerator is a chip built to do that one job, many thousands of times in parallel.
Think of a kitchen with thousands of cooks. Hiring more cooks is easy. Getting ingredients to all of them fast enough is the hard part. AI chip designers spend most of their effort on that delivery problem: memory, wiring, power, and heat.
These chips are often as large as a factory can print in one piece, a ceiling called the . To go further, designers place several chips and towers of stacked memory side by side inside one package so they work as a single unit.
Big cloud companies now design their own AI chips. Google started its TPU project after estimating that people using voice search for three minutes a day would force it to double its datacenters if it kept using ordinary processors.1
Neural-network training and inference are dominated by matrix multiplication, so accelerators spend their area on dense arrays of units and on SRAM to keep them fed. Google’s first TPU had a 256 × 256 array of 8-bit MACs (65,536 in total) running at 700 MHz, with 28 MiB of on-chip memory.1
Four limits shape every design in this market:
- Die size. A single die cannot exceed the scanner’s exposure field, roughly 800 mm². NVIDIA’s H100 is 814 mm².2
- Memory bandwidth. Off-chip DRAM is usually , stacked next to the processor.
- Power and heat. Delivering current without voltage sag and removing the heat both get harder each generation.
- Yield and cost. Bigger dies yield worse, and leading-edge masks are expensive.
Accelerator design is a balance between MAC throughput, on-chip SRAM, off-chip bandwidth, and the power budget, set under a hard area ceiling. Yield is the hidden fifth constraint. In a textbook defect model with fixed defect density, quadrupling die area from 60 to 240 mm² raises cost per good die by 7.28× because there are fewer candidate dies per wafer and a lower fraction of them work.3 Near the reticle limit that curve is steep.
The industry’s answers are architectural and physical at once: split the system into on a or a 3D stack, put HBM beside the compute die, add redundancy so defects cost a small fraction of the die, and move power delivery to the wafer backside. Each answer adds a flow step that a single-die SoC never needed: package co-design, known-good-die test, die-to-die timing, and package-level thermal signoff.
- Size limits. One chip can only be so big, so big systems are built from several chips packed tightly together.
- Memory hunger. AI models are huge. Memory is stacked like a tower of pancakes and placed right next to the processor.
- Power and heat. These chips draw large currents at very low voltage, and all of that energy leaves as heat through a small area.
- Speed to market. AI changes fast. A chip that arrives a year late may be built for yesterday’s models.
Memory. HBM stacks DRAM dies connected by through-silicon vias and connects them to the processor through a substrate such as a silicon interposer.5 The JEDEC HBM4 standard runs up to 8 Gb/s across a 2,048-bit interface, for up to 2 TB/s per stack.6 Several stacks surround the compute die, so their PHYs claim most of the die edge.
Packaging. 2.5D packaging places dies side by side on an interposer. 3D packaging stacks them. Solder microbumps have pitches in the tens of micrometers. joins copper pads directly, and research has shown pitches down to 400 nm.7 The standard defines die-to-die links for both cases. Its advanced-package profile uses 25–55 µm bump pitch and reaches up to 2 mm, and the consortium pitches chiplets as a way to build SoCs larger than the reticle.8
Time to market. Google designed, verified, built, and deployed TPU v1 in 15 months.1
Power delivery. Front-side power competes with signals for metal. imec reports that power interconnect takes at least 20% of routing resources, and that with buried power rails cut by 7× in its simulations.9 Intel’s PowerVia test chip showed over 6% frequency gain and 30% less power loss. Intel also had to write new thermal design rules and invent new debug methods, because the transistors now sit between two interconnect stacks.10 At the extreme, a wafer-scale part cannot be fed from its edges at all. Cerebras delivers power vertically through more than 300 voltage regulator modules.2
Design for yield. Granularity of redundancy decides how much a defect costs. In the WSE-3 each core is about 0.05 mm², against about 6 mm² for an H100 streaming multiprocessor, and a reconfigurable fabric routes around bad cores.2 Smaller repair units mean each defect costs less silicon.
Tapeout cost. Treat headline figures with care. A widely quoted 2018 estimate put a 5 nm design at $542.2 million. Semiconductor Engineering argues such numbers appear inflated and fall as a node matures.11 Masks, IP, verification, and software are all real costs, and they push teams toward reusable chiplets and fewer respins.
- Plan with real AI programs. Before anything is drawn, engineers model how real neural networks would run on the proposed chip.
- Build from repeated tiles. Most of the chip is the same small block copied thousands of times, which makes it easier to design and to repair.
- Test on stand-in hardware. The design runs on racks of reprogrammable chips so software teams can start early.
- Design the package with the chip. Memory towers, connections, power, and cooling are planned at the same time as the chip itself.
- Test each piece before assembly. One bad chip can ruin an expensive package, so every chip is checked first.
Architecture and RTL. Teams build cycle-level performance models and run real networks through them. RTL tends to be generated: one processing element is written once and instantiated into an array, often a that passes operands between neighbors so each value fetched from SRAM is reused many times.1
Verification. Software simulation is too slow to run a compiler, a runtime, and a full model. Teams add and FPGA prototyping. The open-source FireSim project, for example, ran cycle-exact RTL of a 1,024-node cluster on cloud FPGAs at a 3.4 MHz simulated clock, under 1,000× slower than real time.12
Physical design. Implementation is hierarchical: one tile is placed, routed, and closed, then replicated. The floorplan is set by HBM and die-to-die PHYs on the die edge, and power planning is a first-class task.
DFT and tapeout. Every die going into a multi-die package must be a . Redundant lanes, cores, and memory rows are added so test can repair defects.
Macro-dominated floorplans. In TPU v1 the 24 MiB unified buffer took almost a third of the die and the matrix unit a quarter, while control was just 2%. The buffer size was chosen partly to match the matrix unit’s pitch.1 Expect floorplans where SRAM and datapath placement is decided by hand and the tools mostly fill in around them.
Die-to-die closure. Interfaces get their own timing budgets, SI/PI analysis, and lane repair. UCIe’s advanced-package profile includes spare lanes and targets 165–1,317 GB/s per millimeter of shoreline, depending on data rate.8 That number bounds how much die edge each link consumes.
Schedule versus PPA. Time-to-market pressure leaves performance on the table. The TPU v1 authors estimate that more aggressive logic synthesis and block design could have raised the clock by 50% had they had more than 15 months.1 Teams choose which PPA to give up to hit the market window.
Package-level signoff. Thermal, warpage, and power integrity are analyzed for the full stack of die, interposer, HBM, and substrate. With backside power, failure analysis and debug flows must also change.10
Around 2013, Google realized that speech recognition and other neural networks would soon need far more computing than its datacenters could supply with ordinary processors. It started a crash project to build its own chip, the Tensor Processing Unit.
The team kept the design simple to move fast. The TPU plugged into existing servers like an add-in card, and the main computer sent it instructions. Most of the chip was one giant grid of small multipliers plus a large block of memory. Control logic was a tiny fraction.
It went from start to running in Google’s datacenters in 15 months, and it ran Google’s AI workloads roughly 15 to 30 times faster than the processors and graphics chips of the time.1
The TPU v1 paper documents a design built for speed of delivery as much as for speed of computation.1
- Coprocessor on PCIe. To reduce the chance of delaying deployment, it sat on the PCIe bus. The host sent instructions, so the TPU never fetched its own.
- One big matrix unit. 256 × 256 8-bit MACs, 92 TOPS peak, fed by a 24 MiB unified buffer.
- Systolic dataflow. Reading a large SRAM costs more energy than arithmetic, so data flows through the array and is reused instead of being re-read.
- Result. About 15–30× faster than contemporary GPUs and CPUs on Google’s inference workloads, with 30–80× better TOPS per watt.
Two lessons from the paper matter for practitioners. First, the chip was memory-bound on most workloads. The authors estimate that GDDR5 memory in place of the DDR3 weight memory it shipped with would have tripled achieved TOPS and raised TOPS per watt to nearly 70× the GPU.1 Memory bandwidth set delivered performance, and that is why later accelerators moved to HBM.
Second, risk reduction shaped the architecture. A coprocessor with host-issued CISC instructions, a single large matrix unit, and minimal control logic kept verification tractable on a 15-month schedule. The authors note that a longer schedule could have bought a higher clock.1
Later generations moved toward system scale. TPU v4 adds SparseCores, dataflow processors that speed up embedding-heavy models 5–7× while using about 5% of die area and power, and it reaches 2.1× the performance and 2.7× the performance per watt of TPU v3.4 The design target shifted from one chip to a 4,096-chip machine.
Sources
- In-Datacenter Performance Analysis of a Tensor Processing UnitTPU v1: 65,536 MACs, 700 MHz, 28 MiB on-chip memory, 15-month schedule, systolic matrix unit, die floorplan.
- A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial IntelligenceReticle limit near 800 mm², H100 die size, WSE-3 size, core-level redundancy, vertical power delivery.
- IC Manufacturing, Cost, Power, and Dependability (COE 501 lecture slides)Die yield and die cost formulas with a worked example of cost versus die area.
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings4,096-chip pods, optical circuit switches, SparseCores, 2.1× performance over TPU v3.
- High Bandwidth MemoryHBM is stacked DRAM joined by TSVs and connected to the processor through a substrate such as a silicon interposer.
- JEDEC and Industry Leaders Collaborate to Release JESD270-4 HBM4 Standard, Advancing Bandwidth, Efficiency, and Capacity for AI and HPCHBM4: up to 8 Gb/s across a 2,048-bit interface, up to 2 TB/s per stack.
- Hybrid Bonding Plays Starring Role in 3D ChipsHybrid bonding vs microbumps; pitch and connection density; use in stacked cache and HBM.
- The UCIe 1.1 Specification: Future Applications of ChipletsChiplets exceed the reticle limit; standard vs advanced package bump pitch, reach, shoreline bandwidth, spare lanes.
- Backside power deliveryPower interconnect takes at least 20% of routing resources; backside delivery with buried rails cut IR drop 7× in simulation.
- Intel Is All-In on Backside Power DeliveryPowerVia test chip: over 6% frequency gain, 30% less power loss; new thermal rules and debug methods.
- What Will That Chip Cost?Headline advanced-node design-cost estimates vary widely and may be inflated.
- FireSim: FPGA-Accelerated Cycle-Exact Scale-Out System Simulation in the Public CloudFPGA-accelerated RTL simulation of a 1,024-node cluster at a 3.4 MHz simulated clock.