Stage 9 of 13 · Back end

Placement

Software finds a spot for each of hundreds of thousands of building blocks so the connected ones sit close together.

Standard cells are placed into rows to keep wires short while meeting timing and avoiding congestion. Global placement, legalization and detailed placement run in sequence.

Analytical and electrostatic global placement, timing- and congestion-driven optimization, legalization, detailed placement, and in-placement buffering, sizing and Vt swapping.

Builds The cell layer

By this point the chip has a floor plan. The big blocks, like memories, are in place, and a grid of power wires covers the whole area. What’s left is the bulk of the logic: hundreds of thousands or millions of tiny building blocks called . Placement decides exactly where each one goes.

Every connection between two cells becomes a real wire. Long wires slow signals down and burn power, so the software tries to keep connected cells near each other. It also has to avoid jamming too many cells into one spot, because then the wires won’t fit. A good placement makes every later step easier. A bad one can’t be rescued by routing.

Placement takes a synthesized, scan-inserted gate-level netlist and a floorplan, and assigns every a legal location and orientation in the core’s . It sits after floorplanning and power planning and before clock tree synthesis and routing.

The placer balances several goals at once: short total wirelength, met timing on critical paths, a layout the router can complete without , an even cell density, and low switching power on busy nets. Placement is also where the netlist starts changing for physical reasons. Buffers get inserted, gates get resized, and scan chains get reordered, all based on where cells actually landed.

Every commercial and open flow has a placement engine: Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler, and the open-source OpenROAD project with its global placer (gpl, derived from RePlAce) and detailed placer (dpl, derived from OpenDP).

Placement is the stage where wirelength, and with it most of the remaining timing and routability, gets decided. Clock tree synthesis, routing and signoff mostly inherit the placement’s quality. The flow runs a global placer to minimize a smooth wirelength proxy under a density constraint, legalizes onto the site grid, refines with detailed placement, and interleaves physical synthesis: buffering, sizing, VT swaps and restructuring on estimated parasitics.

Commercial tools (Cadence Innovus, Synopsys IC Compiler II and Fusion Compiler) run these steps inside their own placement-optimization flows. OpenROAD exposes them as separate commands, which makes it a good model for what happens inside. Placement sub-problems are NP-hard, so every engine is a stack of heuristics and continuous relaxations. The academic state of the art for global placement is nonlinear electrostatic placement (ePlace, RePlAce, DREAMPlace), and OpenROAD’s gpl uses it.

Default target density, OpenROAD global_placement
0.70
Default stopping overflow, OpenROAD global_placement
0.10
DREAMPlace global placement speedup on GPU vs. multi-threaded RePlAce
>30×

Those numbers come from the OpenROAD gpl documentation and the DREAMPlace paper.

Placement starts with two things: a parts list that says which cells exist and how they connect, and the floor plan that shows the rows, the big blocks and the reserved areas. It also needs the “spec sheet” for each kind of cell: how big it is and how fast it is.

It finishes with the same parts list, slightly edited, plus an exact spot for every cell. Nothing overlaps, and every cell sits neatly in a row. The engineer also gets reports that say how long the wires are likely to be, whether signals will arrive on time, and where wiring might get crowded.

DirectionWhatFormat
InGate-level netlist after synthesis and DFT insertionVerilog
InFloorplan: die and core area, rows, placed macros, I/O pins, power grid, tap and endcap cells, blockages, regionsDEF or a tool database (OpenDB .odb)
InTechnology and cell geometry: layers, sites, cell sizes, pin shapes, obstructionsLEF (tech LEF and cell LEF)
InCell timing, power and areaLiberty (.lib)
InClocks, I/O delays, timing exceptionsSDC
InScan chain description (which flops may be reordered)DEF SCANCHAINS section
OutPlaced design: every cell PLACED or FIXED with location and orientationDEF or .odb
OutUpdated netlist with new buffers, resized cells and reordered scan chainsVerilog
OutReports: utilization, timing on estimated wires, congestion, legalityText logs and GUI heat maps

LEF defines the , the smallest placement step, and each cell’s width as a multiple of it. DEF ROW statements lay those sites out across the core. Liberty supplies the delay and power tables the placer’s timing engine needs. The output DEF marks each cell PLACED, meaning tools may still move it, or FIXED, meaning automatic tools must leave it alone.

Two inputs are easy to forget. The first is per-layer RC (OpenROAD’s set_wire_rc, or the tech file’s layer parasitics), which drives estimate_parasitics -placement and therefore every timing decision made before routing. The second is the scan DEF: SCANCHAINS separates FLOATING segments, which may be reordered, from ORDERED segments, which may not. Physical-only cells (taps, endcaps, fillers, decaps) carry SOURCE DIST in DEF and never appear in the logical netlist, so a netlist-vs-layout check has to account for them.

Placement software works in three passes, from rough to precise.

  1. Rough layout. The software treats connections like rubber bands pulling connected cells together, and treats crowding like a pressure pushing cells apart. It lets the cells slide around until those forces balance. At this point cells can still overlap a little and don’t sit in proper slots.
  2. Snap into rows. Each cell moves to the nearest free slot in a row, so nothing overlaps. This step is called . The goal is to move everything as little as possible.
  3. Polish. The software tries small swaps between neighbors and keeps the ones that shorten the wires.

To score a placement, the software uses a quick wire-length estimate called : draw the smallest box around all the cells on one connection and add its width and height. Adding that up over every connection gives one number to shrink.

Along the way the tool keeps an eye on two more things. It checks timing, and gives the slowest paths extra pull so their cells end up closer. It also checks crowding, and spreads cells out where too many wires would need to squeeze through. When a signal has to travel far, the tool adds small booster cells called buffers, or swaps in a stronger version of a gate. So the parts list keeps changing a little during placement.

Three phases

finds approximate (x, y) positions for all movable cells. It minimizes total , the sum over nets of bounding-box width plus height. It also keeps the cell area in every bin of a coarse grid under a target density. Modern global placers are analytical: they treat positions as continuous variables and use gradients. In OpenROAD’s gpl, wires pull connected cells together while an pushes cells out of crowded bins. The engine stops when the leftover overlap, called , drops below a target (0.1 by default).

moves each cell to a legal site in a row with no overlaps, minimizing displacement from its global position. then improves locally, for example by swapping cells, reordering a few neighbors in a row, or mirroring a cell so its pins face its connections.

A worked HPWL example

Take a net with pins at (2, 1), (5, 4) and (3, 7) µm. Its bounding box runs from x = 2 to 5 and from y = 1 to 7, so its HPWL is 3 + 6 = 9 µm. Move the third cell to (3, 4) and the box shrinks to y = 1 to 4, so HPWL drops to 3 + 3 = 6 µm. For a two-pin net, HPWL equals the length of the shortest Manhattan route. For nets with more pins the routed wire is at least that long and usually longer, so HPWL is an optimistic but consistent score.

Why rough first, then snap

Choosing a site for each of a million cells directly is an enormous discrete search. Treating positions as continuous numbers lets the placer move all cells at once along a gradient, which scales to millions of cells. Snapping to the grid afterwards costs a little wirelength, and that cost stays small only if global placement left enough free space near every cell.

What it optimizes

  • Wirelength. HPWL is the stand-in for routed length, because it is cheap and correlates with both timing and routability.
  • Timing. runs timing analysis during placement and gives the worst-slack nets higher weight, so the placer shortens them first. In OpenROAD the top 10% of nets by slack get weights up to 5× by default.
  • Routability. A congestion estimate runs inside the loop. OpenROAD uses , which spreads each net’s wire evenly over its bounding box. Cells in congested tiles are temporarily “inflated” so they push neighbors away and make room for wires.
  • Density. The target density caps how full any local area gets. Overall is usually well under 100%, because buffering, gate sizing and detailed placement all need free space spread across the die.
  • Power. Keeping high-activity nets short cuts switching capacitance, which cuts dynamic power.

Optimizing while placing

Once cells have locations, the tool can estimate wire parasitics and run real timing. OpenROAD’s resizer shows the usual moves. repair_design inserts buffers to fix max slew, max capacitance, max fanout and long wires, and resizes gates. repair_timing upsizes gates, swaps pins, clones gates, buffers, and swaps cells to a lower threshold voltage (VT) for speed. Low-VT cells are faster but leak more. Logic restructuring can also rebuild a cone of gates for delay or area. Each change gets re-legalized.

Scan-chain reordering

DFT insertion stitched flip-flops into scan chains before anyone knew where they would land. After placement, re-stitches each chain so consecutive flops are physically close. The classic formulation is a traveling-salesman problem over flop locations.

The objective in practice

Global placement solves min HPWL subject to bin density ρ_b ≤ ρ_t, relaxed into minimizing W + λN. Here W is a smooth wirelength, N is a density penalty, and λ sets the balance between them. HPWL is exact for two-pin nets and a lower bound for multi-pin nets. It ignores detours, layer assignment and via cost. That is why routability-driven modes overlay a congestion estimate. OpenROAD gpl runs at set overflow points, inflates cells in tiles whose RC metric exceeds the target (1.01 by default), and keeps iterating until the metric stops improving. RePlAce introduced the same routability extension, with congestion from a global router or RUDY.

mode in gpl runs a virtual repair_design at overflow checkpoints (64% and 20% by default), computes slack, and scales net weights from the maximum (5) on the worst net down to 1.0 at the 10% point of nets sorted by slack. It keeps resizer changes only below a set overflow. Net weighting is coarse. It cannot tell a net that is critical through one path from one that is critical through thousands, and it only sees timing at a few checkpoints.

Utilization and whitespace

Placement density and are different numbers. Core utilization is total cell area over available area. Target density is the per-bin cap. gpl defaults to 0.7. Whitespace has to be local, because buffers and upsized gates belong next to the nets they fix and cannot wait for a distant buffer island. Reported whitespace in real blocks ranges from about 20% to about 70%, depending on methodology and how much congestion guardband the team wants. Many logic blocks start in the 60–75% utilization range and adjust after the first congestion and timing look. Macro-heavy, pin-dense or high-fanout designs sit lower.

Physical synthesis and scan

In-placement optimization is the same move set as post-route ECO but cheaper, because cells can still move. OpenROAD’s resizer handles buffering for slew, cap, fanout and wire length, and gate sizing to normalize slews. Its repair_timing adds pin swapping, gate cloning, load splitting and VT swaps, with -recover_power to downsize on paths that have positive slack. Restructuring cuts out logic cones and re-maps them through ABC for delay or area.

runs on the DEF SCANCHAINS FLOATING lists. ORDERED segments stay intact. Classic reorderers minimize Manhattan distance between flops as a TSP. Routing-aware variants count the incremental connection to the existing net and respect timing slack at the affected sinks. After reordering, the ATPG patterns must be regenerated from the new chain order.

Advanced-node effects

Below roughly 10 nm, non-overlap is no longer enough for a placement to be legal. Drain-to-drain abutment, minimum implant width and area, and oxide-diffusion jog rules depend on which cells sit side by side. Padding every library cell to make all adjacencies safe costs too much area, so flows add a rule-aware final legalization. Multi-row-height cells (large flops, multiplexers, high-drive gates) must land on rows whose power/ground rails match. They couple rows during legalization and bring edge-spacing and pin-short rules. is the other big one. Neighboring and boundary pins can block each other’s access points, so detailed routing hits DRCs that placement never saw. Mitigations include cell padding, pin-access-aware detailed placement and refinement during routing.

The placer doesn’t get a blank floor. Some areas are marked “no parts here,” such as the gap around a big memory block, so wires can reach it. Some groups of parts are told to stay inside a fenced yard. And a few special helper cells are laid down first, in a regular pattern, like the fixed posts in a parking lot. One kind of helper, a , ties the silicon to power and ground so the chip stays electrically stable. Engineers also scatter a few unused spare parts around, so a late bug fix can be wired in without making new masks for the transistor layers.

Blockages and halos

A removes area from the placer. In DEF, a plain PLACEMENT blockage is hard: no cells at all. + SOFT keeps initial placement out but lets later timing optimization and clock tree synthesis use the area. + PARTIAL caps standard-cell density inside the rectangle at a percentage. A is a keep-out margin attached to a macro that moves with it.

Regions

A ties a group of cells to an area. A GUIDE region is a preference that wirelength or timing can override. A FENCE region is exclusive: its cells must stay inside, and no other cells may enter. Voltage areas and hierarchical sub-blocks commonly use fences.

Pre-placed physical cells

  • connect the n-well and p-substrate to the supply. They prevent latch-up, and fabs set a maximum distance from any transistor to the nearest tap. OpenROAD’s tapcell command inserts them in a checkerboard at a set distance.
  • (boundary cells) sit at both ends of every row and around macros.
  • are unconnected gates spread across the die during placement. A later metal-only engineering change order (ECO) can wire them in.

All of these are FIXED before standard-cell placement starts, so the placer works around them.

DEF encodes the constraint vocabulary directly. BLOCKAGES - PLACEMENT is hard. + SOFT is honored only in initial placement. + PARTIAL maxDensity limits standard-cell area in the rectangle and is ignored by later clock buffers and timing buffers. Component halos use + HALO [SOFT] left bottom right top. REGIONS take TYPE FENCE or GUIDE. Partial blockages (some tools call them density screens) are the standard fix for a bin that keeps overflowing next to a macro corner or in a narrow channel. Cell padding (set_placement_padding, or global_placement -pad_left/-pad_right) is the per-cell version, used to give pin-dense or high-fanout cells room for access.

pitch comes from the latch-up rule: the maximum distance from any active area to a tap. tapcell -distance sets it, and -halo_width_x/-y keep taps and endcaps out of the margin when rows are cut around macros. Taps and are LEF CLASS CORE WELLTAP and CLASS ENDCAP, and they appear in DEF with SOURCE DIST. Because they are FIXED at a regular pitch, they chop rows into segments, and legalization must treat each segment separately. follow the same logic: spread uniformly so that a metal-only ECO always has a gate within reach. A patch built from spares that are too far away becomes a timing and congestion problem of its own.

Below is a tiny chip with 18 cells, six input and output pins, and wires drawn between connected cells. Drag cells around and watch “total wire length” change. Try pulling each colored group into a line between its pins. Then press Congestion map to see where wires pile up, and Run annealer to let the computer try thousands of random swaps. The chart shows the wire length dropping over time. Early on, the computer accepts some bad swaps to escape dead ends. Later it gets pickier, the way metal settles into place as it cools.

The grid is 12 × 8 sites with 18 cells and 6 fixed I/O pins. The readout is total HPWL in µm. Drag cells to beat the starting HPWL by hand, then run the annealer. It makes random moves and swaps, and accepts a worse move with probability exp(−ΔHPWL/T). The chart plots HPWL and temperature over cooling epochs. Watch HPWL bounce around early, while T is high, and settle as T drops. Toggle the congestion map (a RUDY estimate) and notice that the lowest-HPWL arrangement can still have a hot spot where several bounding boxes overlap. Press Scatter for a new random start.

Drag any of the 18 cells on the 12 × 8 site grid and total HPWL updates live. Toggle a RUDY congestion map, or press Run annealer to watch simulated annealing work, with HPWL and temperature plotted per cooling epoch. The annealer is a toy TimberWolf: Metropolis acceptance, geometric cooling, and a move window that shrinks with temperature, the same idea as TimberWolf’s range limiter. Watch the accept-rate readout fall as T drops. If the cooling is too fast, the run freezes in a local minimum. Run it from several scatters and compare final HPWL. The spread you see is run-to-run variance, and placement stability matters in a timing-closure loop. Overlay the RUDY map after annealing. Pure HPWL minimization often packs the cross-linked nets together and builds a demand peak, which is exactly what routability-driven inflation is meant to break up.

Loading simulation…

After a placement run, an engineer looks at four things, in roughly this order.

  1. Did it finish cleanly? Every cell must sit in a legal slot with no overlaps. The tool reports this as a simple pass or fail.
  2. How full is it? says what share of the floor the cells use. Too high and wires won’t fit; too low and the chip is bigger, and more expensive, than it needs to be.
  3. Where is it crowded? A map colors the chip like a weather radar. A few warm spots are normal. A bright red blob next to a memory block means trouble for routing.
  4. Is it fast enough? A timing report estimates whether signals arrive before the clock ticks. Wires don’t exist yet, so the tool guesses their length from the placement.

The script below is an illustrative OpenROAD-style placement step. It reads the floorplanned database, runs global placement, repairs electrical rules on estimated wires, legalizes, refines, and checks timing. Real flows such as OpenROAD-flow-scripts split these across several scripts and set the knobs per design.

Below are an illustrative OpenROAD-style script, the DEF that results, and an annotated log. The log is modeled on OpenROAD output but is not verbatim. Read the log for three things: the overflow curve (HPWL should rise as cells spread, then flatten), the legalization cost (displacement and HPWL delta), and where congestion overflow sits relative to macros and fences.

place.tcl (illustrative)tcl
# place.tcl: illustrative OpenROAD-style placement step, not a complete flow
read_db results/2_floorplan.odb   ;# rows, macros, IO pins, PDN, tap and endcap cells
read_liberty lib/stdcells_typ.lib
read_sdc constraints/top.sdc
set_wire_rc -signal -layer M3
set_wire_rc -clock -layer M5

# 1. Global placement: analytical, timing- and routability-driven
global_placement -timing_driven -routability_driven \
    -density 0.65 -overflow 0.10 -pad_left 1 -pad_right 1

# 2. Fix electrical rules using placement-based wire estimates
estimate_parasitics -placement
repair_design -max_wire_length 400
report_design_area

# 3. Legalize, then refine
set_placement_padding -global -left 1 -right 1
detailed_placement
improve_placement
optimize_mirroring
check_placement -verbose

# 4. Check timing before clock tree synthesis
estimate_parasitics -placement
report_worst_slack -max
report_tns
write_db results/3_place.odb
write_def results/3_place.def
  1. 1L2The floorplan database already holds rows, macros, the power grid and the FIXED tap and endcap cells, so the placer only moves standard cells.
  2. 2L5Per-unit-length RC for estimated wires. Every timing number before routing depends on this guess.
  3. 3L9Timing-driven reweights the worst-slack nets. Routability-driven inflates cells in tiles that RUDY flags as congested.
  4. 4L10Target density 0.65 per bin; stop when overflow reaches 10%. One site of padding each side gives pins room for access.
  5. 5L14Buffers long wires and fixes max slew, capacitance and fanout. New buffers land wherever there is free space, which is why utilization matters.
  6. 6L19Legalization: snap every cell to a site in a row, minimizing displacement.
  7. 7L20Detailed placement: local moves that cut wirelength without breaking legality.
  8. 8L22Fails on overlaps, off-site cells, cells in blockages or outside fences.
  9. 9L26Setup slack on placement-based parasitics. The clock is still ideal here, so a small negative slack is often fixed later in CTS and routing optimization.

Here is the matching part of the output DEF. Each component line gives a cell instance, its library cell, its status and its location in database units (1000 per µm here) with orientation.

3_place.def (excerpt, illustrative)text
VERSION 5.8 ;
DESIGN top ;
UNITS DISTANCE MICRONS 1000 ;
DIEAREA ( 0 0 ) ( 200000 200000 ) ;
ROW ROW_0 CoreSite 10000 10000 N DO 900 BY 1 STEP 200 0 ;
ROW ROW_1 CoreSite 10000 12000 FS DO 900 BY 1 STEP 200 0 ;
# ... 88 more rows ...
REGIONS 1 ;
- alu_fence ( 110000 40000 ) ( 180000 90000 ) + TYPE FENCE ;
END REGIONS
COMPONENTS 7 ;
- u_sram SRAM_1KX32 + FIXED ( 20000 120000 ) N + HALO 5000 5000 5000 5000 ;
- ENDCAP_0 ENDCAP_X1 + SOURCE DIST + FIXED ( 10000 10000 ) N ;
- TAP_0 TAPCELL_X1 + SOURCE DIST + FIXED ( 30000 10000 ) N ;
- U1023 NAND2_X2 + PLACED ( 48200 10000 ) N ;
- U1024 INV_X1 + PLACED ( 48800 12000 ) FS ;
- u_alu/U88 DFF_X1 + REGION alu_fence + PLACED ( 130000 60000 ) FS ;
- spare_17 NOR2_X1 + SOURCE USER + PLACED ( 90000 34000 ) N ;
END COMPONENTS
BLOCKAGES 3 ;
- PLACEMENT RECT ( 10000 100000 ) ( 18000 190000 ) ;
- PLACEMENT + SOFT RECT ( 60000 112000 ) ( 66000 160000 ) ;
- PLACEMENT + PARTIAL 40.0 RECT ( 140000 140000 ) ( 170000 170000 ) ;
END BLOCKAGES
  1. 1L5A row of 900 sites, each 0.2 µm wide, starting at (10, 10) µm. The row is one cell tall.
  2. 2L6The next row is 2 µm higher and flipped (FS), so neighboring rows share a power or ground rail.
  3. 3L9A fence: cells assigned to alu_fence must stay inside, and no other cells may enter.
  4. 4L12A FIXED macro with a 5 µm halo on every side. The placer treats the halo as a blockage that moves with the macro.
  5. 5L13SOURCE DIST marks physical-only cells (taps, endcaps, fillers) that are not in the logical netlist.
  6. 6L14A well tap, FIXED before placement. The next tap in this row is a set distance to the right.
  7. 7L15A normal standard cell. PLACED means later tools may still move it, for example during CTS or ECO.
  8. 8L17This flop belongs to the fence region. Its y of 60 µm is an odd row, hence FS orientation.
  9. 9L18A spare gate, unconnected, waiting for a possible metal-only ECO.
  10. 10L21Hard blockage: no standard cells at all.
  11. 11L22Soft blockage in a macro channel: kept empty in initial placement, available to buffers and clock cells later.
  12. 12L23Partial blockage: standard cells may fill at most 40% of this area, which relieves a congestion hot spot.

The log below follows one run. Overflow is how much cell area still sits in over-full bins. It starts near 1, with every cell clumped together, and the run stops at 0.1. HPWL rises while the cells spread, which is expected. Legalization should add only a percent or two.

If you only read three lines, read the converged overflow (did global placement finish spreading?), the legalization displacement and HPWL change (did snapping undo much of it?), and the congestion table (will the router cope?). A hotspot named next to a fence or a macro usually points back to a floorplan or constraint choice.

place.log (illustrative, modeled on OpenROAD output)log
[gpl] Target density 0.650, placeable area 15,220.4 um^2, core util 58.6%
[gpl] Instances 48,211 (movable 45,903, fixed 2,308), nets 47,960
[gpl] Initial placement done, HPWL 121,330 um
[gpl]  Iter   Overflow   HPWL(um)
[gpl]     1      0.974    120,815
[gpl]   100      0.712    402,660
[gpl]   250      0.301    641,027
[gpl] Timing-driven: reweighted worst 10% of nets, max weight 5
[gpl] Routability: RUDY RC metric 1.18 > target 1.01, inflating cells in 1,412 tiles (+3.1% area)
[gpl]   420      0.198    688,904
[gpl]   515      0.099    702,013
[gpl] Converged: overflow 0.099 at iteration 515
[rsz] Inserted 412 buffers on 287 nets, resized 1,086 instances
[dpl] Legalized 46,315 instances
[dpl] Displacement: average 0.96 um, max 14.2 um (U8817)
[dpl] HPWL 702,013 -> 714,583 um (+1.8%)
[dpl] check_placement: 0 overlaps, 0 off-site, 0 fence violations
[grt] Estimated congestion (GCell = 15 M3 pitches)
Layer   Usage   Max H/V overflow   Total overflow
M2      60.9%   0 / 0               0
M3      75.7%   2 / 0              14
M4      71.2%   0 / 3               9
M5      48.3%   0 / 0               0
[grt] Hotspot: 3 GCells over capacity near (142.0, 61.5) um, inside alu_fence
  1. 1L1Average utilization is 58.6%, but the per-bin cap is 65%. Local density can be higher than the average, never higher than the target.
  2. 2L5Overflow near 1: the initial placement puts cells on top of each other near the center of their connections.
  3. 3L7HPWL rises as density forces spread cells. That is normal. A placer that keeps HPWL flat here is not spreading.
  4. 4L8At an overflow checkpoint, slack from a virtual repair raises net weights on the most critical nets.
  5. 5L9RUDY found tiles over the target RC metric. Cells there are inflated so neighbors are pushed out and wire room opens up.
  6. 6L12Stop at the overflow target (0.1). Lower targets spread more evenly but take longer and raise HPWL.
  7. 7L13Physical synthesis changes the netlist: 412 new instances must now be legalized too (45,903 + 412).
  8. 8L15Average under one micron is healthy. Large maximum displacement usually means a cell was pushed out of a full fence or a blocked area.
  9. 9L16Legalization cost. A jump of 5% or more suggests global placement ended too dense.
  10. 10L21M3 has horizontal overflow in two GCells. The router may detour around it, but it will cost timing or DRCs.
  11. 11L24The hotspot is inside the fence. The fence is too small for its logic: enlarge it, add a partial blockage, or lower density there.
  • Too crowded. If cells are packed too tightly, the wires can’t all fit later. The fix is to give the cells more room or move crowded groups apart.
  • Too spread out. If cells are far apart, signals travel too far and the chip runs slower. More empty space also means a bigger, more expensive chip.
  • Traffic jams near big blocks. The edges and corners of memory blocks attract wires. Without a clear border around them, wiring gets stuck there.
  • Problems found late. A placement that looks fine can still fail during wiring. Teams run a quick trial routing right after placement to catch this early.
  • Congestion hot spots. High local density, high-pin-count cells clustered together, and narrow channels between macros all overload the routing. Teams catch it with a global-route congestion report and fix it with partial blockages, padding, halos or a lower target density.
  • Timing that looked fine before placement. Synthesis only guessed at wire delay. Real distances between placed cells can make a path that passed synthesis fail. Run timing-driven placement and check slack on placement-based parasitics before moving to CTS.
  • No room to optimize. At very high utilization, repair_design has nowhere to put buffers, and legalization pushes cells far from their ideal spots.
  • Large legalization moves. If global placement ends too dense, legalization has to push cells a long way to find free sites, which undoes the wirelength and timing work. Check the average and maximum displacement in the detailed placement report.
  • Over-tight fences. A fence that is too small for its logic packs it to the limit and creates a congestion island.
  • Bad I/O or macro placement. Placement can’t fix a floorplan that puts connected macros on opposite sides of the die. The fix is to go back to floorplanning.
  • Optimistic HPWL, pessimistic routing. HPWL ignores detours and layer limits. A placement can win on HPWL and lose on routed DRCs. Always compare a global route’s overflow per layer as well as RUDY.
  • Pin-density congestion. Local cell density can look fine while pin density is not, for example with clusters of complex gates or multi-bit flops. These areas fail pin access in detailed routing. Pad those masters, or put partial blockages over the cluster.
  • Net-weight overshoot. Aggressive timing-driven weights pull critical clusters tight. That can create density peaks and long detours on non-critical nets, so total negative slack gets worse even as worst slack improves. Tune the weight cap and net percentage (OpenROAD exposes both) for each design.
  • Legalization damage. Large maximum displacement usually means a full fence, fragmented rows from tap and endcap cuts, or multi-height cells with no matching row. Check the maximum and the distribution of displacement as well as the average.
  • Advanced-node adjacency violations. Drain-drain abutment, implant width and OD jog rules surface only after placement, and late fixes perturb timing.
  • Non-determinism and instability. Small netlist changes can produce very different placements, which breaks run-to-run comparison and ECO convergence. Stability is a first-class requirement in physical synthesis loops.
  • Scan reorder surprises. Reordering changes which flop drives which scan input, so it changes scan-path timing and the test patterns. Keep ORDERED segments and partitions intact, check hold on the new scan connections, and rerun ATPG on the reordered chains.

This part covers the algorithms inside the tools. It’s written for the Expert tier.

The problem

Model the netlist as a hypergraph G = (V, E): cells are vertices and nets are hyperedges. Find (x, y) for each movable cell to minimize total HPWL, where HPWL_e = max|xᵢ − xⱼ| + max|yᵢ − yⱼ| over pins of net e. The constraints: every cell is in enough free sites, aligned to a row, and overlaps nothing. Most placement sub-problems are at least NP-hard. Even the simpler min-cut bipartitioning problem has no known polynomial-time exact algorithm, so practical tools rely on heuristics. The standard decomposition is global placement, which solves a relaxation with bin density ρ_b ≤ ρ_t, then legalization, then detailed placement.

Historically there are four families: stochastic (simulated annealing), min-cut partitioning, quadratic, and nonlinear analytical.

Simulated annealing and why it lost at scale

TimberWolf (1985) placed standard cells with . It proposed cell displacements and pairwise swaps, accepted a cost increase Δc with probability min(1, exp(−Δc/T)), and cooled with T_new = α·T_old, where α is usually 0.8–0.95. A range limiter shrank the move window with the logarithm of T, so late moves became local. On circuits of 800–2,700 cells it cut estimated wirelength by 45–66% against the placer it was compared with. Annealing gives good quality, but it needs a huge number of moves and converges slowly. That poor scalability pushed it out of global placement once designs reached millions of cells.

Min-cut partitioning

Top-down placers recursively bisect the netlist and the die, assigning each half of the cells to one half of the region and minimizing the nets cut. Fiduccia–Mattheyses (FM) refines a bipartition by moving one cell at a time under a balance constraint. Careful data structures make each pass linear in netlist size. Multilevel partitioning (hMETIS) coarsens the hypergraph, bisects the smallest version, then projects and refines level by level. It produced cuts 6–23% better than earlier tools and ran several times faster. Min-cut placers such as Capo were state of the art in the 2000s. A bad early cut can’t be undone later, though, and whitespace handling across levels is difficult.

Quadratic and force-directed placement

Model each net as springs, so wirelength becomes a quadratic function of cell coordinates. Minimizing it means solving a large sparse linear system, typically with conjugate gradient. Unconstrained, every cell collapses toward the center, so force-directed placers add pseudo-pins and pseudo-nets that pull cells out of dense regions. SimPL keeps two placements. The lower-bound one is quadratic, solved with CG on a Bound2Bound net model. The upper-bound one comes from fast look-ahead legalization. Upper-bound positions become anchors that pull the next lower-bound solve, and the two converge. ComPLx generalizes this as a primal-dual Lagrange optimization. Quadratic placers are fast and stable, but their quality usually trails nonlinear placers.

Smooth wirelength: log-sum-exp and weighted-average

Nonlinear placers need a differentiable wirelength. Log-sum-exp (LSE) approximates the x-span of a net as γ·(ln Σ exp(xᵢ/γ) + ln Σ exp(−xᵢ/γ)), with error at most γ·ln n for an n-pin net. The weighted-average (WA) model takes the difference between a softmax-weighted mean of the pin coordinates (weights exp(xᵢ/γ)) and a softmin-weighted mean (weights exp(−xᵢ/γ)). Its error bound is roughly half that of LSE. γ trades accuracy against smoothness and can’t be set arbitrarily small because of floating-point limits. ePlace, RePlAce and DREAMPlace all use WA.

Electrostatic placement: ePlace, RePlAce, DREAMPlace

In ePlace’s model, each cell is a positive charge with q equal to its area. The density penalty N is the system’s potential energy, and the spreading force on a cell is q·ξ, where ξ is the local electric field. The potential comes from Poisson’s equation, solved spectrally with FFT in O(n log n), with Neumann boundary conditions that keep cells inside the region. The global objective is min W(v) + λN(v), and the whole thing is placed flat, with no clustering.

Instead of nonlinear conjugate gradient with line search, ePlace uses Nesterov’s accelerated gradient method. Its step length is the inverse of a Lipschitz constant predicted in closed form, which avoids line search and is more than 2× faster than CG. A preconditioner approximates the Hessian. RePlAce adds a per-bin local density multiplier that grows exponentially with that bin’s overflow, plus dynamic step-size control. It reports 2.00% lower HPWL than the best published results on ISPD 2005/2006, and it extends to routability through congestion-driven cell inflation. OpenROAD’s gpl is built on RePlAce.

DREAMPlace recasts the same problem as neural-network training. Nets play the role of data samples, the WA wirelength plus density penalty is the loss, the forward pass computes the objective and the backward pass computes gradients. It is built on PyTorch with custom GPU kernels for wirelength and density. It reports over 30× speedup in global placement against multi-threaded RePlAce with no quality loss, about one minute for a million-cell design, and near-linear scaling to 10 million cells.

Legalization: Tetris and Abacus

Tetris-style legalization sorts cells by x and drops each one greedily into the nearest free position. Abacus also sorts by x and legalizes one cell at a time, trying it in nearby rows. For each trial row, its PlaceRow step re-places the cells already legalized in that row to minimize their total movement, and it keeps the row with the lowest cost. That gives lower total displacement than Tetris. OpenROAD’s dpl started from OpenDP’s diamond search and now defaults to a negotiation-based legalizer. It handles 1×–4× multi-height cells, fence regions and fragmented rows. Mixed-cell-height legalizers also have to match each multi-row cell to rows with the right power/ground rails, and they use window-based insertion plus network-flow cleanup for displacement, edge spacing and pin access. At 10 nm and below, a MILP-based final pass can fix adjacency rules (drain-drain abutment, implant width, OD jogs) in independent windows. One such pass fixed 99% of violations on an abstracted 7 nm library.

Detailed placement moves

FastPlace’s detailed placer shows the standard move set. Global swap computes each cell’s optimal region (where its HPWL would be smallest with everything else fixed) and swaps it with a cell or an empty space in that region if the gain is positive. Vertical swap moves a cell up or down a row toward its optimal region. Local reordering tries every order of three consecutive cells in a segment. Single-segment clustering shifts cells within a row segment while keeping their order. The passes repeat until improvement stalls. OpenROAD adds improve_placement and optimize_mirroring, which flips cells so their pins face their connections. Pin-access-aware refinement makes small local cell moves during routing to clear access conflicts that placement models miss.

Novice · 0 of 4 correct
  1. Q1A three-pin net has pins at (1, 1), (4, 3) and (2, 6) µm. What is its HPWL?

  2. Q2What does legalization do?

  3. Q3A soft placement blockage is placed in a channel between two macros. What happens?

  4. Q4Why can a tool reorder the flip-flops in a scan chain after placement?

Sources

  1. Placement (electronic design automation)Wikipedia contributors · WikipediaOverview of placement goals, global vs. detailed placement, NP-hardness and algorithm families.
  2. Global Placement (gpl)The OpenROAD Project · OpenROAD documentationRePlAce-based Nesterov placer; -density default 0.7, -overflow default 0.1; timing-driven net reweighting; RUDY-based cell inflation.
  3. Detailed Placement (dpl)The OpenROAD Project · OpenROAD documentationLegalization to sites/rows minimizing displacement; OpenDP origin; mixed-cell-height and fence support; padding, fillers, improve_placement.
  4. RePlAce: Advancing Solution Quality and Routability Validation in Global PlacementChung-Kuan Cheng, Andrew B. Kahng, Ilgweon Kang, Lutong Wang · IEEE TCAD (author-hosted, UCSD VLSI CAD Lab) · 2018Local density penalty per bin, dynamic step size, routability extension; 2.00% HPWL gain on ISPD 2005/2006.
  5. Gate Resizer (rsz)The OpenROAD Project · OpenROAD documentationrepair_design buffering and sizing, repair_timing with sizing, pin swap, cloning, buffering and VT swap; estimate_parasitics -placement.
  6. ePlace: Electrostatics-Based Placement Using Fast Fourier Transform and Nesterov’s MethodJingwei Lu, Pengwen Chen, Chin-Chih Chang, Lu Sha, Dennis Jen-Hsin Huang, Chin-Chi Teng, Chung-Kuan Cheng · ACM TODAES (copy on UCSD CSE 248 course page) · 2015Problem formulation, HPWL, LSE and WA models, eDensity via Poisson/FFT, Nesterov’s method, survey of four placer families.
  7. DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI PlacementYibo Lin, Shounak Dhar, Wuxi Li, Haoxing Ren, Brucek Khailany, David Z. Pan · DAC 2019 (author version, UT Austin) · 2019Placement cast as neural-network training in PyTorch on GPUs; over 30× faster than multi-threaded RePlAce.
  8. LEF/DEF 5.8 Language Reference: LEF SyntaxCadence Design Systems (open LEF/DEF standard) · Coriolis project, LIP6 (public mirror) · 2016SITE definitions and MACRO classes including WELLTAP and ENDCAP.
  9. LEF/DEF 5.8 Language Reference: DEF SyntaxCadence Design Systems (open LEF/DEF standard) · Coriolis project, LIP6 (public mirror) · 2016ROWS, COMPONENTS (PLACED/FIXED, HALO, REGION, SOURCE DIST), BLOCKAGES (SOFT, PARTIAL), REGIONS (FENCE/GUIDE), SCANCHAINS.
  10. FastPlace 2.0: An Efficient Analytical Placer for Mixed-Mode Designs (slides)Natarajan Viswanathan, Min Pan, Chris Chu · ASP-DAC 2006 · 2006Detailed placement moves: global swap, vertical swap, local re-ordering, single-segment clustering.
  11. On Robustness and Generalization of ML-Based Congestion Predictors to Valid and Imperceptible PerturbationsChester Holtz, Yucheng Wang, Chung-Kuan Cheng, Bill Lin · arXiv · 2024Defines RUDY: each net’s wire volume spread uniformly over its bounding box, used as a congestion indicator.
  12. On Whitespace and Stability in Mixed-Size Placement and Physical SynthesisSaurabh N. Adya, Igor L. Markov, Paul G. Villarrubia · ICCAD 2003 (author-hosted, University of Michigan) · 2003Gate sizing, buffering and detailed placement need local whitespace in every region of the die.
  13. Restructure (rmp)The OpenROAD Project · OpenROAD documentationLogic restructuring for area or delay by handing logic cones to ABC.
  14. A Proposal for Routing-Based Timing-Driven Scan Chain OrderingPuneet Gupta, Andrew B. Kahng, Stefanus Mantik · ISQED 2003 (author-hosted, UCLA NanoCAD Lab) · 2003Scan chain ordering from placement data, cast as a traveling-salesman problem; affects routability, wirelength and timing.
  15. Hierarchical Whitespace Allocation in Top-Down PlacementAndrew E. Caldwell, Andrew B. Kahng, Igor L. Markov · IEEE TCAD (author-hosted, University of Michigan) · 2003Whitespace in real designs varies from about 20% to about 70% with methodology; congestion guardbanding drives it.
  16. Scalable Detailed Placement Legalization for Complex Sub-14nm ConstraintsKwangsoo Han, Andrew B. Kahng, Hyein Lee · UCSD VLSI CAD Lab (author-hosted)Drain-drain abutment, minimum implant width and OD jog rules break correct-by-construction placement at 10 nm and below.
  17. Pin-Accessible Legalization for Mixed-Cell-Height CircuitsHaocheng Li, Wing-Kai Chow, Gengjie Chen, Bei Yu, Evangeline F. Y. Young · IEEE TCAD (author-hosted, CUHK) · 2022Multi-row-height cells, power/ground rail alignment, edge spacing and pin access in legalization.
  18. In-Route Pin Access-Driven Placement Refinement for Improved Detailed Routing ConvergenceAndrew B. Kahng, Jian Kuang, Wen-Hao Liu, Bangqi Xu · IEEE TCAD (author-hosted, UCSD VLSI CAD Lab) · 2021Neighboring and cell-boundary pins degrade pin access and cause routing DRCs at advanced nodes.
  19. Latch-upWikipedia contributors · WikipediaSubstrate taps lower well resistance to prevent latch-up; fabs set maximum distance to the nearest tap.
  20. Tapcell (tap)The OpenROAD Project · OpenROAD documentationTap cell and endcap/boundary cell insertion, tap distance, macro halos.
  21. Resource-Aware Functional ECO Patch GenerationAn-Che Cheng, Iris Hui-Ru Jiang, Jing-Yang Jou · DATE 2016 (open proceedings archive) · 2016Spare cells are spread over a design during placement so later metal-only ECOs can rewire them.
  22. The TimberWolf Placement and Routing PackageCarl Sechen, Alberto Sangiovanni-Vincentelli · IEEE Journal of Solid-State Circuits (copy on University of Toronto course page) · 1985Simulated annealing placement: acceptance rule, geometric cooling with α of 0.8–0.95, range limiter, 800–2,700-cell circuits.
  23. Global Routing (grt)The OpenROAD Project · OpenROAD documentationGCells, capacity, overflow, congestion reports viewable in the GUI.
  24. A Linear-Time Heuristic for Improving Network PartitionsC. M. Fiduccia, R. M. Mattheyses · 19th Design Automation Conference (copy on Utrecht University course page) · 1982Iterative min-cut heuristic, linear time per pass, moves one cell at a time under a balance constraint.
  25. Multilevel Hypergraph Partitioning: Applications in VLSI DomainGeorge Karypis, Rajat Aggarwal, Vipin Kumar, Shashi Shekhar · IEEE Transactions on VLSI Systems (copy on Georgia Tech course page) · 1999Multilevel coarsen, bisect, project and refine; the algorithm behind hMETIS.
  26. SimPL: An Effective Placement AlgorithmMyung-Chul Kim, Dong-Jin Lee, Igor L. Markov · IEEE TCAD (author-hosted, University of Michigan) · 2012Force-directed quadratic placement; lower-bound and upper-bound placements with look-ahead legalization and anchors.
  27. ComPLx: A Competitive Primal-dual Lagrange Optimization for Global PlacementMyung-Chul Kim, Igor L. Markov · DAC 2012 (author-hosted, University of Michigan) · 2012Generalizes SimPL as a primal-dual Lagrange optimization.
  28. Abacus: Fast Legalization of Standard Cell Circuits with Minimal Movement (slides)Peter Spindler, Ulf Schlichtmann, Frank M. Johannes · ISPD 2008 (ispd.cc slide archive) · 2008Tetris-style greedy legalization vs. Abacus, which re-places already-legal cells in a row to minimize total movement.