By this point the chip exists as a written description. Engineers have spelled out what it should do, in a special programming-like language, and checked it carefully. A factory can’t build from a description, though. It needs to know exactly which tiny parts to use and how to wire them together.
Logic synthesis bridges that gap. A synthesis tool reads the description and produces a : a long list of parts and the wires between them. The parts come from a catalog of pre-built building blocks called . Each one is a small, tested circuit, such as a gate that outputs 1 only when both its inputs are 1, or a one-bit memory.
The tool also has goals. It must make the circuit fast enough, keep it small and keep its power use low. Engineers call this trio : power, performance and area. Synthesis is the first point in the flow where these get real numbers.
Logic synthesis translates register-transfer level (RTL) code in Verilog, SystemVerilog or VHDL into a gate-level built from a specific library. It sits after RTL design and verification and before design-for-test (DFT) insertion and floorplanning.
Most tools work in three broad phases:
- Elaboration: parse the RTL and build a generic, technology-independent circuit of operators, multiplexers and registers.
- Technology-independent optimization: simplify that logic, share hardware, re-encode state machines.
- Technology mapping and timing optimization: cover the logic with real library cells and then resize, restructure and buffer until the timing constraints are met.
The commercial tools are Synopsys Design Compiler and Fusion Compiler and Cadence Genus. The open-source path is Yosys, which handles the front end and calls the ABC logic optimizer from UC Berkeley for optimization and mapping.15
Synthesis is where a design first gets measured numbers (power, performance, area). Those numbers are estimates, because nothing has been placed yet, but they tell you early whether the architecture can hit its targets.
Synthesis is the first point where RTL meets a real library, real constraints and a real cost function. Its job is to produce a gate-level that is functionally equivalent to the RTL, meets the constraints with margin, and leaves placement and routing a tractable problem. Synopsys Design Compiler and Fusion Compiler, Cadence Genus, and open-source Yosys with ABC all follow the same broad shape: elaborate, optimize independently of technology, map, then run timing- and power-driven incremental optimization.1
The quality of results (QoR) you get here sets the ceiling for the back end. Place-and-route can size, buffer and restructure locally, but it rarely recovers from a poorly structured datapath or a bad FSM encoding. The bigger risks at this stage are about inputs and estimates:
- SDC that under- or over-constrains the design.
- Wire estimates that disagree with what placement later sees, the problem physical-aware synthesis tries to fix.11
- Transformations such as retiming and FSM recoding that make harder.
Synthesis takes in three things:
- The checked design description that says what the chip does.
- The catalog of building blocks, with each block’s size, speed and power use.
- The goals: how fast the clock must tick and how much time signals have at the chip’s edges.
It hands back the parts-and-wires list plus reports: how fast the result is, how big it is and how much power it’s expected to use. The next teams add test circuitry and start planning the physical layout.
| Direction | What | Format |
|---|---|---|
| In | Design source | Verilog / SystemVerilog / VHDL RTL, plus file lists and defines |
| In | Cell library: function, area, timing, power | Liberty .lib, one per process/voltage/temperature corner7 |
| In | Timing constraints: clocks, I/O delays, exceptions | SDC (Tcl-based)8 |
| In (physical-aware only) | Cell and technology abstracts, rough floorplan | LEF, DEF |
| Out | Mapped design | Gate-level Verilog netlist |
| Out | Constraints for downstream tools | SDC written back against the netlist’s names |
| Out | Quality of results | Timing, area, power and QoR reports (text) |
The netlist goes to DFT, which replaces flops with scan flops and stitches scan chains, and then to floorplanning. A static timing analyzer such as OpenSTA can read the same three things synthesis produced or used (Verilog netlist, Liberty, SDC) and give an independent timing check.9
In practice the inputs also include a dont_use list, a list of modules or instances to keep, and power intent if the design has multiple voltage domains. Library sets usually cover several Vt flavors and at least one slow corner for setup. Physical-aware runs add the technology LEF, cell LEFs and a floorplan DEF (die size, macro placement, blockages) so the tool can run a coarse placement while it optimizes.
Downstream consumers care about more than the netlist. DFT needs to know which flops are clock-gated and whether ICG test-enable pins are tied off or exposed. Floorplanning wants hierarchy preserved for big blocks. LEC needs the RTL, the netlist and any guidance the tool wrote about renamed, merged or retimed registers.
Synthesis happens in a few big steps.
- Read and understand. The tool reads the description and builds its own rough sketch of the circuit: adders here, choosers there, memory bits where values must be kept.
- Simplify. Like simplifying an algebra problem, the tool removes anything unused, merges repeated work and rewrites logic into a tidier form.
- Pick real parts. It covers the simplified sketch with building blocks from the catalog, choosing between small-and-slow and big-and-fast versions of each.
- Check the timing and fix it. It measures every route a signal can take between memory points. If any route is too slow, it swaps in faster parts or reshuffles the logic, then checks again.
- Prove nothing changed. A separate tool runs an , a mathematical proof that the parts list behaves exactly like the original description for every possible input.
Along the way the tool makes trade-offs. Faster usually means bigger parts and more power. To save power, it can add : it stops the clock from reaching parts of the chip that have nothing to do this moment, like turning off the lights in empty rooms.12
Elaboration
resolves parameters, expands generate blocks, builds the module hierarchy and converts always blocks into a generic netlist. Yosys, for example, represents this as coarse-grain word-level cells (adders, comparators, multiplexers, registers) that later break down into fine-grain single-bit gates.2 This is where the tool infers flip-flops from clocked blocks and, if your code leaves a signal unassigned in some branch of a combinational block, a latch.
Technology-independent optimization
Before any library cell appears, the tool cleans up the generic logic. It propagates constants, removes logic that drives nothing and narrows operators to the bit widths actually used. It also applies a few structural transformations:
- uses one adder for two additions that never happen in the same cycle, with a multiplexer in front. Yosys’s
sharepass uses a SAT solver to prove the two operations are mutually exclusive.3 - FSM extraction and encoding finds state registers and may re-encode them. Yosys extracts the state machine, optimizes it and can recode it as , with one flop per state.4 One-hot spends flops to make next-state logic shallow; binary encoding does the opposite.
- Boolean optimization rewrites the logic into fewer, shallower gates. Modern tools do this on an And-Inverter Graph, a network of two-input ANDs and inverters.22
Technology mapping with Liberty
covers the optimized logic with cells from the target library. The tool learns about the cells from the file, which lists each cell’s logic function, area, leakage power, pin capacitances, delay tables and internal energy.7 Delay is stored as an table: rows indexed by input transition time, columns by output load capacitance. The tool interpolates between entries. A trimmed, illustrative entry:
cell (NAND2_X1) {
area : 0.80;
cell_leakage_power : 12.4;
pin (A1) { direction : input; capacitance : 0.0016; }
pin (A2) { direction : input; capacitance : 0.0017; }
pin (ZN) {
direction : output;
function : "!(A1 & A2)";
timing () {
related_pin : "A1";
cell_rise (delay_template_3x3) {
index_1 ("0.01, 0.05, 0.20");
index_2 ("0.001, 0.005, 0.020");
values ("0.012, 0.025, 0.071", \
"0.018, 0.032, 0.079", \
"0.035, 0.050, 0.098");
}
}
}
}- 1L2Cell area, usually in µm². Synthesis sums these for its area report.
- 2L4Input pin capacitance. It becomes part of the load on whatever gate drives this pin.
- 3L8The Boolean function. The mapper matches logic against this.
- 4L12index_1: input transition times (ns) for each row.
- 5L13index_2: output load capacitances (pF) for each column.
- 6L14Delay grows to the right (more load) and downward (slower input edge).
Flip-flops get mapped separately from combinational logic. In Yosys, dfflibmap maps registers to the library’s flops and abc -liberty maps the logic.5
Timing-driven optimization with SDC
Without constraints, a tool has no idea how fast the design must be. supplies that.8
create_clocknames a clock and gives its period.set_input_delayandset_output_delaysay how much of each cycle is used outside the block, treating each I/O as if it connected to a register on the given clock.set_false_pathremoves paths that never need to meet single-cycle timing, such as some crossings between unrelated clocks.set_multicycle_pathgives a path more than one cycle when the design only samples the result every N cycles.
For every path the tool computes : required time minus arrival time. The path with the worst slack is the . Optimization then upsizes cells on that path, restructures logic, inserts buffers and downsizes cells elsewhere to win back area.
Area, power and timing trade-offs
Every one of these choices moves . A tighter clock produces larger cells, more buffers and more leakage. Two techniques target power specifically:
- Clock gating. When a group of flops shares an enable, synthesis can replace the enable multiplexers with one integrated clock-gating (ICG) cell that stops their clock. Yosys’s
clockgatepass does exactly this and can pick the ICG from the Liberty file.6 Because the flops see no clock edges while idle, their switching power drops.12 - Multibit flops. Packing several one-bit registers into one shares the clock buffering inside the cell and lowers the load on the clock network.13
moves registers across combinational logic to balance stage delays while keeping the circuit’s cycle-by-cycle behavior at its ports.18 If one pipeline stage takes 1.2 ns and the next takes 0.6 ns, moving some logic from the first stage across the register into the second can bring both under a 1.0 ns clock without adding a cycle of latency.
Power estimates at this stage have three parts. Leakage comes straight from each cell’s Liberty leakage value. Internal power comes from the Liberty energy tables, charged each time a cell switches. Switching power depends on net capacitance and on how often each net toggles.7 Without simulation data, the tool assumes a default toggle rate, so treat the number as a rough guess until you feed it activity from real tests. Clock gating works by cutting the toggle rate of the clock pins, which are the busiest nets in the design.12
Hierarchy: flatten or keep
Flattening dissolves module boundaries so optimization can cross them, which usually improves timing and area. Keeping hierarchy makes the netlist easier to debug, to floorplan by block and to patch late. Flows often mix the two: OpenROAD-flow-scripts, for example, can synthesize hierarchically, keeping modules above an area threshold and flattening smaller ones.10
Proving it’s still the same design
Synthesis rewrites almost everything, so teams prove the result with (LEC) between RTL and netlist. LEC is formal: it proves equality for all inputs instead of sampling them with tests. Open-source EQY does this with Yosys.14
Elaboration and the generic netlist
resolves parameters and generates, binds the hierarchy and turns processes into word-level operators, multiplexer trees and storage elements. Yosys keeps this as coarse RTLIL cells and lowers them to single-bit gates later.2 Its default script shows the order most tools follow:
- Process conversion (
proc). - FSM extraction (
fsm), word-width reduction (wreduce) and arithmetic extraction (alumacc). - SAT-based resource sharing (
share) and memory inference. - Generic gate mapping (
techmap), followed by ABC.1
Commercial tools add datapath synthesis that picks adder and multiplier architectures from the timing context. Review elaboration warnings before anything else. Inferred latches, truncated widths, multiply driven nets and black boxes all show up here first.
Technology-independent optimization
The logic is restructured on an or similar network. Rewriting and refactoring reduce node count without increasing depth, and balancing reduces depth without increasing node count.22 Structural transforms also run at this level:
- merges mutually exclusive operators. It saves area but puts a mux in front of the shared operator, so a timing-driven tool may refuse to share or may un-share on critical paths. Yosys checks exclusivity with SAT.3
- FSM re-encoding to (the recoding target Yosys documents) or another encoding.4 Recoding changes the register set, which matters for LEC.
Mapping against Liberty
covers the subject graph with library cells to minimize delay on critical paths and area elsewhere. Cell data comes from : area, leakage, pin capacitance, delay and transition tables indexed by input slew and output load, internal-power tables and timing checks.7 Design-rule limits such as max_transition, max_capacitance and max_fanout are hard constraints, and tools give them priority over slack, so a high-fanout net that blows max_transition shows up as a buffer tree. Large high-fanout nets like reset and scan enable are usually marked ideal in synthesis and left for placement-aware buffering.
Timing-driven synthesis
The tool times every path against the , with clocks treated as ideal because no clock tree exists yet. set_clock_uncertainty therefore carries margin for future skew and jitter.8 Exceptions need care:
- A removes a path from both optimization and checking.
- A with
-setup Nalso moves the default hold check. You pair it with-hold N−1to move hold back to the original edge.8
Optimization works on WNS and TNS per path group. On paths with negative the tool upsizes cells, swaps to a faster Vt flavor, restructures logic, duplicates high-fanout drivers and buffers. On paths with positive slack it downsizes cells and reclaims area and leakage. Yosys’s ABC integration shows the open-source version of this: a delay target (-D) adds retiming toward that target, and a constraint file adds buffering and up/down-sizing steps.6
Power-oriented transformations
- . Flops that share an enable lose their recirculating muxes and get one latch-based ICG instead.12 Yosys’s
clockgatepass groups flops by clock and enable, can choose ICGs from Liberty, sets a minimum group size, and ties a named ICG pin low for later connection to scan enable.6 The cost: the enable path now has to arrive before the ICG’s clock edge, which is earlier than a flop’s D-pin setup. That turns some enables into new critical paths. - Multibit banking. share internal clock buffering and present a smaller clock load. Merging them needs physical proximity, so many flows bank at or after placement instead of, or as well as, in synthesis. OpenROAD-flow-scripts clusters flops during placement.1310
Retiming
moves registers across logic to minimize clock period or register count while keeping cycle-by-cycle I/O behavior.17 It is powerful on deep, unbalanced pipelines, but it renames, splits and merges flops, which complicates LEC, debug and any later change that refers to register names. Open flows treat it with caution. OpenROAD-flow-scripts marks its module retiming option experimental, notes that the option does not check equivalence, and notes that its objective ignores the SDC.10
Hierarchy
Flattening enables boundary optimization: constants propagate into blocks, logic moves across ports, unused outputs disappear. Keeping hierarchy preserves block boundaries for hierarchical floorplanning, limits the reach of a late change and makes LEC and timing debug easier. A common compromise is to keep large blocks and flatten small ones; OpenROAD-flow-scripts exposes exactly that threshold.10
Physical-aware synthesis
Classic synthesis estimated wire load from a , a table that maps fanout to length and RC.7 Once interconnect dominates delay, that breaks. A WLM predicts the average net reasonably well, but the variation between nets is so large that timing on the worst paths is grossly mispredicted.11
Placement-aware synthesis replaces the table with estimates from an actual, if rough, placement. Commercial tools offer this as physical-aware modes (Synopsys calls its version topographical): they load LEF and a floorplan and estimate each net from a fast placement. Gluing synthesis to placement has its own convergence trap: synthesis upsizes to fix timing, placement spreads the resulting congestion, and new timing problems appear.11
Equivalence checking
matches state points (flops, ports, black boxes) between the RTL reference and the netlist, then proves each pair of combinational cones equivalent. Anything that breaks one-to-one state correspondence needs extra handling: retiming, FSM recoding, merged constant or duplicate flops. EQY, for instance, lets you declare the new FSM encoding in a recode section or exclude state registers from matching and use a sequential strategy.15 Commercial flows rely on guidance files that the synthesis tool writes for the LEC tool.
Type a short logic rule with up to four inputs (a, b, c, d). In the examples here, & means AND, | means OR, ^ means XOR (one or the other but not both) and ! means NOT. The simulator shows every possible combination of inputs, a tidied-up version of your rule and the parts a tool would pick to build it, with total size and a rough speed. Try (a & b) | (a & c) and see whether the tool spots that a appears in both halves. Then try a ^ b ^ c ^ d, a rule whose tidied-up list is long but whose parts list is short.
Enter a Boolean expression of up to four inputs. The simulator prints its truth table, a minimized , and a netlist mapped onto a tiny generic library (INV, NAND2, NOR2, AND2, OR2, XOR2, AOI21, OAI21) with total area and a delay estimate. Three to try:
(a & b) | (a & c): the SOP has two terms, but the factored form a·(b + c) needs fewer gates. Does the mapped netlist find it?!((a & b) | c): this is exactly one AOI21 cell. Compare its area and depth with an AND2 + OR2 + INV version.a ^ b ^ c ^ d: parity has eight 4-literal product terms and nothing to merge, yet maps to three XOR2 cells.
The simulator is a toy version of the front of the flow: truth table, two-level minimization, then a cover with a generic library (INV, NAND2, NOR2, AND2, OR2, XOR2, AOI21, OAI21) costed by area and .
- Use
a ^ b ^ c ^ dto see why two-level minimization lost to multi-level synthesis: parity has no adjacent minterms, so the SOP is eight 4-literal cubes, while a balanced XOR tree has depth two. - Use
!((a & b) | c)to watch a complex gate absorb three levels of generic logic. - Use
(a & b) | (a & c) | (b & c)(majority) to compare area-oriented and depth-oriented covers.
Remember that depth ignores load and slew; real mappers use the NLDM tables.
After a synthesis run, an engineer opens a handful of reports and checks three numbers before anything else.
- Is it fast enough? The timing report lists the slowest routes. Each one shows how much spare time it has. A negative number means that route is late, and the engineer has to change something.
- How big is it? The area report adds up the size of every part. If it’s much bigger than planned, the chip may not fit or may cost too much.
- How much power? The power report is an estimate at this stage, but a big surprise here is easier to fix now than after layout.
Then they read the warnings. A tool that silently removed a block because nothing used its output, or that built an accidental memory where the designer wanted plain logic, usually says so in a warning first. Finally, the must pass before the netlist goes to the next team.
Below is a small open-source flow, a matching SDC, a fragment of the resulting netlist and one timing path. All are illustrative and use generic cell names.
The artifacts below follow an open-source flow (Yosys + ABC, OpenSTA-style reports). Commercial reports differ in layout but carry the same information. All numbers are illustrative.
read_verilog -sv rtl/alu.sv rtl/top.sv
hierarchy -check -top top
synth -top top -flatten
dfflibmap -liberty lib/generic_typ.lib
abc -liberty lib/generic_typ.lib -D 1000
opt_clean
stat -liberty lib/generic_typ.lib
write_verilog -noattr out/top_netlist.v- 1L2Elaborate and check the hierarchy under the chosen top module.
- 2L3Runs Yosys’s generic script (proc, fsm, share, techmap, ABC) and flattens the design first.
- 3L4Maps generic flip-flops onto the library’s flop cells.
- 4L5ABC optimizes and maps logic onto the Liberty cells. -D sets a 1000 ps delay target and enables retiming toward it.
- 5L7Prints cell counts and total area using the Liberty areas.
- 6L8Writes the gate-level netlist for DFT and floorplanning.
create_clock -name clk -period 1.0 [get_ports clk]
set_clock_uncertainty 0.08 [get_clocks clk]
set_input_delay 0.40 -clock clk [get_ports {a_in[*] b_in[*] en}]
set_output_delay 0.35 -clock clk [get_ports {sum_q[*]}]
set_driving_cell -lib_cell BUF_X2 [get_ports {a_in[*] b_in[*] en}]
set_load 0.005 [all_outputs]
set_false_path -from [get_ports rst_n]
set_multicycle_path 2 -setup -from [get_cells u_div/*] -to [get_cells u_acc/*]
set_multicycle_path 1 -hold -from [get_cells u_div/*] -to [get_cells u_acc/*]- 1L1A 1.0 ns clock (1 GHz) on port clk. Units follow the Liberty file, here ns.
- 2L2Margin for future clock skew and jitter, because clocks are ideal during synthesis.
- 3L30.40 ns of each cycle is spent outside the block before these inputs arrive.
- 4L4The receiving logic outside needs the outputs 0.35 ns before the next edge.
- 5L5Inputs are driven as if by a BUF_X2, so input slews are realistic.
- 6L6Each output drives 0.005 pF, here assumed to be pF from the library units.
- 7L7Excludes the raw asynchronous reset input from timing. Reset is usually synchronized on-chip, and the synchronized reset’s recovery and removal checks are still timed.
- 8L8The divider’s result is only sampled every second cycle, so setup gets two cycles.
- 9L9Moves the hold check back to the original edge. Without it, hold would be checked one cycle late.
module top (clk, rst_n, en, a_in, b_in, sum_q);
input clk, rst_n, en;
input [1:0] a_in, b_in;
output [1:0] sum_q;
wire n1, n2, gclk;
wire [1:0] sum_d;
XOR2_X1 U1 (.A(a_in[0]), .B(b_in[0]), .Z(sum_d[0]));
NAND2_X1 U2 (.A1(a_in[0]), .A2(b_in[0]), .ZN(n1));
XOR2_X1 U3 (.A(a_in[1]), .B(b_in[1]), .Z(n2));
XNOR2_X1 U4 (.A(n2), .B(n1), .ZN(sum_d[1]));
ICG_X1 clk_gate_sum_q (.CK(clk), .E(en), .SE(1'b0), .GCK(gclk));
DFFR_X1 sum_q_reg_0_ (.D(sum_d[0]), .CK(gclk), .RN(rst_n), .Q(sum_q[0]));
DFFR_X1 sum_q_reg_1_ (.D(sum_d[1]), .CK(gclk), .RN(rst_n), .Q(sum_q[1]));
endmodule- 1L8Bit 0 of a + b is a XOR b.
- 2L9NAND2 gives the inverted carry out of bit 0. The mapper chose it because NAND is cheaper than AND.
- 3L11XNOR with the inverted carry equals XOR with the true carry, so no separate inverter is needed.
- 4L13Clock-gating cell inserted because both flops load only when en is high. SE (scan/test enable) is tied low until DFT connects it.
- 5L14The flops lost their enable mux. They are clocked by the gated clock and reset asynchronously by rst_n.
Startpoint: u_alu/op_a_reg_3_ (rising edge-triggered flip-flop clocked by clk)
Endpoint: u_alu/acc_reg_31_ (rising edge-triggered flip-flop clocked by clk)
Path Group: clk
Path Type: max
Delay Time Description
---------------------------------------------------------------
0.000 0.000 clock clk (rise edge)
0.000 0.000 clock network delay (ideal)
0.000 0.000 ^ u_alu/op_a_reg_3_/CK (DFF_X1)
0.112 0.112 v u_alu/op_a_reg_3_/Q (DFF_X1)
0.046 0.158 ^ u_alu/U812/ZN (NAND2_X1)
0.071 0.229 v u_alu/U813/ZN (AOI21_X1)
0.064 0.293 ^ u_alu/U840/ZN (OAI21_X1)
0.069 0.362 v u_alu/U841/ZN (AOI21_X1)
0.058 0.420 ^ u_alu/U902/ZN (OAI21_X2)
0.073 0.493 v u_alu/U903/ZN (AOI21_X1)
0.066 0.559 ^ u_alu/U955/ZN (OAI21_X1)
0.070 0.629 v u_alu/U956/ZN (AOI21_X1)
0.067 0.696 ^ u_alu/U1010/ZN (OAI21_X1)
0.072 0.768 v u_alu/U1011/ZN (AOI21_X1)
0.088 0.856 ^ u_alu/U1102/ZN (XNOR2_X1)
0.034 0.890 v u_alu/U1150/ZN (NOR2_X1)
0.000 0.890 v u_alu/acc_reg_31_/D (DFF_X1)
0.890 data arrival time
1.000 1.000 clock clk (rise edge)
0.000 1.000 clock network delay (ideal)
-0.080 0.920 clock uncertainty
0.920 ^ u_alu/acc_reg_31_/CK (DFF_X1)
-0.052 0.868 library setup time
0.868 data required time
---------------------------------------------------------------
0.868 data required time
-0.890 data arrival time
---------------------------------------------------------------
-0.022 slack (VIOLATED)- 1L1The path starts at a flop’s clock pin (launch).
- 2L2It ends at the D pin of another flop (capture).
- 3L9Ideal clock: no clock tree exists yet, so arrival at every flop is assumed simultaneous.
- 4L11Clock-to-Q delay of the launching flop, from its Liberty table.
- 5L13Alternating AOI21/OAI21 stages are a typical mapped carry chain.
- 6L16The tool already upsized one stage to X2 trying to speed this path up.
- 7L25Total time for data to reach the capturing flop.
- 8L29Uncertainty from the SDC is subtracted from the time available.
- 9L31The capturing flop needs its data this long before the clock edge.
- 10L37Negative slack: data arrives 22 ps late. Fix by restructuring the adder, retiming, or relaxing the constraint.
The netlist is ordinary structural Verilog with no always blocks or operators left. Each line is one cell instance: the library cell name, an instance name and the pin connections. The two-bit adder became four gates, and the enable on the output register became a clock-gating cell feeding both flops. Instance names such as U2 are tool-generated, while register names such as sum_q_reg_0_ keep a trace of the RTL signal, which helps debugging and equivalence checking.
Read a timing path top to bottom. The upper half adds up delays to get the arrival time. The lower half starts from the next clock edge and subtracts uncertainty and setup time to get the required time. Slack is the difference. Here a 32-bit accumulator’s carry chain misses by 22 ps.
Typical responses:
- Let the tool pick a faster adder architecture.
- Enable retiming.
- Move part of the addition into the previous pipeline stage in the RTL.
- Accept a slower clock.
A QoR summary condenses the whole run. Format varies by tool; the fields don’t.
Design : top
Corner : slow (setup)
Clock clk period : 1.000 ns
WNS (setup) : -0.022 ns
TNS (setup) : -0.311 ns
Violating endpoints : 27
Max transition violations : 3
Combinational cells : 41,206
Flip-flop bits : 9,730 (4,914 single-bit, 1,204 x 4-bit multibit)
Clock-gating cells : 212 (gated flop bits: 88.4%)
Total cell area : 48,930.6 µm²
Leakage power : 1.84 mW
Dynamic power (estimate) : 63.2 mW (default toggle rate)- 1L2Setup is judged at the slow corner. Hold is mostly fixed after CTS, so synthesis rarely reports it seriously.
- 2L4Worst negative slack: the single worst path, from the report above.
- 3L5Total negative slack: the sum over failing endpoints. A small WNS with a large TNS means many paths fail, which usually points to architecture or constraints.
- 4L7Design-rule violations get fixed before slack. Non-zero here usually means an unbuffered high-fanout net.
- 5L9Multibit banking packed most bits into 4-bit cells, lowering clock-pin load.
- 6L10Fraction of flop bits behind an ICG. Higher usually means lower clock power, if the enables are actually idle often.
- 7L13With default toggle rates, dynamic power is a rough guess. Feed simulation activity (SAIF or VCD) for a real number.
Check the slack distribution before chasing the worst path. Thousands of endpoints a few picoseconds negative usually means the clock target, uncertainty or a missing exception is wrong. A handful of deep failures points at specific RTL structures. Confirm the same netlist and SDC in a standalone STA run, and run LEC, before the hand-off.
- Wrong goals. If the speed target is missing or mistyped, the tool may build something far too slow, or far too big and power-hungry. Teams review constraint files as carefully as the design itself.
- Code that means one thing in a test and another in hardware. Some ways of writing a design behave differently in simulation than in the synthesized result. Equivalence checks and re-running tests on the netlist catch this.
- Missing pieces. If part of the design isn’t connected to anything useful, the tool deletes it. That is usually a sign of a bug in the design, so engineers read the warnings.
- Optimistic guesses about wires. At this stage nothing is laid out, so wire delays are guesses. A design that just barely passes here may fail later.
- Unconstrained paths. A missing
create_clockor I/O delay leaves paths untimed, and the tool doesn’t optimize what it doesn’t time. Every flow should report unconstrained endpoints and treat a non-zero count as an error. - Inferred latches. A combinational
alwaysblock that doesn’t assign a signal on every path infers a latch. Read elaboration warnings, or usealways_comband lint rules. - Simulation/synthesis mismatch. Initial blocks, delays (
#5), incomplete sensitivity lists and X-assignments can simulate one way and synthesize another. Gate-level simulation of key tests plus LEC catches most cases. - Logic optimized away. If an output is unconnected or tied off, synthesis removes the entire cone that feeds it. Treat a sudden drop in area or flop count as a red flag.
- Over-constraining. Setting the clock tighter “for margin” makes the tool upsize and buffer everywhere, which costs area and power and can make placement harder. Use clock uncertainty for margin instead.
- Wire estimates. Timing with no placement information is optimistic or pessimistic in unpredictable ways.11 Compare synthesis timing with post-placement timing early, while there’s still time to adjust.
Most teams automate these checks. A synthesis run fails the regression if it finds inferred latches, unconstrained endpoints, unmapped cells or a failed equivalence check, and the reports are compared against the previous run so that a jump in area, cell count or negative slack gets investigated the same day. Catching a problem here costs minutes. Catching it after placement and routing costs days.
- Exceptions that hide real paths. A wildcard
set_false_pathor a clock-group declaration that covers a real synchronous crossing removes it from optimization and signoff. Lint SDC, and review every exception with the RTL owner. - Multicycle hold mistakes. A setup multicycle without the matching hold adjustment moves the hold check a cycle late, producing large false hold violations that later stages may try to “fix” with delay cells.8
- Correlation gaps. Wire-load or zero-wire synthesis followed by a real placement can shift slack a lot, worst on long nets, because per-net wire load is mispredicted.11 Physical-aware synthesis narrows the gap. Its convergence still depends on a floorplan that resembles the final one.
- LEC aborts and non-equivalences. Large multipliers and deep arithmetic can make cone proofs hard.20 Retiming, FSM recoding and register merging break state matching.15 Keep the synthesis tool’s guidance, run LEC with datapath-aware settings, and avoid stacking aggressive sequential transforms without a plan to verify them.
- Clock-gating side effects. Enables must meet the ICG’s setup to the clock edge, which is stricter than a flop’s D-pin setup, and the test-enable pin must be connected or tied correctly for scan.6 Too many tiny gating groups waste area and complicate CTS; set a minimum group size.
- Boundary optimization and late changes. Flattened or boundary-optimized blocks lose their port-level meaning, which makes engineering change orders and partial re-synthesis harder. Decide which blocks to keep before the first full run.10
- Library hygiene. Missing
dont_useentries let the mapper pick cells that are hard to route or not meant for signal use (for example clock-only buffers or delay cells).10 A wrong corner or Vt set skews every result.
This part covers the algorithms inside the tools. It’s written for the Expert tier.
- NPN classes of 4-input functions
- 222
- Typical 5-input cuts per AIG node
- 20–30
- Node types in an AIG
- AND2 + inverted edges
Sources: Mishchenko et al. 2006 for the NPN class count, and Mishchenko et al. 2005 for the typical cut count per node.2223
From two-level to multi-level logic
Early logic minimization targeted forms for PLAs. Quine–McCluskey finds all prime implicants and then solves a covering problem. Its runtime and memory grow exponentially with input count. Espresso (Brayton and colleagues at IBM, later refined at Berkeley) avoids expanding the function into minterms. It manipulates cubes (product terms) of the ON-, don’t-care and OFF-sets heuristically, and its results are close to minimal and always free of redundancy.19
Two-level forms are the wrong target for standard cells. Common functions such as parity need exponentially many cubes, and SOP is not canonical.20 MIS and then SIS moved to multi-level Boolean networks. Each node holds a small SOP, simplified with Espresso against local don’t-cares, and algebraic operations (extraction, kerneling, resubstitution) find shared logic between nodes.1821 The SIS scripts that drove this worked well but were slow and needed hand-tuning.22
And-Inverter Graphs
An is a DAG in which every internal node is a two-input AND and edges may be complemented. Primary inputs have no fanin, and registers are cut into input/output pairs. Structural hashing during construction guarantees no two AND nodes have the same pair of fanins, so trivially duplicated logic merges for free.22 AIGs grew out of formal verification and replaced SOP- and BDD-based representations in ABC’s synthesis flow. ABC now scales to designs with millions of nodes that SIS could not finish.21
Rewriting, refactoring and balancing
Rewriting is a greedy local search. For each node it enumerates 4-feasible cuts (sets of at most four nodes that separate it from the inputs). It computes each cut’s function as a 16-bit truth table and classifies it into one of the 222 NPN classes of 4-input functions. It then tries the precomputed AIG subgraphs stored for that class. A replacement is accepted if it reduces node count without adding levels. It is DAG-aware: it counts the nodes that would become dead and the existing nodes it can reuse.22
Refactoring takes one larger cut per node, collapses it and re-factors it algebraically. Balancing reduces depth by algebraic tree-height reduction without adding area. ABC’s resyn2 script interleaves these: balance first, then alternate rewrite/refactor (area without delay) with balance (delay without area). Each pass is cheap, so many passes compound into a global effect. The authors report it is orders of magnitude faster than SIS and MVSIS scripts with comparable or better mapped quality.22
Cut enumeration and cut-based technology mapping
Classical mappers decomposed logic into NAND2/INV subject graphs and covered them with library patterns using tree covering (as in SIS) or DAG covering.18 Because matching was structural, the mapped result inherited the subject graph’s shape. This is the “structural bias” problem.23
Cut-based mapping with Boolean matching, from the same Berkeley group, works in five steps:
- Enumerate k-feasible cuts for every AIG node.
- Compute each cut’s truth table.
- Look the truth table up in a hash table of library functions (Boolean matching, which also chooses input permutations and phases).
- Compute best arrival times in topological order.
- Select the cover in reverse topological order, then recover area on non-critical paths.
Boolean matching is faster and more complete than structural matching for libraries with large complex gates. Two extensions attack structural bias further:
- Supergates are small precomputed networks of library gates treated as single gates.
- Choice nodes store several functionally equivalent structures in the AIG so the mapper can pick among them.23
Cut counts per node grow quickly with k. For LUT mapping, priority cuts keep only a few good cuts per node, making memory and runtime linear in circuit size while preserving quality; the authors note that similar cut-based methods exist for standard cells. The same paper extends mapping to search combined mapping-and-retiming solutions.24
Gate sizing and logical effort
After mapping, the cover is fixed but cell sizes are not.
Logical effort gives the intuition: stage delay is d = gh + p. Here g is the gate’s logical effort (1 for an inverter, 4/3 for NAND2, 5/3 for NOR2), h is the electrical effort (load over input capacitance) and p is the parasitic delay.16 Path delay is minimized when every stage bears equal effort. This explains why NAND-heavy covers win over NOR-heavy ones and why a high-fanout net wants a tapered buffer tree.
Gain-based (constant-delay) synthesis builds on this. It assigns each stage a fixed gain (load over input capacitance), so stage delay is fixed during mapping, and then sizes cells and nets to keep those gains through placement and routing.11 Its limits are real:
- It assumes continuous sizing, while real libraries offer discrete drive strengths.
- It ignores slew and rise/fall asymmetry.
- It needs re-buffering once real wires appear.
Mapping a continuous-sizing solution onto a discrete library can leave a sub-optimal netlist.11 That is why mapped netlists still go through discrete sizing against the library’s NLDM tables.
Retiming
Leiserson and Saxe model a circuit as a graph whose vertices are combinational elements and whose edges carry register counts. Retiming assigns each vertex an integer lag, which moves registers from its inputs to its outputs or back, and preserves I/O behavior. They give algorithms for minimum clock period and show minimum-register retiming is polynomial-time solvable, with a mixed-integer LP characterization of optimal retiming.17 In a full synthesis flow, retiming interacts with mapping. Combining the two searches a larger space than mapping followed by retiming.24
Equivalence checking: BDDs vs SAT
Combinational equivalence checking builds a miter: pair the inputs, XOR each pair of outputs and OR the XORs. The two circuits are equivalent exactly when the miter output is constant 0.25
Reduced ordered BDDs are canonical, so equivalence is a pointer comparison once both sides are built. Their size depends heavily on variable order, though, and for integer multipliers it grows exponentially for every order.20
Modern checkers instead:
- Convert the miter to an AIG and structurally hash it.
- Simulate random and guided vectors to find candidate equivalent internal nodes.
- Prove or refute candidates in topological order with SAT, or with BDDs under small resource limits.
- Merge proven nodes, interleaving light AIG rewriting to shrink the problem before attacking the outputs.
Structural similarity between RTL and netlist is what makes this work. Each proven internal equivalence simplifies the next SAT call.25 ABC’s equivalence checker grew out of exactly this need to verify its own synthesis results.21 It also explains the LEC pain points. Retiming and FSM recoding remove the internal correspondences these methods feed on, and deep arithmetic has few internal equivalences to exploit.
Q1Which input tells the synthesis tool each cell’s area, delay and power?
Q2The clock period is 1.0 ns and the SDC says set_input_delay 0.4 -clock clk on an input. What does that mean for logic inside the block?
Q3A synthesis timing report shows slack −0.022 ns (VIOLATED) on a path. What does that mean?
Q4What does retiming do?
Sources
- synth – generic synthesis script (Yosys command reference)The begin/coarse/fine/check steps of Yosys’s default synthesis script, including proc, fsm, alumacc, share, techmap, abc and the -flatten option.
- Internal cell libraryRTLIL represents a design with coarse-grain word-level cells and fine-grain gate-level cells.
- FSM handlingFSM detection, extraction, optimization and one-hot recoding in Yosys.
- Mapping to cell librariesExample flow: dfflibmap and abc -liberty map a design onto a Liberty cell library.
- Technology mapping commands: abc, clockgate, dfflibmapDefault ABC scripts, the -D delay target (which adds retiming), -constr buffering and sizing, and clock-gating insertion with ICG cells.
- Chapter 8: Cell Characterization (figures), Digital VLSI Chip Design with Cadence and Synopsys CAD ToolsLiberty file structure: area, leakage, lookup-table templates indexed by input transition and output load, cell_rise tables, internal power and wire-load models.
- SDC CommandsMeaning of create_clock (including virtual clocks), set_input_delay/set_output_delay, set_false_path, set_multicycle_path and set_clock_uncertainty.
- OpenSTA: Parallax Static Timing AnalyzerA gate-level static timing verifier that reads Verilog netlists, Liberty libraries and SDC, and supports false-path and multicycle exceptions.
- Flow variablesSynthesis options in an open flow: hierarchical vs flat synthesis, keep-size threshold, ABC area/speed strategy, experimental retiming, and multibit flop clustering at placement.
- Timing and Design Closure in Physical Design FlowsWhy statistical wire-load models mispredict timing, the limits of constant-delay (gain-based) synthesis, and placement-aware synthesis.
- Clock gatingClock gating removes the clock from idle logic to cut dynamic power; ICG cells use an internal latch for a glitch-free gated clock.
- FF-Bond: Multi-bit Flip-flop Bonding at PlacementMultibit flip-flops present a smaller clock load through shared clock logic; merging is often done at or after placement.
- EQY: Getting StartedFormal equivalence checking to ensure a synthesis tool has not changed a design’s function.
- Reference for .eqy file formatGold vs gate designs, match rules for net names, and recode sections for FSM state encodings changed by synthesis.
- Lecture 6: Logical Effort (CMOS VLSI Design, 4th ed. slides)Delay model d = gh + p; logical effort 1 for an inverter, 4/3 for NAND2, 5/3 for NOR2.
- Retiming Synchronous Circuitry (MIT-LCS-TM-309)Retiming as a graph problem: algorithms for minimum clock period and polynomial-time minimum register count.
- SIS: A System for Sequential Circuit Synthesis (Memorandum UCB/ERL M92/41)Multi-level synthesis with Espresso for node simplification, state assignment, tree-covering technology mapping and retiming.
- Espresso heuristic logic minimizerHistory of Espresso (IBM, then UC Berkeley) and why Quine–McCluskey’s exponential growth made a heuristic necessary.
- Graph-Based Algorithms for Boolean Function ManipulationOrdered BDDs are canonical; size depends on variable ordering; multiplier outputs grow exponentially for every ordering.
- ABC: An Academic Industrial-Strength Verification ToolABC’s origins in SIS and MVSIS, its AIG-based synthesis that replaced SIS scripts, and its equivalence checker.
- DAG-Aware AIG Rewriting: A Fresh Look at Combinational Logic SynthesisAIG definition, 4-feasible cuts, 222 NPN classes, rewrite/refactor/balance and the resyn2 script.
- Technology Mapping with Boolean Matching, Supergates and ChoicesCut-based standard-cell mapping with Boolean matching, structural bias, supergates, choice nodes and area recovery.
- Combinational and Sequential Mapping with Priority CutsKeeping a few priority cuts per node instead of all K-cuts gives linear memory and runtime; sequential mapping combines mapping with retiming.
- Improvements to Combinational Equivalence CheckingMiter construction, simulation plus BDD/SAT sweeping of internal equivalences, and interleaving SAT with AIG rewriting.