A chip is full of tiny memory cells called flip-flops, and they all act on the beat of one clock. On every tick, each flip-flop grabs a new value. If the tick reaches some flip-flops much earlier than others, they grab values at the wrong moment and the chip computes garbage.
One wire can’t deliver that tick to a hundred thousand flip-flops. It would be like one person shouting the start signal to every runner in a city marathon: the sound arrives faint and late at the back. builds a branching network of small amplifiers, so every flip-flop gets a crisp tick at nearly the same instant. The gap in arrival times is called , and keeping it small is the main job.
This step comes after placement, because the tree has to reach flip-flops whose positions are now known, and before the rest of the wiring is drawn. It matters for power too: the clock switches on every single tick, so its network is one of the hungriest parts of a chip.1
Before CTS, synthesis and placement time the design against an that reaches every sink instantly or with a guessed delay. replaces that fiction with a real network: it clusters the sinks (flip-flop and latch clock pins, macro clock pins, clock-gating cells), inserts and places clock buffers or inverters, and connects them from the clock root so that every sink sees an acceptable , the between related sinks stays small, and the and power stay low.
A tree is needed for three reasons. Fanout: a block can have tens of thousands of sinks. RC delay: a wire’s resistance and capacitance both grow with length, so unbuffered wire delay grows roughly with length squared.17 Transition: a weak driver on a huge load gives a slow edge, which hurts flop timing and burns short-circuit power. After CTS the clocks become , real skew appears in every timing path, and a post-CTS optimization pass fixes the hold violations that skew creates.7
CTS is where the clock stops being an SDC abstraction and becomes the largest, most active net on the die. The tool builds each clock network against per- targets for skew, latency and max transition, then optimizes with propagated clocks: useful skew or concurrent clock and data (CCD) optimization for setup, buffer insertion for hold, and sizing for slew and power. The same step decides how much of each critical launch/capture pair shares a common clock path, which later determines how much credit signoff gives back under derating.
Clock power is the other half of the objective. Each clock node switches every cycle, while data nodes switch only when logic values change. Published processor breakdowns put the clock network above a quarter of chip power, and at 40–44% across several Alpha generations.21 Production tools: Cadence Innovus (CCOpt), Synopsys IC Compiler II and Fusion Compiler, and open-source OpenROAD TritonCTS, which builds a generalized H-tree over k-means sink clusters.8
- Clock network share of chip power, DEC Alpha 21164
- 40%
- Clock trees’ share of power, Motorola MCORE
- 36%
- Extra power of a clock mesh vs. a conventional tree (one estimate)
- +20–40%
- Sinks per leaf clock buffer in this page’s example (illustrative)
- ~20
In: the placed chip, with every flip-flop at a known spot; the clock’s required speed; and a catalog of amplifier parts the tool may use. Out: the same chip with the clock network added, plus a report card listing how evenly the tick arrives, how long it takes to travel, and how much power the network will use.
| Direction | What | Typical format |
|---|---|---|
| In | Placed design (cells at legal locations) | DEF or an OpenDB .odb, plus Verilog netlist |
| In | Cell timing/power, including clock buffers, inverters and ICGs | Liberty .lib |
| In | Technology and cell abstracts, routing layers, NDR definitions | LEF |
| In | Clock definitions, generated clocks, uncertainty, transition | SDC |
| In | CTS settings: buffer list, skew/latency/transition targets, NDR choice | Tool Tcl (e.g. clock_tree_synthesis options) |
| Out | Netlist with clock buffers, placed and legalized | Verilog, DEF/.odb |
| Out | Constraints switched to propagated clocks | SDC |
| Out | Skew, latency, transition, buffer-count and timing reports | Text reports and logs |
Beyond the basic files, CTS consumes multi-mode multi-corner scenarios (which corner the tree is balanced in matters), skew-group and exclude/ignore-pin definitions, and macro clock-pin insertion delays. Liberty can describe a macro’s internal clock tree with min/max_clock_tree_path timing groups, which STA includes in latency and skew reports on request.5 TritonCTS balances macro and register trees by inserting delay buffers and adds dummy loads to even out leaf capacitance; its inserted instances carry the prefixes clkbuf_regs, delaybuf and clkload, which makes them easy to find in the netlist.6 The clock nets often leave CTS with routing guides or pre-routes and NDRs attached, and are routed before signal nets.
Think of the tool working in four moves:
- Group neighbors. It gathers flip-flops that sit close together into small clusters.
- Feed each group. It places a small amplifier, a , next to each cluster.
- Branch back to the source. Groups of buffers are fed by bigger buffers, level by level, until one trunk reaches the clock’s entry point.
- Balance. Branches that would deliver the tick too early get a little extra delay so everyone lines up.
Timing has two deadlines. The data must arrive a little before the tick () and must not change too soon after it (). Skew moves the tick, so it can help one deadline and hurt the other. If the tick reaches a receiving flip-flop late, slow data has more time to arrive, but fast data has a better chance of racing through and spoiling the old value.4 A slow-arriving problem can be fixed by running the clock slower; a racing problem cannot, so engineers watch hold closely.4
The tree also saves power with : special switch cells on the branches stop the tick from reaching parts of the chip that have nothing to do this cycle.10
The build, step by step
- Find the clock networks. Trace each clock from its SDC source through buffers, gates and muxes to every sink, and note where generated clocks start.
- Cluster the sinks. Group nearby sinks so that one leaf buffer can drive each group within the max-transition and max-capacitance limits. TritonCTS uses k-means clustering with a cap on sinks per cluster and cluster diameter.6
- Build the upper tree. Connect the leaf buffers to the root with a branching structure (an H-tree in TritonCTS), inserting buffers wherever a wire segment gets too long to keep the edge sharp.8
- Balance. Add delay buffers, dummy loads or extra wire on fast branches until latencies match within the target.6
- Legalize and optimize. Snap the new cells onto legal placement sites, switch to propagated clocks, then repair setup and hold.7
Ideal vs. propagated clocks
Pre-CTS, the SDC tells STA what to assume: set_clock_latency and set_clock_transition describe the expected tree when analyzing with ideal clocks. After CTS, set_propagated_clock switches STA to calculated gate and interconnect delays through the real network.5
What CTS measures
- Skew. Global skew is the spread of arrival times over all sinks; local skew is the difference between sinks connected by a data path, and only local skew affects timing.3
- Insertion delay (latency). Source-to-sink delay. Lower is better for power and variation.
- Transition. Clock slew at each pin, held under a max-transition limit.
- Duty-cycle distortion. Buffers with unequal rise and fall delays shift the high/low ratio a little at every level, and a 50% clock can drift toward stuck-high through a long chain.9
Setup, hold and the sign of skew
Take a launch flop and a capture flop. With clock arrivals L and C, data delay d (including clock-to-Q) and period T, define skew s = C − L:
- Setup = (T + C − tsetup) − (L + dmax) = T + s − dmax − tsetup
- Hold slack = (L + dmin) − (C + thold) = dmin − s − thold
Positive skew (capture later) adds to setup slack and subtracts the same amount from hold slack; negative skew does the reverse. The period appears only in the setup check, which is why hold violations cannot be fixed by slowing the clock.4
Useful skew
Zero skew is a convenience. Timing only needs each pair of related flops to meet its checks, so a tool can delay the capture clock of a failing setup path, taking time from the next stage, which shares that flop as its launch.2 This is . Modern tools apply it after CTS as concurrent clock and data optimization, adjusting clock latencies and data-path cells together.15
Clock cells
Libraries include dedicated clock buffers and inverters (here CLKBUF_X4, CLKINV_X8) with strong drive and close rise/fall delays, plus integrated clock-gating cells (ICGs). TritonCTS identifies clock buffers by a name substring (default CLKBUF) or a Liberty footprint, and takes an explicit buffer list.6
Topologies
| Topology | Idea | Trade-off |
|---|---|---|
| Recursive symmetric branching, equal path length to every leaf | Zero skew by symmetry; blockages and uneven sinks break it; used for top levels | |
| Balanced (clustered) tree | Cluster sinks, buffer each cluster, balance upward | Lowest power and wire; most exposed to variation on long uncommon branches |
| Fishbone / spine | Wide trunk with perpendicular ribs feeding local trees | Regular and easy to balance across a wide block; more wire than a tree |
| Grid of shorted wires driven by many buffers | Lowest skew and variation; highest power and routing cost | |
| Multisource CTS | H-tree to a sparse mesh, local trees from tap points | Middle ground: much of the mesh’s robustness at near-tree power |
The H-tree’s exact zero skew makes it the usual choice for top-level distribution.3 Toyama estimates a full mesh at 20–40% more power than conventional CTS, with multisource designs much closer to the tree.11
Routing the clock
Clock nets get a : double spacing to reduce coupling to neighbors and crosstalk, often double width to reduce resistance. TritonCTS applies a 2× spacing rule to non-leaf clock nets, by default on the first half of the levels counted from the root.6 Critical trunks may also be shielded with power or ground wires alongside, as on the Alpha 21264’s global clock grid, which costs tracks.2 The trunk usually runs on upper metal layers, which are thicker and less resistive; using coarser lines for global clock distribution is an old and effective technique.2
Clock gating, generated clocks and domains
An ICG contains a latch, so its enable can only change while the clock is inactive and the gated clock never glitches.10 To CTS, an ICG is a node in the tree that must be balanced through, and its enable pin gets a clock-gating setup/hold check (set_clock_gating_check).5 A clock divider is described with create_generated_clock; its flop’s clock pin is a sink of the master tree and its output roots a new tree. Unrelated clocks are separated with set_clock_groups -asynchronous, and signals between them are made safe by synchronizers designed into the RTL, so CTS does not balance across them.5
After the tree: hold fixing
With real skew in place, short paths into late-clocked flops fail hold. OpenROAD’s repair_timing runs after CTS with propagated clocks, repairs setup first and then hold, and by default won’t insert a hold buffer that breaks setup.7
What the targets really mean
Global skew is a proxy. Correct timing depends only on local skew between sequentially adjacent sinks, which is why useful-skew formulations work on local constraints.3 Tools still minimize skew within each because a tight tree is predictable and leaves room for CCD to add deliberate offsets. The CTS outcome that matters at signoff is slack after variation, which depends on three things: latency, the uncommon fraction of each critical pair’s clock paths, and transition.
Variation, derates and CPPR
STA models by derating early and late paths differently, and clock paths are derated too (set_timing_derate -clock -early / -late).5 Flat derates are too pessimistic for deep paths and too optimistic for shallow ones. AOCV indexes the derate by logic depth and distance; POCV gives each cell a sigma (from LVF data in the library) and combines them statistically along the path.14 For a setup check, the launch clock is late and the capture clock early. On the segment both share, that is physically impossible, so credits back the early/late difference on the common path.13 The design consequence: two flops that split at the last buffer lose almost nothing to clock derates, while two that split at the root lose the derate on the whole tree. Shallow, low-latency trees and clustering of tightly connected flops under one branch both buy real slack.
Uncertainty before and after CTS
covers jitter and, before CTS, the skew the tree will add.5 Jitter varies cycle to cycle at one sink, while skew is a spatial difference.4 A typical flow reduces setup uncertainty after CTS to jitter plus margin, and deletes the ideal-mode set_clock_latency and set_clock_transition, since they apply only to ideal clocks.5 Leaving the pre-CTS numbers in place double-counts skew and sends the optimizer chasing phantom violations.
Topology trade-offs in practice
An H-tree gives balance at the cost of capacitance: its total wirelength, and therefore power, exceeds a sink-driven buffered tree’s, and symmetric structures add latency.2 A puts branch resistances in parallel and shares the path above the mesh among all loads, so only the short stubs below it see uncorrelated variation.211 Multisource CTS drives a mesh fabric that Toyama describes as one to two orders of magnitude sparser than a full clock mesh, fed by H-tree pre-routes, with conventional subtrees hanging off tap points.11 The Alpha 21264 shows the hybrid style at processor scale: a trunk and X/H-trees fed 16 drivers of a gridded global clock, every grid wire shielded by power or ground, local gated clocks below, and 65 ps of skew measured on silicon.2
Clock gating inside the tree
An ICG’s clock pin sits upstream of the flops it gates, so its clock arrives earlier than theirs. The enable is launched by ordinary flops at full latency and captured at that earlier ICG clock, which shortens the effective window for the gating check. Pushing ICGs toward the root gates more of the tree and saves more power, but tightens enable timing; cloning an ICG to sit near its fanout does the opposite. TritonCTS’s -balance_levels keeps a similar number of levels across non-register cells such as clock gates and inverters.6
Generated clocks, macros and skew groups
A divided clock inherits the master’s latency up to the divider flop plus the divider’s clock-to-Q and its own tree, so a div2 domain talking to its master needs the two trees balanced at the crossing flops. Skew groups express exactly which sink sets must align: include the divider’s fanout and the master flops it exchanges data with, exclude scan-only or asynchronous sinks. Macros bring their own internal latency; TritonCTS inserts delay buffers to balance macro and register trees, with a derate knob to insert only part of the needed delay.6
Hold repair and the useful-skew budget
Post-CTS hold repair is where a bad tree shows. OpenROAD caps hold buffers at 20% of the instance count by default and refuses buffers that create setup violations unless told otherwise.7 Every picosecond of useful skew given to a setup path is a picosecond taken from hold on that pair and from setup on the downstream stage, so CCD engines search for latency adjustments that improve worst and total negative slack together.15
Power and EM
Clock nets are long, heavily loaded and switch every cycle.1 Switching activity also drives RMS (Joule-heating) electromigration: time to failure falls as switching rate rises.16 Clock nets have the highest activity on the die, so clock drivers and trunk wires need EM checks with their true switching rate. The usual fixes are wider NDR wires, extra vias on trunk connections, and splitting large loads across several drivers.
The simulator shows two flip-flops with a calculation between them. Sliders move the moment the clock tick reaches the sender (launch) and the receiver (capture), change how long the calculation takes, and set the tick rate. Two scores light up green or red: “setup” (did the data arrive in time?) and “hold” (did it stay put long enough?).
- Make the calculation slow (about 950 ps with a 1000 ps clock) until setup turns red. Now drag the capture tick later by 60 ps. Setup turns green, and the hold score shrinks by exactly 60 ps.
- Make the calculation very short (about 150 ps) and slide the capture tick much later. Hold turns red. Try slowing the clock: hold stays red, because the hold check never looks at the clock rate.
The model is one launch flop, one combinational path and one capture flop. Setup slack = (period + capture arrival − setup) − (launch arrival + data delay). Hold slack = (launch arrival + data delay) − (capture arrival + hold). The data delay includes clock-to-Q.
- Useful skew fixes setup. Period 1000 ps, launch 300 ps, capture 300 ps, data 950 ps, setup 80 ps, hold 40 ps. Setup slack is −30 ps. Move capture to 360 ps: setup slack becomes +30 ps and hold slack drops from 910 ps to 850 ps.
- Hold ignores the period. Set data to 150 ps with launch and capture at 300 ps: hold slack is +110 ps. Move capture to 430 ps: hold slack is −20 ps. Now change the period from 1000 ps to 2000 ps and watch hold slack stay at −20 ps.
- Find the window. With data 950 ps you needed capture at 330 ps or later. Suppose the same capture flop also receives a 150 ps path from the same launch flop. Set data to 150 ps to stand in for it, and slide capture up from 330 ps until hold fails (above 410 ps). Any capture time between 330 and 410 ps satisfies both paths; outside that window, only a data-path change helps.
Use the sim to reason about margins the slider set doesn’t show directly. Fold uncertainty and derates into the sliders and watch how fast the useful-skew window closes.
- Add uncertainty. Take period 1000 ps, launch 300, capture 360, data 950: setup slack +30 ps. Model 50 ps of setup uncertainty by raising setup from 80 to 130 ps: slack −20 ps. The skew that fixed the ideal-mode path is not enough with post-CTS margins.
- Add OCV. Derate the data path late by 5% (950 → 998 ps) and the uncommon part of the capture clock early by 5% (if 200 of its 360 ps is uncommon, capture becomes 350 ps). Setup slack falls by about 58 ps, before even counting the launch clock’s own late derate. The CPPR credit applies only to the shared segment.
- Borrowing from the next stage. Suppose the downstream path from the capture flop has data 900 ps and launches at 300 ps into a flop clocked at 300 ps: setup slack +20 ps. After you moved this flop’s clock to 360 ps, that stage launches at 360 ps and its slack is −40 ps. Around any loop of stages the skews cancel, which is the core of Fishburn’s formulation (see Under the hood).
In the sim, a green slack number means the check passes with time to spare; red means it fails by that many picoseconds.
After the tool builds the tree, an engineer opens a few reports and a picture of the chip with the clock network drawn over it. They check five things: the worst gap in tick arrival between connected flip-flops (skew), the longest travel time (latency), whether any tick edge is too slow, how much power the clock network will use, and how many “too fast” (hold) problems appeared. A branch that is far longer than its neighbors usually points to a big block in the way, and the fix is often to move some flip-flops or give the tool a hint about which parts belong together.
Constraints first. This illustrative SDC (times in ns) shows the pre-CTS clock setup and what changes afterward.
# constraints.sdc (illustrative; times in ns)
create_clock -name core_clk -period 1.000 [get_ports clk]
create_clock -name io_clk -period 4.000 [get_ports io_clk_in]
create_generated_clock -name div2_clk -source [get_ports clk] \
-divide_by 2 [get_pins u_clkdiv/q_reg/Q]
# Pre-CTS only: estimates for a tree that does not exist yet
set_clock_latency 0.400 [get_clocks core_clk]
set_clock_transition 0.060 [get_clocks core_clk]
set_clock_uncertainty -setup 0.120 [get_clocks core_clk]
set_clock_uncertainty -hold 0.040 [get_clocks core_clk]
# Board-level delay before the clock reaches the chip pin
set_clock_latency -source 0.150 [get_clocks core_clk]
# Clock-gating cells must see a stable enable around the edge
set_clock_gating_check -setup 0.040 -hold 0.010
# io_clk is unrelated; CDC is handled by synchronizers in RTL
set_clock_groups -asynchronous -group {core_clk div2_clk} -group {io_clk}
# Post-CTS version replaces lines 8-11 with:
# set_propagated_clock [all_clocks]
# set_clock_uncertainty -setup 0.050 [get_clocks core_clk]
# set_clock_uncertainty -hold 0.020 [get_clocks core_clk]- 1L21 GHz clock on the clk port. Rises at 0, falls at 0.5 ns by default.
- 2L4A divide-by-2 clock at the divider flop’s output. The flop’s CK pin is a sink of core_clk’s tree; its Q roots a new tree.
- 3L8Ideal-mode network latency. Ignored once the clock is propagated; leaving it in is a common bug.
- 4L9Ideal-mode clock slew so pre-CTS flop timing uses a realistic edge.
- 5L10Pre-CTS setup uncertainty = jitter + expected skew + margin. It shrinks after CTS because real skew is then computed.
- 6L14Source latency (outside the chip) stays after CTS; propagated mode only replaces the on-chip network part.
- 7L17Adds setup/hold checks on ICG enable pins against the clock at the gate.
- 8L20Paths between these groups are not timed, and CTS does not need to balance across them.
- 9L23Switch STA to calculated delays through the built tree.
Then the CTS run itself, in illustrative OpenROAD syntax:
# cts.tcl (illustrative, OpenROAD syntax)
read_db results/3_place.odb
read_sdc constraints_cts.sdc
# RC per micron for clock and signal wires (clock trunk on upper metal)
set_wire_rc -clock -layer M7
set_wire_rc -signal -layer M3
clock_tree_synthesis \
-root_buf CLKBUF_X16 \
-buf_list {CLKBUF_X4 CLKBUF_X8 CLKBUF_X16} \
-sink_clustering_enable \
-sink_clustering_size 20 \
-sink_clustering_max_diameter 60 \
-balance_levels \
-apply_ndr half \
-repair_clock_nets
set_propagated_clock [all_clocks]
estimate_parasitics -placement
report_cts
report_clock_skew -setup -digits 3
report_clock_skew -hold -digits 3
# Fix setup first, then hold, then re-legalize the new cells
repair_timing -setup
repair_timing -hold -hold_margin 0.010
detailed_placement
report_checks -path_delay max -digits 3
report_checks -path_delay min -digits 3- 1L6set_wire_rc sets the layer whose RC the tool uses to estimate clock wires; the clock is assumed to route on thick upper metal.
- 2L10Root buffer: the strongest cell, driving the trunk.
- 3L11Candidate buffers for the tree. Clock-specific cells, not generic BUF_X* cells.
- 4L12Pre-cluster nearby sinks; each cluster gets a leaf buffer that becomes an H-tree endpoint.
- 5L13At most 20 sinks per leaf buffer, within a 60 µm diameter.
- 6L15Keep a similar number of levels through clock gates and inverters.
- 7L162× spacing NDR on the first half of the levels counted from the root; leaf nets never get it.
- 8L17Buffer the long wire from the clock pin to the root before latency balancing.
- 9L19TritonCTS already propagates the clocks it builds; repeating it keeps the script correct if the SDC is reloaded.
- 10L26Setup repair before hold, so hold buffers don’t undo setup fixes.
- 11L27Hold repair with 10 ps extra margin. Buffers that would break setup are refused by default.
- 12L28Legalize the buffers inserted by CTS and repair.
Two artifacts tell you whether the tree is good: the CTS summary with a skew report, and the worst setup and hold paths with the clock network expanded. Both are illustrative, in OpenROAD/OpenSTA format, times in ns.
[INFO CTS-0098] Clock net "core_clk"
[INFO CTS-0099] Sinks 18432
[INFO CTS-0100] Leaf buffers 912
[INFO CTS-0101] Average sink wire length 41.27 um
[INFO CTS-0102] Path depth 6 - 8
[INFO CTS-0207] Dummy loads inserted 37
Total number of Clock Roots: 1.
Total number of Buffers Inserted: 1206.
Total number of Clock Subnets: 1206.
Total number of Sinks: 18432.
> report_clock_skew -setup -digits 3
Clock core_clk
0.461 source latency u_lsu/addr_reg[3]/CK ^
-0.389 target latency u_lsu/tag_reg[3]/CK ^
0.050 clock uncertainty
-0.018 CRPR
--------------
0.104 setup skew- 1L2Every clock pin reached by core_clk: flops, latches, ICG clock pins and macro clock pins.
- 2L3About 20 sinks per leaf buffer, matching -sink_clustering_size.
- 3L4Average leaf wire length. Long leaf wires mean slow slew and more coupling; watch for outliers near macros.
- 4L5Depth varies from 6 to 8 levels. Usually extra levels through ICGs or into macro subtrees; each extra level adds uncommon latency.
- 5L6Dummy loads (clkload*) even out leaf capacitance so sibling branches balance.
- 6L14Launch-side latency of the worst pair. This includes source latency and the propagated tree.
- 7L15Capture-side latency, subtracted. Here the capture clock arrives 72 ps before the launch clock, which hurts setup.
- 8L16Setup uncertainty is added to the reported skew so it reflects the effective penalty.
- 9L17CRPR credit for the shared part of the two clock paths reduces the effective skew.
- 10L19Note the sign: this format reports launch minus capture, the opposite of the capture-minus-launch convention used in the sim.
Startpoint: u_alu/acc_reg[12] (rising edge-triggered flip-flop clocked by core_clk)
Endpoint: u_alu/res_reg[31] (rising edge-triggered flip-flop clocked by core_clk)
Path Group: core_clk
Path Type: max
Delay Time Description
---------------------------------------------------------
0.000 0.000 clock core_clk (rise edge)
0.458 0.458 clock network delay (propagated)
0.000 0.458 ^ u_alu/acc_reg[12]/CK (DFF_X1)
0.121 0.579 ^ u_alu/acc_reg[12]/Q (DFF_X1)
0.143 0.722 v u_alu/U812/ZN (NAND2_X2)
0.208 0.930 ^ u_alu/U1033/ZN (AOI22_X1)
0.236 1.166 v u_alu/U1207/ZN (OAI21_X2)
0.198 1.364 ^ u_alu/U1301/ZN (XNOR2_X1)
0.028 1.392 ^ u_alu/res_reg[31]/D (DFF_X1)
1.392 data arrival time
1.000 1.000 clock core_clk (rise edge)
0.443 1.443 clock network delay (propagated)
0.011 1.454 clock reconvergence pessimism
0.000 1.454 ^ u_alu/res_reg[31]/CK (DFF_X1)
-0.050 1.404 clock uncertainty
-0.042 1.362 library setup time
1.362 data required time
---------------------------------------------------------
1.362 data required time
-1.392 data arrival time
---------------------------------------------------------
-0.030 slack (VIOLATED)- 1L9Launch clock latency, late-derated: 0.150 ns source latency + 0.308 ns through the tree.
- 2L11Data path starts at clock-to-Q. Everything from here to line 16 is data delay.
- 3L20Capture clock latency, early-derated: 0.150 + 0.293 ns. The capture clock is 15 ps earlier than launch, so skew costs this path 15 ps.
- 4L21CPPR: the two flops share the root and second-level buffers. Late minus early delay on that shared segment (11 ps) is credited back.
- 5L23Post-CTS setup uncertainty (jitter + margin), no longer a skew estimate.
- 6L24Setup time from the Liberty constraint table, looked up with the actual clock and data slews.
- 7L30−30 ps. Options: CCD delays res_reg[31]’s leaf by ≥30 ps if its downstream paths and hold allow; resize the data path; or regroup the two flops under a deeper common branch to gain CPPR.
- One branch runs late. A big block in the way forces a detour. The engineer spots the outlier in the skew report and moves logic or gives the tool hints.
- Too many “too fast” problems. Once the real tree exists, some short connections get their data too early. Fixing them means adding small delay cells, which cost area and power.
- Noisy neighbors. A busy signal wire running right next to the clock can nudge the tick. Giving the clock extra spacing prevents it.
- Wasted power. A tree with too many strong amplifiers, or without clock gating, burns power every tick even when the chip is idle.
- Pre-CTS constraints left in place. Ideal-mode latency and inflated uncertainty after CTS double-count skew. Caught by reviewing the post-CTS SDC and comparing slack before and after
set_propagated_clock. - Hold explosion. Skew creates thousands of hold violations, and repair adds thousands of buffers. Caught by tracking hold-buffer count; the root cause is often an unbalanced tree or a missing skew-group definition.
- Missing generated clock. Without
create_generated_clock, STA treats the divider output as data, and the downstream flops are either unclocked or wrongly timed. Caught by checks for unclocked sinks. - Max-transition violations at the leaves. Clusters too large or too spread out. Caught in the CTS summary and fixed by smaller cluster size or diameter.
- Crosstalk on clock nets. Default-rule clock wires next to busy signals pick up delta-delay that shows up as skew only after routing. Prevented with NDRs and caught by signal-integrity-aware STA.
- Balanced in the wrong corner. A tree balanced at one corner skews at another, because buffer delay and wire delay scale differently with voltage and temperature. Teams choose the balancing corner deliberately and verify skew across all signoff scenarios.
- Long uncommon paths. Critical pairs that split near the root pay the full clock derate with little CPPR credit. Caught by sorting critical paths by uncommon clock latency; fixed by clustering or skew-group constraints.
- ICG enable timing. Gates placed near the root save power but leave their enable paths failing against an early clock. Look for clock-gating check violations, then clone or push ICGs toward their loads.
- Duty-cycle drift. Long chains of mismatched buffers move the duty cycle at every stage, and the minimum-pulse-width check catches it late.9 Use clock cells with matched edges, or inverter pairs.
- Scan mode surprises. Scan chains reconnect flops along paths with little logic, so skew that is harmless in functional mode can cause races in test mode.2 Time shift mode as its own scenario.
- EM on clock drivers. High activity makes clock trunks the first to hit RMS current limits.16 Run EM checks on clock nets with their true switching activity, not a data-net default.
This part covers the algorithms inside the tools. It’s written for the Expert tier.
The problem
Given sink locations and loads, find a routing tree from the source that meets a skew target at minimum cost. Three formulations: zero-skew trees (ZST), bounded-skew trees (BST) when signoff only needs skew under a bound, and useful skew, which constrains only local skews between related sinks.3 Most algorithms split the work into topology generation (which sinks merge with which) and geometric embedding (where the internal nodes go).
H-trees and the method of means and medians
An is exactly zero-skew by symmetry but only for regular sink arrays, so it serves top-level distribution.3 The method of means and medians (Jackson, Srinivasan and Kuh, 1990) handles arbitrary sink positions top-down: split the sinks at the median into two equal halves, connect the center of mass of the whole set to the centers of mass of each half, and recurse with the split direction alternating.183 It balances sink counts and is fast, with no skew guarantee. Bottom-up recursive geometric matching (Kahng, Cong and Robins) pairs nearby subtrees and achieves near-zero pathlength skew in practice, still without a guarantee.18
Tsay’s exact zero skew
Tsay (1991) was the first to guarantee exact zero skew, under the Elmore model. It merges two zero-skew subtrees bottom-up at a tapping point on the wire joining their roots.18 With subtree delays t1, t2, downstream capacitances C1, C2, a connecting wire of length L, and per-unit resistance r and capacitance c, equalizing the Elmore delay from the tapping point to both sides gives the fraction z of the wire on subtree 1’s side:
z = [t2 − t1 + rL(C2 + cL/2)] / [rL(C1 + C2 + cL)]
If 0 ≤ z ≤ 1 the tap lies on the wire. Otherwise one subtree is too slow, and the algorithm elongates (snakes) the wire on the faster side until delays match.19 Buffered trees apply the same merge per stage, since a buffer hides its subtree’s capacitance from upstream.19
Deferred-merge embedding (DME)
Earlier methods fixed each internal node’s location as soon as it was computed. DME, found independently by three groups in 1992, defers that choice. Given a topology, a bottom-up pass computes for each internal node a merging segment: the locus of all points where the two child subtrees can join with minimum added wire and zero skew. Under Manhattan distance these loci are Manhattan arcs (segments at 45°), and merging two arcs yields another arc. A top-down pass then picks the root anywhere on its segment and places each child at the point of its segment nearest the parent.1819 Kahng and Tsao summarize DME as a linear-time algorithm that optimally embeds any given topology with exact zero skew and minimum total wirelength. Its limitation is that it needs the topology as input, so it is paired with topology generators such as matching or greedy merging.183
Bounded-skew trees (BST/DME)
Zero skew wastes wire when signoff tolerates some skew. BST/DME replaces merging segments with merging regions, the loci of feasible embedding points that keep skew within a bound B, then embeds top-down as before. It works under both pathlength and Elmore delay, and sweeping B gives a smooth trade-off between skew and wirelength; at B = 0 it reduces to DME and at unbounded skew it approaches a Steiner tree.20
Elmore delay and its limits
For an RC tree, the to a node sums, over each resistor on the path from the source, that resistance times all capacitance downstream of it. It is the first moment of the impulse response, which makes it additive along a path and cheap enough to evaluate inside merge loops.17 Gupta, Tutuianu and Pileggi showed that it acts as a delay bound for RC trees, even with general input signals.17 It is often inaccurate but faithful: reducing Elmore delay almost always reduces true delay, so it ranks choices well.17 It knows nothing about input slew, nonlinear driver behavior or inductance, and skew is a small difference of two large delays, so computing it needs far more accuracy than computing either delay alone.2 Modern CTS therefore builds with characterized lookup tables. TritonCTS characterizes buffered wire segments on the fly from Liberty data and selects segments from that table.6
Buffer insertion: van Ginneken’s dynamic program
Van Ginneken (1990) finds optimal buffer positions on a fixed tree under Elmore wire delay and a linear buffer model. Working from sinks to source, each candidate solution at a node is a pair (Q, C): required arrival time and downstream capacitance. At each legal buffer position the algorithm also considers a buffered option, merges candidate lists at branch points, and discards dominated candidates, those with more load and less required time than another. Run time is O(n²) in candidate positions. Lillis, Cheng and Lin extended it to b buffer types in O(b²n²), and Li and Shi reduced that to O(bn²) by showing the useful candidates lie on a convex hull in the (Q, C) plane.21 For clocks the objective changes from maximum slack to balance and power, but the candidate-and-prune structure carries over. TritonCTS uses dynamic programming over its generalized H-tree to choose the minimum-power topology and buffering that meets latency and skew targets, with capacitated k-means for sink clustering.8
Clock mesh analysis
A mesh breaks the assumptions STA delay calculation relies on: many drivers short together, paths reconverge, and the skew between mesh-buffer inputs feeds into the output. SPICE handles this accurately but slowly, and simplified models miss slew and input-skew effects, which is why research on fast mesh timing (including learned models trained on SPICE) continues.12 Design flows typically simulate the mesh and its drivers in SPICE at a few corners, then annotate the resulting stub-input arrivals and slews back into STA for the local trees below.
Useful-skew scheduling as an LP
Fishburn (1990) posed skew scheduling as a linear program. Assign each register i a latency xi. For every path i → j with delays dmin and dmax:
- setup: xi + dmax + tsetup ≤ xj + T
- hold: xi + dmin ≥ xj + thold
Minimize T, or for a fixed T maximize the smallest margin across all constraints. Fishburn named the two hazards zero-clocking (setup) and double-clocking (hold), and on a 4-bit ripple-carry adder with accumulator cut the minimum period from 9.5 ns with zero skew to 7.5 ns with scheduled skew.2 Every constraint bounds a difference xj − xi, so for fixed T the system is a set of difference constraints on a constraint graph: feasible exactly when the graph has no negative cycle, which Bellman–Ford checks. Summing setup constraints around any cycle of k registers cancels every x, leaving T ≥ (Σ(dmax + tsetup)) / k. Skew can redistribute time along a loop but cannot shrink the loop’s average. Tools solve incremental, physically constrained versions of this problem after CTS, because each latency change must be realized with real buffers on a real tree.15
Q1Period 1000 ps, launch clock arrival 300 ps, capture clock arrival 340 ps, data-path delay (including clock-to-Q) 950 ps, setup 80 ps. What is the setup slack?
Q2After CTS, a short path has −25 ps hold slack. Why won’t lowering the clock frequency fix it?
Q3What changes when you run set_propagated_clock after CTS?
Q4Why are clock nets often routed with a double-spacing non-default rule?
Sources
- Power-Aware Clock Tree PlanningClock nets are long, heavily loaded and switch constantly; Alpha 21164 clock network at 40% of chip power, Motorola MCORE clock trees at 36%.
- Clock Distribution Networks in Synchronous Digital Integrated CircuitsSkew vs. min/max path constraints, localized (useful) skew, H-tree vs. buffered tree, mesh, clock power, Fishburn’s LP skew scheduling.
- VLSI Physical Design: From Graph Partitioning to Timing Closure, Chapter 7 slides (Specialized Routing)Global vs. local skew, zero/bounded/useful skew formulations, H-tree, MMM, exact zero skew, DME, buffering under variation.
- Clock skewPositive skew = receiving register clocked later; hold violations cannot be fixed by lengthening the period; skew vs. jitter.
- OpenSTA CommandsSemantics of set_propagated_clock, set_clock_latency, set_clock_transition, set_clock_uncertainty, create_generated_clock, set_clock_groups, set_clock_gating_check, set_timing_derate, report_clock_skew.
- Clock Tree Synthesis (TritonCTS 2.0) documentationclock_tree_synthesis options: H-tree, CKMeans sink clustering, buffer lists, 2X-spacing NDR strategies, dummy loads, delay buffers for macros, report_cts.
- Gate Resizer documentation (repair_timing, repair_clock_nets)repair_timing runs after CTS with propagated clocks; setup before hold; hold buffer cap defaults to 20% of instances.
- Toward an Open-Source Digital Flow: First Learnings from the OpenROAD ProjectTritonCTS: generalized H-tree, dynamic programming for minimum-power topology under latency/skew targets, capacitated k-means sink clustering.
- Duty cycle distortion correction circuitry (US 9,048,823 B2)Unequal rise/fall delays in clock buffers push the duty cycle further from 50% at each stage.
- Clock gatingGating stops flops from switching to save power; integrated clock-gating cells contain a latch for a glitch-free gated clock.
- What’s The Difference Between CTS, Multisource CTS, And Clock Mesh?Mesh: near-zero skew at the fabric, 20–40% more power than conventional CTS; multisource CTS uses tap points on a sparse mesh fed by an H-tree.
- GATMesh: Clock Mesh Timing Analysis using Graph Neural NetworksMesh analysis is hard because of reconvergent paths, multi-source driving and mesh-buffer input skew; SPICE is accurate but slow.
- The TAU 2014 Contest: Removing Pessimism during Timing Analysis (slides)Early/late delay bounds from derating; a signal cannot be both early and late on the common clock path; CPPR credit.
- Parametric on-chip variation: A step towards accurate timing analysisFlat OCV derates vs. AOCV (depth/distance-based derates) vs. POCV (per-cell sigma, LVF).
- DiffCCD: Differentiable Concurrent Clock and Data OptimizationTraditional CTS minimizes skew; concurrent clock and data optimization treats skew as a timing resource post-CTS.
- Electromigration-Aware Architecture for Modern MicroprocessorsPeak, average and RMS (Joule-heating) EM; RMS-EM time to failure falls as switching rate rises.
- Elmore delayElmore delay as the first moment of the impulse response; RC-ladder formula; inaccurate but faithful; bound result of Gupta, Tutuianu and Pileggi (1997).
- Planar-DME: A Single-Layer Zero-Skew Clock Tree RouterReviews MMM, matching-based trees, Tsay’s exact zero skew with wire snaking, and the DME algorithm (merging segments, linear time).
- EE695K VLSI Interconnect, Lecture 9: High-Speed Clock RoutingMMM, Tsay’s Elmore-based merge and tapping-point equation, DME, BST/DME, wire sizing for skew.
- Bounded-Skew Clock and Steiner RoutingBST/DME: merging regions generalize DME’s merging segments; smooth trade-off between skew bound and wirelength.
- An O(bn²) Time Algorithm for Optimal Buffer Insertion with b Buffer TypesVan Ginneken’s O(n²) dynamic program, (Q, C) candidates and dominance pruning, Elmore wire model; extension to b buffer types.