What is the difference between fine-grain and coarse-grain clock gating, and what does an integrated clock gating (ICG) cell do that a plain AND gate cannot?
From PDVerse Low-Power Physical Design Mentor Guide, part of the pdVerse Mentor Guide
Short Answer
Fine-grain clock gating stops the clock to a small group of registers that share an enable, usually inserted by synthesis from RTL enable logic. Coarse-grain gating stops the clock at a branch or block root, so the buffers under it stop toggling too. Both use an ICG cell, which latches the enable while the clock is low so the gated clock never glitches, something a bare AND gate cannot guarantee.
Technical Explanation
- Fine-grain gating: synthesis turns a register bank with a load enable into registers fed by a gated clock, typically for banks above a minimum width.
- Coarse-grain gating: one enable at a clock-tree branch or block input stops every register and clock buffer beneath it.
- Coarse gating saves more per gate because the clock buffers stop too; fine gating catches idle cycles that block-level enables miss.
- ICG cell: a level-sensitive latch, transparent while CLK is low, holds EN during the high phase and feeds an AND gate.
- A plain AND gate passes any EN change during the high phase straight through, producing runt pulses that can clock flops falsely.
- ICGs carry a test enable input ORed with EN so scan shift can clock every flop regardless of functional enables.
- Placing ICGs far from their sinks adds enable-path delay; CTS and placement must keep the enable timing check met.
Common Mistake
The Trap: Building a clock gate from an AND gate and a flop, or letting synthesis use one, to save area.
- An enable that changes while CLK is high cuts the pulse short and creates a glitch that clocks some flops and not others.
- The failure is data dependent and hard to catch in simulation, so it tends to show up only on silicon.
Follow-up Question & Model Response
"Where would you place an ICG physically, near the root or near the registers?"
Candidate Model Response: It is a trade-off. An ICG near the registers keeps the enable path short and easy to time, but the clock buffers upstream of it keep toggling. An ICG near the root saves those buffers too, but the enable must travel far and arrive before the clock rises at the ICG. Most flows let CTS decide by cloning or merging ICGs, and you guide it with the power and timing reports. Check the enable setup at the ICG after CTS, because that is where a far-away gate fails first.
Practical Example
Design Scenario: (illustrative) A 32-bit config register in PD_CPU is written once every 10,000 cycles. Synthesis inserts one fine-grain ICG for the bank, cutting its clock power by over 99% on idle cycles. At block level, the whole PD_DSP clock branch sits behind one coarse ICG driven by dsp_active, so when the DSP idles its 400 flops and 30 clock buffers stop together. Both ICGs have TE tied to the scan-enable network so test mode can shift through every flop. Clock tree power in the idle audio mode falls from 12 mW to 3 mW; what remains is the ungated tree above the ICGs plus leakage, and only power gating removes the leakage.
Low-Power & UPF Handbook
Master Low-Power VLSI & Multivoltage Design
Read the complete low-power guide library covering power domains, level shifters, isolation clamps, state retention, and UPF signoff verification.
Offline PDF Bundle
Want all 1109 questions offline?
Get the complete 4-book PDF bundle (PnR, STA, MMMC, Low Power) with a clickable table of contents - no ads, no internet needed.

Continue practising