CMOS switching, delay and power¶
About this chapter Semiconductor devices and electrical logic
In this chapter
- CMOS switching, delay and power
- Prerequisites and scope
- The CMOS inverter as a switching system
- Effective load capacitance
- Charging a capacitive node
- Dynamic switching power
- Numerical example
- Transition energy and supply current
- First-order RC delay
- Rise and fall delay
- Input slew matters
- Short-circuit current
- Leakage power
- Temperature feedback
- Fan-out
- Transistor sizing
- Buffer chains
- Logical effort as an abstraction
- Elmore delay intuition
- Series transistor stacks
- Internal node capacitance
- Glitching and hazards
- Activity factor
- Clock power
- Data gating and operand isolation
- Power gating
- Dynamic voltage and frequency scaling
- Energy per operation
- Power-delay product
- Critical paths
- Process, voltage and temperature variation
- Noise and delay interaction
- Energy conservation and the PDN
- State storage and switching
- CMOS power is not source-code operation count
- ChrisOS source reconciliation
- ChrisArchitectureState
- cpu_run
- gfx_rgb
- Initialization boundary
- State and data structures
- Algorithms and complexity
- Memory ownership
- ABI boundary
- Concurrency
- Failure and recovery
- Security and privilege
- Performance trade-offs
- Validation evidence for this chapter
- Current limitations
- Roadmap boundary
- Revision provenance
Prerequisites and scope¶
This chapter assumes:
- MOS capacitor accumulation, depletion and inversion;
- MOSFET and CMOS inverter operation;
- capacitance and stored energy;
- resistance and RC transients;
- power-delivery impedance and supply droop;
- Boolean logic and propagation delay.
The central abstraction stack is:
Boolean transition
↓
transistor network changes conduction
↓
node capacitances charge or discharge
↓
analog waveform crosses receiver threshold
↓
new logical value becomes valid
Static truth tables describe only the endpoints.
Delay and power require the continuous transition between those endpoints.
The CMOS inverter as a switching system¶
A CMOS inverter has:
- a PMOS pull-up path to V_DD;
- an NMOS pull-down path to ground;
- an output node with effective load capacitance C_L.
Conceptually:
When the output goes low-to-high, the PMOS network supplies charge to C_L.
When the output goes high-to-low, the NMOS network removes charge from C_L.
The ideal Boolean result is NOT(input), but the physical transition takes finite time.
Effective load capacitance¶
C_L is not one literal capacitor.
It can include:
- input capacitance of driven gates;
- transistor drain junction capacitance;
- interconnect capacitance;
- coupling capacitance;
- package or I/O capacitance where relevant.
A first-order lumped model writes:
The decomposition depends on physical design.
At sufficiently high speed or long interconnect, a distributed RC or transmission-line model can be more appropriate than one lumped C_L.
Charging a capacitive node¶
For an ideal capacitor charged from 0 to V_DD:
The final stored energy is:
An ideal voltage source charging through a resistive path delivers:
The other half:
is dissipated in the charging path in the simple RC model.
When the node later discharges to ground, the stored half is dissipated in the pull-down path.
Therefore one complete 0→1→0 cycle draws approximately:
from the supply in the ideal first-order CMOS dynamic model.
Dynamic switching power¶
If a node has probability/activity factor α of making a 0→1 transition per clock opportunity and opportunities occur at frequency f:
Conventions for α differ across texts.
Some define activity in terms of all transitions rather than 0→1 charging events.
Therefore any numeric use of α must state its definition.
The robust dependencies are:
The quadratic voltage term makes supply-voltage reduction especially powerful for dynamic energy.
Numerical example¶
Suppose:
Then:
for that modeled node.
A real chip contains many nodes with different capacitances and activities, so total dynamic power is a sum.
This example is illustrative and is not a power estimate for any processor running ChrisOS.
Transition energy and supply current¶
A low-to-high transition requires charge:
If it occurs in transition time Δt, average charging current magnitude is approximately:
Faster transitions therefore demand larger transient current for the same C and voltage swing.
That current flows through the PDN.
Thus logic activity couples directly to:
- local droop;
- package inductance;
- ground bounce;
- regulator transient response.
CMOS timing and power integrity cannot be separated completely.
First-order RC delay¶
Treat the conducting transistor network as effective resistance R_eq charging or discharging C_L.
For a first-order RC response:
The time to reach 50 percent V_DD is:
A common first-order propagation-delay estimate is therefore:
The proportionality constant depends on the threshold definition and waveform.
Rise and fall delay¶
Pull-up and pull-down paths need not have identical effective resistance.
Define approximately:
Then average propagation delay is often summarized as:
This is a pedagogical RC model.
Modern standard-cell timing is characterized with richer nonlinear models over input slew, output load, voltage, process and temperature.
Input slew matters¶
The input itself does not switch infinitely fast.
A slow input causes pull-up and pull-down devices to spend longer in intermediate conduction states.
Consequences can include:
- larger propagation delay;
- larger short-circuit current;
- greater timing uncertainty;
- different output slew.
Therefore delay is better described as a function:
rather than one fixed number per gate.
Short-circuit current¶
During an input transition, NMOS and PMOS can both conduct simultaneously for a finite interval.
That creates a direct current path:
The resulting short-circuit or crowbar power is additional to ideal capacitive switching power.
It depends on:
- input slew;
- transistor sizing;
- supply voltage;
- threshold voltages;
- output loading.
An infinitely fast ideal input would reduce the overlap interval, but real edges are finite.
Leakage power¶
CMOS is not perfectly static.
Leakage mechanisms include:
- subthreshold current;
- reverse-biased junction leakage;
- gate tunneling;
- gate-induced drain leakage and related device effects in scaled technologies.
A first-order total static-power relation is:
I_leak depends strongly on:
- temperature;
- threshold;
- process;
- device state;
- geometry.
At advanced nodes, leakage can be a substantial part of total power.
Temperature feedback¶
Electrical power becomes heat.
Rising temperature can increase leakage.
That produces a feedback path:
Thermal management and circuit design must keep the operating point stable.
The exact relation is technology-specific.
Fan-out¶
A gate driving more gate inputs sees larger total input capacitance.
If each load contributes approximately C_in:
for fan-out N in a simplified model.
Larger fan-out therefore increases:
- delay;
- charging energy;
- transient current.
A Boolean net can have one logical value and still be physically expensive to distribute.
Transistor sizing¶
Increasing transistor width generally increases drive capability, reducing effective on resistance.
But larger width also increases:
- gate capacitance presented to the preceding stage;
- diffusion capacitance;
- area;
- dynamic energy.
Thus:
Sizing is an optimization across the entire path, not one gate in isolation.
Buffer chains¶
A very small gate should not always drive a huge load directly.
A chain of progressively larger buffers can distribute the capacitance ratio.
Conceptually:
Each stage adds intrinsic delay but reduces the extreme load seen by the preceding stage.
The optimum staging depends on the delay model, parasitics and physical library.
Logical effort as an abstraction¶
Logical effort separates a gate's topology-dependent difficulty from its electrical fan-out.
In a simplified model, stage delay is written conceptually as:
where:
- g is logical effort;
- h is electrical effort or capacitance ratio;
- p is parasitic delay.
This model is useful for reasoning about path sizing.
It is not a transistor-level timing signoff method.
Elmore delay intuition¶
For a distributed RC network, delay depends on where capacitance is located relative to upstream resistance.
The Elmore first-moment approximation can be written conceptually as:
where each capacitance is weighted by resistance common to the source-to-capacitor path.
This explains why long resistive interconnect can dominate gate delay.
Physical timing tools use richer extraction and models, but Elmore delay provides useful intuition.
Series transistor stacks¶
A CMOS NAND pull-down network may place NMOS devices in series.
Series conduction raises effective resistance compared with one device of the same size.
A NOR pull-up can similarly contain series PMOS devices.
Consequences include topology-dependent:
- delay;
- sizing;
- logical effort;
- parasitic capacitance.
A truth table alone does not expose this cost.
Internal node capacitance¶
Complex gates can contain internal diffusion nodes.
Those nodes can charge or discharge even when the output does not make a full rail-to-rail transition.
Internal switching contributes energy and delay.
A gate-level model that counts only output transitions may therefore underestimate physical activity.
Glitching and hazards¶
Combinational paths rarely have exactly equal delay.
Suppose two logically related inputs reach a gate at different times.
The output can briefly change even though its final Boolean value should remain unchanged.
That transient is a glitch.
A glitch can:
- consume dynamic energy;
- propagate downstream;
- reduce timing margin;
- create unwanted pulses if captured.
The final truth table does not describe this switching activity.
Activity factor¶
For power estimation, each node can be assigned an activity factor.
Conceptually:
Then:
Activity depends on:
- workload;
- data statistics;
- clock gating;
- logic topology;
- glitches;
- architectural state transitions.
It cannot generally be inferred from static source code alone.
Clock power¶
Clock networks switch regularly and drive many sequential elements.
They can therefore contribute substantial dynamic power.
A clock-distribution tree contains:
- buffers;
- wire capacitance;
- local clock pins;
- gating structures.
Clock activity is often close to deterministic while enabled.
Clock gating reduces unnecessary switching by preventing parts of the tree or downstream sequential logic from toggling when no state update is required.
Data gating and operand isolation¶
Unnecessary combinational activity can be reduced by preventing irrelevant inputs from toggling internal logic.
Techniques include:
- operand isolation;
- enable-based gating;
- architectural clock gating;
- power gating at larger granularity.
These techniques require correctness conditions.
Saving energy cannot violate state-update semantics.
Power gating¶
Power gating disconnects a block from its supply using sleep transistors or equivalent physical mechanisms.
Potential benefits:
- greatly reduced leakage in an inactive block.
Costs and constraints include:
- wake-up latency;
- state retention or loss;
- inrush current;
- area;
- power-domain isolation;
- sequencing.
Power gating is a physical design feature, not an operating-system assumption unless the hardware exposes a controlled interface.
Dynamic voltage and frequency scaling¶
A simplified dynamic-power model is:
Reducing V decreases dynamic power quadratically.
Reducing f decreases it approximately linearly.
However maximum safe frequency generally depends on voltage because lower voltage reduces transistor overdrive and current drive.
Thus DVFS couples performance and voltage:
The actual voltage-frequency table belongs to the processor/platform.
Energy per operation¶
If an operation causes a set of nodes to switch, a first-order energy model is:
E_operation
≈
Σ_i transitions_i · C_i V_DD²
+
short-circuit energy
+
leakage energy during execution
This is more informative than power alone when comparing operations that finish at different times.
But mapping software instructions to transistor-level node transitions requires microarchitectural and circuit knowledge.
Power-delay product¶
One simple metric is:
which has units of energy.
Another is energy-delay product:
These metrics weight efficiency and performance differently.
No single metric is universally optimal.
Battery life, thermal density, throughput and latency can demand different objectives.
Critical paths¶
A synchronous pipeline clock period must exceed the worst relevant path delay plus timing margins.
A simplified relation is:
The logic term is built from transistor/gate/interconnect delays.
Reducing average power does not automatically improve worst-case timing.
Likewise, a path that rarely toggles can still determine maximum clock frequency.
Process, voltage and temperature variation¶
Gate delay and leakage depend on PVT:
Examples:
- slower device corner increases delay;
- lower supply often increases delay;
- higher temperature can alter mobility and leakage.
Timing and power signoff therefore use characterized operating corners and statistical models.
One nominal RC estimate is educational, not signoff evidence.
Noise and delay interaction¶
Supply droop can reduce transistor drive during a transition.
That increases delay and can alter output slew.
Ground bounce can shift local thresholds.
Crosstalk can either speed or slow a victim transition depending on aggressor direction and timing.
Thus:
are coupled.
Energy conservation and the PDN¶
Every 0→1 charge event draws energy from the supply network.
At chip scale, many simultaneous transitions create a current transient:
The regulator and decoupling hierarchy must support this current while maintaining rail voltage.
This directly connects the previous power-delivery chapter to gate-level activity.
State storage and switching¶
Latches and flip-flops contain feedback nodes and clocked transistor networks.
Their power includes:
- internal clock switching;
- data-dependent internal switching;
- output load;
- leakage.
A register that holds its logical value can still consume clock-tree energy unless the relevant clock is gated.
The later sequential-logic chapters develop the state semantics; this chapter explains the physical cost of transitions.
CMOS power is not source-code operation count¶
A software expression such as:
can compile into a sequence of machine instructions.
The processor may execute those instructions using:
- pipelines;
- renamed physical registers;
- caches;
- branch prediction;
- multiple execution units;
- clock gating;
- speculative activity.
Therefore source-level operators do not map one-to-one to:
- CMOS gates;
- transistor transitions;
- capacitance charged;
- joules consumed.
The abstraction boundary must remain explicit.
ChrisOS source reconciliation¶
Current ChrisOS revision da3df29cb397932c43d32373871fb9380e688ade exposes architectural and software state, not transistor-level power state.
The reviewed sources are:
The concrete symbols are:
These sources establish the upper abstraction boundary.
ChrisArchitectureState¶
ChrisArchitectureState records software-visible architectural quantities such as:
- general-purpose registers;
- RIP;
- RFLAGS;
- control registers;
- segment state;
- selected MSRs;
- XMM state;
- a modeled TSC.
It does not record:
- transistor capacitance;
- gate delay;
- rail current;
- clock-tree activity;
- die temperature;
- leakage.
An architectural state vector is not a physical switching-state vector.
cpu_run¶
The ChrisCPU cpu_run loop:
- fetches bytes;
- decodes an instruction;
- executes modeled instruction semantics;
- updates RIP when appropriate;
- handles modeled interrupts;
- increments step and modeled TSC counters.
The increments:
do not represent physical clock-tree transitions or joules consumed by a real CPU.
They are emulator-model counters.
This distinction is essential when connecting software execution to CMOS energy.
gfx_rgb¶
gfx_rgb packs three 8-bit channels using shifts and OR operations.
It establishes a software representation rule:
It does not establish how many transistor gates a compiler or CPU uses to execute the expression.
Different processors or compiled instruction sequences can realize the same software result with different physical switching activity.
Initialization boundary¶
There is no CMOS power-model initialization in the cited ChrisOS source.
ChrisCPU initializes architectural emulator state.
Graphics initializes buffers and framebuffer metadata.
Neither path initializes:
- transistor capacitance tables;
- standard-cell timing libraries;
- voltage-frequency curves;
- leakage models.
Therefore no transistor-power initialization is inferred.
State and data structures¶
The physical state relevant to CMOS switching includes:
node voltages
node charges
transistor conduction states
supply current
local temperature
clock phase
Those quantities are not stored in ChrisArchitectureState.
ChrisArchitectureState is a software model of ISA-visible state.
The invariant is:
Algorithms and complexity¶
Computing the simple dynamic-power equation for one node is O(1).
Summing a known list of N node activities is O(N).
Detailed circuit simulation can be far more expensive because it solves nonlinear differential/algebraic equations over many devices and time points.
ChrisCPU instruction emulation complexity is a different problem.
Its step loop cannot be used as evidence of transistor simulation complexity.
Memory ownership¶
Current ChrisOS has no owned data structure for physical CMOS power.
The cited emulator owns architectural state and machine-model structures.
Graphics owns framebuffer/backbuffer state.
None of those buffers is a power waveform or gate-level netlist.
A future power-estimation tool would need explicit ownership for:
- event traces;
- capacitance models;
- activity counters;
- voltage/frequency state;
- thermal state.
ABI boundary¶
The ISA and device ABIs expose functional behavior.
They generally do not expose every internal transistor transition.
Even hardware performance counters, when available, are aggregate microarchitectural events rather than direct transistor counts.
Current cited ChrisOS sources define no ABI for:
- measured watts;
- joules per instruction;
- node capacitance;
- rail current;
- transistor leakage.
Concurrency¶
Real silicon switches many nodes concurrently.
ChrisCPU, in the reviewed emulator path, executes its modeled CPU loop according to software control.
Physical electrical concurrency and software-thread concurrency are different concepts.
No lock in the kernel can serialize transistor switching inside the host CPU.
Conversely, multiple software threads can increase physical switching indirectly by increasing workload.
Failure and recovery¶
CMOS-level failure modes include:
| Mechanism | Possible effect |
|---|---|
| excessive path delay | timing violation |
| excessive supply droop | slower transition or logic error |
| high leakage | thermal/power budget violation |
| overheating | throttling, fault or damage |
| excessive electric field | reliability degradation |
| clock instability | sampling failure |
| excessive crosstalk | timing/noise failure |
Software may observe:
- incorrect computation;
- reset;
- machine-check behavior;
- device timeout;
- performance throttling.
Those symptoms are not sufficient to identify one physical mechanism.
Security and privilege¶
Dynamic power and timing can leak information about computation.
Examples of physical-security research include:
- timing side channels;
- power analysis;
- electromagnetic analysis;
- voltage/clock fault injection.
This chapter does not assert a specific ChrisOS exploit.
The security lesson is architectural: software-visible behavior can modulate lower-level physical activity even though the software does not directly address transistors.
Performance trade-offs¶
CMOS optimization is multiobjective.
| Decision | Speed effect | Power/energy effect |
|---|---|---|
| increase width | stronger drive | larger capacitance/leakage |
| raise V_DD | faster switching | quadratic dynamic-energy increase |
| raise frequency | more throughput | roughly linear dynamic-power increase |
| add buffering | reduces extreme fan-out delay | extra internal capacitance |
| clock gate | little active-path effect when enabled | lowers idle switching |
| power gate | wake-up penalty | reduces leakage |
| slower edge | can reduce noise/short-circuit interactions | may increase delay |
The optimum depends on workload, physical library and product constraints.
Validation evidence for this chapter¶
The deterministic checker validates illustrative first-order relationships:
stored capacitor energy:
E = 1/2 C V²
full charge-discharge supply energy:
E_cycle = C V²
dynamic power:
P = α C V² f
charge:
Q = C V
average transition current:
I = ΔQ/Δt
RC 50-percent delay:
t_50 = ln(2) R C
fan-out capacitance:
C_load = N C_in + C_wire
PDN coupling:
ΔV = Z_PDN ΔI
The source-contract portion checks the current ChrisOS symbols and explicitly verifies that architecture/emulator sources do not claim transistor-level power state.
The checker is not silicon power characterization.
Current limitations¶
This chapter does not provide:
- a foundry standard-cell library;
- SPICE netlists;
- extracted parasitic networks;
- BSIM parameters;
- measured host-CPU power;
- hardware performance-counter energy calibration;
- processor DVFS tables;
- transistor-level ChrisCPU simulation;
- timing signoff.
Those require implementation/process/hardware evidence beyond the current repository.
Roadmap boundary¶
The conceptual progression is:
MOS electrostatics
↓
MOSFET conduction
↓
CMOS pull-up/pull-down logic
↓
node capacitance and RC delay
↓
dynamic + short-circuit + leakage power
↓
fan-out and path timing
↓
logic thresholds and noise margins
↓
sequential timing and clocking
The next missing curriculum chapter, logic-levels-noise-margins, formalizes logic-family voltage thresholds, input/output loading and restoration on top of the device-level switching behavior developed here.
Revision provenance¶
Implementation-facing claims were reconciled against ChrisOS main revision da3df29cb397932c43d32373871fb9380e688ade.
The exact reviewed sources are chrisvm/chris_arch.h, chrisvm/cpu/emulator/chriscpu.c and kernel/gfx/graphics.c, with symbols ChrisArchitectureState, cpu_run and gfx_rgb.
Physical theory was cross-checked against MIT 6.012 material on MOS capacitors, CMOS inverter behavior and CMOS scaling. SI quantities follow the BIPM SI Brochure 9th edition version 4.01.
No mapping from ChrisCPU step count or kernel source operators to real transistor energy is claimed.