22-Elec-B6 Integrated Circuit Engineering · May 2017
Question 4 of 6: Body Effect, Charge Sharing and the C 2 MOS Flip-Flop
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
Paper format. National Exams May 2017 — 16-Elec-B6,
Integrated Circuit Engineering. Three hours, closed book, scientific
calculator permitted. Six questions are printed; any five constitute a
complete paper and the total is 100 marks. Formulae and constants are supplied
on the last page of the examination. All six questions are solved here, so
the paper can be used as a complete study resource.
Reference texts.
J. M. Rabaey, A. Chandrakasan and B. Nikolic, Digital Integrated
Circuits: A Design Perspective, 2nd ed. — the standard reference for
this exam code (complementary CMOS, dynamic and NP-domino logic, C2MOS
and TSPC sequential elements, interconnect and supply-noise chapters).
N. Weste and D. Harris, CMOS VLSI Design: A Circuits and Systems
Perspective, 4th ed. — layout extraction, junction capacitance and
adder structures.
S.-M. Kang and Y. Leblebici, CMOS Digital Integrated Circuits: Analysis
and Design, 3rd ed. — MOS capacitance models and transmission-line
effects on chip.
A. S. Sedra and K. C. Smith, Microelectronic Circuits, 8th ed.
— device-level background (body effect, pass-transistor levels).
In Fig. Q4.2 the gate lead of M4 ends at its own input terminal, labelled C, and is not connected to the internal node between M2 and M3. Q4-2 is solved on that reading.
Question 4: Body Effect, Charge Sharing and the C2MOS Flip-Flop (20 marks)
Given. Fig.Q4.1 is a two-input NAND: pMOS devices M1 (gate
$B$) and M2 (gate $A$) in parallel between $V_{DD}$ and $Y$; nMOS devices M3
(gate $A$, drain at $Y$) and M4 (gate $B$, source at ground) in series. All nMOS
bulks are tied to ground and all pMOS bulks to $V_{DD}$. Fig.Q4.2 is a footed
dynamic gate: precharge pMOS M1 and foot nMOS M5 both gated by CLK; the
pull-down network is M2 (gate $A$) in series with M3 (gate $B$), in parallel
with M4 (gate $C$). Fig.Q4.3 is two cascaded clocked inverters: stage 1
transparent when $\text{CLK}=1$, stage 2 transparent when $\text{CLK}=0$.
Find. The body-effect device and its consequence, and a
design change; the dynamic gate's output in both clock phases and the input
pattern that causes charge sharing; the logic at nodes A and B of the
C2MOS cell in each clock phase, and a proof that it tolerates clock
skew.
Approach. Identify, for each circuit, which node is not tied
to a rail — the internal source node in the NAND stack, the internal
pull-down node in the dynamic gate, and the two dynamic storage nodes in the
flip-flop — because every part of this question turns on the behaviour of
those nodes.
Part 1(a) — the device subject to body effect.
Locate the terminal that is not on a rail. M1 and M2 have
their sources at $V_{DD}$ and their bulks at $V_{DD}$, so
$V_{SB}=0$ for both pMOS devices. M4 has its source at ground and its bulk at
ground, so $V_{SB}=0$ as well. Only M3 has its source at the internal node $X$
between the two series nMOS devices, and $X$ is free to sit above ground.
Therefore only M3 is subject to body effect.
State the effect quantitatively. The threshold voltage of
a device with a reverse-biased source-to-bulk junction is
$$V_{TN}=V_{TN0}+\gamma\left(\sqrt{|2\phi_F|+V_{SB}}-\sqrt{|2\phi_F|}\right),$$
where $\gamma$ is the body-effect coefficient. With $V_{SB}=V_X>0$ the
threshold of M3 rises, its gate overdrive $V_{GS}-V_{TN}$ falls, its drain
current falls and its equivalent on-resistance rises. The consequences are a
longer $t_{pHL}$ at $Y$, a slightly higher switching threshold for the gate as
seen from input $A$, and reduced noise margin at the low output level.
Figure 4.1 — the NAND of Fig.Q4.1 with the internal node $X$ marked. Only M3 has a source that can float above ground, so only M3 sees a non-zero $V_{SB}$.
Part 1(b) — the design change.
Apply the input-ordering rule first, because it is free.
Body effect only bites if the internal node $X$ is charged when M3 turns on. If
the earlier signal drives the device nearer ground, that device
discharges $X$ before the later signal arrives, so M3 switches with
$V_{SB}\approx 0$. Since input $B$ leads input $A$, the required assignment is
$B$ on M4 (nearest ground) and $A$ on M3 (nearest the output) — which is
exactly how Fig.Q4.1 is already wired. Confirm this before changing anything:
had the connections been the other way round, $X$ would charge to
$V_{DD}-V_{TN}$ while $A$ alone was high, and M3 would then switch with a
strongly elevated threshold.
Then make the change that actually removes the mechanism.
With the ordering already correct, the remaining modification available to the
designer is to stop tying M3's bulk to ground: place M3 in its own isolated
p-well (a triple-well or deep-n-well process) and connect that well to M3's
source, so $V_{SB}=0$ under every input sequence, including the case where $A$
happens to arrive first:
$$\boxed{\text{tie the bulk of M3 to its own source rather than to ground.}}$$
Note the cheaper fallback. If the process offers only a
common p-substrate, widen M3 relative to M4 to compensate for its raised
threshold, or add a small nMOS from $X$ to ground gated by
$\overline{Y}$ to keep $X$ discharged while the gate is idle. Both cost area;
neither is as clean as the separate well.
Part 2(a) — the dynamic gate's output.
Precharge phase, $\text{CLK}=0$. M1 conducts and M5 is off,
so the pull-down network is disconnected from ground and the output node is
charged to $V_{DD}$ regardless of $A$, $B$ and $C$:
$$Y=V_{DD}\quad\text{(logic 1), independent of the inputs.}$$
No current path from $V_{DD}$ to ground exists, which is why the foot device is
present.
Evaluation phase, $\text{CLK}=1$. M1 is off and M5
conducts, so the output falls whenever the pull-down network conducts. That
network is M2 in series with M3 (conducting when $A\cdot B$) in parallel with M4
(conducting when $C$), so
$$\boxed{Y=\overline{A\cdot B+C}\;=\;\overline{A\cdot B}\cdot\overline{C},}$$
an AND-OR-INVERT (AOI21) gate. For any input pattern that does not satisfy
$A\cdot B+C$, the output holds its precharged value dynamically on the
capacitance of node $Y$.
Part 2(b) — the charge-sharing inputs.
Identify the internal node. Node $X$ between M2 and M3
carries its own parasitic capacitance $C_X$ (the junction capacitances of M2 and
M3 plus their overlap capacitances). It is not precharged, so at the end of a
cycle in which it was discharged it sits at 0 V.
Find the pattern that connects $X$ to $Y$ without discharging
$Y$. M2 conducts when $A=1$, joining $X$ to $Y$. For the output to
merely droop rather than discharge fully, no complete path to ground may exist,
which requires $B=0$ and $C=0$. Checking all eight patterns confirms this is the
only one:
$$\boxed{A=1,\quad B=0,\quad C=0.}$$
If $B=1$ as well, or $C=1$, the node discharges to ground — that is the
intended logic function, not charge sharing. If $A=0$, M2 is off and $X$ is
never connected to $Y$.
Quantify the droop. Charge is conserved across the two
capacitors when M2 turns on, so
$$C_YV_{DD}=(C_Y+C_X)V_Y\quad\Longrightarrow\quad
V_Y=V_{DD}\frac{C_Y}{C_Y+C_X},\qquad
\Delta V=V_{DD}\frac{C_X}{C_Y+C_X}.$$
With representative values $C_Y=20$ fF and $C_X=5$ fF at $V_{DD}=1.2$ V, the
output settles at $V_Y=1.2\times20/25=0.960$ V, a droop of $0.240$ V. That is
comfortably enough to eat the noise margin of a following gate, and is why real
designs precharge the internal nodes as well, or add a keeper.
Figure 4.2 — charge redistribution when $A=1$, $B=0$, $C=0$. The precharged output shares its charge with the undriven internal node, dropping from 1.200 V to 0.960 V for the values shown.
Part 3(a) — the C2MOS node expressions.
Read the enabling conditions off the clock devices. In
stage 1 the pull-up path contains a pMOS gated by $\overline{\text{CLK}}$ and
the pull-down path an nMOS gated by CLK, so both paths are enabled together only
when $\text{CLK}=1$; when $\text{CLK}=0$ both are cut and node A is in a high
impedance state. Stage 2 has the clock phases interchanged, so it is enabled
when $\text{CLK}=0$ and high impedance when $\text{CLK}=1$.
State the two phases. When $\text{CLK}=1$, stage 1 acts as
an ordinary inverter and stage 2 holds:
$$A=\overline{D},\qquad B=B_{\text{previous}}\ \text{(held dynamically on the node capacitance).}$$
When $\text{CLK}=0$, stage 1 holds and stage 2 acts as an inverter:
$$A=A_{\text{held}}=\overline{D}\big|_{\text{at the falling edge}},\qquad
\boxed{B=\overline{A}=D\big|_{\text{at the falling edge}}.}$$
The value that appears at $B$ is therefore the value $D$ had when CLK last fell:
this is a negative edge-triggered D flip-flop, with the master formed by
stage 1 and the slave by stage 2.
Figure 4.3 — timing of the C2MOS cell. Node A tracks $\overline{D}$ while CLK is high and freezes when CLK falls; node B updates only on the falling edge.
Part 3(b) — skew insensitivity. Clock skew means CLK
and $\overline{\text{CLK}}$ are not exact complements: for a short interval they
can both be high (a 1-1 overlap) or both be low (a 0-0 overlap). Race-through
means new data at $D$ reaching node B within the same clock event, and it must be
shown to be impossible in both overlap cases.
Case 1-1 overlap: both clock phases high. Every pMOS gated
by a clock phase is off, so stage 1 reduces to a pull-down-only network and so
does stage 2. Stage 1 can therefore only drive node A low. But stage 2's
pull-down path requires its input device, gated by node A, to conduct —
that is, it requires $A=1$. Since A can only fall, it can never rise to enable
that path during the overlap. A new value of $D$ can propagate at most one stage,
and node B is untouched.
Case 0-0 overlap: both clock phases low. Now every nMOS
gated by a clock phase is off, so both stages reduce to pull-up-only networks.
Stage 1 can only drive node A high; stage 2's pull-up path needs its
input device gated by node A to conduct, which requires $A=0$. Again the
required condition is the opposite of the only motion available, so no path from
$D$ to node B exists.
Conclude. In both overlap cases the two stages are reduced
to networks of the same polarity, and a two-stage inverting chain of one
polarity cannot propagate a signal — each stage passes a transition in the
one direction the next stage cannot accept:
$$\boxed{\text{no race-through path exists for either clock overlap, so the cell is skew tolerant.}}$$
The tolerance holds provided the overlap is shorter than the sum of the
propagation delays of the two stages; beyond that, and in the presence of slow
clock edges, the pseudo-static version with feedback keepers is required.