1. Introduction
Minimizing the number of CNOT gates is a central problem in exact quantum circuit synthesis and reversible-circuit optimization [
1,
2,
3]. In most architectures, one-qubit gates are far cheaper and more reliable than two-qubit entangling gates, so the CNOT count is a natural and practically meaningful cost measure. The two-qubit SWAP gate is the standard benchmark: the standard optimal construction uses three CNOT gates, and three are necessary in the exact two-qubit synthesis theory [
4,
5]. For larger systems, upper bounds are easy to obtain from explicit constructions, but matching lower bounds are often subtle once arbitrary one-qubit gates are allowed.
Wire permutations are also a basic data-movement primitive. They appear when logical qubits must be routed to physical locations, when swap networks rearrange registers, and when permutation subroutines are used inside larger quantum algorithms. An exact lower bound for these gates therefore separates unavoidable two-qubit cost from artefacts of a particular compiler or circuit decomposition.
This paper develops a binary-shadow framework for exact CNOT counting in the CNOT+local model and uses it to treat the general
n-qubit problem. The target gate is the cyclic wire permutation
For
, this is the ordinary SWAP and costs 3 CNOT gates. For
, the cost is 6. The natural question is whether there is a closed formula for general
n, and more generally, for arbitrary wire permutations.
The binary-shadow method shows that the quantum problem is equivalent to a classical one. Working in the Heisenberg picture, we allow the local Pauli
Z axis on each qubit to rotate under the surrounding one-qubit gates. After this change in viewpoint, each CNOT contributes one elementary transvection over
, so every circuit with
k CNOT gates has a binary shadow that is a product of
k transvections. For a wire permutation, the shadow is rigid: it must coincide with the associated permutation matrix. Consequently, for every wire permutation
, one has
where
is the permutation matrix of
and
is its transvection length over
.
Thus, the quantum question becomes the classical question of computing the transvection length of permutation matrices. This places the problem in the theory of linear reversible circuits, where CNOT-only synthesis is Gaussian elimination over
, and optimization is governed by word length with respect to elementary transvections [
3,
6,
7]. The decisive concept here is the graph-theoretic link-middle-cut decomposition introduced by Bu, Fan, and Joo in their study of exact CNOT synthesis for linear reversible circuits [
8]. Their result shows that an
m-cycle permutation matrix requires exactly
CNOT gates in the CNOT-only model. Related permutation-circuit realizations have also been studied recently from a group-theoretic quantum-circuit perspective [
9]. Combined with the binary-shadow reduction, this immediately suggests a complete answer for cyclic SWAP gates. In fact, by adding a short induction on the number of cycles, one obtains a formula for every permutation matrix.
The main result of this paper is the following.
Theorem 1. Let , and let denote the number of disjoint cycles in its cycle decomposition, including fixed points. Then,In particular, if is the cyclic SWAP gate from (1), then Comparison with Existing Work and Novelty
The CNOT-only part of (
3) is not claimed as new. Bataille studied the algebraic structure and optimization of CNOT circuits and, in compatible notation, reports the cost
for a permutation/cycling with
p components [
10]. Bu, Fan, and Joo later proved, using the link-middle-cut method, that
n-cycle matrices require
CNOT gates and also generalized the statement to permutation circuits [
8]. Liu et al. studied related realizations of permutation groups by CNOT circuits, including exact results for the three-qubit cycle and upper-bound constructions for larger cases [
9]. Recent work on sparse amplitude permutation gates further shows that permutation-type primitives remain relevant in state preparation and circuit decomposition [
11].
The contribution of the present paper is different: it proves that these CNOT-only lower bounds remain valid even in the larger CNOT+local model, where arbitrary one-qubit gates may be inserted before, after, and between the CNOT gates for free. A priori, those local gates can rotate the Pauli axes and might conceivably create a shorter entangling-gate implementation. The binary-shadow rigidity theorem rules this out for wire permutations. Thus, this paper gives a model-lifting argument from CNOT-only transvection length to exact CNOT complexity with free local gates.
Equation (
4) settles the exact values for the first unresolved cases:
More generally, every time the cyclic shift gains one additional qubit, the optimal CNOT count increases by exactly three.
The proof is deliberately phrased in the Heisenberg picture, following a standard viewpoint in quantum computation and stabilizer-circuit analysis [
12,
13]. Unlike stabilizer simulation, however, the binary shadow used here tracks only the support pattern of rotated local
Z axes; this coarse invariant is weak for general unitaries but becomes exact for wire permutations.
This paper is organized as follows.
Section 2 fixes the notation and records the conventions used throughout.
Section 3 develops binary shadows of CNOT+local circuits.
Section 4 proves the rigidity theorem for wire permutations and the reduction (
2).
Section 5 computes the transvection length of permutation matrices using cycle structure and the classical exact
m-cycle theorem.
Section 6 specializes the formula to
n-qubit cyclic SWAP gates, including explicit optimal factorizations for
and
.
Section 7 discusses applications and implications for circuit compilation and resource accounting.
Section 8 closes with a more explicit conclusion and open directions.
2. Notation and Setup
We work on the Hilbert space
of
n qubits. The computational basis on qubit
r is denoted by
. If
then
will also be regarded as a vector in
. Throughout this paper, all addition in
is modulo 2, and products of gates are read from right to left.
If
, then
denotes the controlled-NOT gate with control qubit
i and target qubit
j. Explicitly,
with all other qubits unchanged. The arrow in
therefore points from the control qubit to the target qubit.
A local unitary is a tensor product
where each
is a one-qubit unitary acting on qubit
r only.
For permutation
, we write
for the corresponding wire permutation:
Thus, output wire
r receives the input state that originally sat on wire
. Equivalently,
where
denotes the Pauli
Z operator acting on qubit
r.
We now introduce the binary objects that will encode the CNOT pattern.
Definition 1. A local axis
on qubit r is an observable of the formfor some one-qubit unitary . Equivalently, is a traceless Hermitian unitary on qubit r; in particular, and the eigenvalues of are . In Bloch-sphere language, a local axis is simply a rotated Pauli Z direction. Allowing these axes to vary is what lets the one-qubit gates be absorbed into the formalism.
More explicitly, if
are the Pauli matrices, then every local axis can be written as
so it is a measurement direction on the Bloch sphere obtained from the standard
Z axis by a one-qubit change in basis. The word “local” means that the operator acts on one tensor factor only; the word “axis” refers to the corresponding Bloch-sphere direction.
Definition 2. Let be a family of local axes, and let . We writeIf, for example, , then . Because the factors in (
8) act on different qubits, they commute, and one has
For a bit vector
, we write
for its support.
Definition 3. For , let be the elementary transvection defined bywhere is the standard basis of . The notation in (
9) is chosen to mirror the circuit notation: the CNOT gate
contributes the transvection
on the Heisenberg-side labels.
For clarity, is the only elementary transvection notation used below. When a two-wire swap is expanded, it will be written as the three directed transvections , not as an undefined bidirectional symbol.
Definition 4. For , the associated permutation matrix is defined byWe write for the minimal number of elementary transvections needed to factor . If
has disjoint cycle decomposition
we write
for the number of cycles, counting fixed points as 1-cycles. If
is the length of
, then
Finally, we fix the specific cyclic shift permutation used throughout this paper.
Definition 5. Let be the cycle defined byThe corresponding wire permutation will be denoted by and called the n-qubit cyclic SWAP gate. Thus, The matrix
is the standard cyclic permutation matrix
3. Binary Shadows of CNOT+Local Circuits
The construction below is a support-level version of the Heisenberg evolution of Pauli observables [
12,
13]. The surrounding one-qubit gates are absorbed by allowing the local
Z axes to rotate. We begin with the basic one-CNOT calculation.
Lemma 1. Letbe an n-qubit circuit consisting of one CNOT gate and arbitrary local unitaries before and after it. Then, there exist families of local axes on the input side and on the output side such thatConsequently, Proof. Write
Define local axes on the input and output sides by
For qubits
, the CNOT acts trivially, so
For the two active qubits, we use the standard Heisenberg identities
Hence,
This is precisely the elementary transvection
on exponent vectors. The basis vector
labels the output axis
, and the identity
says that its input-side exponent vector is still
. The basis vector
labels
, and the identity
says that its input-side exponent vector is
. All other basis vectors are fixed. Equation (
15) then follows from multiplicativity: conjugation preserves products, and the observables
multiply by bitwise addition of exponents. □
The previous lemma motivates the central definition.
Definition 6. Let U be an n-qubit unitary. A matrix is called a binary shadow
of U if there exist families of local axes on the input side and on the output side such that A binary shadow need not be unique for a general unitary, because the families of local axes in (
17) may not be unique. The notation “a binary shadow of
U” should therefore be read as “a binary shadow of
U with respect to some chosen input and output local axes.” The axis families
and
may depend on
U and on the particular CNOT+local decomposition used to construct the shadow. The next theorem shows, however, that the existence of a short shadow is automatic for every circuit with few CNOT gates.
Remark 1 (Terminology). The term “binary shadow” is not meant to invoke classical shadow tomography. Classical shadows are randomized measurement data used to estimate properties of quantum states. Here, the object is instead a deterministic matrix over extracted from the Heisenberg evolution of rotated local Z axes. The word “shadow” is used only in the elementary sense that this matrix is a coarse binary projection of the full operator evolution.
Theorem 2. Let U be an n-qubit circuit containing exactly k CNOT gates and any number of one-qubit gates. If the CNOT gates, in temporal order, arethen U has a binary shadowIn particular, every k-CNOT circuit has a binary shadow that is a product of exactly k elementary transvections. Proof. We argue by induction on
k. If
, then
U is local, say
. Choose any family of output axes
and define
Then,
for every
, so the identity matrix is a binary shadow of
U.
Now, suppose the statement is true for circuits with
CNOT gates, and write
where
is the last CNOT block together with the local gates immediately adjacent to it, and
V contains the first
CNOT gates. By Lemma 1, there exist an intermediate family of local axes
and an output family
such that
By the induction hypothesis applied to
V, there exist an input family
and a binary shadow
for which
Combining the two identities gives
Therefore,
is a binary shadow of
U. □
4. Wire Permutations and Transvection Length
For a wire permutation , the binary shadow is in fact unique.
Theorem 3. Let U be an n-qubit circuit built from CNOT gates and one-qubit gates, and suppose that is a wire permutation. Then, every binary shadow of U is equal to . Consequently, Proof. Let
M be any binary shadow of
U. By definition, there exist families of local axes
and
such that
In particular,
where
is a local observable on output qubit
r.
Because
only permutes tensor factors, the operator
is local on input qubit
. On the other hand, the operator
acts nontrivially, exactly on the qubits in
, because every local axis
is non-identity. Therefore,
must contain exactly one index, namely
. Equivalently,
Thus,
.
If
U is realized with
k CNOT gates, then by Theorem 2, it has a binary shadow that is a product of
k transvections. Since every binary shadow is
, the matrix
admits a factorization of length
k. Hence,
, which proves (
19). □
Corollary 1. For every wire permutation ,Equivalently, allowing arbitrary one-qubit gates does not lower the optimal CNOT count of a wire permutation below its optimal CNOT-only count. Proof. The lower bound is exactly Theorem 3. For the reverse inequality, choose a shortest transvection factorization
Consider the CNOT-only circuit
so that the CNOT gates appear in temporal order as
Applying Theorem 2 with no surrounding one-qubit gates shows that
V has binary shadow
relative to the standard
Z axes. Hence,
For computational-basis states, this means that the rth output bit is exactly the th input bit, so V implements the wire permutation . Therefore, . □
Remark 2. Readers who prefer the usual Schrödinger-side bit-update convention may transpose all matrices. Sincetransvection length is unchanged by transposition, and the results above are independent of this convention. Remark 3. For any transposition , one hasso . By Corollary 1, this is exactly the CNOT cost of the corresponding two-wire swap. The exact three-CNOT lower bound for two-qubit SWAP follows from the standard optimal two-qubit synthesis results [4,5]. Hence, Section 4 is the quantum reduction step: every exact CNOT-counting problem for a wire permutation in the CNOT+local model becomes a transvection-length problem for a permutation matrix over
. We now solve that classical problem for arbitrary permutations.
5. Transvection Length of Permutation Matrices
CNOT-only circuits on computational basis states implement invertible linear maps over
, and elementary CNOT gates correspond exactly to elementary transvections. This viewpoint is standard in linear reversible circuit synthesis and optimization [
3,
6,
7]. In this section, we derive the exact formula
for every permutation
. The key external input is the exact length of a single cycle matrix.
5.1. The Cycle Case and the Link-Middle-Cut Method
The relevant classical lower-bound technique is the link-middle-cut (LMC) decomposition of a CNOT synthesis, introduced by Bu, Fan, and Joo [
8]. Their method associates several graphs with a CNOT circuit and classifies each gate as a link, middle, or cut gate. For a cycle matrix, they prove that each of these three classes must contain at least
gates, yielding an exact lower bound of
. The same paper also shows that optimal factorizations of a cycle matrix split into three spanning trees of size
.
This single-cycle theorem is closely aligned with earlier CNOT-only work of Bataille, who analyzed CNOT circuits through their underlying algebraic structure and stated the corresponding
cost for permutation/cycling structures with
p components [
10]. The LMC theorem is used here as the quoted exact transvection-length input, while the binary-shadow argument supplies the additional quantum reduction needed when arbitrary local gates are free.
For the purposes of the present paper, we need only the following theorem.
Theorem 4 (Bu–Fan–Joo)
. Let be an m-cycle, and let be its permutation matrix over . Then, Remark 4. Theorem 4 is the exact classical statement needed to extend the binary-shadow method from the three-qubit cyclic SWAP to all n. In the notation of the introduction, it says precisely that the permutation matrix of the m-qubit cyclic shift has CNOT-only cost .
5.2. Cycle Decomposition and the General Permutation Formula
We now record two elementary lemmas that convert the single-cycle theorem into a formula for arbitrary permutations.
Lemma 2. Let , and let a and b lie in distinct cycles of σ. Then, the permutation has exactly one fewer cycle than σ: Proof. Write the two cycles containing
a and
b as
and let the remaining cycles of
be unchanged. Since the transposition
acts first in the composition
, we compute
and then,
Therefore, the two cycles merge into the single cycle
while all other cycles are unchanged. Hence, the total number of cycles drops by exactly one. □
Proposition 1. For every permutation , Proof. Let
be the disjoint cycle decomposition of
, with
. Each cycle
can be written as a product of
transpositions:
where
. Since each transposition matrix has transvection length 3 by (
22), we obtain
Because the cycles are disjoint, their permutation matrices commute and multiply to
. By subadditivity of transvection length,
This proves (
24). □
We can now prove the exact formula.
Theorem 5. For every permutation , Proof. The upper bound is Proposition 1. For the lower bound, we induct on .
If , then is an n-cycle, and the claim is exactly Theorem 4.
Now, assume the formula is known for all permutations on
n letters with fewer than
cycles, and let
have
c cycles. Suppose for contradiction that
Choose
a and
b from two distinct cycles of
, and set
By Lemma 2,
has
cycles. Moreover, our convention
gives
for all permutations
. Hence, for
,
There is no extra sign or phase issue: these are ordinary 0–1 permutation matrices, now regarded over
. So, by subadditivity of transvection length and (
22), we have
This contradicts the induction hypothesis applied to
, since
. Therefore, no such
can exist, and the lower bound
holds for every
. Together with Proposition 1, this proves (
25). □
Remark 5. Theorem 5 is consistent with the general permutation result proved by Bu, Fan, and Joo [8]. The point of the argument above is that, once the exact single-cycle length is known, the full cycle-structure formula follows by a short induction that fits naturally with the binary-shadow reduction. The quantum consequence is now immediate.
Corollary 2. For every wire permutation , Proof. Combine Corollary 1 with Theorem 5. □
6. Exact CNOT Cost of the -Qubit Cyclic SWAP
We now specialize Corollary 2 to the cyclic SWAP gate
from Definition 5. Since
is a single
n-cycle, we have
, and therefore,
We first record an explicit optimal family of circuits.
Proposition 2. For every ,Consequently,At the level of transvections, Proof. Starting from the right-hand side of (
28), the rightmost swap
moves the state on wire
n one position to the left, the next swap
moves it one step further, and so on. After the full product is applied, the state originally on wire
n has moved to wire 1, while each of the states originally on wires
has shifted one place to the right. This is exactly the action of
.
Each two-wire swap uses three CNOT gates, so (
29) follows immediately. Replacing each swap by its standard three-transvection factorization gives (
30). □
We can now state the exact theorem.
Theorem 6. For every , the n-qubit cyclic SWAP gate satisfiesEquivalently, the cyclic permutation matrix (13) has transvection length Proof. The upper bound is Proposition 2. The lower bound follows from Corollary 2, because . □
The first unresolved cases asked for in the problem statement now follow immediately.
Corollary 3. The exact CNOT costs of the four- and five-qubit cyclic SWAP gates areMore generally, 6.1. Explicit Optimal Factorizations for and
For completeness, we write out the first two new cases explicitly.
Example 1 (The four-qubit cyclic SWAP)
. The four-qubit cyclic SWAP isIts permutation matrix isand by Theorem 6, one hasOne optimal factorization isThus, no exact implementation of can use fewer than nine CNOT gates, even with arbitrary one-qubit gates inserted anywhere in the circuit. Example 2 (The five-qubit cyclic SWAP)
. The five-qubit cyclic SWAP isIts permutation matrix isand Theorem 6 givesAn optimal factorization isHence, twelve CNOT gates are necessary and sufficient. 6.2. First Exact Values
Table 1 lists the first exact values for cyclic SWAP gates.
The table exhibits the simplest qualitative consequence of the theorem: the exact cost grows linearly with slope 3. In particular, extending a cyclic shift from n to wires increases the optimal CNOT count by exactly three.
7. Applications and Implications
The theorem has three immediate implications.
First, it gives a certified benchmark for qubit-routing and register-rearrangement subroutines. If an algorithm or compiler must implement a pure wire permutation in an unconstrained architecture, then CNOT gates are unavoidable. In particular, the adjacent-swap construction for an n-cycle is not merely natural; it is globally optimal even after arbitrary one-qubit gates are allowed at zero cost.
Second, the result clarifies the role of local gates in exact synthesis. In many Clifford, Clifford, and reversible-linear compilation tasks, local gates are treated as comparatively cheap. The binary-shadow rigidity theorem shows that, for wire permutations, those local degrees of freedom cannot be exploited to reduce the two-qubit cost. This provides a reusable proof strategy: if a target family has a rigid support-level Heisenberg shadow, then CNOT-only lower bounds may transfer to the CNOT+local setting.
Third, the result is relevant as a low-level resource-accounting primitive for larger permutation-based constructions. Recent work on sparse amplitude permutation gates uses permutation subroutines in state-preparation problems [
11]. Recent quantum-information and quantum-secure communication protocols also emphasize the need to account carefully for quantum resources, including semi-quantum private comparison protocols [
14,
15,
16] and quantum-based authentication schemes for smart-grid settings [
17]. The present theorem does not optimize those complete protocols; rather, it supplies an exact lower bound for the pure wire-permutation/data-movement component whenever such a component occurs inside a larger construction.
8. Conclusions and Future Work
The binary-shadow method reduces an exact CNOT-counting problem for quantum circuits to a transvection-length problem over . For wire permutations, this reduction is complete: the binary shadow is rigid, and the CNOT complexity in the CNOT+local model is exactly the same as in the CNOT-only model. In this sense, arbitrary one-qubit gates do not create any hidden shortcut for permuting qubit wires.
The present paper extends the three-qubit cyclic SWAP analysis to arbitrary
n. The exact formula
shows that the linear upper bound coming from adjacent swaps is already optimal. More generally, the formula
expresses the exact CNOT cost of a wire permutation entirely in terms of its cycle structure.
Conceptually, three mathematical ideas are doing the work:
The Heisenberg-picture binary shadow, which extracts a matrix over from a CNOT+local circuit;
Rigidity for wire permutations, which forces that shadow to be the permutation matrix itself;
The graph-theoretic link-middle-cut theory of CNOT syntheses, which computes the transvection length of cycle matrices and hence, of arbitrary permutation matrices.
Together, these ideas turn a quantum lower-bound problem into a concrete question about word length in with respect to elementary transvections.
There are several natural directions for further work. First, it would be interesting to identify other classes of target unitaries for which the binary shadow is similarly rigid. Second, one can ask for analogous exact formulas under architectural constraints; for example, when only nearest-neighbour CNOT gates are allowed. Third, the permutation case suggests investigating whether the binary-shadow method can be fused with other classical lower-bound techniques beyond the LMC framework, potentially yielding new exact CNOT counts for larger families of unitaries.
The comparison with the existing CNOT-only literature is now transparent. The formula is already present at the level of CNOT-only permutation circuits; what is proved here is that the same value is the exact CNOT cost in the larger CNOT+local model. Thus, this paper’s main role is not to replace the LMC theory or earlier CNOT-only analyses, but to explain why their permutation lower bounds survive the addition of arbitrary one-qubit gates.
The limitations are also clear. The binary shadow records only the support pattern of rotated local Z axes, so it is generally too coarse to classify arbitrary unitaries. It becomes decisive here because a wire permutation maps each local output observable to a single local input observable, forcing the shadow matrix to be a permutation matrix. Extending the method therefore requires finding other target families with comparable rigidity.
Finally, it would be useful to turn the binary-shadow obstruction into an automated compiler certificate: given a proposed short circuit for a wire permutation, the certificate would show directly that its CNOT count is below the transvection length and hence, impossible.