From 7935ba09dcc1f5787dee3e3cebc7ef5f13002504 Mon Sep 17 00:00:00 2001 From: Ruge Lin Date: Sat, 3 Oct 2026 18:26:46 +0800 Subject: [PATCH 1/2] Prove linear fixed-accuracy Hopf T-depth with unary phase sources --- README.md | 10 +- WORKSPACE.md | 75 +-- docs/GROUPED_PROGRAM_PREFETCH.md | 6 + docs/OPEN_PROBLEM.md | 25 +- docs/README.md | 1 + docs/RELATED_WORK.md | 87 +++- docs/SOURCE_MAP.md | 28 +- docs/UNARY_PHASE_GRADIENT.md | 720 +++++++++++++++++++++++++++++ docs/VERIFICATION.md | 25 + tests/README.md | 1 + tests/test_unary_phase_gradient.py | 293 ++++++++++++ 11 files changed, 1217 insertions(+), 54 deletions(-) create mode 100644 docs/UNARY_PHASE_GRADIENT.md create mode 100644 tests/test_unary_phase_gradient.py diff --git a/README.md b/README.md index 45fb2c2..85223ba 100644 --- a/README.md +++ b/README.md @@ -144,18 +144,18 @@ For $`L=\Theta(n)`$, sufficient $`b=\Theta(n)`$ gives optimal worst-case $`T=\Theta(N)`$ and $`D_T=\Theta(N/n)`$ in one complete real-frame circuit. -At fixed accuracy, [grouped programs](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) +At fixed accuracy, [unary phase-source groups](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) and chunked queries give, with two clean flags and sufficient $`b=\Theta(\sqrt N)`$, one complete real-frame circuit with ```math T=O(\sqrt N),\qquad G=O(N),\qquad -D_T=O\!\left(n\log\log(n+2)\right). +D_T=O(n). ``` -T-count is optimal in worst-case order. Depth optimality and the -high-precision endpoint remain open; the -[variable-precision theorem](docs/OPEN_PROBLEM.md) is unchanged. +Phase-source preparation, reuse, and return are charged. T-count is +optimal in order; depth optimality and the high-precision endpoint +remain open. The [variable-precision theorem](docs/OPEN_PROBLEM.md) is unchanged. Beyond Hopf frames, **literal diagonals and general one-target U(2) multiplexors** attain $`\Theta(\sqrt{NL}+L+NL/b)`$ with one clean diff --git a/WORKSPACE.md b/WORKSPACE.md index 4a07751..38aad56 100644 --- a/WORKSPACE.md +++ b/WORKSPACE.md @@ -3,9 +3,9 @@ This is the entry point when a previous conversation or execution workspace is unavailable. Proofs and decisions live in the repository. -The 2026-10-02 group-selector and source-reuse pass starts from verified main -`9f7906f57ba379bd2cd934d8c625ba9bb06dcca4`, after the grouped-program -fixed-accuracy depth theorem. Check later commits before +The 2026-10-03 unary phase-source pass starts from verified main +`005debe2ccf7a0b2e7cbc032a0f1ae285b83d2ab`, after the incremental +group-selector and common-source audit. Check later commits before continuing. The selected state-based Hopf QBP construction and its bounded-input audit are complete; the [consolidated theorem](docs/STATE_BASED_QBP_THEOREM.md) is their entry point. @@ -25,7 +25,8 @@ and finite checks, without large simulations, QRAM, resets inside a compiler execution, supplied catalysts, or hidden initialized work. 1. Begin active depth work with the - [grouped complete-frame theorem](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy), + [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy), + then the [grouped complete-frame theorem](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy), [chunked dirty indicator](docs/CHUNKED_DIRTY_INDICATOR.md), [conditional geometric source](docs/CONDITIONAL_GEOMETRIC_SOURCE.md), [two-layer obstruction](docs/SHALLOW_SOURCE_OBSTRUCTION.md), @@ -78,7 +79,7 @@ compiler execution, supplied catalysts, or hidden initialized work. | Additional dirty banks | Improve the state preparation T bound while charging the exact coarse circuit; [banked proof](docs/COMPLEX_COARSE_COMPILER.md#8-additional-dirty-banks-improve-fine-state-preparation) | | State-based T-depth | Two complete schedules, one retaining the sharper count at a stronger dirty reservation; [depth proof](docs/STATE_QBP_DEPTH.md) and [fair comparison](docs/QBP_COST_COMPARISON.md#7-state-based-t-depth-comparison) | | Complete-frame T-count and T-depth | Same-circuit bounds at every accuracy; matching in an explicit workspace range, including inverse-polynomial error; [amortized tradeoff](docs/AMORTIZED_DIRTY_LOOKUP.md) | -| Fixed-accuracy large-width depth | Two external flags, sufficient square-root-scale dirty width, optimal-order square-root T-count, and O(n log log n) depth; [grouped program theorem](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) | +| Fixed-accuracy large-width depth | Two external flags, sufficient square-root-scale dirty width, optimal-order square-root T-count, and O(n) depth; [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) | | Complete-frame error accumulation | Sharp ideal-angle stability and finite relative spectra; coherent linear leakage in actual shared-flag source layers; [scoped error audit](docs/HOPF_ERROR_ACCUMULATION.md) | | Filtered complete-frame source | Quadratic radial error on the same two flags, charged native selective phases, and a smaller source-precision cap; [filter proof](docs/HOPF_RADIAL_FILTER.md) | | Conditional precision depth | Logarithmic source/reflection depth using an active zero suffix and two external flags; [conditional source](docs/CONDITIONAL_GEOMETRIC_SOURCE.md). Query/predicate depth remains charged | @@ -152,18 +153,21 @@ The original routed schedule retains an additive $`n^2`$ term. The [dirty-counter hybrid](docs/PARALLEL_DIRTY_LOOKUP.md#6-a-polylogarithmic-depth-indicator-using-dirty-counters) improves fixed-accuracy large-workspace depth to $`O(n\chi(n))`$ while retaining optimal-order T-count. -The [grouped-program theorem](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) +The [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) now improves this fixed-accuracy, sufficient-square-root-width bound to ```math T=O_\eta(\sqrt N),\qquad G=O_\eta(N),\qquad -D_T=O_\eta\!\left(n\log\log(n+2)\right), +D_T=O_\eta(n), \qquad b\ge C_\eta\sqrt N. ``` These are simultaneous bounds on one complete real-frame circuit using -two external clean flags. Early groups store their programs in conditional -logical zeros; chunked dirty indicators control the late-query depth. +two external clean flags. Early groups store one-hot programs and a charged +unary phase source in conditional logical zeros. Bilinear cyclic shifts +have constant T-depth; source preparation and its actual inverse occur +once per group. Chunked dirty indicators and capped source precision +control the entire remaining tail. The worst-case count is optimal in order, but the available depth lower bound remains only Omega(1). Fixed accuracy is essential to this theorem; the general-precision matching interval and endpoint remain unchanged. @@ -277,10 +281,10 @@ claim and its proof obligation. A general compiler is not required to close this selected Hopf-QBP task. The guarded-batch logarithm is now amortized, and the variable-accuracy composition is complete; do not repeat either task. -The next bounded depth target is fixed accuracy at sufficiently large -$`b=\Theta(\sqrt N)`$, retaining optimal-order $`T=\Theta(\sqrt N)`$. -Its available depth bounds are now $`\Omega(1)`$ and -$`O(n\log\log(n+2))`$ by the grouped-program schedule. +The selected fixed-accuracy depth milestone at sufficiently large +$`b=\Theta(\sqrt N)`$ now has $`D_T=O(n)`$, retaining +optimal-order $`T=\Theta(\sqrt N)`$. Its available depth bounds +are $`\Omega(1)`$ and $`O(n)`$; depth optimality remains open. The capped precision allocation has removed the accumulated source widths as a quadratic contribution: sources and suffix predicates now cost $`O(n\log(n+1))`$ depth at fixed L. The retained per-layer routing @@ -448,7 +452,7 @@ leaves, use ell=min(address-half length,2^floor(k/12)). Its count is summable within O(sqrt N), while all final O(log n) queries have O(n) total depth. No unknown dirty program is treated as an initialized cache. -The global theorem uses the unchanged unfiltered additive precision: +The earlier geometric grouped theorem uses the unfiltered additive precision: m=L+4+ceil log2(8n), cutoff R=min(n,256m), and early height g=floor log2(k/(64m)). These groups fit the literal suffix reservation, have O(n/log n) prefetches, and cost O(n log log n) depth internally. @@ -460,7 +464,8 @@ and erase the retained prefix tree in reverse. It has O(g) group depth, fits the existing 16m2^g suffix reservation, and returns its work exactly through source leakage and on arbitrary inactive inputs. Across early groups this contribution is O(n). Source/reflection depth still retains -the log log n factor, so the complete-frame frontier is unchanged. +the log log n factor in that geometric construction. The unary phase +source below gives the current complete-frame frontier. The [common-source audit](docs/GROUPED_PROGRAM_PREFETCH.md#11-a-common-source-identity-and-the-remaining-reflection) gives an exact two-layer conjugation identity, but two conjugated success @@ -470,22 +475,30 @@ by one-use fresh banks also leaves constant rejection unless a later operation mixes that bank's success and failure sectors. These are scoped failures of specified substitutions, not general depth lower bounds. -The next bounded target is a literal implementation of the conjugated -success reflection, or a maintained encoding that updates its monitor -through the programmed word. Aim for O(g+log(m+2)) group depth at -O(m2^g) native count and the stated suffix reservation; a larger charged -reservation needs a revised global allocation. Start with two unequal -legal rows and carry the same rule through a third stage, preserving -literal phases, all rejected action, and exact inactive identity. -An unpriced conjugated reflection or a fresh bank per Q is not a solution. -If no rule meets these obligations, record that there is no selected -construction rather than starting a larger fixture. Keep the high-precision -endpoint separate. Do not repeat the completed selector schedule, -common-source audit, carry pipeline, short echoes, conditional preparation, -grouped program identity, or chunked-query proof; do not apply ideal-angle -stability to unfiltered source leakage. -A matching unrestricted large-width frame-depth lower bound remains -separate. +The [unary phase source](docs/UNARY_PHASE_GRADIENT.md) supplies a different +route. A Karatsuba rank decomposition implements a coherent one-hot cyclic +shift with private conditional work and literal full inactive identity. +The ideal Fourier source supplies signed target phases; its charged native +preparation and actual inverse cost at most twice the preparation error +for the entire group. No geometric success reflection is retained. +The natural suffix cutoff reserves the convolution pool. Early groups +have O(log n) depth and number O(n/log n). The entire remaining tail uses +chunked queries at the original capped source precision, giving O(n) +total depth while preserving O(sqrt N) T-count and O(N) Clifford count. +The six bounded tests cover the rank identity, native guard, arbitrary +source shift, preparation/inversion, unequal rows, and full source-return +error. They do not constitute a scalable native group compiler. + +The next bounded research question is whether this fixed-accuracy route +extends to a width-dependent same-circuit bound of O(N/b²+n) depth with +O(sqrt N+N/b) T-count. That extension is not established here. Start with +a complete shared-work reservation and a summed prefetch schedule; do not +assume the large-width parallel query bound survives reduced dirty width. +A stronger unrestricted depth lower bound and the high-precision endpoint +remain separate. Do not repeat the completed selector, common-source, +unary phase-source, or chunked-query proofs. Any alternative must retain +literal phases, actual inverses, full work return, and charged source +preparation; ideal-angle stability does not bound unfiltered source leakage. The completed modest-width matching theorem does not require solving the high-precision endpoint. For any new component, keep literal phases and actual inverses, declare diff --git a/docs/GROUPED_PROGRAM_PREFETCH.md b/docs/GROUPED_PROGRAM_PREFETCH.md index 7c6ec1b..42159c5 100644 --- a/docs/GROUPED_PROGRAM_PREFETCH.md +++ b/docs/GROUPED_PROGRAM_PREFETCH.md @@ -18,6 +18,12 @@ Both ingredients are needed: retaining logarithmically many old late queries would retain the previous $`O(n\log(n+2))`$ depth allowance. Depth optimality and the high-precision endpoint remain open. +The later [unary phase-source construction](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) +changes the source interface and improves the fixed-accuracy depth to +$`O(n)`$, with the same asymptotic count and external-workspace orders. +The geometric-source schedule proved here remains a valid construction; +its individual conjugated reflections are not resynthesized by that result. + ## 1. Group, program, and local statement Fix a group of heights $`d,\ldots,d+g-1`$, with $`g\ge1`$. diff --git a/docs/OPEN_PROBLEM.md b/docs/OPEN_PROBLEM.md index 94ad3f8..49a3e91 100644 --- a/docs/OPEN_PROBLEM.md +++ b/docs/OPEN_PROBLEM.md @@ -41,6 +41,7 @@ hybrid depth bounds, write | Simultaneous T-count and T-depth | $`T=O(\sqrt{NL}+L\ell_*(n))`$, $`D_T=O(\min\{nL+n^2,L\ell_*(n)+n^3\})`$, $`G=O(NL)`$, at $`a=2`$, $`b\ge C(L+n+7+\sqrt{NL})`$ | Same real-frame circuit, for sufficiently large fixed C; [parallel dirty lookup](PARALLEL_DIRTY_LOOKUP.md); T-depth optimality remains open | | Fixed-accuracy count and depth at modest width | $`T=O(\sqrt N+N/b)`$, $`D_T=O(N/b^2+n\chi(n))`$, $`G=O(N)`$, at $`a=2`$, fixed L, $`b\ge17(L+n+7)`$ | Same complete real-frame circuit; count is optimal in order, and depth is matching for $`b\le\sqrt N/n`$; [amortized dirty lookup](AMORTIZED_DIRTY_LOOKUP.md) | | Variable-accuracy count and depth | $`T=O(\sqrt{NL}+NL/b+nL)`$, $`D_T=O(NL/b^2+nL+n\chi(n))`$, $`G=O(NL)`$, at $`a=2`$, $`L\ge6`$, $`b\ge17(L+n+7)`$ | Same complete real-frame circuit; both are matching when $`b\le\sqrt{NL/(nL+n\chi(n))}`$; [hybrid composition](PARALLEL_DIRTY_LOOKUP.md#every-eligible-width-and-precision) | +| Fixed-accuracy large-width depth | $`T=O_\eta(\sqrt N)`$, $`G=O_\eta(N)`$, $`D_T=O_\eta(n)`$, at $`a=2`$, $`b\ge C_\eta\sqrt N`$ | Same complete real-frame circuit with charged unary source preparation/return; [unary theorem](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy); T-count is optimal in order, depth lower bound remains $`\Omega(1)`$ | Take the best applicable construction. For fixed L, $`a=2`$ and $`b=L+n+7=\Theta(n)`$, the arbitrary-budget matching splice gives @@ -84,8 +85,8 @@ uses $`a=2`$ throughout and respects each sufficient allocation threshold. | Variable L, $`17B_0\le b\le\sqrt{NL/(nL+n\chi(n))}`$ | $`\Omega(NL/b^2)`$ | $`O(NL/b^2)`$ with $`T=\Theta(NL/b)`$ | Matching throughout this interval when nonempty | | $`L=\Theta(n)`$, sufficient $`b=\Theta(n)`$ | $`\Omega(N/n)`$ | $`O(N/n)`$ with $`T=\Theta(N)`$ | Matching at inverse-polynomial error in N for sufficiently large n | | Fixed L, $`2B_0\le b\lt17B_0`$ | $`\Omega(N/n^2)`$ | $`O(N/n)`$ | Earlier schedule remains the proved fallback at this literal reservation | -| Fixed L, sufficiently large $`b=\Theta(\sqrt N)`$ | $`\Omega(1)`$ | $`O(n\log\log(n+2))`$ with $`T=O(\sqrt N)`$ | [Grouped program reuse](GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy); depth lower bound remains unmatched | -| Fixed L, $`b=\Theta(N)`$ | $`\Omega(1)`$ | $`O(n\log\log(n+2))`$ with $`T=O(\sqrt N)`$ | Extra width is not needed by this schedule; depth optimality remains open | +| Fixed L, sufficiently large $`b=\Theta(\sqrt N)`$ | $`\Omega(1)`$ | $`O(n)`$ with $`T=O(\sqrt N)`$ | [Unary phase-source groups](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy); depth lower bound remains unmatched | +| Fixed L, $`b=\Theta(N)`$ | $`\Omega(1)`$ | $`O(n)`$ with $`T=O(\sqrt N)`$ | Extra width is not needed by this schedule; depth optimality remains open | | $`L=N`$, $`b=\Theta(N)`$ | $`\Omega(1)`$ | $`O(N\ell_*(n))`$ | Serial precision cost remains | | Selected endpoint $`L=N,b=B_0`$ | $`\Omega(1)`$ | $`O(N\ell_*(n))`$ from $`D_T\le T`$ | The larger-bank depth theorem does not apply | @@ -101,9 +102,13 @@ The exact dirty-counter indicator, [masked-sum refinement](DIRTY_SUM_COMPRESSION and bilinear-query hybrid now improve the square-root-width upper bound to $`O(n\chi(n))`$. The [grouped-program refinement](GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) -now gives $`O(n\log\log(n+2))`$ at fixed accuracy, with returned +gives $`O(n\log\log(n+2))`$ at fixed accuracy, with returned work and the same optimal-order T-count. It combines cached conditional programs with a summable chunked-indicator budget for the late layers. +The [unary phase-source refinement](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) +now gives $`O(n)`$ depth under the same external-flag and sufficient +dirty-width orders. It charges source preparation/inversion and the +coherent one-hot program interface. The lower bound is still constant in this regime; depth optimality is open. The lower bound does not assume count optimality; the upper circuit also retains optimal-order T-count. No high-precision endpoint improvement follows. @@ -164,15 +169,21 @@ $`16m2^g`$ suffix reservation. They erase each suffix enable before its controls change and retain prefix nodes until the final reverse traversal. Their work returns exactly through source leakage and on arbitrary inactive inputs. The resulting selector contribution is $`O(n)`$; -the source/reflection contribution still determines the displayed frontier. +the source/reflection contribution retains the old geometric bound. The [common-source identity](GROUPED_PROGRAM_PREFETCH.md#11-a-common-source-identity-and-the-remaining-reflection) moves a shared preparation to the group boundaries but retains two conjugated success reflections per stage. A legal exact row rules out a stale success monitor and a one-use source-bank substitution, with constant rejected norm. These are interface restrictions, not depth lower bounds. -The next bounded task is a charged shallow reflection or monitor-update -identity that closes through two unequal rows and a third stage, including -all rejected action. No such improved source construction is yet selected. +The [unary construction](UNARY_PHASE_GRADIENT.md) instead uses a Fourier +eigenstate and constant-T-depth programmed cyclic shifts. Conditional +bilinear work returns exactly on all source inputs, the inactive sector +is literal identity, and source preparation plus actual inversion costs +at most twice its preparation error for the entire group. Its global +allocation proves the displayed $`O(n)`$ upper bound. The next bounded +question is a width-dependent $`O(N/b^2+n)`$ depth extension retaining +$`O(\sqrt N+N/b)`$ T-count; this remains unproved. Its first obligation +is a complete reduced-width prefetch and workspace schedule. The [two-layer obstruction](SHALLOW_SOURCE_OBSTRUCTION.md) separately allows unrestricted Clifford interlayers: the original source at width diff --git a/docs/README.md b/docs/README.md index 778f12f..a227431 100644 --- a/docs/README.md +++ b/docs/README.md @@ -26,6 +26,7 @@ topic has one primary chapter below. | [Two-layer source obstruction](SHALLOW_SOURCE_OBSTRUCTION.md) | Width-independent approximation gaps for full-input sources with arbitrary Clifford interlayers and returned dirty helpers | | [Conditional geometric source](CONDITIONAL_GEOMETRIC_SOURCE.md) | Logarithmic precision depth using the active logical suffix as temporary clean work; lookup and suffix-predicate costs remain separate | | [Grouped program reuse](GROUPED_PROGRAM_PREFETCH.md) | Complete real-frame depth O(n log log n) at fixed accuracy and sufficient square-root dirty width; O(n) selector maintenance and scoped source-reuse audit | +| [Unary phase-source groups](UNARY_PHASE_GRADIENT.md) | Complete real-frame T-depth O(n) at fixed accuracy and sufficient square-root dirty width, retaining optimal-order T-count; charged preparation and exact guarded cyclic shifts | | [Chunked dirty indicator](CHUNKED_DIRTY_INDICATOR.md) | A tunable exact dirty-tree indicator and a summable late-query budget that removes the late routing bottleneck | | [Hopf error accumulation](HOPF_ERROR_ACCUMULATION.md) | Sharp ideal-angle stability, finite relative spectra, and coherent leakage in the actual shared-flag sources; scoped precision boundaries | | [Flag-echo audit](HOPF_FLAG_ECHO.md) | Exact errors of four diagonal Pauli echoes, their generic linear leakage, and an exact equal-mask exception | diff --git a/docs/RELATED_WORK.md b/docs/RELATED_WORK.md index a1caed4..45fbf58 100644 --- a/docs/RELATED_WORK.md +++ b/docs/RELATED_WORK.md @@ -739,7 +739,7 @@ access model. ## 16. Precision depth and workspace assumptions (2 October 2026) The [source-depth audit](SOURCE_T_DEPTH.md) separates a restriction of the -current source implementation from the unrestricted +geometric source implementation from the unrestricted [Hopf T-depth problem](T_DEPTH_COMPILER.md#4-lower-bounds-and-the-remaining-depth-gap). The relevant denominator technique is already present in [Casas et al., *Matchgate synthesis via Clifford matchgates and T gates*, @@ -768,17 +768,20 @@ dirty work. Its bounded-arity elementary-depth lower bound also does not bound T-depth when unrestricted Clifford circuits between T layers are free. [Kim, *Catalytic z-rotations in constant T-depth*, arXiv:2506.15147v3, -Section 3](https://arxiv.org/pdf/2506.15147v3), published in *Quantum* +Section 3](https://arxiv.org/pdf/2506.15147), published in *Quantum* **10**, 2191 (2026), explicitly leaves constant-T-depth rotation using only clean or dirty ancillas open. The depth-three construction assumes a prepared nonstabilizer catalyst; the supplied-catalyst resource cannot be replaced by arbitrary borrowed qubits. Charging catalyst preparation and its initialized workspace is necessary before composition with the -present compiler. +present compiler. Its final note records subsequent depth-two and +measurement-assisted depth-one refinements. The universal-catalyst +comparison in Section 19 below includes Kim–Laakkonen's later construction. These comparisons identify usable techniques and their workspace -conditions. They yield no asymptotic improvement to the unrestricted -complete-frame depth bound in this pass, and no matching lower bound. +conditions. The 2 October source-restriction pass yielded no asymptotic +improvement to the unrestricted complete-frame depth bound, and no +matching lower bound. Section 19 records the later unary-source compiler. The exact source restriction must therefore remain separate from both optimal approximate Hopf T-depth and the constant-clean T-count endpoint. @@ -854,5 +857,75 @@ give fixed-accuracy $`O(n\log\log(n+2))`$ T-depth with optimal-order $`O(\sqrt N)`$ T-count, $`O(N)`$ Clifford count, two external clean flags, and sufficient $`C_\eta\sqrt N`$ dirty workspace. No generic priority claim is inferred. The variable-precision theorem keeps its -previous bounds, and unrestricted large-width depth optimality and the -constant-clean high-precision endpoint remain open. +previous bounds. This geometric-source schedule remains a valid +predecessor to the linear fixed-accuracy bound below; unrestricted +large-width depth optimality and the constant-clean high-precision +endpoint remain open. + +## 19. Unary phase-source reuse and linear T-depth (3 October 2026) + +The [unary phase-source compiler](UNARY_PHASE_GRADIENT.md) uses established +phase kickback. [Jones et al., arXiv:1204.0567, Section 2.1, +Eqs. (2)–(4), and Section 4.1, Fig. 15](https://arxiv.org/pdf/1204.0567) +describe Fourier eigenstates of modular shifts, programmable phases, +and repeated reference reuse. The local source uses a unary encoding, +coherently selects its cyclic shift from a loaded one-hot program, and +prepares/unprepares it inside the active zero suffix. Neither phase +kickback nor shared phase-reference preparation is claimed as new. + +The cyclic convolution uses the standard three-product Karatsuba +identity; see [Iggy van Hoof, arXiv:1910.02849v2, Section 4.2](https://arxiv.org/html/1910.02849v2). +Recursive scalar products give the bilinear rank bound, and reduction +modulo $`X^q-1`$ is linear. The constant T-depth does not come from +van Hoof's space-efficient reversible multiplication schedule. It comes +from the local guarded trilinear-phase construction with private +conditional work and the retained native Toffoli word. Parallel +parity-phase synthesis with initialized ancillas is already explicit in +[Selinger, arXiv:1210.0974v2, Section 2, Eqs. (5)–(6), and +Theorem 4.1](https://arxiv.org/pdf/1210.0974). The local proof must +additionally establish identity on arbitrary inactive inputs and exact +temporary return on every active source state. + +[Kim–Laakkonen, arXiv:2512.24982v1, Theorems 3, 5 and 6, +and Section 5.1](https://arxiv.org/html/2512.24982v1) already give +constant-depth controlled CNOT/Clifford circuits and a universal +logarithmic-size catalyst for rotations. Their catalytic rotation +chooses an angle-dependent CNOT matrix classically; its depth-one +implementation uses measurement-assisted uncomputation. Their charged +preparation uses measured phase estimation, expected repetitions, and +dynamic compilation after selecting the catalyst eigenvalue. This does +not directly supply the unary compiler's coherently loaded angle table, +unitary source boundary pair, or conditional-work return contract. +The present bound does not settle Kim's generic clean/dirty-only +constant-T-depth rotation question: source preparation is charged and +the complete frame has linear, rather than constant, T-depth. + +A recent comparison is [Wu et al., *Shared Phase Arithmetic for Parallel +Quantum Rotations*, arXiv:2609.36574v1, Sections II.C, III.B and +IV.C–E](https://arxiv.org/html/2609.36574v1), submitted 29 September 2026 +and checked here on 3 October. Their Theorem 1 expresses reversible +phase-function evaluation, one binary addition, and decoding; Proposition 1 +gives Clifford-only encoding for disjoint binary supports. The displayed +ripple adder permits measurement/feedforward and has depth linear in +the phase-register width. Preparation is charged separately, and parallel +batches require separate resources. These are useful shared-arithmetic +precedents, not the constant-depth unary-program interface used here. + +The additional result is the complete charged composition: one unitary +source boundary pair per group, coherent one-hot shift selection, +inactive-sector identity, the Hopf angle-stability bound, and an +early/late workspace and query allocation. For every fixed accuracy +$`\eta`$, it gives one complete real-frame circuit with + +```math +D_T=O_\eta(n),\qquad T=O_\eta(\sqrt N),\qquad G=O_\eta(N), +``` + +two external clean flags, and sufficient $`C_\eta\sqrt N`$ dirty +workspace. The initialized-isometry error includes all returned work +and arbitrary dirty-reference entanglement. It uses no supplied phase +state or intermediate measurement. This improves the fixed-accuracy +depth upper bound; it proves neither an unrestricted matching depth +lower bound nor the constant-clean high-precision endpoint. The sources +above identify inherited ingredients and interface distinctions, not +priority for the composite construction. diff --git a/docs/SOURCE_MAP.md b/docs/SOURCE_MAP.md index fa0ec77..75df7f3 100644 --- a/docs/SOURCE_MAP.md +++ b/docs/SOURCE_MAP.md @@ -142,14 +142,17 @@ from the exact clean-workspace size–depth theorem. | F25 | [Ross–Selinger, arXiv:1403.2975v3](https://arxiv.org/abs/1403.2975v3), abstract and synthesis-runtime discussion | distinguishes short native words from efficient search; optimal synthesis uses a factoring oracle, and the efficient expected runtime without it is conditional | computational context only; the bounded-input proof uses F5 word-length existence and guarded exhaustive coarse enumeration, with no factoring oracle or runtime conjecture | | F26 | [Baur–Strassen, *Theoretical Computer Science* 22(3), 317–330 (1983)](https://www.sciencedirect.com/science/article/pii/030439758390110X) | classical reverse differentiation background | the explicit Pauli baseline derives its Hopf reverse recurrence and dyadic error bound directly; reverse differentiation is not claimed as a new algorithmic principle | | F27 | [Casas et al., arXiv:2602.05425v1](https://arxiv.org/html/2602.05425v1#S3.SS2.SSS2), Section III.2.2, Eq. (29) | exact denominator-exponent depth bound for Clifford-matchgate plus T layers | the [source-depth proof](SOURCE_T_DEPTH.md) applies this existing method to two particular sources, permitting signed-permutation Clifford stages including the parity-odd extension; no unrestricted depth lower bound follows | -| F28 | [Vasconcelos, arXiv:2609.34659v1](https://arxiv.org/html/2609.34659v1#S3.SS4.SSS1), Theorem 8; [Kim, arXiv:2506.15147v3](https://arxiv.org/pdf/2506.15147v3), Section 3 | current precision-depth comparisons | the former uses precision-sized clean work, the latter a prepared catalyst and leaves the clean/dirty-only constant-T-depth question open; neither is used as a theorem premise for the Hopf schedules | -| F29 | [Kim–Laakkonen, arXiv:2512.24982v1](https://arxiv.org/pdf/2512.24982v1), Theorems 3 and 5 | constant non-Clifford-depth control of CNOT and Clifford circuits without ancillas | the local controlled-shear lemma is a rank-sensitive specialization with a literal four-T-layer word and explicit Clifford ledger; constant-depth control is inherited | +| F28 | [Vasconcelos, arXiv:2609.34659v1](https://arxiv.org/html/2609.34659v1#S3.SS4.SSS1), Theorem 8; [Kim, arXiv:2506.15147v3](https://arxiv.org/pdf/2506.15147), Sections 2.2, 2.4 and 3 | precision-depth and catalytic-rotation comparisons | the former uses precision-sized clean work; the latter uses a prepared catalyst, leaves the clean/dirty-only constant-depth question open, and records subsequent improvements in its final note; see F29 for the universal-catalyst refinement | +| F29 | [Kim–Laakkonen, arXiv:2512.24982v1](https://arxiv.org/html/2512.24982v1), Theorems 3, 5 and 6; Section 5.1 | constant-depth control of CNOT/Clifford circuits without ancillas, constant-factor depth overhead for controlled Clifford+T, and universal catalytic rotations | the controlled-shear lemma is a rank-sensitive specialization; the catalytic comparison uses classically compiled angle-dependent matrices and measurement-assisted dynamic preparation, not the unary source's coherent one-hot program and charged unitary boundary pair | | F30 | [Boyd, arXiv:2312.00696v2](https://arxiv.org/html/2312.00696v2), Section III and Appendix A | commuting SELECT/QROM groups and Clifford changes of basis for parallel action | its address copies use initialized registers; the local all-dirty shear echo and workspace allocation are proved separately | -| F31 | [Selinger, arXiv:1210.0974v2](https://arxiv.org/html/1210.0974v2), Proposition 5.1, Eqs. (17)–(21) | one-T-layer Pauli-conjugation normal form | the local rational-transfer specialization excludes T-depth one for full-input nonaffine classical permutations with arbitrary returned dirty helpers; initialized-clean-subspace implementations are outside that claim | +| F31 | [Selinger, arXiv:1210.0974v2](https://arxiv.org/pdf/1210.0974), Section 2, Eqs. (5)–(6), Theorem 4.1; Proposition 5.1, Eqs. (17)–(21) | parallel parity-phase synthesis with initialized work, and one-T-layer Pauli-conjugation normal form | parallel phase synthesis is established; the unary-source proof separately gives guarded trilinear phases on conditional work using the retained native Toffoli word; the rational-transfer specialization concerns full-input nonaffine permutations with returned dirty helpers, not initialized-clean isometries | | F32 | [Takahashi–Tani–Kunihiro, arXiv:0910.2530v1](https://arxiv.org/pdf/0910.2530v1), Sections 2.1–2.3 | linear-size, linear-depth exact ripple-carry addition without initialized work | deleting the two gates targeting the arbitrary carry-output wire gives the modular adder used in the signed dirty increment and baseline sum tree; its literal CNOT/Toffoli word is emitted in the bounded checks; later clean-work/fanout constructions are not used | | F33 | [Remaud–Vandaele, arXiv:2501.16802v2](https://arxiv.org/html/2501.16802v2), Lemmas 2/4, Algorithm 3, Theorem 2 | exact helper-free addition via shallow CNOT and Toffoli ladders | truncate the carry output at the abstract ladder level, then synthesize the shorter ladders; applies only to private counters; bounded checks audit the reduced macro, while the optimized ladder-depth bound is imported analytically | | F34 | [Vandaele, arXiv:2603.12917v1](https://arxiv.org/html/2603.12917v1), Section 5, Theorem 4 and Corollary 7 | exact logarithmic-depth increment and controlled increment with one returned dirty helper | supplies the retained round-based compressor and the separate two-dirty-bit read-only increment; temporary control borrowing stays on private supports; the carry-pipeline refinement instead uses linear TTK arithmetic | | F35 | [Aaronson–Gottesman, arXiv:quant-ph/0406196v5](https://arxiv.org/pdf/quant-ph/0406196v5), Section III; [Zhang–Zhang, arXiv:2409.13809v2](https://arxiv.org/html/2409.13809v2#S3.SS1), Theorem III.1, Eqs. (10)–(11) | stabilizer-overlap quantization and Pauli conjugation by one T layer into a Hermitian Clifford | the [two-layer source obstruction](SHALLOW_SOURCE_OBSTRUCTION.md) derives a full-space transfer alphabet and robust source witnesses; initialized-clean isometries and growing frame-depth lower bounds are excluded | +| F36 | [Jones et al., arXiv:1204.0567](https://arxiv.org/pdf/1204.0567), Section 2.1, Eqs. (2)–(4); Section 4.1, Fig. 15 and Eqs. (19)–(20) | Fourier-state phase kickback, programmed shifts, and reusable phase references | the unary source changes the encoding and implements coherent one-hot-selected shifts in conditional Hopf work; neither phase kickback nor reference reuse is new | +| F37 | [Iggy van Hoof, arXiv:1910.02849v2](https://arxiv.org/html/1910.02849v2), Section 4.2; Section 4.3, Theorem 4.1 | standard three-product Karatsuba recursion and subquadratic binary-polynomial multiplication | the unary-source proof uses the bilinear rank recursion followed by linear cyclic reduction; it does not import constant depth from the space-efficient reversible multiplication schedule | +| F38 | [Wu et al., arXiv:2609.36574v1](https://arxiv.org/html/2609.36574v1), Sections II.C, III.B and IV.C–E; Theorem 1 and Proposition 1 | shared phase arithmetic, reversible encoding, and amortized reference preparation | comparison checked 3 October 2026; its binary ripple-adder construction permits measurement/feedforward and has linear-width depth; no coherent unary-program, two-clean Hopf, or constant-depth convolution interface is imported | Standard Pauli linear combinations, reversible arithmetic, and oblivious amplitude amplification are used with their actual preparations and adjoints. @@ -198,6 +201,7 @@ is inherited. | R40 | banked state preparation and fair QBP cost comparison | [banked state corollary](COMPLEX_COARSE_COMPILER.md#8-additional-dirty-banks-improve-fine-state-preparation) reuses R17's exact whole-word dirty-bank query, inherited from F2, in R36/R39's preparation: $`T=O(\sqrt{NL}+L+NL/b+n\sqrt N)`$, $`G=O(NL)`$, at two clean flags, $`L\ge\max\{6,n\}`$, and $`b\ge2(L+n+7)`$; [gauged borrowed baseline](QBP_COST_COMPARISON.md#2-a-gauged-complex-borrowed-frame-baseline) combines the retained interpreter and F24's phase cascade with weighted row precision to give $`T=O(NK/(n+b)+K\sqrt N)`$, $`G=O(NK)`$, exact dirty return, and zero compiler clean work for $`b\ge2n`$; the [task comparison](QBP_COST_COMPARISON.md) retains separate frame/state precisions, literal workspace thresholds, observable costs, sampling, and preprocessing; it compares constructive upper bounds without a literal common-phase-frame or end-to-end gradient optimality claim | | R41 | bounded-input construction and classical comparison | [algebraic residual coefficients](RESIDUAL_TABLE_PREPROCESSING.md) use standard half-phase identities with one shared root, certified rational intervals, and a finite-radius cutoff to preserve R36/R39's error constants without Euler search; [bounded-input audit](BOUNDED_INPUT_QBP.md) combines F5 existence with coarse enumeration, explicit masks and instruction output to prove polynomial construction for the listed grouped/state alternatives; the banked small-system source uses $`P+13`$ dirty wires. Deterministic and term-sampled classical Pauli baselines are charged; no unconditional efficient fine-word search, generic Euler-runtime theorem, or end-to-end quantum advantage is claimed | | R42 | state-based QBP T-depth composition | [depth proof](STATE_QBP_DEPTH.md) composes R21/R23's exact schedules, inherited from F2, with R36/R39's constant number of residual rotations and the actual exact-return coarse interpreter; with $`B_0=P+n+7`$, $`b\ge2B_0`$ gives $`D_T=O(NP/b+P+n^3)`$, $`T,G=O(NP)`$, while $`b\ge16(B_0+\sqrt{NP})`$ gives one circuit with $`T=O(\sqrt{NP}+P+n\sqrt N)`$, $`G=O(NP)`$, and $`D_T=O(P+n^3)`$; both real/complex task streams and oracle depth are charged; no new lookup primitive, depth optimality, total-runtime gain, or general emitter is claimed | +| R43 | fixed-accuracy linear T-depth complete real frame | [unary phase-source proof](UNARY_PHASE_GRADIENT.md), 3 October 2026: coherent one-hot shifts, guarded bilinear work, unitary source preparation/return, and the Hopf angle-stability/group/query allocation give $`D_T=O_\eta(n)`$, $`T=O_\eta(\sqrt N)`$, $`G=O_\eta(N)`$ with two external clean flags and sufficient $`C_\eta\sqrt N`$ dirty work; F5/F8/F31/F36/F37 are attributed ingredients, F29/F38 are comparisons; no supplied catalyst, generic synthesis priority, matching depth lower bound, or high-precision endpoint follows | The [Hopf error audit](HOPF_ERROR_ACCUMULATION.md) derives a sharp ideal-angle stability recurrence and an exact finite relative-spectrum recursion from @@ -282,7 +286,7 @@ native controlled-H, reflection, and amplification ingredients. By itself it changes the precision component only; it is not all-dirty source resynthesis. -The [grouped program construction](GROUPED_PROGRAM_PREFETCH.md) combines +The retained [grouped program construction](GROUPED_PROGRAM_PREFETCH.md) combines that source with exact conditional prefetch, read-only program copies, private conjunction trees, and an internal-enable phase correction. The [chunked dirty indicator](CHUNKED_DIRTY_INDICATOR.md) uses the existing @@ -308,6 +312,22 @@ return remain charged. The legal zero-angle witness excludes only the stated stale-monitor and one-use-bank substitutions. Neither statement imports a new synthesis premise or improves the global depth order. +The [unary phase-source construction](UNARY_PHASE_GRADIENT.md), added +**3 October 2026**, uses F36's phase-kickback identity with a coherent +one-hot program, F37's Karatsuba rank bound, and a separately proved +guarded bilinear circuit in conditionally initialized suffix work. +Parallel non-Clifford synthesis has the precedent F31. The actual source +preparation and inverse are charged once per group; source error is +bounded for the entire group, including rejected components and work +return. The Hopf angle-stability bound and early/late query allocation +then give R43's $`O_\eta(n)`$ T-depth with optimal-order T count at +fixed accuracy. This improves the preceding geometric-source schedule +without giving a shallow implementation of its conjugated reflection. +F29 and F38 give nearby catalytic/shared-arithmetic interfaces, with +different program, preparation, and workspace contracts. No priority +claim is inferred from these comparisons, and the high-precision +endpoint and unrestricted T-depth optimality remain open. + The [consolidated state-based QBP theorem](STATE_BASED_QBP_THEOREM.md) collects R36 and R38–R42 under one input, precision, workspace, and sampling contract. It introduces no additional compiler bound or diff --git a/docs/UNARY_PHASE_GRADIENT.md b/docs/UNARY_PHASE_GRADIENT.md new file mode 100644 index 0000000..e4c9534 --- /dev/null +++ b/docs/UNARY_PHASE_GRADIENT.md @@ -0,0 +1,720 @@ +# A conditional unary phase source for an entire Hopf group + +[Grouped program interface](GROUPED_PROGRAM_PREFETCH.md) · [Exact angle stability](HOPF_ERROR_ACCUMULATION.md) · [Native depth convention](T_DEPTH_COMPILER.md) + +This is a different local construction from geometric-source amplification. +A prepared unary phase state is an eigenvector of every cyclic wire shift. +A bilinear circuit implements a program-selected shift in constant T-depth, +using conditionally initialized work. The same source serves every height +of a group and is unprepared only at its boundary. Its preparation error +is charged twice for the entire group. + +Sections 1–6 prove the local interface. Section 7 supplies the complete-frame +precision, group sizes, early/tail split, and query schedule, giving +fixed-accuracy T-depth $`O_\eta(n)`$ with optimal-order T-count and +square-root-scale dirty width. This construction does not replace the +previous conjugated geometric reflection by a shallow implementation. + +## 1. Local contract + +Use the group registers of [the grouped interface](GROUPED_PROGRAM_PREFETCH.md#1-group-program-and-local-statement): +a preserved external prefix x, local bits $`t_0,\ldots,t_{g-1}`$, +and an outer suffix of length r. The external flag h records that this +outer suffix was zero. Put + +```math +q=2^\ell,\qquad \ell\ge1,\qquad 1\le g\le\ell,\qquad +R=3^\ell,\qquad 0\lt\delta\le1/4, +\qquad \tau=1+\log_2((\ell+1)/\delta). +``` + +For each local height d and prefix $`p\in\{0,1\}^d`$, the loaded +program contains a q-bit one-hot word $`e_{a_{d,p}(x)}`$, where +$`a_{d,p}(x)\in\mathbb Z_q`$. Thus the program width is +$`w=q(2^g-1)`$. It is loaded by an exact arbitrary-input XOR lookup +with output $`hF(x)`$, using returned query helpers disjoint from the +conditional suffix. Loading and its actual inverse are separately priced. + +Let $`G_a`$ be the exact local Hopf group with angles +$`\theta_{d,p}=2\pi a_{d,p}/q`$. There is an absolute C such that + +```math +r\ge C\bigl(q2^g+R+q\log_2 q\bigr) +``` + +is sufficient to implement this group with initialized-isometry error +at most $`2\delta`$, including return of the source and the outer flag. +Its additional resources, beyond the lookup pair and outer-predicate pair, +are + +```math +T=O\!\left(q2^g+gR+q+\ell\tau\right), +``` + +```math +G=O\!\left(q2^g+gqR+q\ell+\ell\tau\right),\qquad +D_T=O(g+\ell+\tau). +``` + +Here G counts elementary Clifford gates; long Clifford networks are not +assigned constant elementary depth. Two external clean flags suffice; +this local construction only uses h. Every other initialized input is +explicitly reserved in the active zero suffix. The construction has exact +identity on the full h-zero sector, including arbitrary suffix helpers +and their reference entanglement. There are no intermediate measurements, +resets, supplied phase states, or uncharged reflections. + +## 2. A guarded bilinear XOR in sixteen T layers + +Consider a fixed bilinear map over $`\mathbb F_2`$ with a rank-R +decomposition + +```math +B(p,z)=\bigoplus_{\rho=1}^{R} +\gamma_\rho\,\alpha_\rho(p)\beta_\rho(z), +``` + +where the alpha and beta functions are linear forms and each +$`\gamma_\rho`$ is an output word. Inputs p and z and output y are +disjoint. Define $`\Gamma_\rho(y)=\gamma_\rho\cdot y`$ over +$`\mathbb F_2`$. + +Apply Hadamards to every y wire. For each rho, use three private zero +leaves and CNOTs to compute +$`\alpha_\rho(p),\beta_\rho(z),\Gamma_\rho(y)`$. +With two more private zero wires compute their three-way AND: +first the AND of the first two leaves, then its AND with the third. +All first-round Toffolis have disjoint triples, as do all second-round +Toffolis. Apply $`\mathrm{CZ}(h,q_\rho)`$ to every final AND wire. +Reverse both AND rounds and all linear-form computations, then reverse +the output Hadamards. Every inverse is the actual gate inverse. + +On h equal to one and zero private work this has the Fourier-basis phase + +```math +(-1)^{\sum_\rho\alpha_\rho(p)\beta_\rho(z)\Gamma_\rho(y)} +=(-1)^{B(p,z)\cdot y}. +``` + +It is therefore exactly +$`(p,z,y)\mapsto(p,z,y\oplus B(p,z))`$, with all private work +returned. This is a basis identity and extends to arbitrary quantum +inputs and references, including arbitrary y. + +On h equal to zero, the central CZ gates are all identity on the entire +Hilbert space. Everything around them cancels as its actual inverse. +This remains true when the private leaves and AND targets are arbitrary; +no copied control is interpreted as clean on that sector. In particular, +the original h is the only control of the central gates. + +The circuit has exactly four Toffolis per rank term. Using the literal +seven-T, four-layer native word from the existing +[batch lemma](T_DEPTH_COMPILER.md#a-shared-control-fredkin-batch-has-at-most-four-t-layers), + +```math +w_{\rm private}=5R,\qquad T\le28R,\qquad D_T\le16. +``` + +For q-bit operands and output, direct CNOT evaluation of every linear +form gives the conservative Clifford bound $`G=O(qR)`$. The private +copies ensure that no simultaneous native Toffolis share controls. + +## 3. Karatsuba convolution and an in-place cyclic shift + +Represent q-bit words as polynomials of degree below q over +$`\mathbb F_2`$, and let B be multiplication modulo $`X^q-1`$. +Its output coefficients are cyclic convolution. A self-contained +rank bound is $`R=3^\ell`$: split both polynomials into lower and +upper halves, + +```math +A=A_0+X^{q/2}A_1,\qquad B=B_0+X^{q/2}B_1. +``` + +Compute the three half-size products +$`P_0=A_0B_0`$, $`P_2=A_1B_1`$, and +$`P_1=(A_0+A_1)(B_0+B_1)`$. Their reconstruction is + +```math +AB=P_0+X^{q/2}(P_1+P_0+P_2)+X^qP_2. +``` + +All additions are linear. Recurse to scalar products and finally reduce +the polynomial modulo $`X^q-1`$, also linearly. Hence every scalar +product is a product of one linear form in each input, and the preceding +bilinear decomposition has $`3^\ell`$ terms. No field extension or +division is needed. + +For the one-hot program $`p=e_a`$, convolution is the wire permutation + +```math +B(e_a,z)=S_a z,\qquad (S_a z)_j=z_{j-a\bmod q}. +``` + +Let $`p^-_j=p_{-j\bmod q}`$, obtained by merely relabeling the input +wires. Reserve a zero q-bit temporary word y. Apply, chronologically, + +1. The guarded bilinear XOR $`y\leftarrow y\oplus B(p,z)`$. +2. The guarded bilinear XOR $`z\leftarrow z\oplus B(p^-,y)`$. +3. Swap each pair $`(z_j,y_j)`$ controlled by the original h. + +On h equal to one, one-hot p, and y equal to zero, the intermediate +words are $`(z,S_a z)`$ and $`(0,S_a z)`$, followed by +$`(S_a z,0)`$. The identity holds for every computational z, so the +same circuit permutes arbitrary quantum source inputs and returns y +exactly. It also holds coherently for superpositions of one-hot programs. + +On h equal to zero, each bilinear XOR is exactly identity on arbitrary +work, and every controlled swap is identity. Thus the whole shift has +the required full inactive identity even for arbitrary, non-one-hot p. +The one-hot promise is needed only on the active initialized sector. + +The existing shared-control Fredkin batch implements the last step +without clean helpers, with at most four T layers and +$`6q+(q\bmod2)`$ T gates, including its literal phase. Reusing the +same $`5R`$ private pool for the two XORs gives + +```math +T_{\rm shift}\le56R+6q+1,\qquad +D_{T,\rm shift}\le36,\qquad +G_{\rm shift}=O(qR). +``` + +The temporary word y and the source word z are separately counted. + +## 4. Prepare and decode one phase source + +Write $`\omega_q=e^{2\pi i/q}`$ and define the unary source + +```math +|\Phi_q\rangle=q^{-1/2}\sum_{j=0}^{q-1}\omega_q^j|e_j\rangle. +``` + +The exact cyclic permutation above satisfies the literal eigenvalue +identity + +```math +S_a|\Phi_q\rangle=\omega_q^{-a}|\Phi_q\rangle. +``` + +It suffices to prepare this source to norm error delta once per group. +Use an ell-bit index register. Hadamards followed by the one-qubit +phases $`\operatorname{diag}(1,\omega_q^{2^k})`$ prepare the binary +phase state. Replace each phase, up to its fixed scalar, by its +determinant-one diagonal rotation and an actual Clifford+T approximation +of error at most $`\delta/\ell`$. The repository's existing +[single-qubit word-length premise](SOURCE_MAP.md#5-fault-tolerant-sources-and-contribution-boundaries) +supplies length $`O(\tau)`$ for each of these ell words. They act on +disjoint qubits and run in parallel. This imports a word-length bound, +not a new efficient classical synthesis-runtime guarantee. + +Decode the binary index as follows. Compute a balanced binary prefix +tree of its literal predicates into private zero wires, using private +control copies at each level. Copy the q leaf predicates into the zero +unary source word. Reverse the entire prefix computation while the +binary index is unchanged. Finally erase each binary index bit by the +CNOT parity of the unary positions whose label has that bit set. + +For every binary basis label j, this maps +$`|j\rangle|0^q\rangle|0^{\rm work}\rangle`$ to +$`|0^\ell\rangle|e_j\rangle|0^{\rm work}\rangle`$. +The prefix computation has $`O(q)`$ Toffolis, $`O(\log q)`$ +T-depth, and $`O(q)`$ initialized work. The final parity erasure has +$`O(q\ell)`$ CNOTs. All of these maps are exact circuits. + +Call the resulting approximate native preparation U. Its initialized +state lies in the unary subspace exactly, even if its synthesized +single-qubit words are not diagonal. By telescoping the ell approximants, +there is a fixed scalar $`\lambda`$, independent of all logical inputs, +such that + +```math +\|U|0\rangle-\lambda|\Phi_q\rangle|0^{\rm work}\rangle\| +\le\delta. +``` + +The boundary inverse will be the actual $`U^\dagger`$. The scalar +lambda cancels in the resulting logical action; it is not discarded as +an input-dependent phase. The preparation and its inverse have + +```math +T=O(q+\ell\tau),\qquad +G=O(q\ell+\ell\tau),\qquad +D_T=O(\ell+\tau). +``` + +## 5. Signed row selection and the literal rotation + +At height d retain the exact prefix selectors +$`\pi_p=[t_0\cdots t_{d-1}=p]`$ and enable +$`u_d=[t_{d+1}\cdots t_{g-1}=0]`$. Use the +[incremental selector and consume-before-change enable schedule](GROUPED_PROGRAM_PREFETCH.md#10-amortized-local-selectors-and-suffix-enables), +which has $`O(g)`$ total T-depth and $`O(2^g+g)`$ count and width. +Its validity requires only that a completed stage changes its current +target and preserves every other local bit, including on source leakage. +The stage below has precisely that property. + +Put $`D=SH`$, so $`DZD^\dagger=Y`$. Apply $`D^\dagger`$ to +the current target, and call its computational value z during this +stage. In a zero q-bit selected-program word p compute + +```math +p_j=[j=0]\oplus u_d[j=0]\oplus +\bigoplus_{v\in\{0,1\}^d}\pi_v u_d +\bigl((1-z)W_{d,v,j}\oplus zW_{d,v,-j}\bigr). +``` + +On the active sector, this is exactly $`e_0`$ if the inner enable +is false, and $`e_{a_{d,v}}`$ or $`e_{-a_{d,v}}`$ for the selected +row when z is zero or one. Each degree-four product is computed with +four private control copies and three private AND targets; its output +is CNOTed into p, then the computation is reversed. All terms run in +parallel within each of their three Toffoli rounds. Open controls use +literal X conjugations on private copies. The first two terms use an X +and a CNOT on $`p_0`$. + +This selected-word computation and its actual inverse have constant +T-depth, $`O(q2^d)`$ T/Clifford count, and $`O(q2^d)`$ private +conditional-zero work. Their Clifford fanout gates may share controls +or targets; all simultaneous non-Clifford gates use private triples. + +Apply the guarded cyclic shift from Section 3 to the common source, then +reverse the selected-word computation, and apply D to the target. The +shift preserves every selected-word control and p itself. Consequently +p and its private work erase exactly, including on an arbitrary source +state. On the ideal phase source, the target diagonal in the z basis is + +```math +\operatorname{diag}(\omega_q^{-a},\omega_q^{a}) +=e^{-i(2\pi a/q)Z}. +``` + +The complete target word is therefore exactly +$`D e^{-i(2\pi a/q)Z}D^\dagger=R_y(2\pi a/q)`$. +There is no half-angle factor or additional branch phase. An inner +disabled stage selects the identity shift, so its two target basis +changes cancel on every source input in the active zero-work sector. + +On h equal to zero the guarded shift is identity on its entire input +space, including arbitrary selected programs and private work. The +selected-word compute/inverse and target basis changes then cancel +exactly. Thus every completed stage has full inactive identity without +assuming clean selector copies there. + +## 6. Group return, error, and resources + +Load the program, prepare the source with U, and perform the g stages +with the retained-selector schedule. Erase the selectors and source +with their stated actual inverses, unload the original program with +the original lookup inverse, and finally reverse the outer predicate. +The original program and external prefix are never modified. + +On h equal to one, every stage is exact on the ideal phase source, +returns its selected program and shift work, and preserves that source. +This holds simultaneously for all logical columns. Let V denote the +exact interior group circuit, including selector computation and cleanup. +Let $`J_F`$ append the valid program $`F(x)`$ and the zero interior +work. Since the group preserves x, this programmed subspace is invariant. +Its ideal-source identity is + +```math +V\bigl(|\Phi_q\rangle\otimes J_F|v\rangle\bigr) +=|\Phi_q\rangle\otimes J_FG_a|v\rangle. +``` + +All stages are unitary on rejected source components; they are neither +discarded nor separately approximated. Comparing the actual prepared +state to $`\lambda\Phi_q`$ before V costs delta. Comparing its +unpreparation to zero after V costs another delta, because the actual +inverse is used. On this valid-program embedding the bound is + +```math +\|(U^\dagger VU)J_F-J_FG_a\|\le2\delta, +``` + +where the source-zero input is implicit in $`J_F`$ in this display. +The original lookup inverse erases $`F(x)`$ exactly and converts this +to the same bound with every program, source, and helper input zero. +The bound is uniform over logical inputs and arbitrary external +references. In particular it is not $`2g\delta`$. + +On h equal to zero, every completed stage is exactly identity. The +selector/enable words cancel by the same preserved-control argument as +the incremental schedule, and U cancels with its actual inverse on +arbitrary source and preparation work. Loading the zero inactive row is +identity as well. Thus the inactive identity is exact. The final actual +outer-predicate inverse preserves the combined active/inactive norm +bound, including approximate return of h. + +The program occupies $`q(2^g-1)`$ wires; the largest selected-word +computation uses $`O(q2^{g-1})`$ private wires. Source, temporary shift +word, selected program, and source preparation use $`O(q\ell)`$ +wires conservatively. The convolution pool occupies $`5R`$ and is +reused after each completed XOR. Selectors and enables use +$`O(2^g+g)`$ more wires. Their disjoint union fits the Section 1 +reservation for a fixed absolute C; no suffix wire is simultaneously +counted as a dirty query helper. + +Summing selected-word costs gives $`O(q2^g)`$. The g shifts contribute +$`O(gR)`$ T gates and $`O(gqR)`$ Clifford gates at $`O(g)`$ +T-depth. The boundary source and retained selectors supply the remaining +terms in the local ledger. The one-hot table queries, outer predicate, +and any angle-grid approximation retain their separate costs. + +## 7. Complete-frame theorem at fixed accuracy + +**Theorem.** Fix $`0\lt\eta\le1/64`$. There is a constant +$`C_\eta`$ such that, for every $`n\ge1`$, $`N=2^n`$, and +$`b\ge C_\eta\sqrt N`$, every prescribed complete real Hopf frame W +has one coherent Clifford+T circuit V using two external clean flags and +at most b arbitrary dirty qubits, with + +```math +\|VJ_2-J_2(W\otimes I_b)\|\le\eta, +\qquad +T=O_\eta(\sqrt N),\quad G=O_\eta(N),\quad D_T=O_\eta(n). +``` + +The error is the full initialized-isometry norm, including returned +work and arbitrary references. The circuit has no measurements, resets, +supplied phase states, QRAM, or uncharged quantum oracles. Its phase state +is prepared and unprepared by charged native words. As elsewhere in this +repository, T-depth permits arbitrary intervening Clifford circuits; +their elementary gate count is included in G, and no matching bound on +total elementary depth is claimed. + +The existing worst-case T-count lower bound is +$`\Omega_\eta(\sqrt N)`$, so the count has optimal order. The available +unrestricted depth lower bound at this width remains $`\Omega(1)`$. +This theorem improves the depth upper bound, not depth optimality or the +high-precision endpoint. Its constants may depend on the fixed accuracy; +it does not replace the all-precision, all-width resource theorem. + +### Early groups and their conditional suffix reservation + +Write $`\rho=\log_2 3`$. Choose q to be the least power of two with + +```math +q\ge \frac{4\pi\sqrt{2n}}\eta, +\qquad +\delta=\frac\eta{8n}. +``` + +Thus $`q=\Theta_\eta(\sqrt n)`$. Let A be an absolute sufficient +constant from the local unary-group construction, so its simultaneously +live conditional suffix registers fit within + +```math +A\left(q2^g+q^\rho+q\log_2 q+g\right) +``` + +bits. This includes the program, phase source, multiplication buffer, +all bilinear work and copies, and selector/enable maintenance. Increase +A to at least one if necessary, and set + +```math +K=\left\lceil16A\bigl(q^\rho+q\log_2q+\log_2q+1\bigr)\right\rceil, +\qquad C=16A. +``` + +This cutoff directly reserves the fixed source and convolution pools. +It satisfies + +```math +K=O_\eta(n^{\rho/2}+\sqrt n\log(n+2)), +\qquad K\log(n+2)=o_\eta(n). +``` + +For sufficiently large n, depending only on eta and the fixed circuit +constants, $`K\lt n`$. At a group start +$`d=n-k`$, while $`k>K`$, take + +```math +g=\left\lfloor\log_2\frac{k}{Cq}\right\rfloor, +\qquad +w=q(2^g-1), +\qquad +r=k-g. +``` + +Each row stores the q-bit one-hot translation program; w is the full +program width for the group. Since $`q^2\ge n`$ and $`C\ge1`$, +all early groups have $`g\le\log_2q`$, as required by the local lemma. +The cutoff also gives + +```math +g\ge\left\lfloor(\rho-1)\log_2q\right\rfloor +=\Omega_\eta(\log(n+2)). +``` + +For sufficiently large n this is positive. Moreover, + +```math +Aq2^g\le k/16, +\qquad +A(q^\rho+q\log_2q+g)\le k/16, +\qquad g\le k/16. +``` + +These inequalities follow directly from the group-size choice and the +cutoff, using $`g\le\log_2q`$. The sufficient reservation is at most +$`k/8`$, while $`r=k-g\ge15k/16`$. These are logical suffix bits, zero only on the group's active +sector; they are not counted again as arbitrary dirty query work. The +external dirty query pools and two returned outer-predicate helpers are +reserved separately. All inactive-sector statements use the local +construction's literal guarded identities on arbitrary work. + +There are $`O_\eta(n/\log(n+2))`$ early groups. The last group may +cross the cutoff by fewer than g layers; its remaining height +$`K'\le K`$ is the start of the individual-layer tail. No rounded early +angle is used by that tail. + +### Full error, including the phase-source return + +Round every early angle to its nearest $`2\pi/q`$ grid point, choosing +a real representative of the difference of magnitude at most +$`\pi/q`$. Leave all other ideal angles unchanged. The +[complete-frame angular bound](HOPF_ERROR_ACCUMULATION.md#1-a-sharp-angle-error-bound-on-the-full-frame) +gives + +```math +\|W_{\rm grid}-W\| +\le\frac{\pi\sqrt{2n}}q\le\eta/4. +``` + +With the exact unary Fourier eigenstate, every early group is exactly +its prescribed grid-angle group and returns that state. On the active +sector, let $`J_F`$ append the valid loaded program $`F(x)`$ and zero +source, preparation, selector, and shift work. Let P be an ideal source +preparation, with its fixed scalar chosen as in Section 4, and +$`\widetilde P`$ its native replacement. Both act only on source and +preparation work, so Section 4 gives +$`\|\widetilde PJ_F-PJ_F\|\le\delta`$ uniformly over logical +inputs. If U is the exact guarded middle word of the complete group, +including selector computation and cleanup, then + +```math +UPJ_F=PJ_FW_{\rm group}. +``` + +Consequently the charged prepare-use-unprepare word obeys + +```math +\|\widetilde P^\dagger U\widetilde PJ_F-J_FW_{\rm group}\| +\le2\delta. +``` + +The first delta bounds the input-preparation difference. For the second, +use the exact return identity and + +```math +(\widetilde P^\dagger-P^\dagger)PJ_F +=\widetilde P^\dagger(P-\widetilde P)J_F. +``` + +The grid-angle group preserves the external prefix x, so its action +also preserves the valid-program embedding $`J_F`$. +Thus this estimate requires neither an operator approximation on arbitrary +source inputs nor a fresh source at each height. It includes the final +source and preparation-work leakage. Exact program and selector return +remain valid through the middle word. On the outer inactive sector the +whole completed group is exactly identity even on arbitrary suffix work, +by the local guarded-cancellation proof. The original lookup inverse +erases $`F(x)`$ exactly, converting the active estimate to the contract +with zero program input and output. Exact load/unload and the actual +outer-predicate inverse preserve the norm estimate. + +Prepare the binary Fourier product state with its +$`\log_2q`$ one-qubit factors approximated to error at most +$`\delta/\log_2q`$, then use the exact unary conversion. The established +single-qubit native approximation bound permits these factors to run on +disjoint wires in parallel. Their T-depth is +$`O(\log(\log(q)/\delta))=O_\eta(\log(n+2))`$; their count and the +charged exact conversion fit the local ledger. The actual circuit +inverse is used on exit. Any consistent global phase in the prepared +state cancels against that actual inverse. + +There are at most n early groups, so unitary telescoping charges at most +$`2n\delta\le\eta/4`$ for their native phase preparations. For the +remaining layers, use the original capped additive source certificate +with accuracy parameter eta divided by four. Restricting that +nonnegative error sum to the tail costs at most $`\eta/4`$. Thus the +total error is at most $`3\eta/4\le\eta`$. Telescoping compares each +stage on its ideal initialized input and uses unitarity on preceding +leakage; it never resets actual work between groups. The estimate is +therefore a full-frame initialized-isometry bound, not a state-preparation +or accepted-block-only estimate. + +### Early prefetches, native counts, and depth + +Use the existing parallel single-output dirty-counter bilinear queries +to prefetch the w program bits from the $`d+1`$ address bits +$`(h,x)`$. Their dirty pools are separate, and the shared address enters +through Clifford controls. With $`Q=2^{d+1}`$, + +```math +T_{\rm load},w_{\rm dirty} +=O\!\left(w\sqrt Q(n+2)^3\right), +\qquad +D_{T,\rm load}=O(\log(n+2)), +``` + +```math +G_{\rm load} +=O\!\left(wQ+w\sqrt Q(n+2)^3\right). +``` + +The $`wQ`$ Clifford cost of arbitrary table coordinate changes is +retained. The program outputs are the separately reserved logical +suffix wires. Actual unload has the same costs. Since $`w\le k/C`$ +and early starts have distinct remaining heights k, + +```math +\sum_{\rm early}T_{\rm load} +\le O\!\left(\sqrt N(n+2)^3 + \sum_{k>K}k2^{-k/2}\right) +=O_\eta(\sqrt N), +``` + +```math +\sum_{\rm early}G_{\rm load} +\le O\!\left(N\sum_{k>K}k2^{-k}\right)+O_\eta(\sqrt N) +=O_\eta(N). +``` + +The first bound also bounds every peak private dirty pool. The large +cutoff absorbs these polynomial factors with ample exponential margin. +The charged local selection, translation, source preparation, and outer +predicate costs are polynomial in n. The local translation and selection +cost $`O(q2^g+gq^\rho)`$ T gates and the conservative +$`O(q2^g+gq^{\rho+1})`$ Clifford gates. Source preparation, conversion, +and predicates also have polynomial cost. Their sum over at most n +groups therefore fits both stated exponential gate-count bounds. No +per-group linear-in-k count claim is needed at this cutoff. + +Each group's internal T-depth is +$`O(g+\log(q)+\log(\log(q)/\delta))=O_\eta(\log(n+2))`$. +Its prefetch, unload, and outer predicate/inverse also cost +$`O(\log(n+2))`$ depth. Multiplying by the number of groups gives +$`O_\eta(n)`$ total early depth. + +### One chunked schedule for the entire remaining tail + +Set + +```math +L'=\max\{6,\lceil\log_2(4/\eta)\rceil\}, +\qquad h_n=\lceil\log_2(8n)\rceil, +\qquad m=L'+4+h_n. +``` + +On every remaining layer $`1\le k\le K'\le K`$, use the original +full-input operator source and literal amplification at the original +capped width + +```math +m_k=L'+4+\min\{k,h_n\}\le m=O_\eta(\log(n+2)). +``` + +Use the already proved +[chunked-query schedule](CHUNKED_DIRTY_INDICATOR.md#5-a-summable-budget-for-the-final-hopf-layers) +for this entire tail. The coefficient query has +$`Q_k=4N2^{-k}`$ rows and $`n-k+2`$ address bits. Split that full +address into balanced halves and use chunk cap +$`a_k=2^{\lfloor k/12\rfloor}`$, truncated to each half's length. +The zero-bit half is a Clifford X. The established resource bounds give + +```math +T_k,w_k +=O\!\left(\sqrt N\,2^{-k/2}[m_k+2^{k/4}]\right), +``` + +```math +G_k=O\!\left(N2^{-k}m_k+\sqrt N\,2^{-k/4}\right), +``` + +```math +D_{T,k} +=O\!\left(m_k+n(k+1)2^{-k/12}+\log(n+2)\right). +``` + +Here the sequential output-bit loop of the bilinear oracle costs +$`O(m_k)`$ depth; the two indicators are shared by those output bits. +The last bound follows directly from the truncated chunk sizes: if a +chunk cap is below its address-half length, its indicator depth is +$`O(n(k+1)2^{-k/12})`$; otherwise it is $`O(\log(n+2))`$. +The bounded number of queries and actual inverses in each amplified +layer changes only constants. + +The count and workspace sums converge uniformly for every tail length +at most n, since $`m_k\le L'+4+k`$: + +```math +\sum_{k=1}^{K'}T_k=O_\eta(\sqrt N), +\qquad +\sum_{k=1}^{K'}G_k=O_\eta(N), +\qquad +\max_{1\le k\le K'}w_k=O_\eta(\sqrt N). +``` + +For depth, retain the sharper cap $`m_k\le m`$, rather than the +loose bound $`L'+4+k`$. Because +$`\sum_{k\ge1}(k+1)2^{-k/12}\lt\infty`$, + +```math +\sum_{k=1}^{K'}D_{T,k} +=O_\eta\!\left(n+K\log(n+2)\right)=O_\eta(n). +``` + +The original full-input sources have $`O(m_k)`$ serial T-depth. +Their reflections and the separately priced suffix predicates fit +$`O_\eta(\log(n+2))`$ additional depth per layer, with polynomial +counts. Thus all non-query tail depth is +$`O_\eta(K\log(n+2))=o_\eta(n)`$. This construction retains the +full n-bit external address and every complete-frame marker column; it +does not reinterpret the tail as a smaller independent frame. + +Keep the original dirty base $`B_0=L'+n+7`$ and the separately +reserved predicate helpers. Reuse each query pool only after its exact +return. Every peak fits $`C_\eta\sqrt N`$ for a sufficiently large +constant. The early and tail ledgers now give all three bounds on the +same circuit. The finitely many n below the stated asymptotic thresholds +are covered by the previously proved fixed-accuracy compiler, increasing +only the eta-dependent constants. + +## 8. Attribution and bounded evidence + +Fourier-state phase kickback and reuse of a prepared phase reference are +established ingredients, as are the three-product Karatsuba identity, +parallel parity-phase synthesis with initialized ancillas, reversible +binary-to-unary conversion, and the single-qubit approximation bound. +The [source map](SOURCE_MAP.md#5-fault-tolerant-sources-and-contribution-boundaries) +records these dependencies, including Jones et al., Iggy van Hoof, +Selinger, and the existing native synthesis toolkit. The Karatsuba +reference supplies the multiplication identity; its reversible schedule +is not the source of this chapter's constant T-depth bound. + +The additional interface proved here is the coherently loaded one-hot +translation with charged conditional work, literal inactive identity, +and exact temporary return on every active source state. Its combination +with a charged source boundary pair, retained local selectors, Hopf angle +stability, and the early/late query allocation gives the complete-frame +theorem. [Related work, Section 19](RELATED_WORK.md#19-unary-phase-source-reuse-and-linear-t-depth-3-october-2026) +compares the catalyst and shared-arithmetic constructions of +Kim–Laakkonen and Wu et al. Their results remain separate from the +specific coherent table and initialized-isometry contract here. These +comparisons identify inherited ingredients and resource distinctions; +they do not establish priority or a generic constant-depth rotation +theorem without supplied resources. + +The [bounded checks](../tests/test_unary_phase_gradient.py) cover exact +small Karatsuba identities and source-bit permutations, a literal native +guarded phase gadget, small native source preparation and its actual +inverse, unequal three-stage row programs, and the retained-source +$`2\delta`$ error bound. The [verification index](VERIFICATION.md) +states their exact ranges and negative controls. The group fixtures use +reduced operators on the invariant unary source space; they do not emit +the complete native bilinear XOR, selected-program network, asymptotic +decoder, or full-frame lookup circuit. Uniform return, workspace, gate +counts, depth, and the global error bound follow from the analytic +construction above, not an extrapolation of those finite checks. diff --git a/docs/VERIFICATION.md b/docs/VERIFICATION.md index 4bd15a5..d932268 100644 --- a/docs/VERIFICATION.md +++ b/docs/VERIFICATION.md @@ -159,6 +159,31 @@ and check constant stale-monitor and one-use-bank leakage on a legal exact zero-angle row. PREP is native; masks, monitors, and reflections are reduced operator fixtures. They supply no new reflection emitter or unrestricted depth lower bound. +The [conditional unary phase-source proof](UNARY_PHASE_GRADIENT.md) has +[six bounded checks](../tests/test_unary_phase_gradient.py). Exact integer +checks establish the Karatsuba cyclic bilinear identities on every basis +pair at $`q=1,2,4,8`$, check their rank counts and dense inputs, and verify +the in-place shift and temporary-word return on every four-bit source +string for both directions of every one-hot program. A literal 28-T +scalar guarded trilinear phase gadget checks its active phase and full +inactive arbitrary-work identity. Native source preparations at q equal +to four and eight include the product phases, reversible binary-to-unary +decode, and actual inverse; these small decode words do not implement +the asymptotic parallel tree schedule. + +The group checks use reduced operators on the invariant unary source +space. Three stages with unequal row programs test the signed-shift +convention, local suffix enables, all logical columns, and exact inactive +action on every reduced source input. A perturbed source is retained +through all three stages and unprepared with its actual inverse; the +full initialized-isometry error obeys the $`2\delta`$ group bound and +includes the residual source component. Negative controls expose a +missing original-h guard, an incorrect inverse shift, a reversed phase +convention, and substitution of the ideal preparation inverse. These +tests do not emit the complete native bilinear XOR, selected-program +network, scalable group, or full-frame lookup schedule. Their uniform +work-return, count, width, depth, and error statements are analytic +proofs, not extrapolations from the finite native and reduced fixtures. The [chunked dirty indicator](CHUNKED_DIRTY_INDICATOR.md) has [three bounded checks](../tests/test_chunked_dirty_indicator.py) for native shared-control phases, actual inverses, arbitrary dirty-tree diff --git a/tests/README.md b/tests/README.md index 53c8a72..8601bba 100644 --- a/tests/README.md +++ b/tests/README.md @@ -48,6 +48,7 @@ theorem by numerical extrapolation. | [`test_shallow_source_obstruction.py`](test_shallow_source_obstruction.py) | Full-input two-layer transfer alphabet, exact dyadic subset grids, optimized robust gaps, and native source witnesses with dirty extensions | | [`test_conditional_geometric_source.py`](test_conditional_geometric_source.py) | Conditional geometric preparation, prefix cleanup, native controlled-H phases, scalar masks, actual inverse, inactive sectors, and amplification with work return | | [`test_grouped_program_prefetch.py`](test_grouped_program_prefetch.py) | Literal masks, native AND/PREP, inactive dirty scratch, coherent program unloading through leakage, enable phases, and adaptive group reservations; reduced group reflections and queries are not a full native emitter | +| [`test_unary_phase_gradient.py`](test_unary_phase_gradient.py) | Karatsuba convolution, guarded native phase gates, exact cyclic-shift return, native small phase sources, and three unequal reduced stages with actual inverse and source error charged once per group | | [`test_grouped_selector_reuse.py`](test_grouped_selector_reuse.py) | Incremental prefix growth, early suffix-enable cleanup, exact rational three-stage leakage, inactive arbitrary-work return, literal private-copy phases, and invalid cleanup orders; reduced stages test the selector interface | | [`test_grouped_source_reuse.py`](test_grouped_source_reuse.py) | Common-source two-stage conjugation, a wrong-reflection negative control, and legal exact-row stale-monitor/one-use leakage; native PREP with reduced masks and reflections | | [`test_chunked_dirty_indicator.py`](test_chunked_dirty_indicator.py) | Native shared-control phases and actual inverses, all-input dirty-tree return, conjugation-order regression, and exact late-query resource sums; asymptotic counter depth is analytic | diff --git a/tests/test_unary_phase_gradient.py b/tests/test_unary_phase_gradient.py new file mode 100644 index 0000000..0878bdd --- /dev/null +++ b/tests/test_unary_phase_gradient.py @@ -0,0 +1,293 @@ +"""Small audits of a unary phase-gradient group, with explicit scope. + +Karatsuba identities and source-bit permutations are exact integer checks. +The scalar guarded phase gadget and q=4,8 PREP use literal Clifford+T +words. Group actions are reduced operators on the invariant unary source +space; they are not a native complete-frame/QROM emitter. Their columns +retain every logical input and the actual source inverse, without resets. +The asymptotic count, workspace and T-depth claims require the proof. +""" +from __future__ import annotations + +import unittest + +import numpy as np + +try: + from .test_operator_source_compiler import ( + H, _adjoint, _apply_native_word, _expand_toffolis, + ) +except ImportError: + from test_operator_source_compiler import ( + H, _adjoint, _apply_native_word, _expand_toffolis, + ) + + +ATOL = 8e-11 + + +def _karatsuba(q): + """(left mask, right mask, output mask), before cyclic reduction.""" + if q == 1: + return [(1, 1, 1)] + half = q // 2 + terms = [] + for a, b, v in _karatsuba(half): + terms.append((a, b, v ^ (v << half))) + terms.append((a << half, b << half, + (v << half) ^ (v << (2 * half)))) + terms.append((a ^ (a << half), b ^ (b << half), v << half)) + return terms + + +def _cyclic_terms(q): + terms = [] + for a, b, v in _karatsuba(q): + folded = 0 + for j in range(2 * q - 1): + if (v >> j) & 1: + folded ^= 1 << (j % q) + terms.append((a, b, folded)) + return terms + + +def _bilinear(terms, a, x): + answer = 0 + for left, right, output in terms: + if (a & left).bit_count() % 2 and (x & right).bit_count() % 2: + answer ^= output + return answer + + +def _convolution(q, a, x): + answer = 0 + for i in range(q): + for j in range(q): + if ((a >> i) & 1) and ((x >> j) & 1): + answer ^= 1 << ((i + j) % q) + return answer + + +def _reverse_program(q, a): + return sum(((a >> j) & 1) << ((-j) % q) for j in range(q)) + + +def _shift_pair(q, a, x, y=0, h=1): + """Two multiply-XORs, then a guarded swap (all-input classical model).""" + if h: + terms = _cyclic_terms(q) + y ^= _bilinear(terms, a, x) + x ^= _bilinear(terms, _reverse_program(q, a), y) + x, y = y, x + return x, y + + +def _native_gradient(q): + """Native product phases and reversible binary-to-unary map, q=4,8.""" + r = q.bit_length() - 1 + binary = list(range(r)) + unary = list(range(r, r + q)) + helper = r + q + width = helper + (r == 3) + word = [('H', bit) for bit in binary] + # q=8: T,S,Z; q=4: S,Z. No floating rotation oracle is used. + for bit in binary: + for _ in range(1 << (bit + 3 - r)): + word.append(('T', bit)) + for label, target in enumerate(unary): + flips = [('X', bit) for bit in binary if not ((label >> bit) & 1)] + word += flips + if r == 2: + word.append(('CCX', binary[0], binary[1], target)) + else: + word += [('CCX', binary[0], binary[1], helper), + ('CCX', helper, binary[2], target), + ('CCX', binary[0], binary[1], helper)] + word += list(reversed(flips)) + # Exactly erase the binary label from the one-hot word. + for bit in binary: + word += [('CX', unary[label], bit) for label in range(q) + if (label >> bit) & 1] + return width, unary, _expand_toffolis(word) + + +def _binary_source(q, phase_errors=None): + r = q.bit_length() - 1 + hadamards = np.array([[1.0]]) + for _ in range(r): + hadamards = np.kron(H, hadamards) + phase_errors = [0.0] * r if phase_errors is None else phase_errors + phases = np.array([ + np.exp(1j * sum(((j >> bit) & 1) * + (2 * np.pi * (1 << bit) / q + phase_errors[bit]) + for bit in range(r))) + for j in range(q) + ]) + return phases[:, None] * hadamards + + +def _target_matrix(state, ell, matrix): + result = state.copy() + for low in range(len(state)): + if (low >> ell) & 1: + continue + high = low | (1 << ell) + result[low] = matrix[0, 0] * state[low] + matrix[0, 1] * state[high] + result[high] = matrix[1, 0] * state[low] + matrix[1, 1] * state[high] + return result + + +def _source_matrix(state, matrix): + return np.einsum('ij,ajc->aic', matrix, state) + + +def _reduced_stages(state, rows, q, h=1, wrong_sign=False): + # B=SH maps the Z eigenbasis to the Y eigenbasis. + basis = np.diag([1, 1j]) @ H + result = state + for ell, table in enumerate(rows): + result = _target_matrix(result, ell, basis.conj().T) + shifted = result.copy() + for local in range(len(result)): + if not h or local >> (ell + 1): + continue + a = table[local & ((1 << ell) - 1)] + direction = -1 if ((local >> ell) & 1) else 1 + if wrong_sign: + direction = -direction + shifted[local] = np.roll(result[local], direction * a, axis=0) + result = _target_matrix(shifted, ell, basis) + return result + + +def _ideal_group(rows, q): + dimension = 1 << len(rows) + result = np.eye(dimension, dtype=complex) + for ell, table in enumerate(rows): + stage = np.eye(dimension, dtype=complex) + for low in range(dimension): + if ((low >> ell) & 1) or low >> (ell + 1): + continue + high = low | (1 << ell) + angle = 2 * np.pi * table[low & ((1 << ell) - 1)] / q + c, s = np.cos(angle), np.sin(angle) + stage[np.ix_([low, high], [low, high])] = [[c, -s], [s, c]] + result = stage @ result + return result + + +def _embedding(g, q): + state = np.zeros((1 << g, q, 1 << g), dtype=complex) + state[:, 0, :] = np.eye(1 << g) + return state + + +class UnaryPhaseGradientTests(unittest.TestCase): + def test_karatsuba_cyclic_bilinear_identity_and_rank(self): + for q in (1, 2, 4, 8): + terms = _cyclic_terms(q) + self.assertEqual(len(terms), 3 ** (q.bit_length() - 1)) + # Agreement on all basis pairs proves this bilinear identity. + for i in range(q): + for j in range(q): + self.assertEqual(_bilinear(terms, 1 << i, 1 << j), + 1 << ((i + j) % q)) + for a in (0, 1, (1 << q) - 1, (1 << q) // 3): + for x in range(1 << q): + self.assertEqual(_bilinear(terms, a, x), + _convolution(q, a, x)) + + def test_native_guarded_trilinear_phase_and_inactive_arbitrary_work(self): + # Controls h,a,b,c; work t=a*b, root=t*c. 28 T gates total. + compute = _expand_toffolis([('CCX', 1, 2, 4), ('CCX', 4, 3, 5)]) + central = [('H', 5), ('CX', 0, 5), ('H', 5)] + word = compute + central + _adjoint(compute) + self.assertEqual(sum(g[0] in ('T', 'TDG') for g in word), 28) + inactive = [j for j in range(64) if not (j & 1)] + active = [1 | (j << 1) for j in range(8)] + columns = np.eye(64, dtype=complex)[:, inactive + active] + expected = columns.copy() + expected[:, -1] *= -1 + np.testing.assert_allclose(_apply_native_word(6, word, columns), + expected, atol=ATOL, rtol=0) + # Without the original h, arbitrary inactive inputs are not protected. + unguarded = compute + [('S', 5), ('S', 5)] + _adjoint(compute) + self.assertGreater(np.linalg.norm( + _apply_native_word(6, unguarded, columns) - expected, 2), 1.9) + + def test_in_place_shift_returns_output_for_every_source_bitstring(self): + q = 4 + for shift in range(q): + for sign in (-1, 1): + a = 1 << ((sign * shift) % q) + for x in range(1 << q): + self.assertEqual(_shift_pair(q, a, x), + (_convolution(q, a, x), 0)) + # No validity or clean-output promise is needed on the h=0 sector. + for a in range(1 << q): + for x in range(1 << q): + for y in range(1 << q): + self.assertEqual(_shift_pair(q, a, x, y, h=0), (x, y)) + # The reverse program is essential: two forward convolutions fail. + a, x = 2, 1 + y = _convolution(q, a, x) + self.assertNotEqual(x ^ _convolution(q, a, y), 0) + + def test_native_product_phases_unary_decode_and_actual_inverse(self): + for q in (4, 8): + width, unary, word = _native_gradient(q) + initial = np.zeros((1 << width, 1), dtype=complex) + initial[0, 0] = 1 + prepared = _apply_native_word(width, word, initial) + expected = np.zeros_like(initial) + for j, bit in enumerate(unary): + expected[1 << bit, 0] = np.exp(2j * np.pi * j / q) / np.sqrt(q) + np.testing.assert_allclose(prepared, expected, atol=ATOL, rtol=0) + np.testing.assert_allclose( + _apply_native_word(width, _adjoint(word), prepared), + initial, atol=ATOL, rtol=0) + + def test_unequal_three_stage_rows_give_literal_rounded_frame(self): + q, rows = 8, [[1], [3, 6], [2, 7, 1, 5]] + source = _binary_source(q) + initial = _embedding(3, q) + prepared = _source_matrix(initial, source) + acted = _reduced_stages(prepared, rows, q) + result = _source_matrix(acted, source.conj().T) + expected = initial.copy() + expected[:, 0, :] = _ideal_group(rows, q) + np.testing.assert_allclose(result, expected, atol=ATOL, rtol=0) + wrong = _source_matrix(_reduced_stages(prepared, rows, q, wrong_sign=True), + source.conj().T) + self.assertGreater(np.linalg.norm((wrong - expected).reshape(64, 8), 2), 1) + # Every source/input column, not merely the prepared eigenstate. + arbitrary = np.eye(64, dtype=complex).reshape(8, 8, 64) + np.testing.assert_allclose(_reduced_stages(arbitrary, rows, q, h=0), + arbitrary, atol=ATOL, rtol=0) + + def test_approximate_source_is_charged_once_per_group_without_reset(self): + q, rows = 8, [[1], [3, 6], [2, 7, 1, 5]] + exact, actual = _binary_source(q), _binary_source(q, [.019, -.027, .013]) + delta = np.linalg.norm(actual[:, 0] - exact[:, 0]) + initial = _embedding(3, q) + result = _source_matrix( + _reduced_stages(_source_matrix(initial, actual), rows, q), + actual.conj().T) + expected = initial.copy() + expected[:, 0, :] = _ideal_group(rows, q) + error = np.linalg.norm((result - expected).reshape(64, 8), 2) + self.assertLessEqual(error, 2 * delta + ATOL) + self.assertGreater(np.linalg.norm(result[:, 1:, :]), 0.01) + # For identity rows actual PREP† gives exact return; the ideal inverse + # incorrectly leaves the preparation error in the source register. + identity_rows = [[0], [0, 0], [0, 0, 0, 0]] + acted = _reduced_stages(_source_matrix(initial, actual), identity_rows, q) + np.testing.assert_allclose(_source_matrix(acted, actual.conj().T), initial, + atol=ATOL, rtol=0) + wrong = _source_matrix(acted, exact.conj().T) + self.assertAlmostEqual(np.linalg.norm((wrong - initial).reshape(64, 8), 2), + delta, places=10) + + +if __name__ == '__main__': + unittest.main() From ae42701f6f2dc5540c86ce4da50321a41c6abde8 Mon Sep 17 00:00:00 2001 From: Ruge Lin Date: Sat, 3 Oct 2026 18:31:38 +0800 Subject: [PATCH 2/2] Keep unary group depth equation within narrow rendered pages --- docs/UNARY_PHASE_GRADIENT.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/UNARY_PHASE_GRADIENT.md b/docs/UNARY_PHASE_GRADIENT.md index e4c9534..1570391 100644 --- a/docs/UNARY_PHASE_GRADIENT.md +++ b/docs/UNARY_PHASE_GRADIENT.md @@ -593,7 +593,11 @@ groups therefore fits both stated exponential gate-count bounds. No per-group linear-in-k count claim is needed at this cutoff. Each group's internal T-depth is -$`O(g+\log(q)+\log(\log(q)/\delta))=O_\eta(\log(n+2))`$. + +```math +O(g+\log(q)+\log(\log(q)/\delta))=O_\eta(\log(n+2)). +``` + Its prefetch, unload, and outer predicate/inverse also cost $`O(\log(n+2))`$ depth. Multiplying by the number of groups gives $`O_\eta(n)`$ total early depth.