diff --git a/README.md b/README.md index 85223ba..d161723 100644 --- a/README.md +++ b/README.md @@ -133,29 +133,24 @@ magnitude frames at $`b\ge L+n+8`$. With $`b\ge2(L+n+8)`$, it gives $`O(\sqrt{NL}+L\ell_*(n)+NL/b)`$ T gates, still $`G=O(NL)`$. Leaf-phase derivatives retain a separate QBP stream. -[Hybrid lookup](docs/PARALLEL_DIRTY_LOOKUP.md) uses two clean flags and -$`b\ge17(L+n+7)`$. Fixed L gives $`T=O(\sqrt N+N/b)`$ and +At fixed accuracy, [blocked bilinear lookup](docs/BLOCKED_BILINEAR_LOOKUP.md) +with charged unary phase-source groups gives one complete real-frame +circuit using two clean flags and $`b\ge17(L+n+7)`$, with ```math -D_T=O(N/b^2+n\log(n+2)). +T=O(\sqrt N+N/b),\qquad G=O(N),\qquad +D_T=O(N/b^2+n). ``` -For $`L=\Theta(n)`$, sufficient -$`b=\Theta(n)`$ gives optimal worst-case $`T=\Theta(N)`$ and -$`D_T=\Theta(N/n)`$ in one complete real-frame circuit. +T-count is optimal in order. Count and depth match simultaneously through +$`b\le\sqrt{N/n}`$ when the interval is nonempty. Larger widths retain +the linear depth upper bound; depth optimality remains open there. +Source preparation, reuse, and return are charged. -At fixed accuracy, [unary phase-source groups](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) -and chunked queries give, with two clean flags and sufficient -$`b=\Theta(\sqrt N)`$, one complete real-frame circuit with - -```math -T=O(\sqrt N),\qquad G=O(N),\qquad -D_T=O(n). -``` - -Phase-source preparation, reuse, and return are charged. T-count is -optimal in order; depth optimality and the high-precision endpoint -remain open. The [variable-precision theorem](docs/OPEN_PROBLEM.md) is unchanged. +The [variable-precision hybrid](docs/PARALLEL_DIRTY_LOOKUP.md) retains its +existing bounds: at $`L=\Theta(n)`$, sufficient $`b=\Theta(n)`$ gives +optimal worst-case $`T=\Theta(N)`$ and $`D_T=\Theta(N/n)`$ in one +complete real-frame circuit. The high-precision endpoint remains open. Beyond Hopf frames, **literal diagonals and general one-target U(2) multiplexors** attain $`\Theta(\sqrt{NL}+L+NL/b)`$ with one clean diff --git a/WORKSPACE.md b/WORKSPACE.md index 38aad56..6708fbc 100644 --- a/WORKSPACE.md +++ b/WORKSPACE.md @@ -3,9 +3,9 @@ This is the entry point when a previous conversation or execution workspace is unavailable. Proofs and decisions live in the repository. -The 2026-10-03 unary phase-source pass starts from verified main -`005debe2ccf7a0b2e7cbc032a0f1ae285b83d2ab`, after the incremental -group-selector and common-source audit. Check later commits before +The 2026-10-03 width-sensitive bilinear pass starts from verified main +`b94c77d8d82881c2bb2efc155a1da1f99f27c92e`, after the charged +unary phase-source theorem. Check later commits before continuing. The selected state-based Hopf QBP construction and its bounded-input audit are complete; the [consolidated theorem](docs/STATE_BASED_QBP_THEOREM.md) is their entry point. @@ -25,7 +25,8 @@ and finite checks, without large simulations, QRAM, resets inside a compiler execution, supplied catalysts, or hidden initialized work. 1. Begin active depth work with the - [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy), + [width-sensitive bilinear theorem](docs/BLOCKED_BILINEAR_LOOKUP.md), + then the [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy), then the [grouped complete-frame theorem](docs/GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy), [chunked dirty indicator](docs/CHUNKED_DIRTY_INDICATOR.md), [conditional geometric source](docs/CONDITIONAL_GEOMETRIC_SOURCE.md), @@ -79,7 +80,7 @@ compiler execution, supplied catalysts, or hidden initialized work. | Additional dirty banks | Improve the state preparation T bound while charging the exact coarse circuit; [banked proof](docs/COMPLEX_COARSE_COMPILER.md#8-additional-dirty-banks-improve-fine-state-preparation) | | State-based T-depth | Two complete schedules, one retaining the sharper count at a stronger dirty reservation; [depth proof](docs/STATE_QBP_DEPTH.md) and [fair comparison](docs/QBP_COST_COMPARISON.md#7-state-based-t-depth-comparison) | | Complete-frame T-count and T-depth | Same-circuit bounds at every accuracy; matching in an explicit workspace range, including inverse-polynomial error; [amortized tradeoff](docs/AMORTIZED_DIRTY_LOOKUP.md) | -| Fixed-accuracy large-width depth | Two external flags, sufficient square-root-scale dirty width, optimal-order square-root T-count, and O(n) depth; [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) | +| Fixed-accuracy width-dependent depth | Two external flags, T-count O(sqrt N+N/b), and depth O(N/b²+n) at b at least 17(L+n+7); both orders match through sqrt(N/n); [blocked bilinear theorem](docs/BLOCKED_BILINEAR_LOOKUP.md) | | Complete-frame error accumulation | Sharp ideal-angle stability and finite relative spectra; coherent linear leakage in actual shared-flag source layers; [scoped error audit](docs/HOPF_ERROR_ACCUMULATION.md) | | Filtered complete-frame source | Quadratic radial error on the same two flags, charged native selective phases, and a smaller source-precision cap; [filter proof](docs/HOPF_RADIAL_FILTER.md) | | Conditional precision depth | Logarithmic source/reflection depth using an active zero suffix and two external flags; [conditional source](docs/CONDITIONAL_GEOMETRIC_SOURCE.md). Query/predicate depth remains charged | @@ -132,13 +133,15 @@ gives $`T^\star=\Theta(N)`$ and $`D_T^\star=\Theta(N/n)`$. Outside the matching range, retain the $`nL`$ term and compare the older grouped schedules before claiming count optimality. -At fixed L, the retained $`T=O(\sqrt N+N/b)`$ and -$`D_T=O(N/b^2+n\chi(n))`$ bounds remain. For sufficiently large +At fixed L, the [blocked bilinear theorem](docs/BLOCKED_BILINEAR_LOOKUP.md) +now gives $`T=O(\sqrt N+N/b)`$, $`G=O(N)`$, and +$`D_T=O(N/b^2+n)`$ at the same threshold. For sufficiently large $`b=\Theta(n)`$ above the threshold, one circuit has optimal-order $`T=\Theta(N/n)`$ and $`D_T=\Theta(N/n^2)`$. -The matching depth tradeoff $`D_T^\star=\Theta(N/b^2)`$ holds throughout -$`17(L+n+7)\le b\le\sqrt{NL/(nL+n\chi(n))}`$ when this interval is nonempty. -The loader places one low-address indicator echo around a multiplexed +The fixed-accuracy matching depth tradeoff +$`D_T^\star=\Theta(N/b^2)`$ now holds throughout +$`17(L+n+7)\le b\le\sqrt{N/n}`$ when this interval is nonempty. +The earlier loader places one low-address indicator echo around a multiplexed family of linear shears. A two-pass dirty traversal and rank-reduced controlled shears amortize both former per-batch logarithms. Dirty selector stacks remain separate from live banks and indicators; @@ -154,7 +157,7 @@ The original routed schedule retains an additive $`n^2`$ term. The improves fixed-accuracy large-workspace depth to $`O(n\chi(n))`$ while retaining optimal-order T-count. The [unary phase-source theorem](docs/UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) -now improves this fixed-accuracy, sufficient-square-root-width bound to +first improved the fixed-accuracy, sufficient-square-root-width bound to ```math T=O_\eta(\sqrt N),\qquad G=O_\eta(N),\qquad @@ -167,9 +170,12 @@ two external clean flags. Early groups store one-hot programs and a charged unary phase source in conditional logical zeros. Bilinear cyclic shifts have constant T-depth; source preparation and its actual inverse occur once per group. Chunked dirty indicators and capped source precision -control the entire remaining tail. -The worst-case count is optimal in order, but the available depth lower -bound remains only Omega(1). Fixed accuracy is essential to this theorem; +control the entire remaining tail. The blocked bilinear refinement removes +the late word-bank route at smaller dirty widths and extends this result +to the displayed width-dependent curve. +The worst-case count is optimal in order. At square-root-scale dirty +width the depth lower bound remains only Omega(1). Fixed accuracy is +essential to this theorem; the general-precision matching interval and endpoint remain unchanged. The older general-precision schedules below remain useful where their source or workspace costs are sharper. The separate state-based and @@ -269,7 +275,7 @@ software extensions below are not prerequisites for the stated theorem. | Variable-size native schedule | Optional software: general tables, predicates, reflections, and banked count/depth scheduling. The bounded component-to-gradient pass is complete | | General guarded decoder | Optional software: replace the general floating-point contractions with the proved certified arithmetic. Exact fixture decoders cover only their fixed target | | Constant-clean complete-frame endpoint | Still open independently of the state-based task. A new candidate must supply an explicit complete native identity and symbolic precision/workspace ledger before another fixture pass | -| Optimal T-depth | Count and depth match for $`17(L+n+7)\le b\le\sqrt{NL/(nL+n\chi(n))}`$, plus the retained fixed-L interval $`b\le\sqrt N/n`$; large-width depth and precision outside these ranges remain unresolved | +| Optimal T-depth | At fixed accuracy both match through $`b\le\sqrt{N/n}`$ above the literal threshold; the variable-L interval remains $`17(L+n+7)\le b\le\sqrt{NL/(nL+n\chi(n))}`$; larger-width optimality and further precision dependence remain unresolved | Do not repeat the completed coefficient-to-row, two- and four-row lookup, one- and two-system-qubit amplification, or coherent residual-selection passes. @@ -489,16 +495,25 @@ The six bounded tests cover the rank identity, native guard, arbitrary source shift, preparation/inversion, unequal rows, and full source-return error. They do not constitute a scalable native group compiler. -The next bounded research question is whether this fixed-accuracy route -extends to a width-dependent same-circuit bound of O(N/b²+n) depth with -O(sqrt N+N/b) T-count. That extension is not established here. Start with -a complete shared-work reservation and a summed prefetch schedule; do not -assume the large-width parallel query bound survives reduced dirty width. -A stronger unrestricted depth lower bound and the high-precision endpoint -remain separate. Do not repeat the completed selector, common-source, -unary phase-source, or chunked-query proofs. Any alternative must retain -literal phases, actual inverses, full work return, and charged source -preparation; ideal-angle stability does not bound unfiltered source leakage. +The [blocked bilinear query](docs/BLOCKED_BILINEAR_LOOKUP.md) now proves +the width-dependent extension: depth O(N/b²+n) and T-count +O(sqrt N+N/b), with two flags and the original 17(L+n+7) threshold. +A naive use of the earlier width-sensitive query fails because its +per-layer route has depth O(n). Instead, one returned dirty helper +implements each selected bilinear leaf in constant depth per output +bit. Two dirty traversals select its matrix block; the four-corner +indicator echo returns all masks. Chunk lengths and block sizes are +chosen together, and every tail resource sum is charged. This extends +the simultaneous fixed-accuracy matching interval to sqrt(N/n). + +The next bounded question is the precision dependence of this route. +First expose the eta-dependent source width, conditional suffix cutoff, +and weighted query sums before asserting any uniform-L improvement. +No such extension is established here. A stronger unrestricted +large-width depth lower bound and the high-precision endpoint remain +separate. Do not repeat the completed selector, source-reuse, unary +source, or blocked-query proofs. Retain literal phases, actual inverses, +full work return, and charged preparation in any new schedule. The completed modest-width matching theorem does not require solving the high-precision endpoint. For any new component, keep literal phases and actual inverses, declare diff --git a/docs/AMORTIZED_DIRTY_LOOKUP.md b/docs/AMORTIZED_DIRTY_LOOKUP.md index e3a5fdf..87c495f 100644 --- a/docs/AMORTIZED_DIRTY_LOOKUP.md +++ b/docs/AMORTIZED_DIRTY_LOOKUP.md @@ -62,12 +62,12 @@ error. Large-workspace optimal depth and the high-precision constant-clean endpoint remain open. T-depth permits arbitrary Clifford circuits between T layers; their elementary depth is not bounded by these theorems, while their gate count remains included in G. -The separate [grouped-program theorem](GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) -improves the fixed-accuracy upper depth to $`O(n\log\log(n+2))`$ -at sufficient $`\Theta(\sqrt N)`$ dirty width, retaining optimal-order -T-count. It uses the same additive precision cap below, conditional -program storage, and a different late-query schedule; it does not extend -this chapter's all-width matching range. +At fixed accuracy, the [blocked bilinear refinement](BLOCKED_BILINEAR_LOOKUP.md) +retains the literal width threshold and optimal-order T-count while giving +$`D_T=O(N/b^2+n)`$. Its simultaneous matching interval extends to +$`b\le\sqrt{N/n}`$ when nonempty. It combines conditional unary groups +with selected bilinear blocks and a width-constrained tail query; the +all-precision theorem above remains unchanged. ## 1. A controlled linear shear has constant T-depth diff --git a/docs/BLOCKED_BILINEAR_LOOKUP.md b/docs/BLOCKED_BILINEAR_LOOKUP.md new file mode 100644 index 0000000..2ce8cea --- /dev/null +++ b/docs/BLOCKED_BILINEAR_LOOKUP.md @@ -0,0 +1,529 @@ +# A blocked bilinear query and the fixed-accuracy width frontier + +[Amortized dirty traversal](AMORTIZED_DIRTY_LOOKUP.md) · [Bilinear queries](PARALLEL_DIRTY_LOOKUP.md#5-a-bilinear-query-reduction) · [Chunked indicators](CHUNKED_DIRTY_INDICATOR.md) · [Unary group compiler](UNARY_PHASE_GRADIENT.md) + +A bilinear table query can process several address blocks without routing +an output bank. Its selected middle operation uses a dirty traversal of +the high address. A constant-depth controlled bilinear operation supplies +each leaf. The two low-address indicators are computed only a constant +number of times around that whole traversal. Their chunk sizes can then +vary with the layer's distance from the leaves. + +Combining this exact query with the conditional unary group compiler +gives the following fixed-accuracy theorem. Fix $`0\lt\eta\le1/64`$, +set $`N=2^n`$, $`L=\max\{6,\lceil\log_2(1/\eta)\rceil\}`$, +and $`B_0=L+n+7`$. For every $`n\ge1`$ and $`b\ge17B_0`$, +every prescribed complete real Hopf frame W has one coherent Clifford+T +circuit V with two external clean flags and at most b arbitrary dirty +qubits satisfying + +```math +\|VJ_2-J_2(W\otimes I_b)\|\le\eta, +``` + +```math +T=O_\eta\!\left(\sqrt N+\frac Nb\right),\qquad +G=O_\eta(N),\qquad +D_T=O_\eta\!\left(\frac N{b^2}+n\right). +``` + +All three bounds describe the same circuit. The error includes every +work return and arbitrary reference correlations. Every query and +inverse below is a literal full-input native circuit. There are no +measurements, resets, QRAM, supplied phase states, or extra external +initialized work. T-depth allows arbitrary intervening Clifford +circuits, whose elementary gate count remains included in G. + +The T-count has the established optimal worst-case order. When the +interval is nonempty, count and depth both match their existing lower +bounds for + +```math +17B_0\le b\le\sqrt{N/n},\qquad +T^\star=\Theta_\eta(N/b),\qquad +D_T^\star=\Theta_\eta(N/b^2). +``` + +The additive n in the upper bound is not an unrestricted depth lower +bound. Depth optimality at larger widths and the variable-accuracy +complete-frame endpoint remain open. Constants may depend on the fixed +accuracy; this theorem does not assert a new all-precision bound. + +## 1. A controlled bilinear output with one arbitrary dirty helper + +Let Y and X be disjoint registers of H and J arbitrary bits, let z be +one arbitrary output bit, and let h be a separate arbitrary control. +For a fixed binary matrix $`D\in\mathbb F_2^{H\times J}`$, we need + +```math +(h,Y,X,z)\longmapsto(h,Y,X,z\oplus hY^{\mathsf T}DX). +``` + +Reserve one arbitrary dirty helper e, disjoint from these registers. +Binary elimination, with actual inverse coordinate changes, puts the +bilinear form into $`\sum_{i=1}^{\rho}Y_iX_i`$, where +$`\rho=\operatorname{rank}(D)`$. The orientation is the same as in +the [existing bilinear lemma](PARALLEL_DIRTY_LOOKUP.md#exact-bilinear-oracle): +if $`PDQ=J_\rho`$, use $`Y'=P^{-\mathsf T}Y`$ and +$`X'=Q^{-1}X`$. These basis changes and their inverses have +$`O(\rho(H+J))`$ Clifford gates and require no extra wire. + +In these coordinates, apply H to z and define the exact diagonal word +and exact Toffoli + +```math +D_e=\prod_{i=1}^{\rho}\mathrm{CCZ}(e,Y_i,X_i),\qquad +E=\mathrm{CCX}(h,z;e). +``` + +Execute chronologically + +```math +D_e,\quad E,\quad D_e^\dagger,\quad E^\dagger, +``` + +then reverse the Hadamard and the coordinate changes. Every dagger is +the actual reversed native word. Writing +$`b(Y,X)=\bigoplus_iY_iX_i`$, the two diagonal actions have combined +phase + +```math +(-1)^{e b(Y,X)+(e\oplus hz)b(Y,X)} +=(-1)^{hz b(Y,X)}. +``` + +The final E inverse restores e. Conjugating z by Hadamards therefore +implements the required controlled XOR, with every helper returned. +This proof permits arbitrary h, z, e, Y, and X. In particular, for h +equal to zero the full completed word is identity on arbitrary dirty +work. The computational-basis identity, including literal phase, +extends to arbitrary superpositions and reference correlations. + +The [shared-control phase schedule](T_DEPTH_COMPILER.md#a-shared-control-fredkin-batch-has-at-most-four-t-layers) +implements $`D_e`$ with at most four T layers and +$`6\rho+(\rho\bmod2)`$ T gates. Its control e may be arbitrary; +no initialized copies are used. Each E has seven T gates and four +T layers. Thus, for $`\rho>0`$, the controlled bilinear output has + +```math +T\le12\rho+2(\rho\bmod2)+14,\qquad +D_T\le16,\qquad G=O(\rho(H+J)). +``` + +For rho equal to zero, emit identity. For an m-bit output Z with +matrices $`D_1,\ldots,D_m`$, process the output bits sequentially, +restoring both coordinate changes after every output. Reuse the same +returned helper e. Writing $`R_D=\sum_j\operatorname{rank}(D_j)`$, +the complete controlled bilinear oracle has + +```math +T=O(R_D),\qquad D_T=O(m),\qquad G=O(mHJ). +``` + +Completed output-bit words preserve Y, X, h, and e. Intermediate +coordinate changes need not preserve individual Y or X bits. + +## 2. Select a bilinear block with dirty traversal + +Split the r-bit table address into three disjoint fields $`(c,a,b)`$, +with $`K=2^p`$ high blocks, H low-a values, and J low-b values, so +$`Q=KHJ=2^r`$. For output bit j in block t, let +$`D_{t,j}\in\mathbb F_2^{H\times J}`$ be its classical table matrix. +Define the desired selected bilinear map + +```math +F_c:(Y,X,Z)\longmapsto +\left(Y,X,\left(Z_j\oplus Y^{\mathsf T}D_{c,j}X\right)_{j=1}^m\right). +``` + +Reserve p arbitrary dirty traversal selectors and the separate helper +e from Section 1. Use the existing +[two-pass dirty traversal](AMORTIZED_DIRTY_LOOKUP.md#2-selecting-a-shear-with-dirty-unary-traversal), +with each leaf now the controlled bilinear oracle for that block. +Its terminal selector is the control h in Section 1. + +For completeness, at leaf t write $`\lambda_i=[c_i=t_i]`$ and let +$`v_1=w_1\oplus\lambda_1`$, +$`v_i=w_i\oplus v_{i-1}\lambda_i`$. Its expanded terminal value is +$`v_p=[c=t]\oplus h_t(c,w)`$. One traversal inserts the leaf +operation controlled by $`v_p`$; the second omits the first selector +update, retaining the control $`h_t`$, and is applied by its actual +inverse. Each traversal restores all its selectors. For p equal to +one, the second traversal still visits its two leaves; only the root +updates are omitted. + +Every completed leaf preserves Y, X, and the selectors. All completed +bilinear maps commute, since they preserve Y and X and XOR their +contributions into Z. The two traversals consequently cancel every +unknown dirty-selector contribution and retain only block c. Each +leaf also returns e. This proves F on the full input space, including +arbitrary selector and helper states. For p equal to zero, use the +existing unguarded bilinear oracle directly. + +The complete depth-first traversals have $`O(K)`$ nodes. Each leaf +has at most $`m\min(H,J)`$ total rank. Therefore + +```math +T(F)=O(Km\min(H,J)),\qquad +D_T(F)=O(Km),\qquad G(F)=O(KmHJ)=O(Qm). +``` + +The additional work is exactly the p selectors and one dirty helper; +Y and X are separately allocated indicator registers. The Clifford +ledger includes all coordinate changes for every block and every +output bit. No classical matrix transformation is treated as a free +quantum oracle. + +## 3. Two indicator echoes complete the exact query + +Let $`I_a:Y\mapsto Y\oplus e_a`$ and +$`I_b:X\mapsto X\oplus e_b`$ be exact dirty indicators, preserving +their addresses and returning all their own helpers. Execute + +```math +F_c,\ I_a,\ F_c^\dagger,\ I_b,\quad +F_c,\ I_a^\dagger,\ F_c^\dagger,\ I_b^\dagger. +``` + +For every fixed c and output bit, its four bilinear contributions sum +over $`\mathbb F_2`$ to + +```math +Y^{\mathsf T}D_cX ++(Y+e_a)^{\mathsf T}D_cX ++(Y+e_a)^{\mathsf T}D_c(X+e_b) ++Y^{\mathsf T}D_c(X+e_b) +=D_c[a,b]. +``` + +Thus the output is XORed with exactly its selected table word. Both +indicators, their helpers, all traversal selectors, and e return +exactly. Every address is preserved. A zero table row gives exact +identity even if the internal words act nontrivially on that sector. +The query works for arbitrary initial outputs and dirty work, and +for coherent addresses and references. + +If the indicators' respective resources are $`T_a,G_a,D_a,w_a`$ +and $`T_b,G_b,D_b,w_b`$, where widths include their outputs, a +sufficient simultaneous query width beyond the original address and +output is + +```math +w_a+w_b+p+1. +``` + +This conservative allocation keeps both private indicator pools +disjoint. Reusing returned helper pools could improve its constant but +is unnecessary. The exact resource ledger is at most four completed F +calls, two a-indicators, and two b-indicators, counting actual inverses +at the same cost. In particular, + +```math +T=O(Km\min(H,J))+2T_a+2T_b, +``` + +```math +G=O(Qm)+2G_a+2G_b, +\qquad D_T=O(Km)+2D_a+2D_b. +``` + +There is no output-bank router. This is the existing bilinear echo +with a new, fully charged selected middle operation. + +## 4. Allocate a truncated balanced query at a prescribed width + +Use the [chunked dirty indicator](CHUNKED_DIRTY_INDICATOR.md) with a +positive integer chunk cap a. Put $`P_a=(a+2)^3`$. Fix an absolute +$`c_I\ge1`$ large enough that an s-output indicator, including its +output word, has T-count, Clifford count, and dirty width each at most +$`c_IsP_a`$. This holds even if its address length is below a, since +its actual chunks are truncated to that length. A zero-bit indicator +is simply X on its sole output. + +Let B be additional dirty width, excluding the query address, output, +and every existing compiler reservation. The following sufficient +condition makes the allocation direct: + +```math +B\ge16c_I(P_a+r+1). +``` + +For $`Q\ge2`$, choose s as the largest power of two at most + +```math +\min\!\left\{2^{\lfloor r/2\rfloor},\frac{B}{8c_IP_a}\right\}. +``` + +Take $`H=J=s`$ and $`K=Q/s^2`$. The unassigned address bits are +the high-block field c. Both H and J divide Q, and K is an integer +power of two; if r is odd the balanced count cap leaves at least one +high-block bit. Each indicator gets its own private dirty pool. Their +combined width is at most B divided by four, and the traversal selectors +plus e fit the stated remaining allocation. Thus the whole query fits B. +For $`Q=1`$, emit its sole row directly as a Clifford X word. + +Power-of-two rounding gives + +```math +\frac1s=O\!\left(\frac1{\sqrt Q}+\frac{P_a}{B}\right), +\qquad +K=O\!\left(1+\frac{QP_a^2}{B^2}\right). +``` + +Since $`s\le\sqrt Q`$, Section 3 yields the uniform upper bounds + +```math +T=O\!\left(m\sqrt Q+\frac{QmP_a}{B}+\sqrt QP_a\right), +``` + +```math +G=O(Qm+\sqrt QP_a), +``` + +```math +D_T=O\!\left( + m\left[1+\frac{QP_a^2}{B^2}\right] + +\left\lceil\frac{\log_2s}{\min\{a,\log_2s\}}\right\rceil + \log_2(\min\{a,\log_2s\}+2)\right). +``` + +The indicator-depth term is defined as zero when s equals one. This +standalone query bound does not assert optimal dependence on arbitrary +word precision m. In the fixed-accuracy Hopf composition, its extra +precision factors occur in geometrically weighted layer sums. + +## 5. Fixed-accuracy composition at every eligible width + +Use the existing theorem's literal threshold $`b\ge17B_0`$ from +the opening statement. Separate a small-width branch from the new +construction. If + +```math +b\le\sqrt N/n, +``` + +then the [amortized fixed-accuracy theorem](AMORTIZED_DIRTY_LOOKUP.md#fixed-accuracy-corollary) +already gives +$`D_T=O_\eta(N/b^2+n^2)=O_\eta(N/b^2)`$ with the required +same-circuit T and Clifford counts. No query or source is changed on +this branch. + +For the remainder suppose $`b>\sqrt N/n`$ and n is sufficiently +large, with the threshold depending only on eta and fixed circuit +constants. Use the unary compiler's +[early groups and error allocation](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy). +In detail, choose the least power of two +$`q\ge4\pi\sqrt{2n}/\eta`$, put $`\rho=\log_2 3`$ and +$`\delta=\eta/(8n)`$, and retain its sufficient local constant A +and cutoff + +```math +K_*=\left\lceil16A(q^\rho+q\log_2q+\log_2q+1)\right\rceil, +\qquad C=16A. +``` + +At remaining height $`k>K_*`$, use +$`g=\lfloor\log_2(k/(Cq))\rfloor`$. Their program width is +$`w=q(2^g-1)\le k/C`$. These groups fit their conditionally zero +logical suffixes, number $`O_\eta(n/\log(n+2))`$, and have +$`O_\eta(n)`$ total internal, prefetch, and predicate depth. +Their local gate counts are polynomial in n. The cutoff satisfies + +```math +K_*=O_\eta(n^{\rho/2}+\sqrt n\log(n+2)), +\qquad K_*\log(n+2)=o_\eta(n). +``` + +The early parallel query helper pools still fit the smaller available +dirty width. Their peak is bounded by + +```math +O\!\left(\sqrt N(n+2)^3 + \sum_{k>K_*}k2^{-k/2}\right) +=o_\eta(\sqrt N/n). +``` + +Their T-count is $`O_\eta(\sqrt N)`$ and their Clifford count is +$`O_\eta(N)`$, exactly as in the square-root-width proof. These +pools are separate from the conditional logical suffix and are returned +before every group begins. + +Set + +```math +L'=\max\{6,\lceil\log_2(4/\eta)\rceil\}, +\qquad B'_0=L'+n+7, +\qquad B=b-B'_0-2. +``` + +Since $`L'\le L+2`$, the original literal threshold ensures +$`B\ge b/2`$. On the large-width branch, +$`B>\sqrt N/(2n)`$, so this pool is larger than any fixed polynomial +in n for sufficiently large n. The source base, two predicate helpers, +and two external clean flags remain separately reserved. The predicate +helpers are arbitrary dirty inputs and return exactly. + +### A width-constrained query on the entire tail + +After the last unary group, the remaining height is +$`K'\le K_*`$. For every $`1\le k\le K'`$, use the old +full-input operator source and literal amplification with its capped +precision + +```math +m_k=L'+4+\min\{k,\lceil\log_2(8n)\rceil\} +\le m=O_\eta(\log(n+2)). +``` + +Its exact table has $`Q_k=4N2^{-k}`$ rows and +$`r_k=n-k+2\le n+1`$ address bits. Use Section 4 with + +```math +a_k=\min\{n+2,2^{\lfloor k/12\rfloor}\}, +\qquad P_k=(a_k+2)^3. +``` + +Both useful estimates hold: + +```math +P_k\le(n+4)^3,\qquad P_k\le27\,2^{k/4}. +``` + +The first ensures the sufficient width condition +$`B\ge16c_I(P_k+r_k+1)`$ for every tail query on this branch, +including the smallest k. The second makes the following resource sums +converge. No source or predicate borrows a live query register. + +The query's indicator depth is + +```math +O\!\left(n(k+1)2^{-k/12}+\log(n+2)\right). +``` + +Indeed, when the untruncated exponential cap is below the indicator +address length, its reciprocal is within a fixed factor of +$`2^{-k/12}`$ and its logarithm is $`O(k+1)`$. Otherwise the +indicator has one chunk and logarithmic address-length depth. This +argument also covers s equal to one by the Clifford convention above. +Thus each completed query obeys + +```math +T_k=O\!\left(m_k\sqrt{Q_k} + +\frac{Q_km_kP_k}{B}+\sqrt{Q_k}P_k\right), +``` + +```math +G_k=O(Q_km_k+\sqrt{Q_k}P_k), +``` + +```math +D_{T,k}=O\!\left(m_k+ + \frac{Q_km_kP_k^2}{B^2} + +n(k+1)2^{-k/12}+\log(n+2)\right). +``` + +The bounded number of queries and actual inverses per amplified stage +only changes constants. These are exact full-input replacements for +the old table words, so the source's accepted block, entire rejected +action, and full error certificate are unchanged. + +### Summing the tail resources + +Use $`m_k\le L'+4+k`$ for the weighted count and chunk-depth +terms. With $`Q_k=4N2^{-k}`$ and $`P_k\le27\,2^{k/4}`$, + +```math +\sum_{k=1}^{K'}m_k\sqrt{Q_k}=O_\eta(\sqrt N),\qquad +\sum_{k=1}^{K'}\sqrt{Q_k}P_k=O(\sqrt N), +``` + +```math +\sum_{k=1}^{K'}Q_km_kP_k=O_\eta(N),\qquad +\sum_{k=1}^{K'}Q_km_kP_k^2=O_\eta(N),\qquad +\sum_{k=1}^{K'}Q_km_k=O_\eta(N). +``` + +For example, the P-squared sum is bounded by a constant times +$`N\sum_{k\ge1}(L'+4+k)2^{-k/2}`$. Hence the whole tail has + +```math +T=O_\eta\!\left(\sqrt N+\frac NB\right),\qquad G=O_\eta(N). +``` + +For the unweighted terms retain the cap $`m_k\le m`$ rather +than the looser bound linear in k. The sum +$`\sum_{k\ge1}(k+1)2^{-k/12}`$ is finite, so + +```math +\sum_{k=1}^{K'}D_{T,k} +=O_\eta\!\left(\frac N{B^2}+n+K_*\log(n+2)\right) +=O_\eta\!\left(\frac N{b^2}+n\right). +``` + +The old full-input source words, reflections, and suffix predicates add +$`O_\eta(K_*\log(n+2))=o_\eta(n)`$ depth and polynomial gate +counts. Every query fits the extra pool B by construction; this pool is +reused only after its exact return. The early-group resources already +fit the same b. Consequently all three advertised bounds hold on the +same circuit. + +### Error, the literal threshold, and the matching interval + +The unary group's unchanged error argument charges at most eta divided +by four for rounding all early angles, using the full-frame angular +square-sum bound; at most eta divided by four for their native source +preparations and actual inverses; and at most eta divided by four for +the remaining old source layers with accuracy parameter eta divided +by four. This gives total initialized-isometry error at most eta. +All work return and reference correlations are included. Exact table +replacement introduces no additional approximation error, and no +intermediate source is assumed reset. + +The finitely many n below the large-n thresholds use the original +amortized theorem at accuracy eta and literal width $`b\ge17B_0`$. +Its depth $`O_\eta(N/b^2+n^2)`$ becomes +$`O_\eta(N/b^2+n)`$ over that finite set by increasing only the +eta-dependent constant. The same fallback already handles the +small-width branch for every n. Thus the global statement retains +$`17(L+n+7)`$ despite using $`L'`$ in the large-width construction; +it does not silently impose $`17(L'+n+7)`$ on the whole theorem. + +Finally, the inherited complete-frame lower bounds give +$`T^\star=\Omega_\eta(\sqrt N+N/b)`$ and +$`D_T^\star=\Omega_\eta(N/b^2)`$ under this width threshold. +For $`b\le\sqrt{N/n}`$, the n term in the new depth bound is at +most $`N/b^2`$, and the square-root count term is at most $`N/b`$. +This proves the simultaneous matching interval in the opening +statement. It makes no lower-bound claim for the additive n outside +that interval. + +## 6. Proof and evidence boundary + +The controlled bilinear echo, high-block selection, and two-indicator +identity are exact algebraic circuit proofs. They use the established +shared-control native phase schedule, rank reduction, dirty traversal, +and chunked dirty indicators; none is an assumed quantum table oracle. +The global theorem combines these identities with the already proved +conditional unary group, angular error bound, and original full-input +source certificate. Its count and depth estimates follow from the +explicit reservations and geometric sums above, not finite-size fits. + +The [five bounded checks](../tests/test_blocked_bilinear_lookup.py) verify +native controlled bilinear leaves on every input, one complete eight-wire +native selected-block word, and its actual inverse. Exact Boolean +polynomials check complete queries, including arbitrary indicator words, +selector stacks, outputs, and helpers. Emitted schedules check literal +T-count, T-depth, and disjoint T targets. Negative controls expose a +missing helper inverse, omitted traversal, and incorrect block labels. +The complete-query checks use a reduced controlled-bilinear gate and the +existing routed indicator; they do not emit asymptotic chunked indicators +or a variable-size complete-frame compiler. Their exact scope is recorded +in [verification](VERIFICATION.md). + +The [source map](SOURCE_MAP.md#5-fault-tolerant-sources-and-contribution-boundaries) +and [related-work comparison](RELATED_WORK.md#20-blocked-bilinear-queries-and-the-dirty-width-tradeoff-3-october-2026) +attribute the dirty traversal, phase synthesis, and bilinear framework. +The additional result is their charged selected-block interface and +fixed-accuracy width composition, not generic priority for phase +polarization or dirty cancellation. No unrestricted depth lower bound +is inferred from the finite checks. diff --git a/docs/OPEN_PROBLEM.md b/docs/OPEN_PROBLEM.md index 49a3e91..4082360 100644 --- a/docs/OPEN_PROBLEM.md +++ b/docs/OPEN_PROBLEM.md @@ -39,7 +39,7 @@ hybrid depth bounds, write | One-clean phase-dressed complex magnitude frame | The same grouped and banked T-counts, at $`b\ge L+n+8`$ and $`b\ge2(L+n+8)`$, respectively | Compose the real compiler and literal phase diagonal in the same workspace; [composition corollary](ONE_CLEAN_COMPILER.md#8-phase-dressed-complex-magnitude-frames) | | T-depth with additional dirty banks | $`D_T=O(NL/b+\min\{nL+n^2,L\ell_*(n)+n^3\})`$ at $`a=2`$, $`b\ge2(L+n+7)`$, with $`T,G=O(NL)`$ | Choose between the layerwise and grouped [schedules](T_DEPTH_COMPILER.md); real frames; optimizing depth may increase T-count; no matching frontier established | | Simultaneous T-count and T-depth | $`T=O(\sqrt{NL}+L\ell_*(n))`$, $`D_T=O(\min\{nL+n^2,L\ell_*(n)+n^3\})`$, $`G=O(NL)`$, at $`a=2`$, $`b\ge C(L+n+7+\sqrt{NL})`$ | Same real-frame circuit, for sufficiently large fixed C; [parallel dirty lookup](PARALLEL_DIRTY_LOOKUP.md); T-depth optimality remains open | -| Fixed-accuracy count and depth at modest width | $`T=O(\sqrt N+N/b)`$, $`D_T=O(N/b^2+n\chi(n))`$, $`G=O(N)`$, at $`a=2`$, fixed L, $`b\ge17(L+n+7)`$ | Same complete real-frame circuit; count is optimal in order, and depth is matching for $`b\le\sqrt N/n`$; [amortized dirty lookup](AMORTIZED_DIRTY_LOOKUP.md) | +| Fixed-accuracy count and depth versus width | $`T=O(\sqrt N+N/b)`$, $`D_T=O(N/b^2+n)`$, $`G=O(N)`$, at $`a=2`$, fixed L, $`b\ge17(L+n+7)`$ | Same complete real-frame circuit; count is optimal in order, and depth is matching through $`b\le\sqrt{N/n}`$; [blocked bilinear lookup](BLOCKED_BILINEAR_LOOKUP.md) | | Variable-accuracy count and depth | $`T=O(\sqrt{NL}+NL/b+nL)`$, $`D_T=O(NL/b^2+nL+n\chi(n))`$, $`G=O(NL)`$, at $`a=2`$, $`L\ge6`$, $`b\ge17(L+n+7)`$ | Same complete real-frame circuit; both are matching when $`b\le\sqrt{NL/(nL+n\chi(n))}`$; [hybrid composition](PARALLEL_DIRTY_LOOKUP.md#every-eligible-width-and-precision) | | Fixed-accuracy large-width depth | $`T=O_\eta(\sqrt N)`$, $`G=O_\eta(N)`$, $`D_T=O_\eta(n)`$, at $`a=2`$, $`b\ge C_\eta\sqrt N`$ | Same complete real-frame circuit with charged unary source preparation/return; [unary theorem](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy); T-count is optimal in order, depth lower bound remains $`\Omega(1)`$ | @@ -81,7 +81,7 @@ uses $`a=2`$ throughout and respects each sufficient allocation threshold. | Regime | Depth lower bound | Available depth upper bound | Remaining issue | |---|---|---|---| | Fixed L, $`b=\Theta(n)`$ and $`b\ge17B_0`$ | $`\Omega(N/n^2)`$ | $`O(N/n^2)`$ with $`T=\Theta(N/n)`$ | Matching count and depth in one circuit; [amortized schedule](AMORTIZED_DIRTY_LOOKUP.md) | -| Fixed L, $`17B_0\le b\le\sqrt N/n`$ | $`\Omega(N/b^2)`$ | $`O(N/b^2)`$ with optimal-order count | Matching throughout this interval when nonempty | +| Fixed L, $`17B_0\le b\le\sqrt{N/n}`$ | $`\Omega(N/b^2)`$ | $`O(N/b^2)`$ with optimal-order count | Matching throughout this interval when nonempty | | Variable L, $`17B_0\le b\le\sqrt{NL/(nL+n\chi(n))}`$ | $`\Omega(NL/b^2)`$ | $`O(NL/b^2)`$ with $`T=\Theta(NL/b)`$ | Matching throughout this interval when nonempty | | $`L=\Theta(n)`$, sufficient $`b=\Theta(n)`$ | $`\Omega(N/n)`$ | $`O(N/n)`$ with $`T=\Theta(N)`$ | Matching at inverse-polynomial error in N for sufficiently large n | | Fixed L, $`2B_0\le b\lt17B_0`$ | $`\Omega(N/n^2)`$ | $`O(N/n)`$ | Earlier schedule remains the proved fallback at this literal reservation | @@ -108,8 +108,13 @@ programs with a summable chunked-indicator budget for the late layers. The [unary phase-source refinement](UNARY_PHASE_GRADIENT.md#7-complete-frame-theorem-at-fixed-accuracy) now gives $`O(n)`$ depth under the same external-flag and sufficient dirty-width orders. It charges source preparation/inversion and the -coherent one-hot program interface. -The lower bound is still constant in this regime; depth optimality is open. +coherent one-hot program interface. The [blocked bilinear refinement](BLOCKED_BILINEAR_LOOKUP.md) +extends this to $`D_T=O(N/b^2+n)`$ throughout +$`b\ge17(L+n+7)`$, retaining the optimal-order count +$`T=O(\sqrt N+N/b)`$. Both orders match through +$`b\le\sqrt{N/n}`$ when the interval is nonempty. +At square-root-scale dirty width, the depth lower bound is still constant, +so depth optimality there remains open. The lower bound does not assume count optimality; the upper circuit also retains optimal-order T-count. No high-precision endpoint improvement follows. For varying L outside the matching interval, the new count bound retains @@ -180,10 +185,13 @@ eigenstate and constant-T-depth programmed cyclic shifts. Conditional bilinear work returns exactly on all source inputs, the inactive sector is literal identity, and source preparation plus actual inversion costs at most twice its preparation error for the entire group. Its global -allocation proves the displayed $`O(n)`$ upper bound. The next bounded -question is a width-dependent $`O(N/b^2+n)`$ depth extension retaining -$`O(\sqrt N+N/b)`$ T-count; this remains unproved. Its first obligation -is a complete reduced-width prefetch and workspace schedule. +allocation proves the displayed $`O(n)`$ upper bound. The +[blocked bilinear theorem](BLOCKED_BILINEAR_LOOKUP.md) now closes the +width-dependent extension. It replaces the residual bank route by +selected bilinear blocks with one returned dirty phase helper, then +charges the coupled block-size/chunk-length allocation. The next bounded +question is the precision dependence of this route; its unary source +width and suffix cutoff must be made explicit before any uniform-L claim. The [two-layer obstruction](SHALLOW_SOURCE_OBSTRUCTION.md) separately allows unrestricted Clifford interlayers: the original source at width diff --git a/docs/PARALLEL_DIRTY_LOOKUP.md b/docs/PARALLEL_DIRTY_LOOKUP.md index 7fd50b1..4264baa 100644 --- a/docs/PARALLEL_DIRTY_LOOKUP.md +++ b/docs/PARALLEL_DIRTY_LOOKUP.md @@ -63,12 +63,12 @@ elementary depth. The selected allocation $`L=N,b=N+n+7`$ is not covered by the new sufficient-width hypothesis; its linear T-count endpoint remains open. -At fixed accuracy, [grouped conditional program reuse](GROUPED_PROGRAM_PREFETCH.md#8-complete-frame-theorem-at-fixed-accuracy) -and [chunked dirty indicators](CHUNKED_DIRTY_INDICATOR.md) further give -$`D_T=O(n\log\log(n+2))`$ with the same optimal-order count at -sufficient square-root-scale dirty width. Their global proof charges -wide prefetches and the entire late-query segment. The uniform hybrid -theorem and matching range above remain valid separately. +At fixed accuracy, the [blocked bilinear refinement](BLOCKED_BILINEAR_LOOKUP.md) +uses conditional unary groups and width-constrained bilinear blocks to +give $`D_T=O(N/b^2+n)`$, retaining optimal-order T-count throughout +$`b\ge17B_0`$. Count and depth match through $`b\le\sqrt{N/n}`$ +when nonempty. Its fixed-accuracy proof does not replace the uniform +precision statements above. ## 1. An exact dirty indicator by conjugated routing diff --git a/docs/README.md b/docs/README.md index a227431..f63e0fd 100644 --- a/docs/README.md +++ b/docs/README.md @@ -27,6 +27,7 @@ topic has one primary chapter below. | [Conditional geometric source](CONDITIONAL_GEOMETRIC_SOURCE.md) | Logarithmic precision depth using the active logical suffix as temporary clean work; lookup and suffix-predicate costs remain separate | | [Grouped program reuse](GROUPED_PROGRAM_PREFETCH.md) | Complete real-frame depth O(n log log n) at fixed accuracy and sufficient square-root dirty width; O(n) selector maintenance and scoped source-reuse audit | | [Unary phase-source groups](UNARY_PHASE_GRADIENT.md) | Complete real-frame T-depth O(n) at fixed accuracy and sufficient square-root dirty width, retaining optimal-order T-count; charged preparation and exact guarded cyclic shifts | +| [Blocked bilinear lookup](BLOCKED_BILINEAR_LOOKUP.md) | Full-input selected bilinear blocks with returned dirty work; fixed-accuracy depth O(N/b²+n), optimal-order T-count, and matching range through sqrt(N/n) | | [Chunked dirty indicator](CHUNKED_DIRTY_INDICATOR.md) | A tunable exact dirty-tree indicator and a summable late-query budget that removes the late routing bottleneck | | [Hopf error accumulation](HOPF_ERROR_ACCUMULATION.md) | Sharp ideal-angle stability, finite relative spectra, and coherent leakage in the actual shared-flag sources; scoped precision boundaries | | [Flag-echo audit](HOPF_FLAG_ECHO.md) | Exact errors of four diagonal Pauli echoes, their generic linear leakage, and an exact equal-mask exception | diff --git a/docs/RELATED_WORK.md b/docs/RELATED_WORK.md index 45fbf58..4a55ad1 100644 --- a/docs/RELATED_WORK.md +++ b/docs/RELATED_WORK.md @@ -928,4 +928,51 @@ state or intermediate measurement. This improves the fixed-accuracy depth upper bound; it proves neither an unrestricted matching depth lower bound nor the constant-clean high-precision endpoint. The sources above identify inherited ingredients and interface distinctions, not -priority for the composite construction. +priority for the composite construction. This square-root-width result +is also a corollary of the broader fixed-accuracy tradeoff below. + +## 20. Blocked bilinear queries and the dirty-width tradeoff (3 October 2026) + +The [blocked bilinear lookup](BLOCKED_BILINEAR_LOOKUP.md) combines the +existing [two-pass dirty traversal](AMORTIZED_DIRTY_LOOKUP.md#2-selecting-a-shear-with-dirty-unary-traversal) +with the [bilinear indicator echo](PARALLEL_DIRTY_LOOKUP.md#5-a-bilinear-query-reduction). +Their lineage remains [LKS, Appendix C](https://arxiv.org/html/1812.00954v2), +[Khattar–Gidney, Sections 4 and 7](https://arxiv.org/html/2407.17966v2), +and the controlled-linear and commuting-basis constructions of +Kim–Laakkonen and Boyd discussed in Section 9. Standard Boolean phase +polarization supplies the controlled leaf: for a bilinear phase P, +apply $`(-1)^{dP}`$, toggle the dirty bit d by hz, apply the phase +again, and undo the toggle. This full four-step word leaves exactly +$`(-1)^{hzP}`$ and returns d. +Neither that algebra nor generic dirty cancellation is a novelty claim. + +The local proof establishes a rank-sensitive native leaf with preserved +controls and exact helper return, a selected-block traversal, and a chunk +allocation whose depth sums across the final Hopf layers. Together with +the unary early groups, it gives, for each fixed $`0\lt\eta\le1/64`$, +$`L=\max\{6,\lceil\log_2(1/\eta)\rceil\}`$ and +$`b\ge17(L+n+7)`$, one complete real-frame circuit with + +```math +T=O_\eta\!\left(\sqrt N+\frac Nb\right),\qquad G=O_\eta(N), +\qquad D_T=O_\eta\!\left(\frac N{b^2}+n\right). +``` + +Two external clean flags suffice. Full initialized-isometry error, +literal phase and arbitrary dirty-reference return retain their existing +contracts. Sources, actual inverses and all queries are charged. + +When the interval is nonempty, the inherited count lower bound and the +physical width $`n+2+b=\Theta(b)`$ give simultaneous worst-case matches: + +```math +17(L+n+7)\le b\le\sqrt{N/n},\qquad +T^\star=\Theta_\eta(N/b),\quad D_T^\star=\Theta_\eta(N/b^2). +``` + +This extends the fixed-accuracy matching window; it introduces no new +external synthesis premise or depth lower-bound method. The preceding +variable-accuracy bounds remain separate and unchanged. The unary +square-root-width schedule remains a valid predecessor and corollary; +unrestricted large-width depth optimality and the constant-clean +high-precision endpoint remain open. No generic lookup priority is claimed. diff --git a/docs/SOURCE_MAP.md b/docs/SOURCE_MAP.md index 75df7f3..7093906 100644 --- a/docs/SOURCE_MAP.md +++ b/docs/SOURCE_MAP.md @@ -202,6 +202,7 @@ is inherited. | R41 | bounded-input construction and classical comparison | [algebraic residual coefficients](RESIDUAL_TABLE_PREPROCESSING.md) use standard half-phase identities with one shared root, certified rational intervals, and a finite-radius cutoff to preserve R36/R39's error constants without Euler search; [bounded-input audit](BOUNDED_INPUT_QBP.md) combines F5 existence with coarse enumeration, explicit masks and instruction output to prove polynomial construction for the listed grouped/state alternatives; the banked small-system source uses $`P+13`$ dirty wires. Deterministic and term-sampled classical Pauli baselines are charged; no unconditional efficient fine-word search, generic Euler-runtime theorem, or end-to-end quantum advantage is claimed | | R42 | state-based QBP T-depth composition | [depth proof](STATE_QBP_DEPTH.md) composes R21/R23's exact schedules, inherited from F2, with R36/R39's constant number of residual rotations and the actual exact-return coarse interpreter; with $`B_0=P+n+7`$, $`b\ge2B_0`$ gives $`D_T=O(NP/b+P+n^3)`$, $`T,G=O(NP)`$, while $`b\ge16(B_0+\sqrt{NP})`$ gives one circuit with $`T=O(\sqrt{NP}+P+n\sqrt N)`$, $`G=O(NP)`$, and $`D_T=O(P+n^3)`$; both real/complex task streams and oracle depth are charged; no new lookup primitive, depth optimality, total-runtime gain, or general emitter is claimed | | R43 | fixed-accuracy linear T-depth complete real frame | [unary phase-source proof](UNARY_PHASE_GRADIENT.md), 3 October 2026: coherent one-hot shifts, guarded bilinear work, unitary source preparation/return, and the Hopf angle-stability/group/query allocation give $`D_T=O_\eta(n)`$, $`T=O_\eta(\sqrt N)`$, $`G=O_\eta(N)`$ with two external clean flags and sufficient $`C_\eta\sqrt N`$ dirty work; F5/F8/F31/F36/F37 are attributed ingredients, F29/F38 are comparisons; no supplied catalyst, generic synthesis priority, matching depth lower bound, or high-precision endpoint follows | +| R44 | fixed-accuracy dirty-width/depth tradeoff | [blocked bilinear lookup](BLOCKED_BILINEAR_LOOKUP.md), composed with R43's early groups: for fixed $`0\lt\eta\le1/64`$, $`L=\max\{6,\lceil\log_2(1/\eta)\rceil\}`$ and $`b\ge17(L+n+7)`$, one two-clean complete real-frame circuit has $`T=O_\eta(\sqrt N+N/b)`$, $`G=O_\eta(N)`$ and $`D_T=O_\eta(N/b^2+n)`$; simultaneous worst-case orders are $`\Theta_\eta(N/b)`$ and $`\Theta_\eta(N/b^2)`$ when $`b\le\sqrt{N/n}`$; the full-input leaf and allocation use existing F2/F8/F29/F30/F31 ingredients, and F3/F4 supply the lower-bound lineage; large-width depth optimality and the high-precision endpoint remain open | The [Hopf error audit](HOPF_ERROR_ACCUMULATION.md) derives a sharp ideal-angle stability recurrence and an exact finite relative-spectrum recursion from @@ -328,6 +329,21 @@ different program, preparation, and workspace contracts. No priority claim is inferred from these comparisons, and the high-precision endpoint and unrestricted T-depth optimality remain open. +The [blocked bilinear query](BLOCKED_BILINEAR_LOOKUP.md), also added +**3 October 2026**, extends that fixed-accuracy conclusion to every +$`b\ge17(L+n+7)`$ through R44. Its controlled leaf polarizes a +bilinear Boolean phase with a returned dirty bit; the completed leaves +then use the existing two-pass dirty traversal and bilinear indicator +echo. F2/F8/F29/F30/F31 retain their established query, traversal, +controlled-linear and phase-synthesis roles. The local proof adds the +literal full-input leaf, charged chunk allocation and complete-frame +composition, without a new external premise or generic priority claim. +R43 remains a valid construction and a square-root-width corollary. +The previous variable-accuracy theorem remains valid separately; R44's +constants depend on fixed eta. Its enlarged matching window follows +from the existing F3/F4 count lower bound divided by physical width, +not a new depth lower-bound method. + The [consolidated state-based QBP theorem](STATE_BASED_QBP_THEOREM.md) collects R36 and R38–R42 under one input, precision, workspace, and sampling contract. It introduces no additional compiler bound or diff --git a/docs/UNARY_PHASE_GRADIENT.md b/docs/UNARY_PHASE_GRADIENT.md index 1570391..8c0df41 100644 --- a/docs/UNARY_PHASE_GRADIENT.md +++ b/docs/UNARY_PHASE_GRADIENT.md @@ -15,6 +15,11 @@ fixed-accuracy T-depth $`O_\eta(n)`$ with optimal-order T-count and square-root-scale dirty width. This construction does not replace the previous conjugated geometric reflection by a shallow implementation. +The [blocked bilinear extension](BLOCKED_BILINEAR_LOOKUP.md) retains this +source construction and proves fixed-accuracy depth $`O(N/b^2+n)`$ with +optimal-order T-count throughout the original sufficient dirty-width +range. The square-root-width theorem below remains valid. + ## 1. Local contract Use the group registers of [the grouped interface](GROUPED_PROGRAM_PREFETCH.md#1-group-program-and-local-statement): diff --git a/docs/VERIFICATION.md b/docs/VERIFICATION.md index d932268..896094b 100644 --- a/docs/VERIFICATION.md +++ b/docs/VERIFICATION.md @@ -205,6 +205,24 @@ rectangular binary basis changes, literal shared-target Toffoli phases, the four-corner echo, actual inverses, and arbitrary dirty-input return. These bounded fixtures use the existing routed indicators. They do not supply the shallow indicator themselves. + +The [blocked bilinear checks](../tests/test_blocked_bilinear_lookup.py) +add five bounded audits of the [selected-block query](BLOCKED_BILINEAR_LOOKUP.md). +Native matrices cover rank-zero, rank-one, and rank-two controlled +bilinear words, rectangular coordinate changes, arbitrary helper return, +and an eight-wire selected block with a dirty traversal selector. Their +emitted schedules check literal phases, actual inverses, T-count, +T-depth, and disjoint targets within every T layer. Complete queries +with three or four address bits and one or two output bits are checked +as exact Boolean polynomials in every input wire, including both dirty +indicator words, the selector stack, the helper, and arbitrary outputs. +These larger queries use a reduced controlled-bilinear action whose +native interface is checked separately. Negative cases remove the +helper inverse or second traversal, or select the wrong block. The +fixtures do not emit the scalable chunked indicators or establish the +width-sensitive frame theorem by extrapolation; that resource and +composition argument remains analytic. + The [dirty-counter checks](../tests/test_counter_dirty_indicator.py) audit the two-adder signed increment, both modular-adder actions, cyclic routing, nested full-input echoes, actual inverses, and parallel native diff --git a/tests/README.md b/tests/README.md index 8601bba..f5f26b1 100644 --- a/tests/README.md +++ b/tests/README.md @@ -57,6 +57,7 @@ theorem by numerical extrapolation. | [`test_radial_filter.py`](test_radial_filter.py) | Fixed-point source filtering with literal phase, full polar error, phase-approximation budgets, actual inverse, and smaller precision allocation; native phase-word costs are analytic | | [`test_parallel_dirty_lookup.py`](test_parallel_dirty_lookup.py) | Scratch-free routed indicators, actual-inverse orientation, literal phases, symbolic all-input return, complete native queries, and disjoint T layers | | [`test_bilinear_dirty_lookup.py`](test_bilinear_dirty_lookup.py) | Rectangular bilinear basis changes, shared-target native phases, the two-indicator query echo, actual inverses, and arbitrary dirty-input return; uses existing routed indicators | +| [`test_blocked_bilinear_lookup.py`](test_blocked_bilinear_lookup.py) | Native dirty-controlled bilinear phases and selected blocks, symbolic blocked queries on every input variable, actual inverses, emitted T-layer ledgers, and missing-helper/selection counterchecks; no scalable chunked-query emitter | | [`test_dirty_indicator_depth.py`](test_dirty_indicator_depth.py) | Exact emitted parity phases, two-layer dirty-helper Toffoli, complete six-wire indicator, actual inverses, and no extra helper | | [`test_counter_dirty_indicator.py`](test_counter_dirty_indicator.py) | Two-adder signed increments, TTK and shortened RV modular-adder actions, cyclic echoes, arbitrary dirty offsets, full indicator return, and shared-address native scheduling; optimized RV depth is analytic | | [`test_readonly_dirty_increment.py`](test_readonly_dirty_increment.py) | Two-dirty-bit involution increment, both literal polarities, full native phases and actual inverses, shared-address scheduling, and wrong-order decrement witness; optimized increment depth is analytic | diff --git a/tests/test_blocked_bilinear_lookup.py b/tests/test_blocked_bilinear_lookup.py new file mode 100644 index 0000000..6b5eb3d --- /dev/null +++ b/tests/test_blocked_bilinear_lookup.py @@ -0,0 +1,271 @@ +"""Bounded audits of a dirty-controlled bilinear block lookup. + +Native small matrices retain literal phases, arbitrary dirty helpers, and +actual inverses. Larger complete queries use exact Boolean-polynomial +gates, with a reduced C3X for the already-audited controlled bilinear +action. Every input wire, including both indicator words and traversal +selectors, remains a symbolic variable. These checks do not emit the +asymptotic chunked indicators or prove the width-sensitive frame bounds. +""" +from __future__ import annotations + +import unittest + +import numpy as np + +try: + from .test_amortized_dirty_lookup import _normal_form + from .test_bilinear_dirty_lookup import _bilinear + from .test_operator_source_compiler import _word_matrix + from .test_parallel_dirty_lookup import _classical_indicator, _expected, _multiply + from .test_t_depth import _ccz_batch, _inverse, _native_schedule, _word +except ImportError: + from test_amortized_dirty_lookup import _normal_form + from test_bilinear_dirty_lookup import _bilinear + from test_operator_source_compiler import _word_matrix + from test_parallel_dirty_lookup import _classical_indicator, _expected, _multiply + from test_t_depth import _ccz_batch, _inverse, _native_schedule, _word + + +ATOL = 8e-11 + + +def _controlled_bilinear(matrix, control, left, right, output, helper, *, native=False): + wires = [control, *left, *right, output, helper] + assert len(wires) == len(set(wires)) + before, rank = _normal_form(matrix, right, left) + if not rank: + return [], 0 + left_set = set(left) + before = [('CX', gate[2], gate[1]) if gate[1] in left_set else gate + for gate in before] + if not native: + # Semantic all-input interface; literal phases/helper return are + # checked independently below using the D,E,D†,E† native word. + middle = [('C3X', control, left[j], right[j], output) for j in range(rank)] + return before + middle + list(reversed(before)), rank + pre = [('C', before + [('H', output)])] + diagonal = _ccz_batch(helper, list(zip(left[:rank], right[:rank]))) + flip = _native_schedule([('CCX', control, output, helper)]) + return (pre + diagonal + flip + _inverse(diagonal) + _inverse(flip) + + _inverse(pre)), rank + + +def _selected_bilinear(address, left, right, output, selectors, helper, blocks, + *, native=False): + """Enabled DFS followed by the actual inverse of root-disabled DFS.""" + wires = address + left + right + output + selectors + [helper] + assert len(wires) == len(set(wires)) and len(selectors) == len(address) + assert len(blocks) == 1 << len(address) + reverse = _inverse if native else lambda word: list(reversed(word)) + if not address: + word, ranks = [], [] + for target, matrix in zip(output, blocks[0]): + leaf, rank = _bilinear(matrix, left, right, target, native=native) + word += leaf + ranks.append(rank) + return word, [ranks], word + leaves, ranks = [], [] + for block in blocks: + leaf, leaf_ranks = [], [] + for target, matrix in zip(output, block): + part, rank = _controlled_bilinear( + matrix, selectors[-1], left, right, target, helper, native=native) + leaf += part + leaf_ranks.append(rank) + leaves.append(leaf) + ranks.append(leaf_ranks) + + def traversal(root_enabled): + def descend(level, prefix): + if level == len(address): + return leaves[prefix] + result = [] + for value in (0, 1): + control, target = address[level], selectors[level] + negative = [('X', control)] if value == 0 else [] + toggle = ('CX', control, target) if level == 0 else ( + 'CCX', selectors[level - 1], control, target) + update = negative + [toggle] + list(reversed(negative)) + if level == 0 and not root_enabled: + update = [] + if native: + update = _native_schedule(update) + result += (update + descend(level + 1, prefix | (value << level)) + + reverse(update)) + return result + return descend(0, 0) + + first, second = traversal(True), traversal(False) + return first + reverse(second), ranks, first + + +def _symbolic(word, initial): + values = [set(value) for value in initial] + for name, *wires in word: + assert name in ('X', 'CX', 'CCX', 'C3X') + product = {0} + for control in wires[:-1]: + product = _multiply(product, values[control]) + values[wires[-1]] ^= product + return values + + +def _expected_selected(initial, address, left, right, output, blocks): + result = [set(value) for value in initial] + for block, matrices in enumerate(blocks): + select = {0} + for bit, wire in enumerate(address): + literal = initial[wire] if (block >> bit) & 1 else initial[wire] ^ {0} + select = _multiply(select, literal) + for target, matrix in zip(output, matrices): + for i, row in enumerate(matrix): + for j, bit in enumerate(row): + if bit: + result[target] ^= _multiply( + select, _multiply(initial[left[i]], initial[right[j]])) + return result + + +def _fixture(a_bits, b_bits, high_bits, word_bits): + bits = a_bits + b_bits + high_bits + height, width = 1 << a_bits, 1 << b_bits + address = list(range(bits)) + output = list(range(bits, bits + word_bits)) + left = list(range(bits + word_bits, bits + word_bits + height)) + right = list(range(left[-1] + 1, left[-1] + 1 + width)) + selectors = list(range(right[-1] + 1, right[-1] + 1 + high_bits)) + helper = right[-1] + 1 + high_bits + total = helper + 1 + # Unequal binary block matrices, with zero rows in the upper half. + table = [(((13 * row) ^ (row >> 1) ^ (row >> 2) ^ 5) + & ((1 << word_bits) - 1)) if row < max(1, 1 << (bits - 1)) else 0 + for row in range(1 << bits)] + blocks = [ + [[[table[i + height * j + height * width * c] >> bit & 1 + for j in range(width)] for i in range(height)] for bit in range(word_bits)] + for c in range(1 << high_bits) + ] + high = address[a_bits + b_bits:] + selected, ranks, first = _selected_bilinear( + high, left, right, output, selectors, helper, blocks) + ia = _classical_indicator(address[:a_bits], left) + ib = _classical_indicator(address[a_bits:a_bits + b_bits], right) + + def echo(middle): + reverse = lambda word: list(reversed(word)) + return (middle + ia + reverse(middle) + ib + middle + reverse(ia) + + reverse(middle) + reverse(ib)) + + return dict(total=total, address=address, output=output, left=left, right=right, + selectors=selectors, helper=helper, high=high, blocks=blocks, + table=table, query=echo(selected), selected=selected, ranks=ranks, + first_only=echo(first), echo=echo) + + +class BlockedBilinearLookupTests(unittest.TestCase): + def test_native_controlled_bilinear_literal_phase_dirty_return_and_counts(self): + matrices = ([[0]], [[1]], [[1, 1], [1, 1]], + [[1, 1], [1, 0]], [[1, 1, 0], [1, 0, 1]]) + for matrix in matrices: + height, width = len(matrix), len(matrix[0]) + control, target = 0, 1 + left = list(range(2, 2 + height)) + right = list(range(2 + height, 2 + height + width)) + helper, total = 2 + height + width, 3 + height + width + schedule, rank = _controlled_bilinear( + matrix, control, left, right, target, helper, native=True) + images = [] + for basis in range(1 << total): + bit = sum(matrix[i][j] * (basis >> left[i] & 1) + * (basis >> right[j] & 1) + for i in range(height) for j in range(width)) % 2 + images.append(basis ^ (((basis & 1) * bit) << target)) + expected = np.eye(1 << total)[:, np.argsort(images)] + np.testing.assert_allclose(_word_matrix(total, _word(schedule)), expected, + atol=ATOL, rtol=0) + np.testing.assert_allclose(_word_matrix(total, _word(_inverse(schedule))), + expected.conj().T, atol=ATOL, rtol=0) + self.assertEqual(sum(kind == 'T' for kind, _ in schedule), 16 if rank else 0) + self.assertEqual(sum(len(gates) for kind, gates in schedule if kind == 'T'), + 12 * rank + 2 * (rank % 2) + 14 if rank else 0) + for kind, gates in schedule: + if kind == 'T': + targets = [gate[1] for gate in gates] + self.assertEqual(len(targets), len(set(targets))) + + def test_native_selected_block_on_every_dirty_selector_and_helper_input(self): + # Eight-wire fixture: one high address, two input bits on each side, + # one output, one selector, one returned dirty helper. + address, left, right, output, selectors, helper = [0], [1, 2], [3, 4], [5], [6], 7 + blocks = [[[[1, 1], [1, 0]]], [[[0, 1], [0, 1]]]] + schedule, _, _ = _selected_bilinear( + address, left, right, output, selectors, helper, blocks, native=True) + images = [] + for basis in range(256): + matrix = blocks[basis & 1][0] + bit = sum(matrix[i][j] * (basis >> left[i] & 1) * (basis >> right[j] & 1) + for i in range(2) for j in range(2)) % 2 + images.append(basis ^ (bit << output[0])) + expected = np.eye(256)[:, np.argsort(images)] + np.testing.assert_allclose(_word_matrix(8, _word(schedule)), expected, + atol=ATOL, rtol=0) + np.testing.assert_allclose(_word_matrix(8, _word(_inverse(schedule))), + expected.conj().T, atol=ATOL, rtol=0) + + def test_complete_query_symbolic_all_input_return_and_actual_inverse(self): + for parameters in ((1, 1, 1, 1), (1, 1, 2, 2), (0, 1, 2, 1), + (1, 0, 2, 2), (1, 1, 0, 2)): + f = _fixture(*parameters) + initial = [{1 << wire} for wire in range(f['total'])] + self.assertEqual(_symbolic(f['selected'], initial), _expected_selected( + initial, f['high'], f['left'], f['right'], f['output'], f['blocks'])) + actual = _symbolic(f['query'], initial) + self.assertEqual(actual, _expected(initial, f['address'], f['output'], f['table'])) + self.assertEqual(_symbolic(list(reversed(f['query'])), actual), initial) + # The designated upper-half rows are zero, with all work arbitrary. + inactive = [set(value) for value in initial] + inactive[f['address'][-1]] = {0} + self.assertEqual(_symbolic(f['query'], inactive), inactive) + + def test_two_pass_emitted_native_ledger_and_disjoint_t_targets(self): + for high_bits in (1, 2, 3): + for word_bits in (1, 2): + f = _fixture(1, 1, high_bits, word_bits) + schedule, ranks, _ = _selected_bilinear( + f['high'], f['left'], f['right'], f['output'], f['selectors'], + f['helper'], f['blocks'], native=True) + positive = [rank for row in ranks for rank in row if rank] + traversal_toffolis = 8 * ((1 << high_bits) - 2) + self.assertEqual(sum(kind == 'T' for kind, _ in schedule), + 32 * len(positive) + 4 * traversal_toffolis) + self.assertEqual(sum(len(gates) for kind, gates in schedule if kind == 'T'), + 2 * sum(12 * r + 2 * (r % 2) + 14 for r in positive) + + 7 * traversal_toffolis) + for kind, gates in schedule: + self.assertTrue(all(0 <= wire < f['total'] + for gate in gates for wire in gate[1:])) + if kind == 'T': + targets = [gate[1] for gate in gates] + self.assertEqual(len(targets), len(set(targets))) + + def test_missing_helper_inverse_and_incorrect_block_selection_are_detected(self): + # Omitting the final E† fails to return the arbitrary helper. + diagonal = _ccz_batch(4, [(2, 3)]) + flip = _native_schedule([('CCX', 0, 1, 4)]) + wrong = diagonal + flip + _inverse(diagonal) + correct = wrong + _inverse(flip) + self.assertGreater(np.linalg.norm(_word_matrix(5, _word(wrong)) + - _word_matrix(5, _word(correct)), 2), 1.9) + f = _fixture(1, 1, 2, 2) + initial = [{1 << wire} for wire in range(f['total'])] + expected = _expected(initial, f['address'], f['output'], f['table']) + self.assertNotEqual(_symbolic(f['first_only'], initial), expected) + wrong_selected, _, _ = _selected_bilinear( + f['high'], f['left'], f['right'], f['output'], f['selectors'], + f['helper'], list(reversed(f['blocks']))) + self.assertNotEqual(_symbolic(f['echo'](wrong_selected), initial), expected) + + +if __name__ == '__main__': + unittest.main()