Search arXiv⌕ Search

arXiv · 2609.29480

Shared-End-First Control Ordering in Clean-Ancilla V-Chain Synthesis of Nested Multi-Controlled Cascades: A Quantum Binary-Addition-Tree Case Study

Abstract

Control order is logically irrelevant for a multi-controlled-X gate but can become physically consequential after ordered ancilla-chain synthesis. This work derives and validates a shared-end-first ordering rule for nested-control cascades, using reversible Binary-Addition-Tree (BAT) successor circuits as the primary case. Under Maslov clean-v-chain synthesis, pre-optimization CX cost is order-invariant, whereas shared-end-first ordering exposes P(m)=(m-4)(m-3)/2 inverse relative-phase-CCX pairs across stage boundaries; standard optimization removes 6P(m) CX, giving 12m-31 for QBAT after an independently derived two-CX terminal identity. At m=20, retaining an ascending control list after mirroring the BAT direction yields 1,025 CX, while mirroring the chain orientation restores 209. A disjoint-target nested cascade reproduces the same cancellation count for count, showing that BAT target geometry is not required; an ordered dirty-v-chain control shows zero orientation advantage, delimiting the mechanism within the tested synthesis families. Independent ripple-carry baselines give 11m-28 CX on FULL connectivity, so the QBAT/RC ratio approaches 12/11, with equal natural width 2m-3 in the tested QBAT-A1 and INC_RC implementations. The contribution is a synthesis-ordering rule that exposes cancellation to an existing optimizer, not a new optimizer pass or a universal MCX rule.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei-Chang Yeh. 2026-08-24. Shared-End-First Control Ordering in Clean-Ancilla V-Chain Synthesis of Nested Multi-Controlled Cascades: A Quantum Binary-Addition-Tree Case Study. https://arxiv.org/abs/2609.29480

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations

In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73-2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime $\propto a^{0.39}$, $r = 0.72$). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free all-to-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads.

cs.ET↗

P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction

AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never observed as a pair, and cells under one condition remain heterogeneous and noisy. Regression on individual cells absorbs sampling variation into the estimated effect, whereas interpretation requires the reproducible population effect. P2P (Perturbation-to-Perturbation) takes a stochastic cell-set view as its supervision unit. Two views drawn from the same condition share a reproducible population effect and differ by view-specific variation. A permutation-invariant set encoder summarizes the control population, a structured encoder represents perturbation tokens, cellular context, dose, and combination interactions, and a gate blends empirical condition-effect memory with a neural residual. A heteroscedastic head predicts the population mean and gene-wise response variance. Under one protocol and five seeds, P2P attains the lowest expression RMSE and the highest Effect Pearson, DEG F1, and DEG average precision on each of Adamson, Norman, Replogle K562, and Replogle RPE1 relative to GenePert, LinearPert, SLIM, Scouter, and scPILOT. On Replogle K562, Effect Pearson rises from 0.643 to 0.702 and DEG F1 rises from 0.067 to 0.178 relative to Scouter, the strongest baseline on both metrics.

cs.ET↗

JFS-CryoMem: A Cryogenic Memory with Voltage-Controlled Superconducting Devices and Femtojoule-Scale Write/Read Energies

Scalable cryogenic systems require memory that combines nonvolatile storage, selective access, low thermal disturbance, and compatibility with superconducting electronics. We present a cryogenic memory architecture that integrates a voltage-controlled Josephson junction field-effect transistor (JJFET) selector with a ferroelectric superconducting quantum interference device (FeSQUID) storage element, hereafter termed JFS-CryoMem. The JJFET provides gate-controlled cell selection, whereas the FeSQUID stores information in stable remanent-polarization states. JFS-CryoMem features separate read and write path mechanisms that support nondestructive readout and independent optimization of programming and sensing conditions. The architecture is evaluated using experimentally calibrated compact models that reproduce the measured electrical characteristics of both constituent devices. We demonstrate selective programming using a half-bias scheme, nonvolatile state retention, and distinguishable readout in a $4 \times 4$ array while accounting for the selected cell and all unselected parallel branches. We then extend the analysis to arrays up to $16 \times 16$ and examine how array scaling alters current distribution, column-equivalent resistance, readout separation, required bitline current, and read energy. The results reveal the principal sensing and energy tradeoffs associated with larger arrays and identify the operating conditions required to preserve read distinguishability as the array grows. JFS-CryoMem provides a device-to-array framework for cryogenic memory in quantum, high-performance, and space-oriented computing systems.

cs.ET↗