Search arXiv⌕ Search

arXiv subjects

Zhenyu Zhang

Publications and source records attributed to Zhenyu Zhang.

At least 19 recordsLinked to original sources

Investigation of Hopping Conduction and Its Impact on the Subthreshold and Transition Region Transport in Oxide Semiconductor Transistors by Magneto-Transport Measurements

In this work, magneto-transport measurements, including Hall effects and magnetoresistance (MR) at various temperatures, are employed to directly probe the electron transport properties in crystalline indium oxide (In2O3) and amorphous indium-zinc oxide (IZO) transistors. For the first time, we develop a subgap density of states (DOS) extraction method based on the MR measurements at low temperature, considering both the interference and orbital shrinkage effects. The sign and magnitude of MR are used as direct evidence to distinguish the dominating transport mechanisms between hopping conduction and free-electron conduction. It is found that crystalline In2O3 exhibits a subgap DOS more than two orders of magnitude lower than that of amorphous IZO, together with a smaller localization radius of hopping sites. As a result, crystalline In2O3 not only enhances the carrier mobility but also significantly reduces the supply voltage (VDD) because of the smaller transition region enabled by the much-suppressed subgap DOS. This study reveals that structural disorder critically influences the device operation in the subthreshold and transition regions of oxide semiconductor transistors.

cond-mat.mtrl-sci↗

Emergence of Chiral Dynamical Multiferroicity in a Ferroelectric Lattice Nonadiabatically Driven by Ultrafast Achiral Electric Fields

It has been shown recently that chiral dynamical multiferroicity can be generated on a ferroelectric lattice whose electric dipoles respond masslessly under chiral optical pumping. Here we demonstrate that, when driven by ultrafast electric pulses, chiral dynamical multiferroicity can also emerge even in situations where the external field is achiral. We reveal this striking phenomenon using a prototypical system of a BaTiO$_3$ moiré ferroelectric lattice, emphasizing the key factor that its chiral electric dipoles inevitably behave massively upon ultrafast driving. At a deeper level, the massive nature is attributed to the anisotropic and nonadiabatic dipolar responses, as captured by the phase difference between the faster longitudinal and slower transverse components of the ferroelectric polarization. Crucially, such a phase difference naturally also gives rise to dynamical magnetization, which exhibits chiral magnetic textures with monopole-like topology coexisting with the ferroelectric chirality. These findings establish the dipolar mass as an enabling and tunable degree of freedom in inducing chiral dynamical multiferroicity, offering new avenues for ultrafast, non-contact magnetic control using pure electric probes.

cond-mat.mtrl-sci↗

Spectral Overfitting in Noisy Linear Probing of Pretrained Representations

Frozen pretrained features are often treated as a safe interface for downstream learning: only a small linear readout is trained, while the backbone is fixed. We show that this readout can still overfit noisy labels in a structured way. A label-blind PCA rank sweep reveals a sharp spectral pattern: under label noise, exposing all pretrained directions can hurt clean accuracy, and intermediate ranks often recover much of the lost performance. Rank-matched random projections help less, and measured between-class signal is strongly concentrated in leading PCs. The pattern appears across three ImageNet-pretrained backbones on CIFAR-10, with gains up to $36.0\pm0.8$ points over the default full-rank probe at 40\% noise. Tuned full-rank probes outperform validation-selected PCA probes, so we present the sweep as a diagnostic of spectral overfitting rather than a competitive noisy-label method.

cs.LG↗

Learning Dynamic Evidence Routes for Vision Transformer Probing

Probing frozen vision transformers typically uses permutation-invariant aggregation (GAP or $\texttt{[CLS]}$), treating patch tokens as an unstructured set. Content-dependent probes such as self-attention are useful accuracy controls, but they do not expose a fixed token schedule or fixed position weights for auditing. We introduce $\textbf{SSMProbe}$, an explicitly inspectable probe that replaces invariant pooling with a Sinkhorn-learned evidence route followed by a diagonal S4 decoder. The S4 decoder is a linear time-invariant (LTI) system whose final state has fixed, position-dependent coefficients, so the probe-induced routed sequence can be audited as a concrete object rather than inferred only from accuracy. Our central measurement is the geometry of routed evidence: which patch tokens are moved to influential positions by this diagnostic, whether those tokens form spatially organized regions or random-like dispersed sets, and how the fixed S4 kernel weights them. Across MAE, BEiT, DINOv2, and supervised ViT, this route geometry separates MAE's dispersed, nearly random-like routes from the more spatially organized routes of BEiT, ViT, and DINOv2, with DINOv2 retaining a distinct strong $\texttt{[CLS]}$ profile. SSMProbe uses the mathematical transparency of state-space models to turn a frozen ViT readout into an auditable evidence-routing analysis.

cs.CV↗

Dense Dump Codec: Error-Controlled, Random-Access Compression of GRMHD Time Series for Slow-Light Radiative Transfer

Black-hole movies provide a unique venue to study the time evolution of accreting plasma. Modeling this evolution with slow-light radiative transfer requires closely spaced simulation outputs, because the plasma changes as light propagates through it. Saving these outputs in full is costly, whereas sparse sampling introduces time-interpolation errors. We present the Dense Dump Codec (DDC), a compression method for dense GRMHD time series. DDC stores exact anchor states and compactly represents the intermediate evolution, while allowing direct access to selected times and variables. Applied to eight production sequences spanning SANE and MAD simulations at several black-hole spins with 0.1M output, DDC archives are about 10 times smaller than the original dense data and even about 50% smaller than the original 0.5M output in aggregate. We further integrate a native DDC reader into polarized slow-light radiative transfer. Tests at 86, 230, and 345 GHz show that DDC may reduce aggregate Stokes-image errors by 54-79% relative to slow-light calculations using original 0.5M input. These results indicate that DDC enables high-cadence slow-light calculations without the prohibitive cost of saving every dense GRMHD state in its original form.

astro-ph.HE↗

Coexistence and Interconversion of Multiple-Order Majorana Modes in Topological Superconductor Films with Varying Thickness

We theoretically investigate the thickness-dependent evolution of Majorana modes in $C_{2h}$-symmetric topological superconductor films (such as the recently discovered 2M-WS$_2$) proximity coupled with magnetic insulators. For sufficiently thick films, two Majorana bound states coexist as end modes at the surface and interface along a vortex line, with the interfacial mode evolving into a chiral Majorana edge mode upon increasing the proximity-induced exchange field. The intrinsic $C_{2h}$ crystalline symmetry selects two chiral Majorana modes circulating along the hinges on two of the four side surfaces of the film. When the penetration depth of the exchange field is sufficiently shallow, the two circulating modes are localized near the interface, but with qualitatively different subsequent evolutions. One of them further collapses to form two Majorana corner modes, while the other merges with the chiral Majorana mode circulating around the interface. Importantly, the corner modes are well decoupled from the interfacial chiral mode, thereby enabling an unprecedented coexistence of first-, second-, and third-order Majorana modes within a single material platform. We further show that such coexistence persists even in the ultrathin-film limit, where the electric-field-controlled two-dimensional $Z_2$ topology offers an extra advantage to readily interconvert the multiple-order Majorana modes. These findings highlight the pivotal role of the proper crystalline symmetry in enabling emergence, manipulation, and potential braiding of Majorana modes for demonstrating non-Abelian statistics and fault-tolerant quantum computation.

cond-mat.supr-con↗

MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple complementary strategies. We propose MOSCOPT, a text-native, parameter-free algorithm that jointly optimizes a pool of $N$ skills and a gating skill $G$ that dynamically selects $K$ skills per step. To effectively optimize the skills, we build the EditAdam with internally maintained dual states. Through the three-phase interleaved updates with EditAdam, the system monotonically improves without gradient or parameter tuning. Extensive experiments and detailed ablations across 5 benchmarks and 3 target LLMs demonstrate that MOSCOPT consistently outperforms all baselines, and confirm that both the mixture-of-skills architecture with selective activation and the collective evolution with three-phase interleaving are essential to its superior performance. Code is released https://github.com/zhangzhenyu13/SummerClaw/tree/master/summerclaw/agent_trainer/algorithms/moscopt.

cs.AI↗

Programming In-Storage Computing with Located, Stateful Dataflow

In-storage computing (ISC) reduces host--storage data movement by executing computation inside computational storage devices (CSDs). For multi-stage applications, realizing these benefits requires coordinating data placement, I/O--compute overlap, and device-resident state across the workflow, yet existing interfaces lack a unified abstraction for these decisions. We present Epic, an NVMe-based ISC stack that provides this abstraction by capturing data residency and lifetime in the program: location types declare logical residency, dataflow derives lifetimes for intermediate values and operation state within an invocation, and a keep primitive extends selected state across invocations. These semantics expose the complete offloaded workflow as a located, stateful dataflow. A storage-aware compiler transforms this workflow, performs movement-aware logical mapping and fusion, and exposes I/O--compute overlap; a runtime completes the plan using execution-time information, asynchronously binding work to physical resources and managing device-resident state. Across 12 file-scanning, database, and machine learning workloads, Epic is 1.6$\times$ faster on average than the strongest of five prior ISC systems, while achieving 4.2$\times$ speedup on average and up to 16.1$\times$ over the corresponding host baselines, and reducing application-side code by up to 14$\times$ in our implementations.

cs.AR↗

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.

cs.AI↗

HintMiner: Automatic Question Hints Mining From Q&A Web Posts with Language Model via Self-Supervised Learning

Users often need ask questions and seek answers online. The Question - Answering (QA) forums such as Stack Overflow cannot always respond to the questions timely and properly. In this paper, we propose HintMiner, a novel automatic question hints mining tool for users to help them find answers. HintMiner leverages the machine comprehension and sequence generation techniques to automatically generate hints for users' questions. It firstly retrieve many web Q\&A posts and then extract some hints from the posts using MiningNet that is built via a language model. Using the huge amount of online Q\&A posts, we design a self-supervised objective to train the MiningNet that is a neural encoder-decoder model based on the transformer and copying mechanisms. We have evaluated HintMiner on 60,000 Stack Overflow questions. The experiment results show that the proposed approach is effective. For example, HintMiner achieves an average BLEU score of 36.17\% and an average ROUGE-2 score of 36.29\%. Our tool and experimental data are publicly available.

cs.LG↗

POSPAN: Position-Constrained Span Masking for Language Model Pre-training

Span-level masked language modeling (MLM) has shown to be advantageous to pre-trained language models over the original single-token MLM, as entities/phrases and their dependencies are critical to language understanding. Previous works only consider span length with some discrete distributions, while the dependencies among spans are ignored, i.e., assuming that the positions of masked spans are uniformly distributed. In this paper, we present POSPAN, a general framework to allow diverse position-constrained span masking strategies via the combination of span length distribution and position constraint distribution, which unifies all existing span-level masking methods. To verify the effectiveness of POSPAN in pre-training, we evaluate it on the datasets from several NLU benchmarks. Experimental results indicate that the position constraint is capable of enhancing span-level masking broadly, and our best POSPAN setting consistently outperforms its span-length-only counterparts and vanilla MLM. We also conduct theoretical analysis for the position constraint in masked language models to shed light on the reason why POSPAN works well, demonstrating the rationality and necessity of POSPAN.

cs.LG↗

Thermodynamics of Kerr-Newman-Bertotti-Robinson black holes

In this work, we extend the thermodynamic analyses of the neutral and specially charged Kerr-Bertotti-Robinson black holes to the general Kerr-Newman-Bertotti-Robinson family, in which the electric and external-field parameters are independent. Using covariant surface charges and canonical integrability methods, we determine the total angular momentum and the canonical mass. The angular momentum follows analytically from the combined gravitational and electromagnetic surface charges. However, the infinitesimal charge associated with coordinate time translations is not integrable in solution space, so the mass must be associated with a more general symmetry generator. Imposing the canonical integrability conditions on this generator, together with the Kerr-Newman mass as the zero-field boundary condition, selects the Christodoulou-Ruffini mass and determines the associated thermodynamic potentials. The first law and Smarr formula take the same form as the ones in usual Kerr-Newman case, and the two previously studied Kerr-Bertotti-Robinson cases are recovered as special limits.

gr-qc↗

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

cs.CV↗

Shadows and Thin-Disk Images of Kerr-Newman Black Holes in a Bertotti-Robinson Magnetic Field

In this paper, we investigate the optical properties of Kerr-Newman-Bertotti-Robinson (KN-BR) black holes. We use the separability of null geodesics to analyze unstable spherical photon orbits and determine the radial extent of the photon shell. Because the spacetime is not asymptotically flat, we construct the critical curve on the screen of a finite-distance zero-angular-momentum observer. We then perform backward ray tracing for a geometrically thin and optically thin disk that extends from the outer region to the event horizon, and examine the resulting images, intensity profiles, critical-curve areas, and inner-shadow areas. We find that the genuine neutral Kerr-BR$_0$ and specially charged Kerr-BR$_s$ configurations have nearly identical optical appearances. It is remarkable that for the KN-BR black holes increasing the electric charge reduces the characteristic image size in the Kerr-Newman limit but enlarges it in the magnetized configurations considered here. We also find that the external magnetic field strongly increases the apparent image scale, while the observer inclination affects the inner-shadow area more significantly than the critical-curve area. These results may provide useful theoretical insight for future observations aimed at identifying such exotic magnetized black holes.

gr-qc↗

ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance

Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a continuous maintenance workflow into a series of independent sessions, ignoring the cumulative dependencies that make real-world bug fixing challenging. To bridge this gap, we introduce ChainSWE, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase. We collect chronological chains of 304 issues across 54 Python projects, mined from six SWE-bench-family datasets. Our evaluation across a range of agents and models reveals a consistent performance drop by up to 70% as the chain length increases.

cs.SE↗

Machine Learning Unveils Finite-volume Energy Shifts in Three-body System

Finite-volume extrapolation (FVE) is essential for extracting physical observables in the lattice calculation. While rigorous FVE formulations are well established for short-range potentials in both two- and three-body systems, long-range interactions with force ranges comparable to the lattice size $L$ remain challenging. Extending a previous data-driven scheme for two-body systems, we apply symbolic regression (PySR) to uncover universal three-body FVE formulae. For short-range potentials, we reproduce the two limiting cases, i.e. $κ_3\ggκ_2$ and $κ_3\simκ_2$. For pure long-range potentials, we obtain a dedicated analytic expression, and after incorporating short-range contributions, we uncover a unified formula consistent with the original PySR solution, which performs excellently in the intermediate force range around 1 fm. This work demonstrates that combining machine learning with physical constraints can yield novel analytical results inaccessible to conventional theoretical tools, advancing data-driven methodologies in hadron physics.

hep-lat↗

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may disrupt logical coherence and degrade performance. We formalize this trade-off as the Context-Generation Substitution Law, where explicit reasoning context substitutes for part of decode-time generation. Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds. Rather than using raw demonstrations, these memories summarize reusable reasoning patterns, key constraints, and critical operations to compensate for information lost during compression. Experiments show that Memory consistently improves prompt-based Chain-of-Draft (CoD) compression across mathematical reasoning, complex reasoning, and science question answering tasks, yielding accuracy gains of 21.4, 28.0, 29.5, and 6.61 points over CoD on GSM8K, MATH, BBH, and MMLU-Sci, while achieving a 1.14-1.49x latency speedup latency speedup over standard CoT. Memory is also compatible with token-level, reasoning-trace-level, and inference-state compression mechanisms.

cs.CL↗

Advantageous Parameter Expansion Training Makes Better Large Language Models

Although scaling up the number of trainable parameters can effectively improve the training performance of large language models, it also leads to increased computational overhead. When delving into the parameter difference, we find that a subset of parameters, termed advantageous parameters, plays a crucial role in determining model performance. Further analysis reveals that stronger models tend to possess more such parameters. In this paper, we propose Advantageous Parameter EXpansion Training (APEX), a method that progressively expands advantageous parameters into the space of disadvantageous ones, thereby increasing their proportion and enhancing training effectiveness, while keeping the total parameter count unchanged. Extensive experiments on both instruction tuning and continued pre-training across five base models demonstrate that, in instruction tuning, APEX outperforms full-parameter tuning while using only 52% of the trainable parameters. In continued pre-training, APEX achieves the same perplexity level as conventional training with only approximately 30% of the training data, and yields significant improvements on downstream tasks.

cs.CL↗