Search arXiv⌕ Search

arXiv subjects

Yifan Zhou

Publications and source records attributed to Yifan Zhou.

At least 19 recordsLinked to original sources

Security Limits of Mining Before Validation in Nakamoto Consensus

Mining before validation allows miners to extend a newly received block before completing its validity checks, giving them a head start in the race for the next block reward. This head start, however, comes with a security risk: rejecting one invalid block also discards the honest work built on it, an effect missed when validation is treated as instantaneous. We quantify this risk in a model of Nakamoto consensus with bounded network delay and a validation-time bound independent of processing load. We establish an explicit threshold on adversarial mining power below which honest miners' fully validated chains continue to grow and agree on a stable history with high probability over any fixed observation period. We also construct an attack that repeatedly draws honest mining onto invalid branches. When the adversary produces more than one block on average during the allowed validation time, the attack can eventually remove a target block at any fixed initial confirmation depth. Combined with ordinary private mining, this attack yields a matching asymptotic bound in the fully decentralized regime, where each honest miner has negligible mining power. The results show how validation latency limits the security of mining before validation, beyond the constraint imposed by network delay.

cs.CR↗

Evaluation of OpenAI o1: Opportunities and Challenges of AGI

This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, including computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. Through rigorous testing, o1-preview demonstrated remarkable capabilities, often achieving human-level or superior performance in areas ranging from coding challenges to scientific reasoning and from language processing to creative problem-solving. Key findings include: -83.3% success rate in solving complex competitive programming problems, surpassing many human experts. -Superior ability in generating coherent and accurate radiology reports, outperforming other evaluated models. -100% accuracy in high school-level mathematical reasoning tasks, providing detailed step-by-step solutions. -Advanced natural language inference capabilities across general and specialized domains like medicine. -Impressive performance in chip design tasks, outperforming specialized models in areas such as EDA script generation and bug analysis. -Remarkable proficiency in anthropology and geology, demonstrating deep understanding and reasoning in these specialized fields. -Strong capabilities in quantitative investing. O1 has comprehensive financial knowledge and statistical modeling skills. -Effective performance in social media analysis, including sentiment analysis and emotion recognition. The model excelled particularly in tasks requiring intricate reasoning and knowledge integration across various fields. While some limitations were observed, including occasional errors on simpler problems and challenges with certain highly specialized concepts, the overall results indicate significant progress towards artificial general intelligence.

cs.CL↗

A Very Big Video Reasoning Suite

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/?v=vbvr .

cs.CV↗

Multi-Epoch Stability in the Rotational Modulation of the Planetary-Mass Companion Ross 458C

Spectroscopic variability is detected in brown dwarfs across all spectral types, suggesting atmospheric patchiness in brown dwarf atmospheres. We present JWST/NIRSpec time series observations of the T8.0 planetary-mass companion Ross 458C. We compared our derived spectral variability and light curve to those obtained using HST/WFC3 observations more than 7.5 yr before and published by Manjavacas et al. (2019). We found that the light curve of Ross 458C is remarkably similar in the two epochs in terms of shape and variability amplitude potentially created by stable weather patterns. We measured a rotational period of 11.57+\-0.09 hr, demonstrating that the previously reported HST/WFC3 period of 6.75+\-1.58 hr is most likely a second-order harmonic produced by an unresolved double-peaked rotational modulation, likely due to the higher uncertainties of the HST/WFC3 light curve and the short time baseline of those observations. We measured the variability amplitude in the white light curve created using only the wavelength range covered by HST/WFC3 (1.10-1.63 micron), obtaining a value of 1.47+\-0.18%, which is consistent with that measured in the HST/WFC3 light curve.

astro-ph.SR↗

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.

stat.ML↗

HBFlex: A Flexible Memory System for Bridging Fine-Grained LLM States and Coarse-Grained HBF Parallel Execution

Large language models (LLMs) require increasing memory capacity to accommodate growing model weights and KV caches. High-Bandwidth Flash (HBF) offers high memory density and aggregate read bandwidth through massive plane-level parallelism, making it an attractive option for LLM serving. However, serving LLMs entirely from HBF introduces three challenges: fine-grained KV reads create placement and access imbalance, incremental writes interfere with foreground reads, and mixed KV lifetimes amplify garbage collection. Hybrid HBM/HBF designs retain HBM to support dynamic KV management, but this allocation reduces the HBF resources available under a fixed packaging budget, limiting aggregate HBF bandwidth. We present HBFlex, a full-HBF memory system with coordinated optimizations for KV reads, writes, and reclamation. HBFlex balances KV placement and attention accesses to improve plane utilization. It aggregates incremental updates and schedules writeback within sufficiently long compute windows to reduce write--read interference. It also combines lifetime-guided block packing with deferred reclamation to reduce valid-page migration. We evaluate HBFlex through trace-driven simulation across different configurations. HBFlex achieves average throughput speedups of up to 1.58$\times$ over FlashAccel and 3.30$\times$ over H3, benefiting from higher HBF bandwidth and more efficient management of dynamic KV-cache reads, writes, and erases.

cs.AR↗

Deep H$α$ Imaging Survey of IC 348 with the Hubble Space Telescope: I. Accretion Properties of Stellar and Substellar Objects

Accretion governs the growth of young stars and the early evolution of their circumstellar disks, yet population-level measurements of accretion are often hampered by heterogeneous diagnostics and by samples preferentially selected toward disk-bearing or accreting objects. This can impact the mass accretion rate-stellar mass ($\dot{M}$-$M_\star$) relation, particularly at substellar masses. We present a uniform analysis of accretion in the $\sim$2 Myr-old star-forming cluster IC 348 based on deep Hubble Space Telescope F656N imaging. Using H$α$ excess as a single, homogeneous accretion tracer, we derive accretion rates and robust upper limits for 200 cluster members spanning the stellar to substellar mass regime ($3\ M_\odot$ to $4\ M_{\rm Jup}$). Accretion is detected in $37\pm3\%$ of the sample, with fractions of $34\pm4\%$ among stellar members and $46\pm6\%$ among substellar objects. For accretors alone, the inferred $\dot{M}$-$M_\star$ relation is consistent with those measured in similarly aged regions such as Lupus. In contrast, including weak accretors and non-detections increases the scatter and flattens the slope while lowering the intercept, demonstrating the strong influence of the low-accretion tail on population-level accretion relations. For free-floating planetary-mass objects in IC 348, extrapolating the accretors-only fit overpredicts $\dot{M}$ by approximately an order of magnitude compared to the fit that includes upper limits, which instead implies mass accretion rates of $<10^{-12}$ $M_{\odot}\ \mathrm{yr}^{-1}$. These results show that sample selection, specifically whether the sample is restricted to disk-bearing or accreting targets, or instead is drawn from a complete membership census, is a dominant factor shaping population-level accretion relations across a wide dynamic range in stellar mass.

astro-ph.SR↗

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.

cs.CV↗

Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, onboard-resource, and cloud-affected availability constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy regulates exploration and exploitation according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO obtains the highest mean observation benefit in every scenario, improves the result of conventional ant colony optimization by 3.40\%--9.40\%, accelerates convergence, and remains stable under different objective-weight settings. These results demonstrate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.

cs.AI↗

A Scalable Path to Astrometric Exomoon Discoveries with the Nautilus Space Observatory

Moons orbiting exoplanets (exomoons) can be detected through the reflex motion they impart to their host planet, which is recoverable in relative star-planet astrometric time series. The signal grows with moon mass and orbital separation and decreases with distance, so the nearest and least massive imaged planets are the most favorable targets. Recovering small (<Earth-mass) moons requires continuous, long-baseline, high-precision monitoring that is only practical with a dedicated or nearly dedicated facility. Building on recent simulations of astrometric exomoon detection and of the resulting population yields, we argue that the scalable, replicable architecture of the Nautilus Space Observatory is uniquely suited to this problem, and we outline a staged campaign. In an initial phase, one or a few small apertures target the nearest imaged giant planets--a high-reward but low-probability search focused on the closest stars. As the array is built out, the astrometric noise floor decreases and the same technique extends the search to the nearest such systems among nearby stars of spectral type K and earlier. This would be performed in parallel with high-contrast imaging and spectral characterization of the host planets and in synergy with a companion starshade concept for imaging Earth-like planets around the same nearby stars. Nautilus thus provides a scalable path from the first detection of a nearby exomoon toward a systematic search for exomoons around the closest stars.

astro-ph.IM↗

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation

Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into an existing PLC project and then runs correctly has been checked only in limited tests. We present \textsc{SemaPLC}, a project-grounded and verification-gated agent harness assembled from conventional tools but governed by a strict completion rule. Rather than stopping when the model judges its own output adequate, \textsc{SemaPLC} declares a task complete only when logged external checks confirm it. Those checks cover the specification, the compilation, and the behavior on a live runtime. On 117 independent-POU tasks matching existing benchmarks, it attains the highest strict verified pass rate on all seven models (72.6\% mean). On a project-context track of 65 tasks whose generated logic must compile and run inside a real project, it attains the highest mean on integrated compilation, static behavior, and dynamic behavior. Of the three layers, dynamic behavior is the most revealing. We measure it by deploying the generated and the reference logic to a live PLC runtime and comparing their executed traces. All methods fall within 10 static points of one another, whereas dynamic scores separate them sharply, from 22.4 to 31.4 for the baselines against 52.2 for \textsc{SemaPLC}. Overall, our verification-gated harness raises the mean at every layer and most sharply at runtime. Execution, not static scoring, is the faithful test of whether generated control logic actually works. \textsc{SemaPLC} is open-sourced at https://github.com/midea-ai/SemaPLC.

cs.SE↗

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.

cs.AI↗

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

cs.CV↗

Revolutionizing Finance with LLMs: An Overview of Applications and Insights

In recent years, Large Language Models (LLMs) like ChatGPT have seen considerable advancements and have been applied in diverse fields. Built on the Transformer architecture, these models are trained on extensive datasets, enabling them to understand and generate human language effectively. In the financial domain, the deployment of LLMs is gaining momentum. These models are being utilized for automating financial report generation, forecasting market trends, analyzing investor sentiment, and offering personalized financial advice. Leveraging their natural language processing capabilities, LLMs can distill key insights from vast financial data, aiding institutions in making informed investment choices and enhancing both operational efficiency and customer satisfaction. In this study, we provide a comprehensive overview of the emerging integration of LLMs into various financial tasks. Additionally, we conducted holistic tests on multiple financial tasks through the combination of natural language instructions. Our findings show that GPT-4 effectively follow prompt instructions across various financial tasks. This survey and evaluation of LLMs in the financial domain aim to deepen the understanding of LLMs' current role in finance for both financial practitioners and LLM researchers, identify new research and application prospects, and highlight how these technologies can be leveraged to solve practical challenges in the finance industry.

cs.CL↗

Same Targets, Different Computation: How Post-Training Divides Work Across Model Layers

A late-layer change learned during post-training may work on the base model's earlier state, or it may depend on earlier computation learned with it. We distinguish these cases with a four-cell diagnostic that crosses base or descendant upstream states with base or descendant late stacks. A large late-stack effect need not imply strong upstream dependence. On math prompts, OpenMath2's late stack changes the target margin by +3.43 logits after base upstream state and +3.28 after its own, giving a near-zero interaction. Several instruction-following descendants of the same Llama-3.1-8B base show greater dependence, while controlled code and biomedical continuation-training runs sit near zero on the common support; seven released descendants span -0.54 to +2.20 logits. Because those checkpoints differ in many ways, we then isolate one training property. Two LoRA fine-tunes learn identical target responses requested either by familiar natural-language instructions or by newly learned nonce codes. Changing only this cue-to-response relation increases upstream dependence by +5.56 logits on Qwen3-4B and +4.18 on Llama-3.1-8B, with a positive paired shift in all six model-by-seed runs. The interaction is also positive in all five released base/instruction pairs we test, and late-stack replacement changes the full-vocabulary argmax in about half of events. Post-training can therefore organize the same target behavior with different dependencies between earlier and later computation. The diagnostic measures local next-token compatibility, not free-running component transplantation.

cs.LG↗

One-Step Epitaxial Access to Rhombohedral Graphene Flat-Band States on Step-Bunched SiC

Rhombohedral graphene multilayers provide a moiré-free platform for correlated and topological flat-band physics, but direct, transfer-free epitaxial access to thickness-tunable multilayers remains limited. Here we report a one-step graphitization route on 4$^\circ$ off-axis 4H-SiC, in which high-temperature flash annealing simultaneously drives self-organized step bunching and multilayer graphene formation. Atomic-resolution cross-sectional scanning transmission electron microscopy identify local ABC registry and distinguish rhombohedral from Bernal stacking. The thickness is tuned from bilayer to more than twenty layers by varying single parameter, the annealing temperature. Angle-resolved photoemission spectroscopy directly tracks the thickness-dependent evolution from interface-dominated low-energy states toward pronounced near-Fermi-level flat-band spectral weight in thick multilayers. Low-temperature scanning tunneling microscopy and spectroscopy on a 17-layer film further reveal a 13.4 meV low-energy spectral reconstruction and a $\sqrt{3} \times \sqrt{3}$ Kekulé-like modulation, providing microscopic signatures consistent with an intervalley-mixed electronic texture. This one-step, transfer-free approach establishes step-bunched SiC as an epitaxial platform that links stacking engineering with moiré-free correlated flat-band electronic states.

cond-mat.mtrl-sci↗

The JWST weather report: Unravelling the atmospheric variability of isolated worlds using principal component analysis

Brown dwarf variability directly probes atmospheric dynamics beyond the Solar System, and recent JWST time-resolved spectroscopy has opened a new window into these processes. Principal component analysis (PCA) offers a data-driven framework to identify the dominant, independent patterns of spectral variability of variable targets without relying on prior atmospheric assumptions. SIMP 0136 is a young, T2.5, brown dwarf at the planetary-mass boundary, making it an ideal analogue for directly imaged exoplanets. We analysed one rotation of JWST/NIRSpec PRISM time-series spectroscopy to investigate the drivers of its variability using PCA. Two principal components are sufficient to reduce the residual spectra to the propagated noise floor, indicating that they capture the detectable coherent spectroscopic variability. The leading principal component captures broadband variability consistent with temperature changes, while the second traces chromatic variability linked to vertical cloud structure. The dominance of two components implies that the spectra can be described as mixtures of three distinct atmospheric states, whose relative contributions we mapped as a function of rotational phase. The observed spectra are described as evolving linear combinations of these states, indicating that the variability arises from the changing visibility of spatially distinct atmospheric regions. By projecting Sonora Diamondback forward models into the same principal component space, we found that the principal components capture a large fraction of the model variance, demonstrating that the same physical processes that govern SIMP-0136's observed variability also capture much of the model grid's variation. Our results establish PCA as a computationally efficient, physically interpretable framework for analysing JWST time-resolved spectroscopy of substellar atmospheres.

astro-ph.EP↗

Photometric Variability and Rotation of Beta Pictoris b from JWST NIRCam Coronagraphic Imaging

We report the detection of photometric variability in the directly imaged super-Jupiter $β$ Pictoris b. Using JWST NIRCam dual-band coronagraphic imaging, we conducted a 16-hour continuous photometric monitoring campaign in the F210M and F410M filters. We developed and validated a time-series photometry framework that combines PSF subtraction, principal component analysis for systematic noise removal, and injection-and-recovery tests to confirm signal fidelity. Both light curves show consistent sinusoidal variability at $\sim$5$σ$ and $\gg 5σ$ significance in the F210M and F410M bands, respectively. A joint sinusoidal fit yields a rotation period of $P_{\rm rot} = 9.00 \pm 0.13$ hr and variability amplitudes of $0.85 \pm 0.07\%$ and $0.89 \pm 0.04\%$ in F210M and F410M, respectively. The near-identical amplitudes and periods in both bands confirm a common astrophysical origin in a heterogeneous atmosphere. Combining $P_{\rm rot}$ with the previously measured projected rotational velocity, we constrain the line-of-sight spin axis inclination of $β$ Pic b. The result favors an equator-on viewing geometry, consistent with line-of-sight spin-orbit alignment: the planetary spin axis, orbital plane, debris disk, and stellar equator are all mutually aligned. This stands in sharp contrast to the large obliquities of wide-separation companions that are likely formed via gravitational fragmentation. Together with the system's young age, this observation provides independent dynamical evidence that $β$ Pic b formed via core accretion. This result constitutes the first detection of rotational modulation in a close-in, high-contrast exoplanet that likely formed via core accretion, demonstrating that time-series coronagraphic imaging with JWST opens a powerful new window onto the rotation, atmospheric dynamics, and spin-orbit architecture of this population.

astro-ph.EP↗