Search arXivSearch

arXiv subjects

Tianshu Wang

Publications and source records attributed to Tianshu Wang.

At least 19 recordsLinked to original sources

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.

cs.CL

Effects of Rotation on 3D Core-Collapse Supernova Models for Low-Mass Progenitors

We explore the dependence upon rotation rate alone of various supernova observables simulated to their saturation for the explosion of a 9-$M_{\odot}$ progenitor. We find that the explosion energy is non-monotonic with, and weakly dependent upon, spin across a broad range of initial spins. The asymmetries of the blast depend weakly on spin, with faster spins leading to only slightly greater asymmetries. There is little significant pole-equator neutrino heating asymmetry during explosion, even for rapid rotation, and only for the fastest rotator does the neutrino heating rate diminish noticeably. Hence, the effects of rotation alone on all salient aspects of supernova dynamics are not large. We find that the recoil kick and spin are clearly aligned only for the most rapidly rotating model. Interestingly, for the fastest rotator, we witness a $T/|W|$ corotation instabilities near a value of 0.05 and spiral arm modes emerge. We find that the nucleosynthetic yields depend little upon the rotation rate and determine that the ratio of initial to final core spin period is near $\sim$4000, implying, given the modest inferred radio pulsar periods at birth, that the initial spin periods of most supernova cores are likely quite long. However, we focus on only one progenitor and do not include magnetic fields. Nevertheless, at least for low-mass progenitors which explode early, we find muted consequences of rotation in most major particulars across a wide range of initial spins.

astro-ph.HE

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite. Recent work explores a paradigm where a main agent decomposes tasks and dispatches subtasks to subagents, which execute and return only summarized results, conserving the main agent's context budget. However, performing this well requires delegation intelligence: the ability to decompose complex tasks, determine when and what to delegate, and integrate returned results into the ongoing workflow. Training data for this capability is scarce in naturally occurring text, and to our knowledge, how to synthesize such data and train models to acquire this capability remains largely unexplored in the open-source community. To bridge this gap, we present a preliminary exploration targeting deep research, a representative long-horizon agent task. Specifically, we design a harness that guides the model toward high-quality task decomposition and delegation, while constraining subagents to return results properly to support the main agent's workflow. The harness-guided trajectories naturally encode correct delegation decisions, which we use as supervised fine-tuning data to internalize delegation intelligence into model weights. Our resulting model, SearchSwarm-30B-A3B, achieves 68.1 on BrowseComp and 73.3 on BrowseComp-ZH, the best results among all models of comparable scale. We will release our harness, model weights, and training data to facilitate future research.

cs.AI

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether models can autonomously develop agent systems. We introduce the Meta-Agent Challenge (MAC), an evaluation framework designed to test the capacity of frontier models for autonomous agent development. Specifically, a code agent (the meta-agent) is given a sandboxed environment, an evaluation API, and a time limitation to iteratively program an agent artifact that maximizes performance on a held-out test set across five domains. To ensure evaluation integrity, this framework is secured by multi-layer defenses against reward hacking. Leveraging this framework, we demonstrate that meta-agents rarely match human-engineered baseline policies, and the few that do are dominated by proprietary frontier models. Moreover, the design process exhibits high variance, and high optimization pressure surfaces emergent adversarial behaviors like ground-truth exfiltration-highlighting critical deficits in both robustness and model alignment. Ultimately, MAC provides a rigorous, open-source benchmark for autonomous AI research and development, offering an empirical proxy for evaluating recursive self-improvement. Benchmark is publicly available at: https://github.com/ant-research/meta-agent-challenge.

cs.AI

An Exploration of the Equation of State Dependence of Core-Collapse Supernova Explosion Outcomes and Signatures

We explore, using a state-of-the-art simulation code in 3D and to late enough times to witness final observables, the dependence of core-collapse supernova explosions on the nuclear equation of state. Going beyond questions of explodability, we compare final explosion energies, nucleosynthetic yields, recoil kicks, and gravitational-wave and neutrino signatures using the SFHo and DD2 nuclear equations of state (EOS) for a 9-$M_{\odot}$/solar-metallicity progenitor star. The DD2 EOS is stiffer and has a lower effective nucleon mass. The result is a more extended protoneutron star (PNS) and lower central densities. As a consequence, the mean neutrino energies, final explosion energy, and recoil kick speed are lower. Moreover, the evolution of PNS convection differs between the two EOS models in significant ways. This translates in part into interestingly altered neutrino ``light" curves and noticeably altered gravitational-wave signal strengths and frequency characteristics that may be diagnostic. The faster exploding model (SFHo) yields slightly more neutron-rich ejecta and more species with atomic weights between 60 and 90 and a weak r-process. However, this is merely a preliminary study. The next step is a more comprehensive and multi-progenitor set of 3D supernova simulations for various EOSes to late times when the observables have asymptoted. Such a future investigation will have a direct bearing on the neutron star and black hole birth mass functions and the quest towards a fully quantitative theory of supernova observables.

astro-ph.HE

The Effect of the Fast-Flavor Instability on Core-Collapse Supernova Models: II. Quasi-Equipartition and the Impact of Various Angular Reconstruction Methods

In this work, we explore in a consistent fashion the effects of fast flavor conversion (FFC) in 1D and 2D core-collapse supernova (CCSN) simulations. In addition, we investigate the impact of various angular reconstruction methods and compare the ``3-species'' and ``4-species'' neutrino transport schemes. We find that the FFC effects are insensitive to the different methods tested and that the FFC alters supernova hydrodynamics is only minor ways. We also present a ``quasi-equipartition'' approximation which can be used to estimate the FFC-altered neutrino properties by post-processing the neutrino signals extracted from no-oscillation CCSN simulations. The relative errors in neutrino number and energy luminosities of this phenomenological method are less than 2\% for 1D models, and less than 10\% for 2D models. This method provides a simple way to include the effects of FFC on neutrino signals without implementing a complex and expensive FFC scheme or redoing simulations.

astro-ph.HE

Simulated 3D $^{56}$Ni Distributions of Type IIp Supernovae

We present the first three-dimensional study of the asymptotic ejecta distributions for a suite of theoretical Type IIp supernovae originating from red supergiant progenitors. We simulate using the radiation-hydrodynamic code F{\sc{ornax}} from core bounce through the first seconds of the neutrino-driven explosion and then follow using a hydrodynamic variant of the code FLASH until shock breakout of the star and through to homologous expansion of the ejecta into the circumstellar environment. Our studied progenitor models range from 9 to 25 M$_{\odot}$, with explosion energies spanning $\sim$0.1$-$1 Bethe. The shock breakout times span the range $\sim$1$-$4 days, with a breakout time spread by direction ranging from hours to over a day. We find that the dipole orientation of the $^{56}$Ni ejecta is well-preserved from the first seconds out to shock breakout. The $^{56}$Ni ejecta penetrates through the initially outer oxygen shell, and its global structure is imprinted with small-scale clumping as the ejecta evolve through the stellar envelope. For the majority of our models, the neutron star kick is anti-aligned with the $^{56}$Ni ejecta. Models with strongly dipolar ejecta morphology and a massive hydrogen/helium envelope with an inner boundary located deep see as much as $\sim$70\% of the $^{56}$Ni ejecta mixed into that outer envelope, reaching asymptotic velocities ranging from $\sim$350 to 3200 km s$^{-1}$. Supernovae arising from red supergiant progenitors and exhibiting prominent nickel features generally display significant $^{56}$Ni mixing into the stellar envelope.

astro-ph.HE

The Effect of the Collisional Flavor Instability on Core-Collapse Supernova Models

We explore the effects of the neutrino collisional flavor instability (CFI) based on 1D and 2D core-collapse supernova (CCSN) simulations done using the sophisticated radiation-hydrodynamic code Fornax. We compare the growth rates of homogeneous CFI (hCFI) modes calculated by numerically solving the multi-group dispersion relation to those calculated using the monochromatic approximation. We find that the widely-used monochromatic approximation leads to incorrect growth rates} when applied in multi-group scenarios. As opposed to the $\sim10^5$ s$^{-1}$ values given by the monochromatic approximation, the actual growth rates of non-resonance multi-group hCFI are at most $\sim$200 s$^{-1}$ in all our models and they are too slow to affect CCSN outcomes. We adopt a BGK flavor conversion scheme in the simulations to include the effects of resonance-like hCFI. We find that the CCSN dynamics and neutrino emission properties are only weakly influenced, and the intrinsic stochasticity due to convection and neutrino-driven turbulence can naturally lead to comparable effects. Hence, our analysis of the non-resonance and resonance-like hCFI into CCSN simulations suggests that the effects of neutrino flavor conversion triggered by hCFI modes are in general small.

astro-ph.HE

ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search

Large language models (LLMs) have demonstrated impressive capabilities and are receiving increasing attention to enhance their reasoning through scaling test--time compute. However, their application in open--ended, knowledge--intensive, complex reasoning scenarios is still limited. Reasoning--oriented methods struggle to generalize to open--ended scenarios due to implicit assumptions of complete world knowledge. Meanwhile, knowledge--augmented reasoning (KAR) methods fail to address two core challenges: 1) error propagation, where errors in early steps cascade through the chain, and 2) verification bottleneck, where the explore--exploit tradeoff arises in multi--branch decision processes. To overcome these limitations, we introduce ARise, a novel framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval--augmented generation (RAG) within a Monte Carlo tree search paradigm. This approach enables effective construction and optimization of reasoning plans across multiple maintained hypothesis branches. Experimental results show that ARise significantly outperforms the state--of--the--art KAR methods by up to 23.10%, and the latest RAG-equipped large reasoning models by up to 25.37%. Our project page is at https://opencausalab.github.io/ARise.

cs.AI

The Effect of the Fast-Flavor Instability on Core-Collapse Supernova Models

Merging our supernova code F{\sc{ornax}} with the Box3D fast-flavor neutrino oscillation formalism, we explore the effects of fast-flavor conversion (FFC) in state-of-the-art 1D and 2D core-collapse supernova simulations. We find that after a few tens of milliseconds after bounce the FFC emerges just interior to and exterior to the stalled shock wave. It does not obtain in the PNS core nor near the average neutrinosphere radii. Interior to the shock, this results in a temporary change in the net neutrino heating rate of $\sim$10\%, due mostly to a hardening of the $\nu_e$ and $\bar{\nu}_e$ neutrino spectra, despite the decrease in their corresponding neutrino number fluxes. In 1D, the hydrodynamic effects are not large, with increases in the stalled shock radius by of order ten to twenty kilometers that abate within a few hundred milliseconds. In 2D, the hydrodynamic effect of the FFC is a bit more noticeable, resulting in slightly earlier explosions for models for lower-mass progenitors, but also potentially inhibiting explosions for some higher-mass progenitors. Fast-flavor conversion continues to operate at larger radii at later times. The net result is a shift upward in the $\nu_{\mu}$ energy and number luminosities and a shift downward in the same quantities for both the $\nu_e$ and $\bar{\nu}_e$ neutrinos. There seems to be a trend at very large radii and later times towards partial species and spectral equipartition. If this is true, it could be an interesting feature of supernova neutrino detection at later times in underground and under-ice facilities.

astro-ph.HE

Channels of Stellar-mass Black Hole Formation

On the basis of a large collection of detailed 3D core-collapse supernova simulations carried to late times, we identify four channels of stellar mass black hole formation. Our examples for Channel 1 involve the formation of lower-gap and above black holes in energetic asymmetric supernova explosions. Our Channel 2 example involves a modest supernova explosion that may leave behind a lower-gap to $\sim$10 $M_{\odot}$ black hole. The latter may not be easily distinguishable from ``standard" supernovae that birth neutron stars. Our Channel 3 example experiences an aborted core-collapse explosion, more often in the context of a low-metallicity progenitor, whose residue is a black hole with a mass perhaps up to $\sim$40 $M_{\odot}$. The latter may be accompanied by a pulsational-pair instability supernova (PPISN). Channel 4 is the only quiescent or ``silent" scenario for which perhaps $\sim$5 to 15 $M_{\odot}$ black holes are left. Where appropriate, we estimate $^{56}$Ni yields, explosion energies, approximate recoil speeds, and residual black hole masses. The progenitor mass density and binding energy profiles at collapse influence the outcome in a systematic way. We speculate that the statistics and prevalence of these various channels depend not only on still evolving supernova theory, but on remaining issues with the theory of massive star evolution, binary interaction, wind mass loss, metallicity, and the nuclear equation of state. Importantly, we suggest, but have not proven, that the silent channel for black hole formation may not be the dominant formation modality.

astro-ph.SR

A 3D Simulation of a Type II-P Supernova: from Core Bounce to Beyond Shock Breakout

In order to better connect core-collapse supernovae (CCSN) theory with its observational signatures, we have developed a simulation pipeline from the onset of core collapse to beyond shock breakout. Using this framework, we present a three-dimensional simulation study following the evolution from five seconds to over five days of a 17-M$_{\odot}$ progenitor that explodes with $\sim$10$^{51}$ erg of energy and $\sim$0.1 M$_{\odot}$ of $^{56}$Ni ejecta. The early explosion is highly asymmetric, expanding most prominently along the southern hemisphere. This early asymmetry is preserved to shock breakout, $\sim$1 day later. Breakout itself evinces strong angle-dependence, with as much a day delay in shock breakout by direction. The nickel ejecta closely tails the forward shock, with velocities at breakout as high as $\sim$7000 km s$^{-1}$. A delayed reverse shock forming at the H/He interface on hour timescales leads to the formation of Rayleigh-Taylor instabilities, fast-moving nickel bullets, and almost complete mixing of the metal core into the hydrogen envelope. For the first time, we illustrate the angle-dependent emergent broadband and bolometric light curves from simulations evolved in three-dimensions in entirety, continuing through hydrodynamic shock breakout a CCSN model of a massive stellar progenitor evolved with detailed, late-time neutrino microphysics and transport. Our case study of a single progenitor suggests that 3D simulations initiated with detailed neutrino heating can begin to generically produce the cornucopia of suggested asymmetries and features in CCSNe observations, while establishing the methodology to study this problem in breadth.

astro-ph.HE

Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation

Recent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs). Most existing studies have focused on converting user behavior logs into textual prompts and leveraging techniques such as prompt tuning to enable LLMs for recommendation tasks. Meanwhile, research interest has recently grown in multimodal recommendation systems that integrate data from images, text, and other sources using modality fusion techniques. This introduces new challenges to the existing LLM-based recommendation paradigm which relies solely on text modality information. Moreover, although Multimodal Large Language Models (MLLMs) capable of processing multi-modal inputs have emerged, how to equip MLLMs with multi-modal recommendation capabilities remains largely unexplored. To this end, in this paper, we propose the Multimodal Large Language Model-enhanced Multimodaln Sequential Recommendation (MLLM-MSR) model. To capture the dynamic user preference, we design a two-stage user preference summarization method. Specifically, we first utilize an MLLM-based item-summarizer to extract image feature given an item and convert the image into text. Then, we employ a recurrent user preference summarization generation paradigm to capture the dynamic changes in user preferences based on an LLM-based user-summarizer. Finally, to enable the MLLM for multi-modal recommendation task, we propose to fine-tune a MLLM-based recommender using Supervised Fine-Tuning (SFT) techniques. Extensive evaluations across various datasets validate the effectiveness of MLLM-MSR, showcasing its superior ability to capture and adapt to the evolving dynamics of user preferences.

cs.IR

Insights into the Production of $^{44}$Ti and Nickel Isotopes in Core-Collapse Supernovae

We report nucleosynthetic results for both $^{44}$Ti and nickel isotopes for eighteen three-dimensional (3D) core-collapse supernova (CCSN) simulations extended to $\sim$20 seconds after bounce. We find that many of our long-term models are able to achieve $^{44}$Ti/$^{56}$Ni ratios similar to that observed in Cassiopeia A, and modern supernova models can synthesize up to $2\times10^{-4}M_\odot$ of $^{44}$Ti. Neutrino-driven winds and the fact that there can be simultaneous accretion and explosion in 3D models of core-collapse supernovae play central roles in its production. We conclude that the $^{44}$Ti underproduction problem in previous CCSN models is no longer an issue. In addition, we discuss the production of both $^{57}$Ni and stable nickel/iron ratios and compare our results to observations of SN1987A and the Crab.

astro-ph.HE

Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching

Entity matching (EM) is a critical step in entity resolution (ER). Recently, entity matching based on large language models (LLMs) has shown great promise. However, current LLM-based entity matching approaches typically follow a binary matching paradigm that ignores the global consistency among record relationships. In this paper, we investigate various methodologies for LLM-based entity matching that incorporate record interactions from different perspectives. Specifically, we comprehensively compare three representative strategies: matching, comparing, and selecting, and analyze their respective advantages and challenges in diverse scenarios. Based on our findings, we further design a compound entity matching framework (ComEM) that leverages the composition of multiple strategies and LLMs. ComEM benefits from the advantages of different sides and achieves improvements in both effectiveness and efficiency. Experimental results on 8 ER datasets and 10 LLMs verify the superiority of incorporating record interactions through the selecting strategy, as well as the further cost-effectiveness brought by ComEM.

cs.CL

Supernova Explosions of the Lowest-Mass Massive Star Progenitors

We here focus on the behavior of supernovae that technically explode in 1D (spherical symmetry). When simulated in 3D, however, the outcomes of representative progenitors of this class are quite different in almost all relevant quantities. In 3D, the explosion energies can be two to ten times higher, and there are correspondingly large differences in the $^{56}$Ni yields. These differences between the 3D and 1D simulations reflect in part the relative delay to explosion of the latter and in the former the presence of proto-neutron star convection that boosts the driving neutrino luminosities by as much as $\sim$50\% at later times. In addition, we find that the ejecta in 3D models are more neutron-rich, resulting in significant weak r-process and $^{48}$Ca yields. Furthermore, we find that in 3D the core is an interesting, though subdominant, source of acoustic power. In summary, we find that though a model might be found theoretically to explode in 1D, one must perform supernova simulations in 3D to capture most of the associated observables. The differences between 1D and 3D models are just too large to ignore.

astro-ph.HE

Guidance Design for Escape Flight Vehicle Using Evolution Strategy Enhanced Deep Reinforcement Learning

Guidance commands of flight vehicles are a series of data sets with fixed time intervals, thus guidance design constitutes a sequential decision problem and satisfies the basic conditions for using deep reinforcement learning (DRL). In this paper, we consider the scenario where the escape flight vehicle (EFV) generates guidance commands based on DRL and the pursuit flight vehicle (PFV) generates guidance commands based on the proportional navigation method. For the EFV, the objective of the guidance design entails progressively maximizing the residual velocity, subject to the constraint imposed by the given evasion distance. Thus an irregular dynamic max-min problem of extremely large-scale is formulated, where the time instant when the optimal solution can be attained is uncertain and the optimum solution depends on all the intermediate guidance commands generated before. For solving this problem, a two-step strategy is conceived. In the first step, we use the proximal policy optimization (PPO) algorithm to generate the guidance commands of the EFV. The results obtained by PPO in the global search space are coarse, despite the fact that the reward function, the neural network parameters and the learning rate are designed elaborately. Therefore, in the second step, we propose to invoke the evolution strategy (ES) based algorithm, which uses the result of PPO as the initial value, to further improve the quality of the solution by searching in the local space. Simulation results demonstrate that the proposed guidance design method based on the PPO algorithm is capable of achieving a residual velocity of 67.24 m/s, higher than the residual velocities achieved by the benchmark soft actor-critic and deep deterministic policy gradient algorithms. Furthermore, the proposed ES-enhanced PPO algorithm outperforms the PPO algorithm by 2.7\%, achieving a residual velocity of 69.04 m/s.

cs.LG

URL: Universal Referential Knowledge Linking via Task-instructed Representation Compression

Linking a claim to grounded references is a critical ability to fulfill human demands for authentic and reliable information. Current studies are limited to specific tasks like information retrieval or semantic matching, where the claim-reference relationships are unique and fixed, while the referential knowledge linking (RKL) in real-world can be much more diverse and complex. In this paper, we propose universal referential knowledge linking (URL), which aims to resolve diversified referential knowledge linking tasks by one unified model. To this end, we propose a LLM-driven task-instructed representation compression, as well as a multi-view learning approach, in order to effectively adapt the instruction following and semantic understanding abilities of LLMs to referential knowledge linking. Furthermore, we also construct a new benchmark to evaluate ability of models on referential knowledge linking tasks across different scenarios. Experiments demonstrate that universal RKL is challenging for existing approaches, while the proposed framework can effectively resolve the task across various scenarios, and therefore outperforms previous approaches by a large margin.

cs.CL