Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 721 records · Page 40Linked to original sources

Keeping the interaction structure explicit: comment on "Graphs are maximally expressive for higher-order interactions"

Peixoto et al. (arXiv:2602.16937v2) stress that graph-based formulations can express any interaction model, a point the literature on higher-order networks (HONs) has often overlooked. We agree with much of their critique. We argue that expressiveness, however, might not be the main reason hypergraphs are often preferred. A network is read from a model rather than assumed, and to properly study the role of structure-one of the most basic questions in complex systems research-one needs a representation that keeps that structure separate from the functional form of the interactions. We show that such representation is in general a directed hypergraph (equivalently, a directed factor graph), with one hyperarc per interaction term, recovering a (directed) graph for pairwise interactions, and an undirected hypergraph for symmetric, higher-order ones. Most of the HONs literature has focused so far on symmetric interactions, which, in light of what established here, might explain why there (undirected) hypergraphs are regularly presented as the natural representation for systems with higher-order interactions.

physics.soc-ph↗

Mitigating prior-volume effects in galaxy clustering by adopting Mpc over Mpc/h units

We investigate prior-volume effects in Bayesian analyses of galaxy clustering based on state-of-the-art models, such as the EFT and the VDG$_\infty$ model. We show that these effects are largely driven by imposing priors on the model parameters defined in ${\rm Mpc}/h$ units, thereby introducing a spurious correlation with the Hubble parameter $h$, and that they can be strongly mitigated by adopting a more physically motivated formulation in ${\rm Mpc}$ units. We test this approach using a set of synthetic data vectors designed to reproduce the clustering signal expected in a Stage-IV survey across a range of cosmological models, including $Λ$CDM, $K$CDM, $w$CDM, and $w_0w_a$CDM. We find that the resulting constraints are largely free from shifts induced during the marginalisation process, even in the challenging $w_0w_a$CDM runs, reducing discrepancies with the fiducial values from several standard deviations to well within $1σ$. The performance of this approach is comparable to that of alternative mitigation strategies, such as the adoption of invariant priors or the reparametrisation of nuisance parameters in terms of growth and geometrical amplitudes. For the latter, we find that using $σ_{12}$ in place of the commonly adopted $σ_8$ to describe the amplitude of density fluctuations improves the performance of the reparametrisation technique, even at high redshifts, where constraints on the dark-energy equation of state are weaker. We further show that these conclusions hold with the inclusion of additional BAO constraints, and for data sets with reduced statistical power. These results add to the benefits of working with ${\rm Mpc}$ units, avoiding an explicit dependence on $h$, with direct relevance for the analyses of Stage-IV galaxy surveys.

astro-ph.CO↗

Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States

Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable challenge. Conventional models rely on historical observations and typically falter when encountering nodes without prior records. To address this, we redefine the problem as a conditional generation task on spatio-temporal graphs and propose GenST, a framework that introduces Large Language Models (LLMs) as a semantic bridge, leveraging a pre-trained LLM fine-tuned to extract rich semantic features from node descriptions, such as functional zones and road network structures, to compensate for missing spatio-temporal signals. Specifically, we design a two-stage generative architecture: a Spatio-Temporal VAE first compresses spatio-temporal dynamics into a latent space, followed by a Generative Transformer (GenT) that reconstructs the future states of unobserved nodes from noise, guided by multi-modal conditions including semantics, geographic coordinates, and neighborhood contexts. Experiments on six traffic and two non-traffic datasets show GenST significantly outperforms existing baselines in zero-shot prediction tasks, demonstrating the practical potential of semantic-guided generation for mitigating spatio-temporal data sparsity.

cs.LG↗

Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News

The spread of fake news may cause severe social consequences. Existing fake news detection methods mainly focus on stylistic variations or incorporate external information such as explanations. However, news articles are often rewritten under different emotional backgrounds while preserving their underlying factual claims, which may affect the robustness of detection models. In this work, we investigate fake news detec- tion under fact-preserving emotional variations. To study this problem, we construct emotion-rewritten test sets and generate explanations from the original news articles as stable background knowledge. We then propose a Gated Cross Attention (GCA) framework that adaptively integrates emotionally rewritten news with the corresponding explanations, enabling the model to focus on informative explanation content while reducing potential mismatches caused by emotional reframing. Experiments on PolitiFact, GossipCop, and LUN demonstrate that the proposed method achieves notable improvements under multiple emotional conditions on PolitiFact and LUN, while maintaining competitive performance on GossipCop. We further analyze the effects of explanation guidance and gating mechanisms under different emotional conditions. Our code and data are available at: https://github.com/Flulike/fakenews gca .

cs.CL↗

An Empirical Study of Agent Skills' Downstream Utility

Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a curated corpus of 37,596 Skills. LLM-assisted analysis of content, execution traces, and final artifacts, followed by author review, relates provided support to actual use and task outcomes. The same Skills help some configurations and hurt others on 36.78\% of tasks, with trajectories showing that recommended procedures can become an execution burden. Relevance rankings overlook more useful candidates. Within the evaluated candidate sets, reranking by support for required operations raises first-choice pass rates by 4.35--5.80 percentage points across the three configurations. We derive 17 authoring practices linking executable procedures to recovery, preservation of task requirements, and checks on final artifacts. Stage Plan and Dependency DAG outperform use order alone, with DAG's additional benefits concentrated in tasks supplied with five or six Skills. These findings guide developers to assess usable operation support, allow procedure adaptation while preserving task requirements, and make artifact dependencies explicit when organizing Skills.

cs.AI↗

Humanize: Judgement Engineering for Agentic Coding

Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates. Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it. We study Humanize through its deployment, 118 public postmortems of real loops, and its applications. Over 68 versions in 108 days, it gathered 1,468 GitHub stars. Applications include a 567-file gem5 build-system migration under upstream review; Kernel Design Agents, which extend the loop with a kernel knowledge base and profiling feedback and placed in the top three of all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; and, through Humanize Olympiad Agents (HOA), full scores in IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, 418.5/437 in IChO 2026 (gold-medal). Humanize also achieves 672/672 on PutnamBench and ranks first (251/303) on Lean-Eval's leaderboard even competiting with professional mathematicians. The postmortems show that independent review catches unsupported builder claims, but stopping remains a key weakness. In reports that separate rounds by phase, two thirds of rounds occurred after implementation was accepted. This evidence is observational, not a controlled comparison of workflows.

cs.AI↗

PEARLS: NuSTAR and XMM-Newton Extragalactic Survey of the JWST North Ecliptic Pole Time-Domain Field V: Unraveling X-ray Spectral Variability

As one of the first deep time-domain surveys in the hard X-ray band, the NuSTAR campaign in the JWST North Ecliptic Pole Time Domain Field (NEP-TDF) offers a prime opportunity to measure variability in a faint population of Active Galactic Nuclei (AGN) in the medium-to-high redshift universe. In the X-ray band, AGN variability is generally caused by changes in the brightness of the corona ($F_{int}$ variability) or changes in the column density of obscuring material near the SMBH ($N_H$ variability). Broad-band spectroscopy ($\sim$0.5-24 keV) is required to reliably disentangle these two phenomena. With simultaneous NuSTAR and XMM-Newton coverage---as well as 1.8 Ms of Chandra monitoring---the X-ray surveys in the NEP-TDF are ideally situated for this purpose. This work presents the first spectral modeling of the 52 NuSTAR-identified sources discovered in the latter half of the campaign, as well as a spectroscopic variability study of the entire 112-source NuSTAR NEP-TDF catalog. Seven variable candidates are identified. We find that $\gtrsim 1000$ total counts split across many ($\gtrsim 10$) epochs are needed in order to detect variability. Additionally, variable sources are compared to the survey sensitivity in order to inform population synthesis models of how variations can affect the detectability of faint AGN, which dominate the SMBH accretion history. Lastly, the timescale of $N_H$ variability is investigated as a probe for the obscurer location. Significant variability is slightly more common on $> 100$ day timescales, suggesting that the torus may have only slight dominance in driving $N_H$ variability in faint AGN.

astro-ph.HE↗

Directed Temporal Representations for Offline Visual Control

Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.

cs.LG↗

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.

cs.AI↗

Convergence of Roberts flow dynamos in Eulerian and Lagrangian codes

Standard benchmark tests for numerical magnetohydrodynamics codes focus on one- and two-dimensional problems that ignore the possibility of dynamos, i.e., an exponential instability converting kinetic energy into magnetic energy. They also tend to ignore magnetic helicity conservation that can have dramatic effects in periodic domains. The existence of magnetic diffusion is crucial for dynamos, as was exemplified by the impossibility of finding dynamo solutions when using Euler potentials, which is a perfectly valid approach in the absence of magnetic diffusion. The Roberts flows~I, II, III, and IV are two-dimensional flows that support three-dimensional dynamos, all of which are also large-scale dynamos in the sense that their planar averages are of significant strengths, even though flow~II is pointwise non-helical. Here we demonstrate the numerical convergence properties of solutions obtained using discretisation schemes of second, sixth, and tenth order in the mesh width in the Pencil Code, compare the results with the SPH-based SWIFT code, and present a list of averaged quantities that facilitate their use in characterising the solutions as benchmarks. In addition, we present a procedure to obtain two-dimensional time-independent magnetic field visualisation for all Roberts flows, and perform per-pixel comparison between the Pencil and SWIFT codes. The rich combination of features makes Roberts flows an excellent dynamo benchmark for the direct numerical simulations of magnetohydrodynamics equations.

physics.comp-ph↗

SPIN: Shadow Predictive Indexer for Sparse Attention

Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.

cs.LG↗

MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence ($A$), read-in ($B$), read-out ($C$), skip ($D$), and discretization ($Δ$) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.

cs.LG↗

The coupler Eve did not monitor: a black-box cavity cheat

Eve receives two average photon counts, while a monitor watching for external leakage may register nothing. Can she conclude that Alice and Bob's opaque cavity devices did not communicate? We give a constructive counterexample: two independently operated controls activate a shared beam-splitter interaction after the inputs arrive. The parties report genuine sample averages, and a fixed local threshold converts these reports into a CHSH game with expected score $S=13/4$ for four repetitions per block, approaching $4$ for large blocks in the ideal model. This exceeds the usual $2\sqrt2$ quantum benchmark because the devices interact during their response. A separate dispersive-bus model explains how substantial exchange can coexist with a small probability of a click in a specified leakage monitor; an illustrative open-system calculation exhibits both effects without postselection. The output marginals nevertheless reveal signalling. The example separates three questions: whether the reports are honest, whether a monitored port is quiet, and whether the devices are physically isolated.

quant-ph↗

Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models

We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we achieve better generalization in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants ($3.73\% \to 24.65\%$ solve rate on Sudoku Hard), and increased diversity of generated solutions ($50.60\% \to 73.79\%$ coverage on 10x10 N-Queens). We also show measurable improvements to natural language generation quality in modest dataset regimes with Gigaword summarization, but notably demonstrate that gains do not transfer to all natural language tasks (e.g open ended dialogue generation). Our method is a single line change to the training objective, requires no additional inference costs by default, and provides the flexibility of classifier-free guidance inspired guided sampling. Our \href{https://github.com/LateralIntelligence/noise-your-prompt}{code} is publicly available.

cs.LG↗

An Accuracy-Information Tradeoff for Loss-Difference Conditional Mutual Information

Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power $r\ge2$, on product distributions over a scaled sign cube in dimension at least linear in $n$, every proper learner with expected excess risk at most $\varepsilon$ on these distributions at the optimal sample size $n\asymp\varepsilon^{-2+2/r}$ has worst-case ld-CMI of order $n$ bits, and $Θ(n/(1+(τ/\varepsilon)^2))$ bits under Gaussian noise of standard deviation $τ$ on the loss differences. The same holds without a regularizer, at $n\asymp\varepsilon^{-2}$. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is $O(n^{-1/2})$. We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.

cs.LG↗

Embeddings of Generalized Weighted Morrey Spaces

We study continuous embeddings from two classes of generalized weighted Morrey spaces into generalized Morrey spaces. The Komori--Shirai-type spaces use the weighted measure of cubes in their normalization, whereas the Samko-type spaces use the Lebesgue measure. We first consider generalized Morrey spaces on bounded domains. In the non-radial setting, we establish embedding results under local radial comparability assumptions on the Morrey functions. As a consequence, we characterize the corresponding embeddings in the radial setting for the full range $0<p_1,p_2<\infty$. We then study embeddings from generalized Komori--Shirai-type weighted Morrey spaces associated with piecewise power weights and general Muckenhoupt weights. For the Samko-type scale, we treat both generalized and classical weighted Morrey spaces, again considering piecewise power weights and general Muckenhoupt weights. Finally, assuming that the weight is comparable to a positive constant on some cube or ball, we show that the continuous embeddings under consideration are not compact.

math.FA↗

Consumption and Investment When Welfare Floors Change over Time

We study consumption and investment over a finite horizon when continuation utility must stay above a welfare floor that changes over time, and the floor at the horizon acts as a terminal wealth floor. A cumulative multiplier reduces the problem to a parabolic obstacle problem. When relative risk aversion exceeds one, the intervention boundary exists at every interior date. When it is below one, the boundary exists exactly at the dates where a discounted transformation of the floor is locally increasing and sets a strict running record, so it can disappear and later reenter. We prove local regularity of the boundary, construct the dual value from the obstacle solution, and verify the optimal cumulative multiplier, whose moments are controlled through its representation as a running supremum. Strong duality identifies the exact lower endpoint of feasible wealth. With stochastic income in a complete market, the floor becomes the value of an outside option, which links the boundary to limited commitment and endogenous borrowing capacity.

math.OC↗

LACE-CRAFT: Robot Co-Design with Actor Inheritance and Blackboard Collaboration

Robot co-design couples morphology search with policy learning, yet training every new design from scratch discards acquired control experience. We present LACE-CRAFT, which compares continued learning on the current robot with policy adaptation to new morphology-reward pairs. LACE resumes the incumbent's full learning state and initializes compatible challengers with its actor parameters and observation statistics. A fixed task metric selects among both branches and the frozen incumbent. CRAFT coordinates Feedback, Morphology, Reward, and Integration roles through shared experimental records and behavioral replays to generate and cross-review paired proposals. A generative extension converts generated meshes into editable articulated models with configured joints, actuator interfaces, and consistently updated simulation assets. Across five locomotion benchmarks, mean scores over three evaluation seeds are 6.4-91.9% higher than D2C. Both methods train 30 new morphology-reward pairs over five rounds; LACE additionally uses four continuation training units. Five-task ablations examine policy inheritance and replay-derived feedback. A fabricated prototype demonstrates indoor walking and illustrates the geometry-to-hardware workflow.

cs.RO↗