Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 109 records · Page 6Linked to original sources

Evolving Idea Graphs with Learnable Edits-and-Commits for Multi-Agent Scientific Ideation

LLM-empowered multi-agent systems offer new potential to accelerate scientific discovery by generating novel research ideas. However, existing methods typically coordinate agents through temporary texts, such as drafts or chat logs; it is difficult to pinpoint the weaknesses in the generated ideas and how the agents refine them. To this end, we introduce \textbf{Evolving Idea Graphs} (EIG), a graph-based multi-agent scientific ideation framework that can generate high-performance research ideas across various benchmark-native metrics, such as novelty, feasibility, and clarity. Instead of coordinating solely through texts, EIG represents a partially formed proposal as an evolving idea graph, where nodes capture scientific claims and edges encode relations (e.g., support and conflict), enabling unresolved weaknesses to remain identifiable throughout the idea evolving process. Specifically, a learned two-head controller operates over the evolving graph to guide the ideation: one head selects graph edits for agents to execute, while the other decides when the graph is ready for commit as final proposal synthesis. On AI Idea Bench 2025 and LiveIdeaBench, EIG outperforms all compared systems on both automatic benchmark scores and blinded expert ratings. Ablations show that the graph-based state, its typed relations, and the learned edit-and-commit control each contribute to final proposal quality.

cs.MA↗

From Dual Tracking to Clipping: Provably Faster Distributionally Robust Multi-Objective Optimization

Multi-objective optimization (MOO) has received growing attention in applications that require learning under multiple criteria. However, most existing MOO formulations do not explicitly account for distributional shifts in the data. We introduce distributionally robust multi-objective optimization (DR-MOO), which minimizes multiple objectives under their respective worst-case distributions. We propose Pareto-type solution concepts for DR-MOO and develop multi-gradient descent algorithms (MGDA) with provable guarantees. Leveraging a Lagrangian dual reformulation, we first design a double-loop MGDA that uses an inner loop to estimate dual variables and achieves a total sample complexity $\mathcal{O}(ε^{-8})$ for reaching an $ε$-Pareto-stationary point. To further improve convergence, we combine large-batch sampling with gradient clipping to accommodate generalized smoothness and control bias in stochastic preference updates, eliminating the need for double sampling. This yields a single-loop double-clip MGDA with substantially improved sample complexity $\mathcal{O}(ε^{-4})$. Our theory applies to nonconvex problems without requiring uniformly bounded gradients of the dual objectives. Experiments demonstrate that our methods are competitive with state-of-the-art MGDA baselines.

cs.LG↗

Stateful Agent Backdoors: Constructing Cross-Session Attack Programs

Multi-step attacks on large language model agents may depend on opportunities distributed across sessions, such as access to target information or the availability of required tools. To combine these opportunities, an attack needs to retain its state and intermediate results, and choose actions based on current conditions. In this work, we study such attacks as cross-session attack programs. We focus on their shared control structures, including sequential execution with conditional waiting, branching and merging, looping, and condition accumulation. We represent these programs as Mealy machines and construct a sub-backdoor for each transition. We build single-session training trajectories for each sub-backdoor, combine them into a training dataset, and fine-tune the model to learn the local behaviors jointly. After a single injection of the initial trigger, the agent uses persistent memory to connect local behaviors into a complete cross-session attack program. We evaluate four instantiations of these control structures across four models in a LangChain-based agent environment. The primary instantiation achieves mean complete-program success rates of 71.7\%--94.7\% in LangChain. In the official OpenClaw runtime under a controlled configuration, it achieves a complete-program success rate of 85\% for each of two evaluated models. These results demonstrate the feasibility of end-to-end execution of cross-session attack programs with different control structures.

cs.CR↗

Enabling Unsupervised Training of Deep EEG Denoisers With Intelligent Partitioning

Denoising electroencephalogram (EEG) is an inherently challenging task, since neural activity is not only subtle but also inseparable from spectrally overlapping noise artifacts. Today, effective EEG denoising is more important than ever, given the rapid adoption of wearables across various applications. Deep learning methods have shown promising results in decomposition-free denoising that handles the time-varying pervasive EEG artifacts. However, training highly expressive neural networks requires artifact-free EEG, which is inherently unobtainable. To address this, we propose Intelligent Partitioning for Self-supervised Denoising (iPSD). Our method eliminates the need for clean references by learning to partition an input EEG segment into independent noisy realizations with the same underlying signal. This enables self-supervision of deep learning denoisers, even in zero-shot settings where only a single EEG segment to be denoised is available. We validate iPSD through extensive experiments, including validations on wearable EEG from in-ear sensors. The results show that iPSD achieves state-of-the-art performance, most notably under extremely low signal-to-noise ratios (down to -10 dB) and challenging artifacts (e.g., EMG), with spectral fidelity orders of magnitude higher than competitive baselines.

cs.LG↗

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

Large language models are widely used in everyday writing, making reliable AI-generated text detection crucial for academic integrity, content moderation, and provenance tracking. Yet high AUROC on clean, in-distribution benchmarks is not sufficient. Practical detectors must resist adversarial rewrites, generalize to unseen generators and writing domains, and maintain low false-positive rates (FPR). A pooled AI-versus-human objective does not explicitly require the model to distinguish among generator families, so it may fail to learn the generator-specific structure needed for generalization and attribution. We introduce MELD (Multi-Task Equilibrated Learning Detector), which instead trains on class-balanced, family-specific AI-versus-human tasks while sharing a common representation of human writing. MELD produces detection, generator-family, rewrite-task, and token-level predictions in a single forward pass. Format normalization further makes its scores invariant to the modeled reformatting operations. MELD ranks first among open-source submissions in the public RAID leaderboard and matches or exceeds supervised baselines on five of six held-out evaluation pools. To evaluate transfer to unseen generators, we introduce MELD-eval, a held-out test pool built from four frontier chat models. Without further fine-tuning, MELD achieves 99.7% TPR at 1% FPR on MELD-eval and 98% TPR at the same FPR on a held-out generator whose family is absent from training. Finally, in a case study of 5.9 million scientific texts from 2016--2026, MELD's prediction scores remain stable through 2022 and increase from 2023 onward, coinciding with the widespread adoption of LLM-based writing tools. The model, MELD-eval pool, source code, and live demo are available.

cs.CL↗

Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency

Recovering continuous-time dynamics from discrete observations is difficult because local supervision loses fidelity as the observation interval grows. We replace local supervision with a global structural constraint: any flow representing autonomous dynamics must satisfy the semi-group property under time translation. We train a time-conditioned secant velocity field whose deviation from this property---which we call Symmetry Rupture---serves two roles: as a training regularizer it confines the hypothesis space to flows that compose consistently across temporal scales; as an inference oracle it guides an adaptive solver to select the largest step size that preserves internal consistency, without relying on local truncation error estimates. On the diffusion-reaction and Navier-Stokes benchmarks, our method achieves the lowest rollout RMSE among all evaluated methods using $\leq 3$ function evaluations per step. On the near-conservative shallow water benchmark, our method matches or exceeds competing flow-based methods at far lower computational cost; the strongest Neural ODE baseline achieves lower absolute RMSE but at the expense of $>100$ function evaluations per step. In the direct auto-regressive setting, our solver maintains stable long-horizon rollouts where flow-matching baselines diverge and Neural ODE requires up to $12\times$ more function evaluations. The core contribution is a model that internalizes temporal consistency, removing the need for high-order external solvers to compensate for structural bias.

cs.LG↗

Towards Effective Theory of LLMs: A Representation Learning Approach

We propose Representational Effective Theory (RET), a framework for describing large language model computation in terms of learned macrostates rather than microscopic details. RET learns these macrostates from hidden-state trajectories using a BYOL/JEPA-style self-supervised objective, coarse-graining activations into macrovariables that preserve higher-level structure relevant for prediction and interpretation. We evaluate whether these macrovariables are practically relevant for interpretability: RET yields temporally consistent states that reveal ``mental-state'' trajectories of reasoning, capture high-level semantic structure, support early prediction of behavioral outcomes such as sycophancy and alignment faking, while providing causal handles for steering generations toward interpretable computational phases. Together, these results suggest that LLM computation admits useful effective descriptions via RET: high-level, dynamically meaningful variables for interpretation, prediction, and control.

cs.LG↗

PoDAR: Power-Decoupled Audio Representation for Generative Modeling

The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor decoupling. We present PoDAR (Power-Decoupled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a 2x acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.

eess.AS↗

FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity

While the overall inference latency of Video Diffusion Transformers (DiTs) can be substantially reduced through model distillation, per-step inference latency remains a critical bottleneck. Existing acceleration paradigms primarily exploit redundancy across the denoising trajectory; however, we identify a limitation where these step-wise strategies encounter diminishing returns in few-step regimes. In such scenarios, fewer intermediate states can limit opportunities for feature reuse or predictive modeling, motivating complementary forms of acceleration. To overcome this, we propose Frame Interleaved Sparsity DiT (FIS-DiT), a training-free and operator-agnostic framework that shifts the optimization focus from the temporal trajectory to the latent frame dimension. Our approach is motivated by an intrinsic duality within this dimension: the existence of frame-wise sparsity that permits reduced computation, coupled with a structural consistency that calls for maintaining coverage of frame positions in the global spatiotemporal context. Leveraging this insight, we implement Frame Interleaved Sparsity (FIS) as an execution strategy that manipulates frame subsets across the model hierarchy, refreshing all latent positions without requiring full-scale block computation. Empirical evaluations on Wan 2.2 and HunyuanVideo 1.5 demonstrate that FIS-DiT reaches 2.4$\times$ speedup while maintaining comparable performance on VBench, providing a scalable and robust pathway toward real-time high-definition video generation.

cs.CV↗

The End of Trust: How Agentic AI Breaks Security Assumptions

For decades, the security of digital interaction has rested on an unacknowledged economic constraint. Attackers faced a tradeoff between the fidelity of a deception and the scale at which it could be deployed. Convincing impersonation required sustained human effort and was confined to a narrow set of high-value targets, while mass-market attacks sacrificed plausibility for reach. Detection systems, verification mechanisms, and user awareness training have all been implicitly calibrated to the artifacts of cheap deception that this tradeoff produced. Agentic AI collapses the tradeoff, allowing high-fidelity, individually tailored deception to be produced at mass-market scale. We argue that this shift exhausts a security paradigm rather than merely intensifying the threat landscape. We introduce the Infinite Impostor, an attack model in which an autonomous agent interposes itself between two parties who already trust each other, hijacking an existing relationship rather than building a new one from scratch. Detection-oriented defenses share an assumption that generative progress is eliminating, that synthetic outputs are distinguishable from authentic ones. We propose a suspect-by-default paradigm that shifts security from authenticating actors to evaluating actions, and examine the governance tensions that arise when platforms become the regulatory substrate of digital interaction.

cs.CR↗

Fidelity Probes for Specification--Code Alignment

Software modernization often relies on natural-language specifications recovered from an existing implementation. Incorrect or missing requirements can lead to defects in the modernized system. We introduce fidelity probes to check a specification against the code and guide its revision. Each probe pairs a question about program behaviour with a reference answer derived from the code. A language model answers the question using only the specification, and the fraction of agreeing answers defines fidelity. Disagreements flag conflicting claims or missing information and guide proposed corrections and additions to the requirements. An LLM generates probes directly from code or phrases facts selected through control-flow, data-flow, and call-graph analysis. We use an observability rule to focus probe questions on behaviour that the modernized system is intended to preserve, including user-visible outputs, changes to stored business data, and interactions with other systems. We revise the specification using fresh probes and evaluate each version on a fixed held-out set. A repair-regression model describes how new failures can offset the gains from repairs. On CardDemo, held-out fidelity improves from 0.59 to 0.89, exceeding free-form revision from source code. We also evaluate the method on industrial systems and public documentation, and assess probe quality and reported defects through human reviews and comparison with an independent requirements audit.

cs.LG↗

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reasoning over temporally complex content and nuanced policy constraints. Existing approaches typically rely on large vision-language models applied uniformly across all inputs, resulting in high inference costs and inefficient allocation of computation. We propose SafeLens, a video guardrail framework that introduces a fast-and-slow inference architecture for efficient and accurate content moderation with variable computational cost across inputs. Additionally, we construct a high-quality dataset by applying influence-guided filtering to the SafeWatch Dataset, retaining only 2.4% of the original data. To further address limitations of training-time scaling, we enable test-time reasoning by augmenting the filtered data with structured Chain-of-Thought traces. Across real-world and AI-generated video benchmarks, SafeLens achieves state-of-the-art performance, outperforming strong open-source video guardrails (e.g., SafeWatch-8B, OmniGuard-7B) and closed-source models (e.g., GPT-5.4, Gemini-3.1-pro) while significantly reducing inference cost, demonstrating that efficient design serves to be more effective than scaling data or model size alone.

cs.CV↗

DCFold: Efficient Protein Structure Generation with Single Forward Pass

AlphaFold3 introduces a diffusion-based architecture that elevates protein structure prediction to all-atom resolution with improved accuracy. This state-of-the-art performance has established AlphaFold3 as a foundation model for diverse generation and design tasks. However, its iterative design substantially increases inference time, limiting practical deployment in downstream settings such as virtual screening and protein design. We propose DCFold, a single-step generative model that attains AlphaFold3-level accuracy. Our Dual Consistency training framework, which incorporates a novel Temporal Geodesic Matching (TGM) scheduler, enables DCFold to achieve a 15x acceleration in inference while maintaining predictive fidelity. We validate its effectiveness across both structure prediction and binder design benchmarks.

cs.LG↗

Wasserstein Saddle-Free Newton: Saddle Escape and Local Spectral Convergence

Extending saddle-free Newton methods to Wasserstein space requires controlling Hessian dependent operators along nonlinear transport trajectories and accounting for a Hessian spectrum that may accumulate at zero. We address these challenges through Wasserstein Saddle-Free Newton (WSFN), a second-order method for nonconvex optimization over Wasserstein space that preconditions the Wasserstein gradient by a regularized inverse square root Hessian operator. The dynamics exploit curvature magnitude while maintaining negative curvature repulsion and positive curvature attraction, avoiding key Wasserstein Newton limitations. Whereas previous Wasserstein saddle escape results rely on perturbed first-order dynamics, WSFN uses curvature both to modify the deterministic transport and to construct the perturbation mechanism. Our analysis combines stability estimates for Hessian operators with a perturbative saddle escape argument to establish polynomial time convergence to approximate second-order stationarity under regularity of landscape assumptions. The contribution of the saddle curvature parameter $δ$ improves from $\tilde O(δ^{-4})$ for prior perturbed first-order Wasserstein methods to $\tilde O(δ^{-3})$. For benign landscapes, this implies proximity to global minimizers. The cubic dependence corresponds to the scale of the objective decrease that can be guaranteed from negative curvature under Lipschitz Hessian regularity. Near a minimizer, the possible absence of a uniform spectral gap prevents a uniform contraction result and instead calls for a spectral characterization. We prove geometric contraction of the linearized dynamics on positive curvature spectral bands bounded away from zero and polynomial decay under spectral source conditions. Together, these results connect quantitative saddle escape with the local spectral behavior of second-order Wasserstein dynamics.

math.OC↗

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large-scale image generation. We identify a key obstacle: NFs are required to learn a single invertible transport over the full ambient space, making them highly sensitive to high-dimensional representations. This leads to a semantic-capacity mismatch in modern visual representation spaces, where semantic information is compact but encoded in overcomplete features. We propose SRC-Flow, which introduces a Semantic Representation Compressor (SRC) to compact high-dimensional RAE features into a low-dimensional semantic space before flow modeling and preserve reconstruction through the frozen RAE decoder. This compact space reduces the modeling burden of NFs and enables effective likelihood-based generation in semantic representation space. We further adopt constant noise regularization tailored to the fixed unconditional bijection learned by flows. On ImageNet 256x256 and 512x512, SRC-Flow achieves state-of-the-art generation quality among normalizing flow methods, with gFID scores of 1.65 and 2.07 under classifier-free guidance, while retaining exact likelihood computation in the compact semantic representation space and deterministic invertible sampling at the flow level. Codes and models will be released in the future.

cs.CV↗

Generalized Functional ANOVA: A Complete Theoretical Framework

The functional ANOVA provides a fundamental representation of square-integrable multivariate functions into main effects and higher-order interactions. For independent inputs, the components belong to mutually orthogonal Hilbert subspaces and admit an explicit representation. For dependent inputs, however, the components are only hierarchically orthogonal: although existence and uniqueness results are available, the Hilbert subspaces underlying the generalized decomposition have remained implicit. We resolve this representation problem for continuous inputs supported on a bounded hyperrectangle whose joint density is bounded above and away from zero. We introduce a distribution-adapted family of functions and prove that it forms a Riesz basis of the $L^2$ space, thereby guaranteeing a unique, stable, and unconditionally convergent representation. We then show that, for every coalition of variables, the corresponding block of this basis exactly characterizes the Hilbert subspace containing the functional ANOVA component. Our construction recovers the classical orthogonal decomposition under input independence. As a direct consequence, computing the generalized functional ANOVA reduces to estimating coefficients in an explicit, distribution-adapted basis. Finally, as a \emph{proof of concept}, we introduce an elementary, fast and model-agnostic estimator based on our theoretical results. Experiments on synthetic and real-world datasets illustrate its connections with established tabular explanation methods and show that low order components often capture most of the signal in the model output.

stat.ML↗

Weak Triplet Models of Neutrino Magnetic Moments

Experimental limits on neutrino magnetic moments remain several orders of magnitude above the predictions of the Standard Model; therefore, any future detection would provide unambiguous evidence for new physics. In models with Dirac neutrinos, however, mechanisms that enhance the magnetic moment typically generate excessively large neutrino masses. Recently, it has been argued that in frameworks where neutrinos mix with weak-triplet Dirac fermions, the magnetic moment can be decoupled from the neutrino mass. In this work, we revisit this possibility and show that sizable enhancements remain highly nontrivial to realize naturally. We demonstrate that, although the minimal realization allows the magnetic moment to be decoupled from the neutrino mass, obtaining an observable enhancement requires a delicate adjustment of the model parameters. Moreover, in extended scenarios, the decoupling no longer persists: the magnetic moment and neutrino mass become intrinsically linked, such that attempts to enhance the former inevitably induce large contributions to the latter.

hep-ph↗

Software Product Line Engineering: Adoption, Tooling and AI Era Challenges

Software Product Line Engineering (SPLE) enables systematic reuse across families of related software-intensive systems. Following an explicit review protocol covering five databases and four research questions, this survey synthesises SPLE foundations, lifecycle concepts, adoption models, tooling and AI-era challenges. We compare seven adoption and evaluation models (BAPO, FEF, PuLSE, SIMPLE, COPLIMO, PROMOTE-PL and APPLIES), trace four eras of SPLE research, and analyse AI applications with their limitations and assurance requirements. Open challenges in interoperability, UVL standardisation, SME adoption, clone-and-own migration, variability-aware DevOps and empirical evidence are consolidated into a compact research agenda.

cs.SE↗