Search arXivSearch

arXiv subjects

Cong Xu

Publications and source records attributed to Cong Xu.

At least 19 recordsLinked to original sources

Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle. It mirrors computing before operating systems, when every program re-implemented basic services. This position paper argues that the field now needs a Foundation Model Operating System (FMOS) -- a system layer that virtualizes FM interactions analogous to how virtual machines abstract physical hardware, giving applications the illusion of dedicated, trustworthy FM instances with effectively unbounded capabilities. Internally, the FMOS orchestrates knowledge across memory tiers, model selection and resource allocation, and verification and policy enforcement. Like the human brain switching between fast intuition and slow deliberation, the FMOS learns when to intervene and when to let inference proceed directly and continuously adapting its policies based on operational experience.

cs.AI

Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection

We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Workshop, ECCV 2026. The detector must be trained only on synthetic data and is evaluated on 50 held-out real LiDAR frames; a separate 50-frame synthetic submission is evaluated for point-cloud realism. Our method addresses the Sim2Real gap at three levels. First, we align synthetic scans to the 50k-point test density and build a 30k-record training pool using UT-LUMPI geometry, RangeLDM-based sampling diversification, rare-class copy-paste, and pedestrian-oriented augmentation. Second, complementary DSVT detectors and Car/Bus PointPillars specialists are trained under the same synthetic-only constraint. Third, predictions are integrated by class-aware routing, asymmetric agreement fusion, constrained residual-recall supplementation, class-coverage auditing, and selective box-size calibration. The realism branch is optimized independently with radial-density matching, weak affine calibration, and calibrated set mixing. The final submission obtains a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and 3D mAP@0.5 of 0.1258.

cs.CV

Solution for UCF UrbanTwin V2X-Real Track: Sim-to-Real Urban LiDAR 3D Object Detection

Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method organizes digital-twin scans, diffusion-redrawn scans, density-stabilized scans, and pedestrian morphology-aligned samples into a unified training pool with complementary roles. Within a common DSVT detection formulation, source-specialized expert branches preserve those roles while optimizing for the same detection objective. At inference, a predefined class-aware fusion pathway integrates geometry-stable and calibration-aware branches for vehicles, sampling-complementary branches for trucks, and morphology-consistent evidence for pedestrians. A label-free point-cloud center blend then refines geometric localization. On the UrbanTwin V2X-Real hidden test set, the unified system achieves a combined score of 0.7421, with 3D mAP@0.5 of 0.4518 and a realism score of 0.8871. The results indicate that a stable, interpretable collaboration among data sources is more valuable than unconstrained aggregation of model outputs.

cs.CV

Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving

Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in which a language-guided intent module reads structured descriptors, computes a smoothed intent-geometry divergence score, and gates the planned maneuver before commitment, upstream of a corridor envelope. On a replayed off-road departure and four crash clips under a frozen, disclosed implementation, gating is the only layer that repairs the plan: on the main case it fires 72 ms after the drift onset but 161 ms before the corridor exit, keeping the trajectory in the corridor in all ten replays. The first calibration draws nine false triggers in 5.9 minutes, each from scoring uncertainty as half a conflict; a preregistered redesign treating uncertainty as abstention cuts this to 0.341 per minute. Two ablations bound the model's contribution: the full score detects fastest on four of five failures under the deployed eligibility, three of five against the unvetoed rule (000871 by one cycle; 000228 by a pre-onset fire on an uncertain stretch that five clips cannot classify as signal or coincidence; dropping the confidence term costs two detections), while on in-domain tracks at equal false positives the geometric rule more than triples its detection. The evidence supports the gating mechanism; the model's demonstrated roles are the fastest detection on these failures and an uncertainty veto on the geometric rule.

cs.RO

Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed, mathematically executable, and unit-consistent, yet contextually ungrounded. Current approaches either rely on formal verifiers that cannot assess semantic intent, or burden Process Reward Models (PRMs) with the dual task of checking both arithmetic and logic. In this paper, we propose a neuro-symbolic framework that cleanly decouples reasoning into two formal dimensions: Symbolic Validity ($V$) and Semantic Groundedness ($G$). We guarantee $V$ by construction using a deterministic symbolic verifier acting as a hard filter. To assess $G$, we train a PRM conditionally on the verifier-accepted manifold. To train this PRM efficiently, we introduce Counterfactual Symbolic Perturbation (CSP), a novel data synthesis strategy that algorithmically generates constraint-preserving hard negatives (steps that perfectly pass the verifier but are logically flawed). At inference, we deploy a verifier-first constrained search that guarantees execution consistency for verifier-covered operations while relying on the PRM solely to rank semantic grounding. By targeting the exact residual error class of strong tool-using LLMs, our method significantly improves reasoning reliability without the sprawling heuristics of prior frameworks.

cs.CL

Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails

A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-critical content when perception degrades, but it presumes observations to populate it; when the primary visual modality is occluded or degraded, those observations may be missing. We address how to sustain the world model from a complementary modality by treating the absence of expected co-evidence as evidence of a hidden cause. The abductive framework is modality-agnostic; this article instantiates it acoustically. A microphone-array front-end estimates the bearing of engine and tire sources and extracts approach-rate evidence (Doppler when a stable tone exists, a broadband looming readout otherwise); the event "signature present, visual co-evidence absent" then triggers abductive inference of a hidden road user, emitting a calibrated risk advisory rather than a control command. Recoverability of the hidden state is analyzed as an identifiability question separating shared from modality-unique information, and cueing is cast as Neyman-Pearson detection under an explicit false-alarm budget. On real occluded-approach recordings at blind junctions, the method warns a mean 1.7 seconds before line-of-sight entry, matches the sustained-window variant of the published acoustic baseline's detection rate with 42% fewer false alarms, localizes to 3.4 degrees median once in view, is well calibrated (expected calibration error 0.034), and keeps hazard awareness above 0.87 under staged vision degradation that collapses a vision-only channel to 0.03. We also measure the method's limits: calibration transfers to an unseen junction almost losslessly, the signature classifier does not, and moving-ego noise is the binding deployment constraint.

cs.RO

Critical Thresholds in Non-Pharmaceutical Interventions for Epidemic Control

Non-pharmaceutical interventions, such as contact tracing and social distancing, are critical for controlling epidemic outbreaks, yet their dynamic interactions remain underexplored. We introduce a probabilistic framework to analyze the synergy between contact tracing speed, quantified by the contact tracing period $τ$, and the average number of close contacts, $\bar{k}_+$, reflecting social distancing measures. We identify critical thresholds ($R=1$) that separate pandemic and contained phases in the $\bar{k}_{+}-τ$ plane, validated using high-resolution data from Shenzhen's 2022 Omicron outbreak (1,187 cases, 86,451 contacts). Our findings show that contact tracing alone can contain diseases with $R_0 < 2.12$ (95% CI 2.07-2.16), covering 43.33% of major infectious diseases, while combining with social distancing extends control to $R_0 < 7.82$ (95% CI 7.70-7.93), encompassing 86.67% of pathogens. These results, supported by empirical data, highlight the efficacy of rapid tracing and targeted social distancing as alternatives to mass PCR testing. Our framework offers actionable insights for optimizing NPI strategies, though challenges in scaling to regions with higher tracing miss rates or weaker infrastructure underscore the need for adaptive, data-driven policies.

physics.soc-ph

Testing the Transverse Scalar Mode of Gravitational Quantum Field Theory with Taiji and LISA

Space-based gravitational-wave (GW) detectors, including LISA and Taiji, offer unprecedented access to regimes where alternative theories of gravity may deviate from General Relativity (GR). Gravitational Quantum Field Theory (GQFT) provides a novel framework in which the Poincaré-type inhomogeneous spin symmetry of Weyl-type fermions in the Standard Model is elevated to a gauge symmetry. Within this construction, the fundamental gravitational field is identified with a gravigauge field which behaves as a Goldstone-type bi-covariant vector field. Unlike GR, GQFT predicts additional polarization states: one transverse scalar (breathing) mode and two vector modes. In this work, we focus on the transverse, isotropic scalar mode and investigate its detectability with Taiji. To isolate this mode, we employ the null-response channel (NRC), a specific interferometric combination designed to suppress contributions from other polarizations. We implement an analytical, dynamic orbital model to realistically simulate a triangular constellation. We compute the response functions and sensitivity curves for various interferometric channels, compare them with the standard Michelson channel, and demonstrate the effectiveness of the NRC approach. Our results show that the NRC provides a reliable, waveform-independent criterion for testing non-GR polarizations, and we anticipate that it will serve as a valuable tool for probing gravitational theories in future space-based GW missions.

gr-qc

Probing Gravitational Quantum Field Theory through Polarization Fingerprints of Gravitational Waves

Gravitational Quantum Field Theory (GQFT) has been proposed as a candidate framework to reconcile general relativity with quantum field theory, and a distinctive imprint on gravitational-wave (GW) polarizations is crucially predicted. While general relativity allows only two tensor modes ($+, \times$), GQFT additionally favors a massless breathing scalar mode, providing a compelling yet largely unexplored observational target for testing quantum gravity. The central challenge is therefore to assess, in a mission-agnostic manner, how well future space-based interferometers can disentangle and detect these tensor and scalar polarization components across the sky. In this work, we develop a model-independent response formalism for LISA- and Taiji-like detectors by incorporating first-order orbital dynamics in the Solar System Barycenter frame. This framework yields three key observational consequences: (1) characteristic interference patterns between tensor and scalar modes, (2) a generalized, model-independent response function for the breathing mode, and (3) sky-position-dependent strategies that optimize detectability. We further translate the formalism into comprehensive polarization maps that provide complete sky coverage and remain fully compatible with existing mission designs, thereby circumventing the need for challenging direct breathing-mode measurements. Overall, our results deliver practical tools for future data analysis and establish a systematic avenue to test fundamental theories of gravity through their GW polarization fingerprints.

gr-qc

A Utility Score Framework for Dose Optimization Studies with Binary Efficacy-Safety Endpoints: Sample Size Determination and Bias Characterization

The FDA's Project Optimus initiative emphasizes patient-centered dose selection in oncology that balances efficacy and safety. We develop a framework for randomized dose optimization studies that uses clinically interpretable utility scores to integrate binary efficacy and safety endpoints and select the optimal dose for a follow-on confirmatory trial. The framework provides: (i) a systematic method for eliciting utility scores that reflect clinical priorities; (ii) closed-form sample size formulas to achieve prespecified Probabilities of Correct Selection (PCS) under clinically relevant scenarios; and (iii) analytical expressions characterizing the propagation of selection-induced bias to confirmatory trials, including time-to-event endpoints correlated with the selection endpoint. Extensive simulations (10^6 replications per scenario) confirm that the sample size methods achieve target PCS and that the bias and Type I error formulas closely match empirical estimates. An R package DoseOptDesign and an interactive Shiny application are publicly available.

stat.AP

HiGR: Industrial-Scale Hierarchical Generative Slate Recommendation Framework in Tencent

Slate recommendation, which presents users with a ranked item list in a single display, is ubiquitous across mainstream online platforms. While recent generative recommendation methods have shown strong potential in modeling item sequences with semantic IDs, directly applying them to industrial-scale slate recommendation faces a fundamental disconnect: entangled SID spaces confound high-level list planning, fine-grained autoregressive decoding over long sequences limits semantic planning efficiency, and token-level objectives misalign with holistic slate quality. In this paper, we propose HiGR, an industrial-scale hierarchical generative framework for slate recommendation that bridges this disconnect through a co-designed pipeline. First, HiGR learns structured SIDs via a Prefix-Contrastive Residual Quantized VAE (PCRQ-VAE). By enforcing high-level prefixes to capture shared semantics, PCRQ-VAE creates a controllable discrete space that acts as a prerequisite for efficient planning. Leveraging this structured space, our Hierarchical Slate Decoder (HSD) shifts autoregressive modeling from entangled token-level decoding to coarse-grained preference embeddings. This design significantly reduces inference latency while allowing explicit global slate structure planning. Finally, this stable planning space enables an ORPO-based listwise alignment mechanism to optimize triple-objective implicit feedback-ranking fidelity, genuine user interest, and diversity. Extensive offline experiments show that HiGR outperforms state-of-the-art baselines by over 10% in offline recommendation quality while achieving a $5\times$ inference speedup. Online A/B tests on Tencent platforms further improve watch time by 1.22% and video plays by 1.73%. HiGR has been deployed on multiple Tencent platform surfaces, serving hundreds of millions of users and proving its industrial-scale applicability.

cs.IR

Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation

Recommender systems in concert with Large Language Models (LLMs) present promising avenues for generating semantically-informed recommendations. However, LLM-based recommenders exhibit a tendency to overemphasize semantic correlations within users' interaction history. When taking pretrained collaborative ID embeddings as input, LLM-based recommenders progressively weaken the inherent collaborative signals as the embeddings propagate through LLM backbones layer by layer, as opposed to traditional Transformer-based sequential models in which collaborative signals are typically preserved or even enhanced for state-of-the-art performance. To address this limitation, we introduce FreLLM4Rec, an approach designed to balance semantic and collaborative information from a spectral perspective. Item embeddings that incorporate both semantic and collaborative information are first purified using a Global Graph Low-Pass Filter (G-LPF) to preliminarily remove irrelevant high-frequency noise. Temporal Frequency Modulation (TFM) then actively preserves collaborative signal layer by layer. Note that the collaborative preservation capability of TFM is theoretically guaranteed by establishing a connection between the optimal but hard-to-implement local graph fourier filters and the suboptimal yet computationally efficient frequency-domain filters. Extensive experiments on four benchmark datasets demonstrate that FreLLM4Rec successfully mitigates collaborative signal attenuation and achieves competitive performance, with improvements of up to 8.00\% in NDCG@10 over the best baseline. Our findings provide insights into how LLMs process collaborative information and offer a principled approach for improving LLM-based recommendation systems.

cs.CL

Bridging the Generalization Gap in Adverse Weather Segmentation: A Training Recipe Perspective

This paper describes our approach for the 8th UG2+ Workshop (CVPR 2026) Track~2, which targets semantic segmentation of outdoor scenes degraded by five weather conditions: blur, darkness, snow, haze, and glare. A central challenge we observe is a severe generalization gap -- models that perform well on the validation set often collapse on the test set. For instance, SegFormer-B5 drops 16.1 mIoU points from validation to test, suggesting that model capacity alone is insufficient for robustness. We investigate whether a carefully designed training recipe, rather than architectural complexity, can address this gap. Starting from a pre-trained SegMAN-S backbone, we systematically study the effects of domain-adaptive fine-tuning, multi-source data mixing, scene-balanced sampling, and synthetic degradation augmentation. Our final system achieves 59.9\% mIoU on the official test set while maintaining a validation-test gap of only 6.5 points -- less than half that of larger models. We analyze negative results from architectural modifications, loss function variants, and model scaling to provide practical insights for weather-robust segmentation under limited data.

cs.CV

Explainable Knowledge Tracing via Probabilistic Embeddings and Pattern-based Reasoning

Knowledge Tracing (KT) models students' knowledge states based on learning interactions to predict performance. While deep learning-based KT models have boosted predictive accuracy, most models rely on deterministic vector embeddings and opaque latent state transitions, limiting interpretability regarding how specific past behaviors influence predictions. To address this limitation, we propose Probabilistic Logical Knowledge Tracing (PLKT), an interpretable KT framework that formulates prediction as a goal-conditioned evidence reasoning process over historical learning behaviors. Instead of representing knowledge states as deterministic vector embeddings, PLKT employs robust Beta-distributed probabilistic embeddings to represent student knowledge states. This probabilistic foundation allows us to model the uncertainty of historical behaviors and perform explicit logical operations (e.g., conjunction), constructing transparent reasoning paths that reveal how specific past interactions contribute to the prediction. Extensive experiments show that PLKT outperforms state-of-the-art KT methods while achieving superior interpretability. Our code is available at https://anonymous.4open.science/r/PLKT-D3CE/.

cs.AI

Quantum average correlations and complementarity relations via metric-adjusted skew information

We investigate quantum average correlations and complementarity relations based on metric-adjusted skew information. Several natural averaging procedures are considered, including complete families of mutually unbiased bases, all orthonormal bases, operator orthonormal bases, and twirling channels induced by the unitary group. All these approaches lead to the same closed expression, which identifies the resulting average correlation as an intrinsic quantity independent of the averaging scheme. By defining measures of wave and particle features via metric-adjusted skew information, we establish complementarity relations among wave and particle features, quantum entropy, and average correlation. These results provide a unified framework for investigating quantum average correlations and complementarity relations in terms of metric-adjusted skew information.

quant-ph

Quantum average correlation based on average coherence

This paper studies the quantification and structural properties of quantum average correlation based on average coherence. Motivated by two mathematically equivalent approaches to define average coherence: one by averaging over complete sets of mutually unbiased bases, and the other by integrating over all orthogonal bases under the Haar measure, we define an average correlation for bipartite systems as the difference between global and local skew information. This correlation measure is shown to satisfy essential properties including non negativity, contractivity under local quantum channels, and local unitary invariance. We further prove the equivalence between the average correlation defined via mutually unbiased bases and that defined via unitary groups. Finally, we derive a complementarity relation that connects wave-particle duality with the average correlation between a system and its environment.

quant-ph

ASPIRE: Make Spectral Graph Collaborative Filtering Great Again via Adaptive Filter Learning

Graph filter design is central to spectral collaborative filtering, yet most existing methods rely on manually tuned hyperparameters rather than fully learnable filters. We show that this challenge stems from a bias in traditional recommendation objectives, which induces a spectral phenomenon termed low-frequency explosion, thereby fundamentally hindering the effective learning of graph filters. To overcome this limitation, we propose a novel adaptive spectral graph collaborative filtering framework (ASPIRE) based on a bi-level optimization objective. Guided by our theoretical analysis, we disentangle the filter learning objective, which in turn leads to excellent recommendation performance, spectral adaptivity, and training stability in practice. Extensive experiments show our learned filters match the performance of carefully engineered task-specific designs. Furthermore, ASPIRE is equally effective in LLM-powered collaborative filtering. Our findings demonstrate that graph filter learning is viable and generalizable, paving the way for more expressive graph neural networks in collaborative filtering.

cs.IR

AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report

This report presents an overview of the AIM 2025 RipSeg Challenge, a competition designed to advance techniques for automatic rip current segmentation in still images. Rip currents are dangerous, fast-moving flows that pose a major risk to beach safety worldwide, making accurate visual detection an important and underexplored research task. The challenge builds on RipVIS, the largest available rip current dataset, and focuses on single-class instance segmentation, where precise delineation is critical to fully capture the extent of rip currents. The dataset spans diverse locations, rip current types, and camera orientations, providing a realistic and challenging benchmark. In total, $75$ participants registered for this first edition, resulting in $5$ valid test submissions. Teams were evaluated on a composite score combining $F_1$, $F_2$, $AP_{50}$, and $AP_{[50:95]}$, ensuring robust and application-relevant rankings. The top-performing methods leveraged deep learning architectures, domain adaptation techniques, pretrained models, and domain generalization strategies to improve performance under diverse conditions. This report outlines the dataset details, competition framework, evaluation metrics, and final results, providing insights into the current state of rip current segmentation. We conclude with a discussion of key challenges, lessons learned from the submissions, and future directions for expanding RipSeg.

cs.CV