Search arXivSearch

arXiv subjects

Zhe Feng

Publications and source records attributed to Zhe Feng.

At least 19 recordsLinked to original sources

On the Coordination of Value-Maximizing Bidders

While the auto-bidding literature predominantly considers independent bidding, we investigate the coordination problem among multiple auto-bidders in online advertising platforms. Two motivating scenarios are: collaborative bidding among multiple bidders managed by a third-party bidding agent, and strategic bid selection for multiple ad campaigns managed by a single advertiser. We formalize this coordination problem as a theoretical model and investigate the coordination mechanism where only the highest-value bidder competes with outside bidders, while other coordinated bidders refrain from competing. We demonstrate that such a coordination mechanism dominates independent bidding, improving both Return-on-Spend (RoS) compliance and the total value accrued for the participating auto-bidders or ad campaigns, for a broad class of auto-bidding algorithms. Additionally, our simulations on synthetic and real-world datasets support the theoretical result that coordination outperforms independent bidding. These findings highlight both the theoretical potential and the practical robustness of coordinated auto-bidding in online auctions.

cs.GT

LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear attention variants often replace the softmax with Gaussian kernels to reduce complexity, but such approximations lack theoretical grounding and tend to oversuppress mid-range token interactions. We propose LaplacianFormer, a Transformer variant that employs a Laplacian kernel as a principled alternative to softmax, motivated by empirical observations and theoretical analysis. To address expressiveness degradation under low-rank approximations, we introduce a provably injective feature map that retains fine-grained token information. For efficient computation, we adopt a Nyström approximation of the kernel matrix and solve the resulting system using Newton--Schulz iteration, avoiding costly matrix inversion and SVD. We further develop custom CUDA implementations for both the kernel and solver, enabling high-throughput forward and backward passes suitable for edge deployment. Experiments on ImageNet show that LaplacianFormer achieves strong performance-efficiency trade-offs while improving attention expressiveness.

cs.CV

MAVEN: A Mesh-Aware Volumetric Encoding Network for Simulating 3D Flexible Deformation

Deep learning-based approaches, particularly graph neural networks (GNNs), have gained prominence in simulating flexible deformations and contacts of solids, due to their ability to handle unstructured physical fields and nonlinear regression on graph structures. However, existing GNNs commonly represent meshes with graphs built solely from vertices and edges. These approaches tend to overlook higher-dimensional spatial features, e.g., 2D facets and 3D cells, from the original geometry. As a result, it is challenging to accurately capture boundary representations and volumetric characteristics, though this information is critically important for modeling contact interactions and internal physical quantity propagation, particularly under sparse mesh discretization. In this paper, we introduce MAVEN, a mesh-aware volumetric encoding network for simulating 3D flexible deformation, which explicitly models geometric mesh elements of higher dimension to achieve a more accurate and natural physical simulation. MAVEN establishes learnable mappings among 3D cells, 2D facets, and vertices, enabling flexible mutual transformations. Explicit geometric features are incorporated into the model to alleviate the burden of implicitly learning geometric patterns. Experimental results show that MAVEN consistently achieves state-of-the-art performance across established datasets and a novel metal stretch-bending task featuring large deformations and prolonged contacts.

cs.LG

One Model, Two Markets: Bid-Aware Generative Recommendation

Generative Recommender Systems using semantic ids, such as TIGER (Rajput et al., 2023), have emerged as a widely adopted competitive paradigm in sequential recommendation. However, existing architectures are designed solely for semantic retrieval and do not address concerns such as monetization via ad revenue and incorporation of bids for commercial retrieval. We propose GEM-Rec, a unified framework that integrates commercial relevance and monetization objectives directly into the generative sequence. We introduce control tokens to decouple the decision of whether to show an ad from which item to show. This allows the model to learn valid placement patterns directly from interaction logs, which inherently reflect past successful ad placements. Complementing this, we devise a Bid-Aware Decoding mechanism that handles real-time pricing, injecting bids directly into the inference process to steer the generation toward high-value items. We prove that this approach guarantees allocation monotonicity, ensuring that higher bids weakly increase an ad's likelihood of being shown without requiring model retraining. Experiments demonstrate that GEM-Rec allows platforms to dynamically optimize for semantic relevance and platform revenue.

cs.IR

HMR-1: Hierarchical Massage Robot with Vision-Language-Model for Embodied Healthcare

The rapid advancement of Embodied Intelligence has opened transformative opportunities in healthcare, particularly in physical therapy and rehabilitation. However, critical challenges remain in developing robust embodied healthcare solutions, such as the lack of standardized evaluation benchmarks and the scarcity of open-source multimodal acupoint massage datasets. To address these gaps, we construct MedMassage-12K - a multimodal dataset containing 12,190 images with 174,177 QA pairs, covering diverse lighting conditions and backgrounds. Furthermore, we propose a hierarchical embodied massage framework, which includes a high-level acupoint grounding module and a low-level control module. The high-level acupoint grounding module uses multimodal large language models to understand human language and identify acupoint locations, while the low-level control module provides the planned trajectory. Based on this, we evaluate existing MLLMs and establish a benchmark for embodied massage tasks. Additionally, we fine-tune the Qwen-VL model, demonstrating the framework's effectiveness. Physical experiments further confirm the practical applicability of the framework.Our dataset and code are publicly available at https://github.com/Xiaofeng-Han-Res/HMR-1.

cs.RO

ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers

Tool calling has become increasingly popular for Large Language Models (LLMs). However, for large tool sets, the resulting tokens would exceed the LLM's context window limit, making it impossible to include every tool. Hence, an external retriever is used to provide LLMs with the most relevant tools for a query. Existing retrieval models rank tools based on the similarity between a user query and a tool description (TD). This leads to suboptimal retrieval as user requests are often poorly aligned with the language of TD. To remedy the issue, we propose ToolDreamer, a framework to condition retriever models to fetch tools based on hypothetical (synthetic) TD generated using an LLM, i.e., description of tools that the LLM feels will be potentially useful for the query. The framework enables a more natural alignment between queries and tools within the language space of TD's. We apply ToolDreamer on the ToolRet dataset and show that our method improves the performance of sparse and dense retrievers with and without training, thus showcasing its flexibility. Through our proposed framework, our aim is to offload a portion of the reasoning burden to the retriever so that the LLM may effectively handle a large collection of tools without inundating its context window.

cs.CL

Neural Latent Arbitrary Lagrangian-Eulerian Grids for Fluid-Solid Interaction

Fluid-solid interaction (FSI) problems are fundamental in many scientific and engineering applications, yet effectively capturing the highly nonlinear two-way interactions remains a significant challenge. Most existing deep learning methods are limited to simplified one-way FSI scenarios, often assuming rigid and static solid to reduce complexity. Even in two-way setups, prevailing approaches struggle to capture dynamic, heterogeneous interactions due to the lack of cross-domain awareness. In this paper, we introduce \textbf{Fisale}, a data-driven framework for handling complex two-way \textbf{FSI} problems. It is inspired by classical numerical methods, namely the Arbitrary Lagrangian-Eulerian (\textbf{ALE}) method and the partitioned coupling algorithm. Fisale explicitly models the coupling interface as a distinct component and leverages multiscale latent ALE grids to provide unified, geometry-aware embeddings across domains. A partitioned coupling module (PCM) further decomposes the problem into structured substeps, enabling progressive modeling of nonlinear interdependencies. Compared to existing models, Fisale introduces a more flexible framework that iteratively handles complex dynamics of solid, fluid and their coupling interface on a unified representation, and enables scalable learning of complex two-way FSI behaviors. Experimentally, Fisale excels in three reality-related challenging FSI scenarios, covering 2D, 3D and various tasks. The code is available at \href{https://github.com/therontau0054/Fisale}.

cs.LG

Scaling Inference-Time Computation via Opponent Simulation: Enabling Online Strategic Adaptation in Repeated Negotiation

While large language models (LLMs) have emerged as powerful decision-makers across a wide range of single-agent and stationary environments, fewer efforts have been devoted to settings where LLMs must engage in \emph{repeated} and \emph{strategic} interactions with unknown or dynamic opponents. In such settings, recipes built upon \emph{offline} pre-training or fine-tuning, though robust against worst-case adversaries, do not fully exploit the capability of LLMs to adapt \emph{online} based on interaction feedback. Instead, we explore the more natural perspective of scaling inference-time computation as a mechanism for adaptation, embedding the principles of a classical game-theoretical learning dynamic, \emph{smooth Fictitious Play (sFP)}, into LLM inference: (i) for belief formation, we employ an auxiliary opponent model that in-context learns to imitate the time-averaged behavior of the opponent; (ii) for best response, we advance best-of-$N$ (BoN) sampling by simulating against the opponent model. Empirical evaluations on two distinct forms of repeated negotiation games demonstrate that our method enables significant performance improvement over repeated online interaction compared to various baselines, offering a scalable and principled approach to repeated strategic decision-making without any parameter updates.

cs.MA

Incentive-Aligned Multi-Source LLM Summaries

Large language models (LLMs) are increasingly used in modern search and answer systems to synthesize multiple, sometimes conflicting, texts into a single response, yet current pipelines offer weak incentives for sources to be accurate and are vulnerable to adversarial content. We introduce Truthful Text Summarization (TTS), an incentive-aligned framework that improves factual robustness without ground-truth labels. TTS (i) decomposes a draft synthesis into atomic claims, (ii) elicits each source's stance on every claim, (iii) scores sources with an adapted multi-task peer-prediction mechanism that rewards informative agreement, and (iv) filters unreliable sources before re-summarizing. We establish formal guarantees that align a source's incentives with informative honesty, making truthful reporting the utility-maximizing strategy. Experiments show that TTS improves factual accuracy and robustness while preserving fluency, aligning exposure with informative corroboration and disincentivizing manipulation.

cs.CL

FilDeep: Learning Large Deformations of Elastic-Plastic Solids with Multi-Fidelity Data

The scientific computation of large deformations in elastic-plastic solids is crucial in various manufacturing applications. Traditional numerical methods exhibit several inherent limitations, prompting Deep Learning (DL) as a promising alternative. The effectiveness of current DL techniques typically depends on the availability of high-quantity and high-accuracy datasets, which are yet difficult to obtain in large deformation problems. During the dataset construction process, a dilemma stands between data quantity and data accuracy, leading to suboptimal performance in the DL models. To address this challenge, we focus on a representative application of large deformations, the stretch bending problem, and propose FilDeep, a Fidelity-based Deep Learning framework for large Deformation of elastic-plastic solids. Our FilDeep aims to resolve the quantity-accuracy dilemma by simultaneously training with both low-fidelity and high-fidelity data, where the former provides greater quantity but lower accuracy, while the latter offers higher accuracy but in less quantity. In FilDeep, we provide meticulous designs for the practical large deformation problem. Particularly, we propose attention-enabled cross-fidelity modules to effectively capture long-range physical interactions across MF data. To the best of our knowledge, our FilDeep presents the first DL framework for large deformation problems using MF data. Extensive experiments demonstrate that our FilDeep consistently achieves state-of-the-art performance and can be efficiently deployed in manufacturing.

cs.AI

A Pilot Kinematic Study on the Forehand Reverse Flick: Feasibility of a Novel Short Return Technique in Table Tennis

Background Following changes in table tennis ball materials, offensive returns have become more important for initiating sustained topspin offense. However, using the backhand flick (BF) to return forehand short balls often increases the difficulty of recovery and continuity, revealing a technical gap. This study preliminarily verified a novel forehand short return technique, the forehand reverse flick (FRF), and analyzed its similarities and differences with the BF. Methods Four elite athletes completed seven consecutive days of FRF specific training. Infrared motion capture and ultra-high-speed cameras were used to collect data on racket kinematics, movement duration, and ball performance. Results The success rate of the FRF increased steadily, reaching 86%. Racket trajectories of the two techniques were highly similar along the X (r = 1) and Y (r = 0.99) axes but differed along the Z (r = -0.04) axis. Racket and ball velocities were comparable between techniques, whereas the FRF showed lower resultant acceleration (approximately 265.57 m/s) and required about 0.03 s more for movement duration. Ball velocity was comparable between techniques, for the ball spin, the FRF generated lower spin (approximately 76.61 r/s) about 64% of the BF value (approximately 120.13 r/s). The highest participant mean spin rate reached 93 r/s, about 77% of the BF mean. Conclusion Overall, the FRF was found to have favorable learnability and training value, with potential for further optimization and competitive application.

physics.med-ph

A Unified Approach to Submodular Maximization Under Noise

We consider the problem of maximizing a submodular function with access to a noisy value oracle for the function instead of an exact value oracle. Similar to prior work, we assume that the noisy oracle is persistent in that multiple calls to the oracle for a specific set always return the same value. In this model, Hassidim and Singer (2017) design a $(1-1/e)$-approximation algorithm for monotone submodular maximization subject to a cardinality constraint, and Huang et al (2022) design a $(1-1/e)/2$-approximation algorithm for monotone submodular maximization subject to any arbitrary matroid constraint. In this paper, we design a meta-algorithm that allows us to take any "robust" algorithm for exact submodular maximization as a black box and transform it into an algorithm for the noisy setting while retaining the approximation guarantee. By using the meta-algorithm with the measured continuous greedy algorithm, we obtain a $(1-1/e)$-approximation (resp. $1/e$-approximation) for monotone (resp. non-monotone) submodular maximization subject to a matroid constraint under noise. Furthermore, by using the meta-algorithm with the double greedy algorithm, we obtain a $1/2$-approximation for unconstrained (non-monotone) submodular maximization under noise.

cs.DS

NIR-II Fluorescence Project Technology for Augmented Reality Surgical Navigation

NIR-II fluorescence imaging provides superior tissue penetration and clarity, yet its clinical use in surgical navigation is hindered by a critical workflow issue. Surgeons must divert their attention between the operative field and external monitors, increasing cognitive load and disrupting procedures. Current strategies have failed to resolve this fundamental problem. Here, we developed a co-axial NIR-II fluorescence projection navigation system to enable real-time, in situ visualization. This system creates an intraoperative augmented reality by directly projecting high-precision, pseudocolored fluorescence images onto the surgical field, spatially integrating functional signals with patient anatomy. Validated through in vitro, in vivo, and clinical patient studies, our system eliminates visual field switching, reduces intraoperative distraction, and preserves natural stereoscopic vision. This approach represents a paradigm shift toward a more coherent, efficient, and ergonomically optimized optical imaging modality for surgical navigation.

physics.optics

Single mode lasing and spectral narrowing in photonic crystal line-defect cavities via spatially selected Bloch modes

The demand for high-efficiency and miniaturized on-chip light sources drives continuous innovation in photonic crystal (PhC) microcavity lasers. The presence of slow-light effects in PhC microcavities leads to the mode competition between Bloch modes resulting in multi-mode lasing, which obstructs the dense integration of PhC lasers. Here, we theoretically verify a technical scheme for the single-mode lasing of PhC line-defect-cavity lasers by spatially pumping a certain Bloch mode via optical interference.We demonstrate the capability to select a specific longitudinal mode to lase with a side mode suppression ratio (SMSR) exceeding 30 dB. The interaction between optical interference fringes and the vacuum electromagnetic field inside the PhC cavity improves the linewidth and noise characteristics of lasers. This scheme of Bloch mode selection provides a novel and viable tool for the manipulation of PhC microcavity lasers.

physics.optics

Lattice Boltzmann Boundary Conditions for Flow, Convection-Diffusion and MHD Simulations

A general derivation is proposed for several boundary conditions arisen in the lattice Boltzmann simulations of various physical problems. Pair-wise moment conservations are proposed to enforce the boundary conditions with given macroscopic quantities, including the velocity and pressure in flow simulations, concentration in convection-diffusion (CD) simulations, as well as magnetic field components in magnetohydrodynamical (MHD) simulations. Additionally, the CD and MHD simulations might involve the Robin boundary condition for surface reactions and a Robin-like boundary condition for thin walls with finite electrical conductivities, respectively, both of which can be written in a form with a variable flux term. In this case, the proposed boundary scheme takes the flux term as an increment to the bounced distribution function and a reference frame transformation is used to obtain a correction term for moving boundaries. Spatial interpolation and extrapolation are used for arbitrary boundary locations between computational grid points. Due to using the same approach in derivations, the obtained boundary schemes for different physical processes in a coupled simulation are compatible for arbitrary boundary-to-grid distances (not limited to the popular half-grid boundary layout) and arbitrary moving speeds. Simulations using half-grid and full-grid boundary layouts for flat boundaries are conducted for demonstrations and validations. Moving boundaries are simulated in hydrodynamic and MHD flows, while static boundaries are used in the CD simulations with surface reactions. The numerical and analytical solutions are in excellent agreement in the studied cases. The proposed boundary schemes are also applied in simulating fully coupled MHD pipe flows of a curved boundary with various boundary-to-grid distances and excellent agreement with analytical solutions is also obtained.

physics.comp-ph

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion methods and VLMs in the field of robot vision. For semantic scene understanding tasks, we categorize fusion approaches into encoder-decoder frameworks, attention-based architectures, and graph neural networks. Meanwhile, we also analyze the architectural characteristics and practical implementations of these fusion strategies in key tasks such as simultaneous localization and mapping (SLAM), 3D object detection, navigation, and manipulation. We compare the evolutionary paths and applicability of VLMs based on large language models (LLMs) with traditional multimodal fusion methods.Additionally, we conduct an in-depth analysis of commonly used datasets, evaluating their applicability and challenges in real-world robotic scenarios. Building on this analysis, we identify key challenges in current research, including cross-modal alignment, efficient fusion, real-time deployment, and domain adaptation. We propose future directions such as self-supervised learning for robust multimodal representations, structured spatial memory and environment modeling to enhance spatial intelligence, and the integration of adversarial robustness and human feedback mechanisms to enable ethically aligned system deployment. Through a comprehensive review, comparative analysis, and forward-looking discussion, we provide a valuable reference for advancing multimodal perception and interaction in robotic vision. A comprehensive list of studies in this survey is available at https://github.com/Xiaofeng-Han-Res/MF-RV.

cs.RO

The Average-Value Allocation Problem

We initiate the study of centralized algorithms for welfare-maximizing allocation of goods to buyers subject to average-value constraints. We show that this problem is NP-hard to approximate beyond a factor of $\frac{e}{e-1}$, and provide a $\frac{4e}{e-1}$-approximate offline algorithm. For the online setting, we show that no non-trivial approximations are achievable under adversarial arrivals. Under i.i.d. arrivals, we present a polytime online algorithm that provides a constant approximation of the optimal (computationally-unbounded) online algorithm. In contrast, we show that no constant approximation of the ex-post optimum is achievable by an online algorithm.

cs.DS

Maximum a Posteriori Probability (MAP) Joint Carrier Frequency Offset (CFO) and Channel Estimation for MIMO Channels with Spatial and Temporal Correlations

We consider time varying MIMO fading channels with known spatial and temporal correlation and solve the problem of joint carrier frequency offset (CFO) and channel estimation with prior distributions. The maximum a posteriori probability (MAP) joint estimation is proved to be equivalent to a separate MAP estimation of the CFO followed by minimum mean square error (MMSE) estimation of the channel while treating the estimated CFO as true. The MAP solution is useful to take advantage of the estimates from the previous data packet. A low complexity universal CFO estimation algorithm is extended from the time invariant case to the time varying case. Unlike past algorithms, the universal algorithm does not need phase unwrapping to take advantage of the full range of symbol correlation and achieves the derived Bayesian Cramér-Rao lower bound (BCRLB) in almost all SNR range. We provide insight on the the relation among the temporal correlation coefficient of the fading, the CFO estimation performance, and the pilot signal structure. An unexpected observation is that the BCRLB is not a monotone function of the temporal correlation and is strongly influenced by the pilot signal structures. A simple rearrangement of the 0's and 1's in the pilot signal matrix will render the BCRLB from being non-monotone to being monotone in certain temporal correlation ranges. Since the BCRLB is shown to be achieved by the proposed algorithm, it provides a guideline for pilot signal design.

cs.IT