Search arXiv⌕ Search

arXiv subjects

Pengfei Li

Publications and source records attributed to Pengfei Li.

At least 19 recordsLinked to original sources

World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild

We introduce World In Your Hands (WIYH), a large-scale open-source ecosystem comprising over 1,000 hours of human manipulation data collected in-the-wild with millimeter-scale motion accuracy. Specifically, WIYH includes (1) the Oracle Suite, a wearable data collection kit with an auto-labeling pipeline for accurate motion capture; (2) the WIYH Dataset, featuring over 1,000 hours of multimodal manipulation data across hundreds of skills in diverse real-world scenarios; and (3) extensive annotations and benchmarks supporting tasks from perception to action. Furthermore, experiments based on the WIYH ecosystem show that integrating WIYH's human-centric data improves robotic manipulation success rates from 8% to 60% in cluttered scenes. World In Your Hands provides a foundation for advancing human-centric data collection and cross-embodiment policy learning. All data and hardware design will be open-source.

cs.RO↗

Charge-4e Superconducting Ground State without Pair Condensation: Exact Quartet Dynamics, Rigorous Order, and a Microscopic Route

A direct charge-\(4e\) superconductor exhibits coherent four-electron order while every charge-\(2e\) pairing channel remains uncondensed. We establish three complementary results. First, building on the \(η\)-clustering states and bipartite parent of Yoshida and Katsura, we formulate and exactly solve a minimal two-term parent on any connected graph. Its fixed-number ground states have quartet off-diagonal long-range order (ODLRO) without charge-\(2e\) ODLRO, while an exact mapping to classical hard-core exclusion dynamics yields the full fixed-sector gap and a branch of quartet-density modes. Nonzero quartet stiffness and vanishing inverse quartet compressibility identify this \(z=2\) parent as a phase-separation boundary. Second, for a finite-range fermionic family with explicit quartet transfer and sufficiently large onsite penalty, we rigorously prove quartet ODLRO without charge-\(2e\) ODLRO at half quartet filling, both at the hypercubic XY point for \(d\geq2\) and throughout a finite XXZ interval on the square lattice. Third, we derive a strong-coupling realization using only electron hopping and two-body interactions. With local gap \(U_0\), pair hopping \(K\) generates quartet motion at order \(K^2/U_0\), whereas electron hopping \(t\) first contributes at order \(t^4/U_0^3\); charge-\(2e\) excitations remain gapped at \(O(U_0)\). On bipartite lattices, positive inverse quartet compressibility opens an asymptotically controlled homogeneous \(z=1\) regime with short-ranged pair correlations, while negative curvature drives phase separation.

cond-mat.str-el↗

Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids

Realizing interactive whole-body control for multi-humanoid systems is critical for unlocking complex collaborative capabilities in shared environments. Although recent advancements have significantly enhanced the agility of individual robots, bridging the gap to physically coupled multi-humanoid interaction remains challenging, primarily due to severe kinematic mismatches and complex contact dynamics. To address this, we introduce Rhythm, the first unified framework enabling real-world deployment of dual-humanoid systems for complex, physically plausible interactions. Our framework integrates three core components: (1) an Interaction-Aware Motion Retargeting (IAMR) module that generates feasible humanoid interaction references from human data; (2) an Interaction-Guided Reinforcement Learning (IGRL) policy that masters coupled dynamics via graph-based rewards; and (3) a real-world deployment system that enables robust transfer of dual-humanoid interaction. Extensive experiments on physical Unitree G1 robots demonstrate that our framework achieves robust interactive whole-body control, successfully transferring diverse behaviors such as hugging and dancing from simulation to reality.

cs.RO↗

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

cs.RO↗

Semiparametric Receiver Operating Characteristic Analysis in the Presence of an Imperfect Reference Standard via a Box-Cox Density Ratio Model

Receiver operating characteristic (ROC) analysis is commonly used to evaluate the diagnostic accuracy of continuous biomarkers. In practice, the true disease status may be unavailable and only a nominal disease status provided by an imperfect reference standard is observed. Existing nonparametric methods have been developed for ROC analysis in this setting, but may suffer from reduced estimation efficiency, numerical instability, or sensitivity to the choice of biomarker scale. We propose a semiparametric method based on a Box-Cox density ratio model, which links the biomarker distributions of the truly healthy and diseased populations while leaving the baseline distribution unspecified. A key feature of the proposed method is that the transformation parameter is estimated from the data rather than specified in advance, allowing the density-ratio structure to adapt to different transformation scales. We develop an empirical likelihood approach for estimation and an expectation-maximization algorithm for computation. We establish the asymptotic distributions of estimators of the ROC curve, area under the curve, Youden's index, and the sensitivity and specificity at the Youden-optimal cutoff, and develop bootstrap confidence intervals and a goodness-of-fit test. Simulation studies demonstrate that the proposed method provides accurate and numerically stable estimation and inference across a range of distributional settings without requiring the transformation scale to be specified in advance. The proposed method is illustrated using data from a malaria study.

stat.ME↗

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited. All datasets, models, and code are available at https://risys-lab.github.io/SAFIRE/.

cs.CV↗

Revisiting dependence in multiple testing: empirical distribution approaches for FDP control

Large-scale multiple hypothesis testing is central to the analysis of high-throughput data, where controlling false discoveries is critical. Classical procedures typically rely on theoretical null distributions and often adjust for dependence among test statistics, but these approaches may be misleading when the empirical distribution of null statistics deviates from theoretical assumptions. Motivated by this observation, we investigate the role of the empirical cumulative distribution function (c.d.f.) of null test statistics in controlling the false discovery proportion (FDP). We first show that, under an oracle scenario where the empirical c.d.f. of the test statistics for all null hypotheses is known, FDP control can be achieved optimally regardless of the dependence structure, highlighting that explicit modeling of dependence may be unnecessary. Building on this insight, we propose an empirical c.d.f.-based FDP control (eFDP) method, implemented via a multivariate mixture model framework and a nonparametric estimation procedure for the empirical c.d.f.s, establish its asymptotic convergence, and construct an FDP control procedure that achieves asymptotic FDP control. Extensive simulations demonstrate that eFDP attains more accurate FDP control and higher power than existing approaches, particularly under strong dependence, and analysis of a high-dimensional breast cancer gene expression dataset confirms its practical utility.

stat.ME↗

Influence of twist direction and large deformation on soft material torsional contact

Shear-induced contact area reduction is widely observed in soft contacts, yet recent torsional experiments have revealed a more complex non-monotonic evolution in which the contact area first increases and then decreases with twist angle. The mechanism responsible for this initial area increase and the role of large deformation in the overall area evolution remain unclear. In this study, we experimentally investigate the torsional contact response of soft Polydimethylsiloxane (PDMS) spheres by combining forward-backward twist tests with a systematic variation of the curing-agent-to-base ratio to tune material softness and deformation level. The loading-unloading tests show that the torsional interface is strongly irreversible: during unloading, the contact area follows a decrease-increase-decrease path rather than retracing the loading branch, and repeatable petal-like edges appear, indicating a wrinkle-induced surface instability. By decreasing the mixing ratio, we find that larger deformation strengthens the area-reduction contribution and eventually suppresses the initial area increase, leading to a monotonic area decrease during loading for sufficiently soft PDMS. Softer PDMS also exhibits lower shear strength, weaker torque oscillations, and improved repeatability. The results provide experimental evidence that large deformation can drive shear-induced contact area reduction, while the origin of the initial area increase remains unresolved. These findings narrow the possible mechanisms (e.g., triboelectrification) responsible for the initial area increase and provide a stringent benchmark for frictional contact models of soft interfaces.

cond-mat.soft↗

Bound states in doped charge transfer insulators

Understanding the physics of doped charge transfer insulators is the most important problem in high-temperature superconductivity. In this work, we show that an in-gap bound state emerges from the localized hole of the doped charge transfer insulator. We propose an approximate ground state wavefunction based on one localized Zhang-Rice singlet and the Neel state. By calculating the excitation states with one hole added and removed from this ground state, we successfully identify the existence of bound states inside the charge transfer gap. This feature is further confirmed by a Lanczos calculation based on matrix product states (MPS) for a system of $4\times4$ CuO$_2$ unit cells. How these bound states evolve into metallic states is further discussed. Our findings identify the key component of recent STM results on lightly doped Ca$_2$CuO$_2$Cl$_2$ and provide a new understanding of hole-doped charge transfer insulators.

cond-mat.supr-con↗

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing TIMT methods, whether cascaded pipelines or end-to-end multimodal large language models (MLLMs),struggle with high-resolution text-rich images due to cluttered layouts, diverse fonts, and non-textual distractions, resulting in text omission, semantic drift, and contextual inconsistency. To address these challenges, we propose GLoTran, a global-local dual visual perception framework for MLLM-based TIMT. GLoTran integrates a low-resolution global image with multi-scale region-level text image slices under an instruction-guided alignment strategy, conditioning MLLMs to maintain scene-level contextual consistency while faithfully capturing fine-grained textual details. Moreover, to realize this dual-perception paradigm, we construct GLoD, a large-scale text-rich TIMT dataset comprising 510K high-resolution global-local image-text pairs covering diverse real-world scenarios. Extensive experiments demonstrate that GLoTran substantially improves translation completeness and accuracy over state-of-the-art MLLMs, offering a new paradigm for fine-grained TIMT under high-resolution and text-rich conditions.

cs.CV↗

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and activities in both humans and Artificial Intelligence (AI). This capability, required for planning and executing daily activities in a task-oriented manner, relies on commonsense knowledge of object physics and functionalities, extending beyond simple object recognition. Current computational models for affordance reasoning from perception lack generalizability, limiting their applicability in novel scenarios. Meanwhile, comprehensive Large Language Models (LLMs) with emerging reasoning capabilities are challenging to deploy on local devices for task-oriented manipulations. Here, we introduce LVIS-Aff, a large-scale dataset comprising 1,496 tasks and 119k images, designed to enhance the generalizability of affordance reasoning from perception. Utilizing this dataset, we develop Afford-X, an end-to-end trainable affordance reasoning model that incorporates Verb Attention and Bi-Fusion modules to improve multi-modal understanding. This model achieves up to a 12.1% performance improvement over the best-reported results from non-LLM methods, while also demonstrating a 1.2% enhancement compared to our previous conference paper. Additionally, it maintains a compact 187M parameter size and infers nearly 50 times faster than the GPT-4V API. Our work demonstrates the potential for efficient, generalizable affordance reasoning models that can be deployed on local devices for task-oriented manipulations. We showcase Afford-X's effectiveness in enabling task-oriented manipulations for robots across various tasks and environments, underscoring its efficiency and broad implications for advancing robotics and AI systems in real-world applications.

cs.CV↗

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

cs.AI↗

Quiescent and traveling solitons in the fractional parametrically driven damped nonlinear Schrödinger equation

We systematically investigate the existence, stability, and dynamics of optical solitons in the framework of the one-dimensional nonlinear Schrödinger equation with the Riesz-fractional diffraction operator, cubic self-focusing, and linear loss, balanced by a linear parametric drive. The model, which can be realized in a laser cavity, produces standing and moving solitons, the latter ones existing below a critical velocity. One of the soliton species is stable in a wide range of parameters, while others are unstable. The fractional diffraction significantly alters the existence conditions and stability thresholds of the solitons. Collision between moving solitons are considered too. The results essentially expand the variety of nonlinear modes in media with fractional diffraction.

nlin.PS↗

Legible Shared Autonomy: Implicit Communication of Robot Belief through Motion

Shared autonomy systems combine user input with autonomous assistance to help users with motor impairments control robot arms to perform everyday manipulation tasks, by inferring user goals and providing appropriate guidance. However, the robot's internal beliefs about user goals cannot be observed by users. Traditional shared autonomy systems provide assistance along efficient shortest paths toward inferred goals, but when multiple objects lie in similar directions, such assistive motion remains ambiguous and fails to reveal the specific goal identified by the robot. This creates two critical problems. First, when the robot correctly infers the goal, users continue controlling because they cannot perceive understanding from ambiguous assistive motion, wasting effort when autonomous completion would suffice. Second, when the robot misunderstands intent, users cannot quickly detect errors until assistive motion diverges significantly, requiring substantial corrective input. We address this by introducing legible motion into shared autonomy, where robot actions must both advance toward the goal and clearly reveal which goal has been inferred, enabling users to understand the robot's beliefs and adjust control accordingly. The robot modulates communication strength through confidence-aware adaptive authority allocation by providing assertive legible assistive actions when confident while increasing user authority when uncertain, transforming shared autonomy into transparent bidirectional collaboration. User studies including simulation, and physical experiments with a six-degree-of-freedom robot arm demonstrate that legible shared autonomy significantly improves users' understanding of robot beliefs and reduces user control effort compared to standard shared autonomy.

cs.RO↗

Analytical and fitting formulae for solutions to Lyman-alpha radiative transfer equations: the effects of geometry, recoil, and velocity gradients

Lyman-alpha (Ly$α$) radiative transfer (RT) is important in many astrophysical environments and governed by multiple physical processes. In this paper, we provide analytical formulae/procedures for the solutions to Ly$α$ RT equations under three simple geometrical symmetries and investigate the effects of atomic recoil and gas bulk motion. We first study Ly$α$ spectra by solving Ly$α$ RT equations for a static, uniform gas cloud under cylindrical geometry. The solution is verified through Ly$α$ Monte Carlo RT simulations, and compared to those under slab and spherical geometries in literature. Second, to characterise the recoil effect, we empirically modify recoil-free Ly$α$ spectra. The method is motivated by Ly$α$ RT equations with recoil and justified by simulations. Finally, we account for constant velocity gradients in Ly$α$ RT equations and obtain series solutions for Ly$α$ spectra. The solutions demonstrate good agreement to Ly$α$ spectra from simulations for small velocity gradients (i.e. edge velocity $v_{\rm E}$ of a cloud being comparable to the thermal velocity $b$) but become less accurate for large ones. To characterise Ly$α$ spectra under large velocity gradients, we empirically extend the functional form of solutions and constrain them from fitting simulated Ly$α$ spectra. The resulting fitting formulae show significant improvement for large velocity gradients ($v_{\rm E}/b \sim 100$) under large optical depths. The analytical study of Ly$α$ spectra in this work completes the set of solutions under simple geometries, provides physical insights for Ly$α$ RT under recoil and velocity gradient, and develops analytical tools for theoretical studies that require inputs from Ly$α$ RT.

astro-ph.GA↗

FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric Mixture

Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level intelligence: at inference time, the server routes each external query to the best-matched client for prediction. Existing approaches, however, typically treat each client as internally homogeneous, overlooking latent subpopulations within local data. For example, patients with the same diagnosis at one hospital may exhibit morphologically distinct disease subtypes. The coexistence of inter-client and intra-client heterogeneity, which we call dual heterogeneity, can impair both routing and prediction. To address this challenge, we propose FedSPM, a routing-enabled semiparametric mixture framework that represents each client using client-specific latent components. Each component combines a predictive distribution for classification with a feature distribution for routing. To flexibly model feature distributions while effectively sharing information across clients, FedSPM models their density ratios relative to a common nonparametric measure estimated via empirical likelihood. We develop a federated expectation-maximization algorithm that optimizes a tractable surrogate and prove convergence of the exact profiled objective at the standard $\mathcal{O}(1/\sqrt{T})$ rate when the surrogate errors are properly controlled. Experiments on controlled benchmarks and real-world medical data demonstrate consistent improvements in routing and prediction under dual heterogeneity. Code is available at https://github.com/zijianwang0510/FedSPM.

cs.LG↗

Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluation

While federated learning enables collaborative modelling on decentralised data, standard methods merely fit historical observations. This purely observational approach is fundamentally insufficient for interventional inference and policy evaluation, as sequential actions dynamically alter future states. We propose \textbf{Fed-CausalDiff}, a federated causal diffusion framework for do-simulation. The architecture decomposes the evolution of the latent state into a global causal score function and a local confounding score function. This design enables \emph{decoupled synchronisation} (DSS), where clients aggregate only the shared causal mechanism while retaining site-specific confounders locally to handle heterogeneity. Experiments on four datasets demonstrate that Fed-CausalDiff achieves better ATE and policy-value estimation accuracy, offering a favorable trade-off between communication cost and inference fidelity.

cs.LG↗

A Water Efficiency Dataset for African Data Centers

Artificial intelligence (AI) computing and data centers consume large amounts of freshwater, both directly for cooling and indirectly for electricity generation. While most attention has been paid to developed countries such as the U.S., this paper presents the first-of-its-kind dataset that combines nation-level weather and electricity generation data to estimate water usage effectiveness for data centers in 41 African countries across five different climate regions. We also use our dataset to evaluate and estimate the water consumption of inference on two large language models (i.e., Llama-3-70B and GPT-4) in 11 selected African countries. Our estimates suggest that writing a 10-page report using Llama-3-70B could consume as much as {0.66 liters} of water, while the water consumption by GPT-4 for the same task may go up to about {59 liters}. For writing a medium-length email of 120-200 words, Llama-3-70B and GPT-4 could consume about {0.13 liters} and {2.9 liters} of water, respectively. All the numbers for generative model inference tasks are based on public information available in 2024, when we initially prepared the analysis. Since then, AI inference systems have improved substantially. For example, recent disclosures suggest that energy efficiency improved by more than 30x between May 2024 and May 2025. Accordingly, our 2024 estimates should be interpreted as historical reference values rather than as representative of current performance. Interestingly, given the same AI model, 9 of the 11 selected African countries consume less water than the global average, mainly because of lower water intensities for electricity generation.

cs.LG↗