Search arXivSearch

arXiv subjects

Vikash Kumar

Publications and source records attributed to Vikash Kumar.

At least 19 recordsLinked to original sources

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

cs.CV

MyoChallenge 2025: A New Benchmark for Human Athletic Intelligence

Athletic performance represents the pinnacle of human motor intelligence, demanding rapid choices, precise control, agility, and coordinated physical execution. Replicating this seamless combination of capabilities remains elusive in current artificial intelligence and robotic systems. Concurrently, understanding the biological mastery of these movements is hindered because complex muscle coordination is rarely measured in vivo due to the limitations of physical equipment. To bridge this fundamental gap in understanding, MyoChallenge at NeurIPS 2025 established a pioneering benchmark for motor control intelligence in sports, leveraging high-fidelity musculoskeletal models within physics simulation combined with machine learning-driven algorithms. The competition introduces two distinct tracks emphasizing either upper or lower limbs control: a table tennis rally task utilizing a biomechanic upper limb composed of an arm with a hand and a trunk; and a soccer penalty kick using a biomechanic model of legs and a trunk. Marking the fourth iteration of the MyoChallenge series, this event attracted almost 70 teams and over 560 submissions globally, uniting a diverse community ranging from physicians and neuroscientists to machine learning experts. The competition facilitated the development of several state-of-the-art control algorithms for a musculoskeletal system capable of sports agility, leveraging techniques such as physics-based motion planners, on-policy behaviour cloning, hierarchical planning, and muscle synergies. By integrating standardized tasks and physiologically realistic models into the open-source framework of MyoSuite, MyoChallenge'25 serves as a reproducible and reusable testbed to accelerate interdisciplinary research across machine learning, biomechanics, sports science, and neuroscience. Project page: https://www.myosuite.org//myochallenge/myochallenge-2025.

cs.RO

Two-point completion of zero interlacing and Wronskian-type bounds, with applications to Jacobi and other orthogonal polynomials

We investigate completed interlacing of zeros for pairs of polynomial sequences that fail to interlace by exactly two points. Using suitable mixed recurrence relations, we identify two additional points, called the completion points, required to obtain completed interlacing and establish general results for polynomials with real zeros. We also show that completed interlacing yields Wronskian-type bounds for the original polynomial pairs. The completed-interlacing results are applied to Jacobi, Meixner-Pollaczek, and Pseudo-Jacobi polynomials. In the Jacobi case, we improve earlier results by determining two explicit additional points that complete the interlacing of $P_n^{(\alpha,\beta)}$ and $P_{n+1}^{(\alpha+1,\beta+1)}$. We also analyze the locations of these points and show that the common-zero cases are non-generic in the Jacobi parameter domain. For this doubly-shifted Jacobi pair, the Wronskian-type bounds yield cross-family Tur\'an-type inequalities and lead to an associated Stieltjes representation and complete-monotonicity consequences. For Meixner-Pollaczek polynomials, we address an open question concerning consecutive degrees with the parameter increased by one, and we obtain analogous completed interlacing results for Pseudo-Jacobi polynomials.

math.CA

Separating zeros of polynomials using an added interlacing point

Following a systematic analysis of existing results, we investigate when complete interlacing between the zeros of distinct polynomial sequences, $\{\mathcal{P}_n\}$ and $\{\mathcal{G}_n\}$ can be achieved by using a naturally arising extra point. Specifically, we analyse several general mixed recurrence relations that ensure the $n+1$ zeros of the polynomial $(x-E)\mathcal{P}_n(x)$ interlace with the $k$ zeros of $\mathcal{G}_k$, where $k=n$ or $n+1$. In addition, we show that imposing specific conditions on the extra point $E$ yields full interlacing between the zeros of $\mathcal{P}_n$ and $\mathcal{G}_k$ for a suitable choice of $n$. The approach provides a consolidated framework broadly applicable to both orthogonal and non-orthogonal polynomials and we illustrate this with new interlacing results for zeros of Krawtchouk, Meixner, and Narayana polynomials. We also illustrate that this general approach can be used to recover and refine existing results regarding the complete interlacing of zeros for classical Jacobi and Laguerre polynomials.

math.CA

A Very Big Video Reasoning Suite

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/?v=vbvr .

cs.CV

Investigation of competing magnetic orders and the associated spin-phonon coupling effect in quasi-2D Cr1+xTe2 (x = 0.22) single crystal

Single crystals of quasi-2D chromium telluride system represented by Cr1+xTe2 (x = 0.22), which crystallizes in the trigonal structure with the c-axis as the growth direction, are synthesized by the flux method. The magnetization measurements revealed the coexistence and competition between ferromagnetic and antiferromagnetic exchange interactions due to the presence of the intercalated Cr-layers. A series of diverse magnetic transitions is exhibited by the crystal. While cooling the crystal, it undergoes a paramagnetic to antiferromagnetic transition at 190 K, followed by a transition into a ferromagnetic state at around 160 K, and a spin-canting at a lower temperature of 75 K. A possible lack of inversion symmetry in the crystal structure, along with the observance of an unusual jump and loop opening in the isothermal magnetization suggests that the crystals Cr1+xTe2 (x = 0.22) may host skyrmions. Furthermore, concomitant to their magnetic transitions, anomalies were observed in the derived Breit-Wigner-Fano fit parameters obtained from the asymmetric temperature-dependent Raman spectra, evidencing a strong spin-phonon coupling effect, intrinsic to the grown crystals. A strong perpendicular magnetic anisotropy along with the robust spin-phonon coupling, makes the system a promising candidate for prospective spintronics applications.

cond-mat.str-el

Hyperparametric solitons in nondegenerate optical parametric oscillators

Dissipative solitons and their associated low-noise, chip-scale frequency combs hold great potential for applications in optical communications, spectroscopy, precision time-keeping, and beyond. These applications drive interest in shifting soliton spectra to frequency bands far detuned from the telecom's C-band pump sources. Recent demonstrations have utilized second-harmonic generation and degenerate optical parametric oscillators (OPOs) to shift soliton combs away from the primary pump. However, these approaches lack the tunability offered by nondegenerate OPOs. This work presents a proof-of-principle demonstration of solitons in a silicon-nitride microresonator-based nondegenerate OPO system with engineered dispersion and optimized coupling rates. By pumping a relatively low-Q resonance in the C-band, we excite a signal soliton comb centred around a far-detuned, high-Q O-band resonance. This process also generates repetition-rate-locked combs at the pump and idler frequencies, with the latter occurring at a wavelength beyond 2$\mu$m. We demonstrate that the solitons supported by this platform are distinct from other families of dissipative solitons and call them - hyperparametric solitons. They emerge when the narrow-band signal mode, phase-matched under negative pump detuning, reaches sufficient power to drive bistability in the parametric signal. We investigate the properties of hyperparametric solitons, including their parametrically generated background and multisoliton states, both experimentally and through theoretical modelling.

physics.optics

Cascaded Diffusion Models for Neural Motion Planning

Robots in the real world need to perceive and move to goals in complex environments without collisions. Avoiding collisions is especially difficult when relying on sensor perception and when goals are among clutter. Diffusion policies and other generative models have shown strong performance in solving local planning problems, but often struggle at avoiding all of the subtle constraint violations that characterize truly challenging global motion planning problems. In this work, we propose an approach for learning global motion planning using diffusion policies, allowing the robot to generate full trajectories through complex scenes and reasoning about multiple obstacles along the path. Our approach uses cascaded hierarchical models which unify global prediction and local refinement together with online plan repair to ensure the trajectories are collision free. Our method outperforms (by ~5%) a wide variety of baselines on challenging tasks in multiple domains including navigation and manipulation.

cs.RO

Towards Efficient Flash Caches with Emerging NVMe Flexible Data Placement SSDs

NVMe Flash-based SSDs are widely deployed in data centers to cache working sets of large-scale web services. As data centers face increasing sustainability demands, such as reduced carbon emissions, efficient management of Flash overprovisioning and endurance has become crucial. Our analysis demonstrates that mixing data with different lifetimes on Flash blocks results in high device garbage collection costs, which either reduce device lifetime or necessitate host overprovisioning. Targeted data placement on Flash to minimize data intermixing and thus device write amplification shows promise for addressing this issue. The NVMe Flexible Data Placement (FDP) proposal is a newly ratified technical proposal aimed at addressing data placement needs while reducing the software engineering costs associated with past storage interfaces, such as ZNS and Open-Channel SSDs. In this study, we explore the feasibility, benefits, and limitations of leveraging NVMe FDP primitives for data placement on Flash media in CacheLib, a popular open-source Flash cache widely deployed and used in Meta's software ecosystem as a caching building block. We demonstrate that targeted data placement in CacheLib using NVMe FDP SSDs helps reduce device write amplification, embodied carbon emissions, and power consumption with almost no overhead to other metrics. Using multiple production traces and their configurations from Meta and Twitter, we show that an ideal device write amplification of ~1 can be achieved with FDP, leading to improved SSD utilization and sustainable Flash cache deployments.

cs.AR

Semantically Controllable Augmentations for Generalizable Robot Learning

Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For robot learning to generalize despite these challenges, it is essential to leverage sources of data or priors beyond the robot's direct experience. In this work, we posit that image-text generative models, which are pre-trained on large corpora of web-scraped data, can serve as such a data source. These generative models encompass a broad range of real-world scenarios beyond a robot's direct experience and can synthesize novel synthetic experiences that expose robotic agents to additional world priors aiding real-world generalization at no extra cost. In particular, our approach leverages pre-trained generative models as an effective tool for data augmentation. We propose a generative augmentation framework for semantically controllable augmentations and rapidly multiplying robot datasets while inducing rich variations that enable real-world generalization. Based on diverse augmentations of robot data, we show how scalable robot manipulation policies can be trained and deployed both in simulation and in unseen real-world environments such as kitchens and table-tops. By demonstrating the effectiveness of image-text generative models in diverse real-world robotic applications, our generative augmentation framework provides a scalable and efficient path for boosting generalization in robot learning at no extra human cost.

cs.RO

Improving Domain Adaptation Through Class Aware Frequency Transformation

In this work, we explore the usage of the Frequency Transformation for reducing the domain shift between the source and target domain (e.g., synthetic image and real image respectively) towards solving the Domain Adaptation task. Most of the Unsupervised Domain Adaptation (UDA) algorithms focus on reducing the global domain shift between labelled source and unlabelled target domains by matching the marginal distributions under a small domain gap assumption. UDA performance degrades for the cases where the domain gap between source and target distribution is large. In order to bring the source and the target domains closer, we propose a novel approach based on traditional image processing technique Class Aware Frequency Transformation (CAFT) that utilizes pseudo label based class consistent low-frequency swapping for improving the overall performance of the existing UDA algorithms. The proposed approach, when compared with the state-of-the-art deep learning based methods, is computationally more efficient and can easily be plugged into any existing UDA algorithm to improve its performance. Additionally, we introduce a novel approach based on absolute difference of top-2 class prediction probabilities (ADT2P) for filtering target pseudo labels into clean and noisy sets. Samples with clean pseudo labels can be used to improve the performance of unsupervised learning algorithms. We name the overall framework as CAFT++. We evaluate the same on the top of different UDA algorithms across many public domain adaptation datasets. Our extensive experiments indicate that CAFT++ is able to achieve significant performance gains across all the popular benchmarks.

cs.CV

Orthogonality of quasi-nature spectral polynomials of Jacobi and Laguerre type

In this work, the explicit expressions of coefficients involved in quasi Christoffel polynomials of order one and quasi-Geronimus polynomials of order one are determined for Jacobi polynomials. These coefficients are responsible for establishing the orthogonality of quasi-spectral polynomials of Jacobi polynomials. Additionally, the orthogonality of quasi-Christoffel Laguerre polynomials of order one is derived. In the process of achieving orthogonality, in both cases, one zero is located on the boundary of the support of the measure. This allows us to derive the chain sequence and minimal parameter sequence at the point lying at the end point of the support of the measure. Furthermore, the interlacing properties among the zeros of quasi-spectral orthogonal Jacobi polynomials and Jacobi polynomials are illustrated. Finally, we define the quasi-Christoffel polynomials of order one on the unit circle and analyze the location of their zeros for specific examples, as well as propose the problem in the general setup.

math.CA

2D Monolayer Molybdenum (IV) Telluride TMD: An Efficient Electrocatalyst for Hydrogen Evolution Reaction

An electrocatalyst is needed to efficiently lower the reaction barriers to produce hydrogen through the H2 evolution reaction (HER). Recently, two-dimensional transition metal dichalcogenides (2D TMDs), such as the pure 2D monolayer MoTe2 TMD, have become attractive materials for HER. Using the first principle-based hybrid DFT-D method, we have computationally designed a pure 2D monolayer MoTe2 TMD and examined its structural and electronic properties and electrocatalytic efficacy towards HER. A non-periodic finite molecular cluster model Mo10Te21 system was employed to explore the feasibility of both the Volmer-Heyrovsky and Volmer-Tafel reaction mechanisms for the HER. The solvent-phase calculations of the HER on the 2D monolayer MoTe2 TMD demonstrate that this material can effectively undergo either Volmer-Heyrovsky or Volmer-Tafel reaction pathways. This conclusion is supported by our determination of low reaction barriers for the H*-migration, Heyrovsky, and Tafel transition states (TSs), which were found to be approximately 9.80, 12.55, and 5.29 kcal.mol-1, respectively. These results highlight the potential utility of MoTe2 TMD as a promising electrocatalyst for HER. The unusual electrocatalytic activity of the pure 2D monolayer MoTe2 TMD is evidenced by its ability to significantly reduce reaction barriers, achieving impressive turnover frequency during the Heyrovsky and Tafel reaction steps, respectively. Additionally, it demonstrates a remarkably low Tafel slope of 29.58 mV.dec-1. Further exploration of its potential applications in electrocatalysis is warranted. The present work provides valuable insights into the atomic modulation of active sites for enhanced electrocatalytic performance towards HER, paving a way for designing advanced non-noble metal free electrocatalysts.

cond-mat.mtrl-sci

Recovering orthogonality from quasi-nature of Spectral transformations

In this contribution, quasi-orthogonality of polynomials generated by Geronimus and Uvarov transformations is analyzed. An attempt is made to discuss the recovery of the source orthogonal polynomial from the quasi-Geronimus and quasi-Uvarov polynomials of order one. Moreover, the discussion on the difference equation satisfied by quasi-Geronimus and quasi-Uvarov polynomials is presented. Furthermore, the orthogonality of quasi-Geronimus and quasi-Uvarov polynomials is achieved through the reduction of the degree of coefficients in the difference equation. During this procedure, alternative representations of the parameters responsible for achieving orthogonality are derived. One of these representations involves the Stieltjes transform of the measure. Finally, the recurrence coefficients ensuring the existence of a measure that makes the quasi-Geronimus Laguerre polynomial of order one an orthogonal polynomial are calculated.

math.CA

Towards Generalizable Zero-Shot Manipulation via Translating Human Interaction Plans

We pursue the goal of developing robots that can interact zero-shot with generic unseen objects via a diverse repertoire of manipulation skills and show how passive human videos can serve as a rich source of data for learning such generalist robots. Unlike typical robot learning approaches which directly learn how a robot should act from interaction data, we adopt a factorized approach that can leverage large-scale human videos to learn how a human would accomplish a desired task (a human plan), followed by translating this plan to the robots embodiment. Specifically, we learn a human plan predictor that, given a current image of a scene and a goal image, predicts the future hand and object configurations. We combine this with a translation module that learns a plan-conditioned robot manipulation policy, and allows following humans plans for generic manipulation tasks in a zero-shot manner with no deployment-time training. Importantly, while the plan predictor can leverage large-scale human videos for learning, the translation module only requires a small amount of in-domain data, and can generalize to tasks not seen during training. We show that our learned system can perform over 16 manipulation skills that generalize to 40 objects, encompassing 100 real-world tasks for table-top manipulation and diverse in-the-wild manipulation. https://homangab.github.io/hopman/

cs.RO

RoboHive: A Unified Framework for Robot Learning

We present RoboHive, a comprehensive software platform and ecosystem for research in the field of Robot Learning and Embodied Artificial Intelligence. Our platform encompasses a diverse range of pre-existing and novel environments, including dexterous manipulation with the Shadow Hand, whole-arm manipulation tasks with Franka and Fetch robots, quadruped locomotion, among others. Included environments are organized within and cover multiple domains such as hand manipulation, locomotion, multi-task, multi-agent, muscles, etc. In comparison to prior works, RoboHive offers a streamlined and unified task interface taking dependency on only a minimal set of well-maintained packages, features tasks with high physics fidelity and rich visual diversity, and supports common hardware drivers for real-world deployment. The unified interface of RoboHive offers a convenient and accessible abstraction for algorithmic research in imitation, reinforcement, multi-task, and hierarchical learning. Furthermore, RoboHive includes expert demonstrations and baseline results for most environments, providing a standard for benchmarking and comparisons. Details: https://sites.google.com/view/robohive

cs.RO

MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation

Robotic systems that aspire to operate in uninstrumented real-world environments must perceive the world directly via onboard sensing. Vision-based learning systems aim to eliminate the need for environment instrumentation by building an implicit understanding of the world based on raw pixels, but navigating the contact-rich high-dimensional search space from solely sparse visual reward signals significantly exacerbates the challenge of exploration. The applicability of such systems is thus typically restricted to simulated or heavily engineered environments since agent exploration in the real-world without the guidance of explicit state estimation and dense rewards can lead to unsafe behavior and safety faults that are catastrophic. In this study, we isolate the root causes behind these limitations to develop a system, called MoDem-V2, capable of learning contact-rich manipulation directly in the uninstrumented real world. Building on the latest algorithmic advancements in model-based reinforcement learning (MBRL), demo-bootstrapping, and effective exploration, MoDem-V2 can acquire contact-rich dexterous manipulation skills directly in the real world. We identify key ingredients for leveraging demonstrations in model learning while respecting real-world safety considerations -- exploration centering, agency handover, and actor-critic ensembles. We empirically demonstrate the contribution of these ingredients in four complex visuo-motor manipulation problems in both simulation and the real world. To the best of our knowledge, our work presents the first successful system for demonstration-augmented visual MBRL trained directly in the real world. Visit https://sites.google.com/view/modem-v2 for videos and more details.

cs.RO

Natural and Robust Walking using Reinforcement Learning without Demonstrations in High-Dimensional Musculoskeletal Models

Humans excel at robust bipedal walking in complex natural environments. In each step, they adequately tune the interaction of biomechanical muscle dynamics and neuronal signals to be robust against uncertainties in ground conditions. However, it is still not fully understood how the nervous system resolves the musculoskeletal redundancy to solve the multi-objective control problem considering stability, robustness, and energy efficiency. In computer simulations, energy minimization has been shown to be a successful optimization target, reproducing natural walking with trajectory optimization or reflex-based control methods. However, these methods focus on particular motions at a time and the resulting controllers are limited when compensating for perturbations. In robotics, reinforcement learning~(RL) methods recently achieved highly stable (and efficient) locomotion on quadruped systems, but the generation of human-like walking with bipedal biomechanical models has required extensive use of expert data sets. This strong reliance on demonstrations often results in brittle policies and limits the application to new behaviors, especially considering the potential variety of movements for high-dimensional musculoskeletal models in 3D. Achieving natural locomotion with RL without sacrificing its incredible robustness might pave the way for a novel approach to studying human walking in complex natural environments. Videos: https://sites.google.com/view/naturalwalkingrl

cs.RO