Search arXiv⌕ Search

arXiv subjects

Dmitrii Maliukov

Publications and source records attributed to Dmitrii Maliukov.

5 recordsLinked to original sources

Kintsugi-VLA: Turning Failed Robot Rollouts into Recovery Data through Interventional Recoverability

Simulation enables scalable training of Vision-Language-Action policies by using privileged experts to generate visual demonstrations without requiring every trajectory to be collected through manual teleoperation. However, such pipelines typically retain successful demonstrations while failed rollouts are discarded, even though they expose precisely the off-nominal states from which recovery must be learned. We introduce Kintsugi-VLA, a framework for converting failed rollouts into targeted synthetic recovery data by exploiting exact state restoration and branching in simulation. For a fixed privileged expert, we define interventional recoverability as the probability of completing the original task after the simulator is restored to a given state, estimate it using adaptive Monte Carlo continuations with pointwise Wilson intervals, and characterize its non-monotonic evolution along failed trajectories. These estimates identify an observed terminal low-recoverability frontier-the point after which measured recoverability remains below a threshold-which is then used to select informative recovery starting states. In a simulated Franka manipulation task, targeted recovery data yield aggregate SmolVLA recovery success of 34.6\% and 38.4\% under difficulty- and frame-budget matching, respectively, 5.8 and 6.7 percentage points above uniform sampling within the same recovery window. The same ordering is observed under disturbed end-to-end execution and shifted clutter and physics conditions, while clean-task success decreases from 76.8\% to 74.7\%. Kintsugi-VLA demonstrates how failed simulator rollouts can be transformed from discarded experience into structured recovery-training data through direct interventional measurement.

cs.RO↗

CognitiveReality: Robot-Agnostic Semantic Gaussian Mapping with an LLM Agent for Immersive Collaborative VR Teleoperation

A photorealistic 3D view tells a teleoperator where a robot is, but not what the scene contains, how well each object has been observed, or how to turn pointing and speech into robot action. CognitiveReality turns a robot's RGB-D stream into a live, semantically indexed Gaussian-TSDF map shared by an operator in virtual reality and a tool-using language agent. One mapper binary serves any platform through configuration alone: it ingests poses from robot SLAM, joint kinematics, motion capture or an inline visual tracker, bridges localization outages with a shadow tracker and keyframe-anchored PnP, and maintains open-vocabulary instance identities with per-object quality at 2 Hz. Speech and controller rays are grounded against persistent scene objects through validated typed tools and operator-confirmed robot actions. In the controlled agent evaluation, the deployed local Qwen3-VL-8B router reaches 81.24\% tool exact match, while merge-aware replay correctly redirects 101 absorbed object identifiers. On robot data CognitiveReality exceeds a Gaussian-plus-SDF baseline by 2-8 dB; pose error through 5-40 s SLAM outages stays within 1-8 cm. Deployed live on two quadrupeds, the agent executed 26 of 30 navigation requests and 20 of 20 re-observation requests, raising object quality by 2-5 dB.

cs.RO↗

GaussAnything: Semantic Intent-Driven Refinement of Evolving Gaussian Scenes for Standalone VR

Deploying reconstructed 3D environments on standalone VR headsets are constrained by limited compute and memory, and conventional level-of-detail policies optimize for visibility without accounting for the user's explicit inspection intent. We present GaussAnything, a native OpenXR system for intent-conditioned reallocation and progressive publication of evolving semantic Gaussian+SDF scenes. GaussAnything resolves class- or instance-level queries to persistent 3D objects and reallocates a fixed Gaussian resident budget toward the selected object while retaining global context, applying incremental, stable-identity updates coordinated with the TSDF-derived mesh through a source-epoch mechanism. Across eight scenes, an object query concentrates 88-90% of the fixed client budget onto the queried object without enlarging it, on-device rendering reproduces the host render to within a small margin (up to 36.7 dB), and the standalone client renders each stereo frame at a steady-state GPU cost of roughly 10 ms within the frame budget of standard standalone panels.

cs.GR↗

AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration

We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, our system combines ORB-SLAM3-based pose estimation with online Gaussian reconstruction for noisy real-world data. A VR pipeline enables immersive exploration of incremental reconstructions; a semantic module transcribes voice commands, generates scene descriptions, and records points of interest. Against state-of-the-art online Gaussian splatting methods, we improve image quality on our dataset (+14.5% PSNR, +8.6% SSIM, -14.3% LPIPS) and TUM-RGBD (+11.7% PSNR, +7.8% SSIM, -21.6% LPIPS), with comparable or superior frame rates via quality-speed configurations. We achieve an 88% VLM object-recognition rate.

cs.CV↗

VersualRL: Closed-Loop Verbal Reinforcement Learning with Visual Execution Feedback for Task-Level Robot Planning

We introduce VersualRL, a closed-loop framework for task-level robot planning that uses visual execution feedback to iteratively refine executable Behavior Trees through structured natural-language critique. VersualRL combines a vision-language model critic, which analyzes visual observations and Behavior Tree execution traces, with a large language model actor that performs discrete and interpretable policy updates. During physical deployment, both foundation models remain frozen, while the executed Behavior Tree is updated at the symbolic level without online gradient-based policy optimization. This enables transparent reasoning, explicit causal feedback, and human-interpretable policy evolution. We validate VersualRL on a real mobile robot performing a multi-stage navigation and manipulation task under execution uncertainty. Experimental results show that the framework supports explainable policy improvements, closed-loop adaptation to execution failures, and successful deployment on physical robotic systems.

cs.RO↗