Search arXiv⌕ Search

EXPLORE CONNECTIONS

Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Follow the relationships that help you find your next source.

Based on 525 indexed works selected for this snapshot; the candidate window is limited. Counts describe this index, not the complete source archives. Prepared from the PostgreSQL corpus; source versions are checked before display. Snapshot 2026-09-26. Up to 32 works or names per graph.

Subjects & research connections

Connections use shared source subjects, names, places, and explicitly mentioned entities. A shared label is not evidence of a citation, experimental result, or verified species identification.

cs.CV · cs.AI · cs.ROcs.RO · cs.AI · cs.CVcs.CV · cs.AIcs.CV · cs.AIeess.IV · cs.CVcs.CV · cs.AIcs.RO · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.RO · cs.CVcs.CV · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.CV · cs.AIcs.RO · cs.AIcs.RO · cs.CVcs.CV · cs.AIcs.RO · cs.AIcs.CV · eess.IVcs.CVcs.CVcs.AIcs.CVcs.AIcs.AIcs.AIcs.AIcs.CVSurgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic SurgerySurgical-LVLM: Learning…VLANeXt: Recipes for Building Strong VLA Models2Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use3Temporal and Contextual Transformer for Multi-Camera Editing of TV Shows4Band-Attention Modulation Network for Robust Face Forgery Detection5Bridging the Inter-Domain Gap through Low-Level Features for Cross-Modal Medical Image Segmentation6Cross-Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation7HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving8MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction9PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment10RotVLA: Rotational Latent Action for Vision-Language-Action Model11Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow12A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring13Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images143D Oral Modelling with Improved Vertex Distribution Using Matching-Based Learning15JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence16Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning17SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction18More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning19Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces20Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning21TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation22TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision23Improving Binary Neural Networks through Fully Utilizing Latent Weights24FlowFace++: Explicit Semantic Flow-supervised End-to-End Face Swapping25Joint Prediction and Denoising for Large-scale Multilingual Self-supervised Learning26To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now27ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs28ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks29Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders30Probing many-body Bell correlation depth with superconducting qubits31AIR: Analytic Imbalance Rectifier for Continual Learning32
Relationships as a list (31)

Citation graph

Arrows run from the citing work to its reference. Only explicit source/provider reference lists are used. Incoming links cover this candidate window; this is not a global citation count.

No supported relationships are available in this snapshot. This does not mean that no relationships exist.

Collaborator network

Shared authorship within up to 200 candidate works (1 examined); up to 32 names shown. Names are matched as supplied, without verified person disambiguation. Shared credit does not necessarily establish personal collaboration.

1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared work1 shared workGuankun WangGuankun WangHongbin LiuHongbin LiuHongliang RenHongliang RenJie WangJie WangJinlin WuJinlin WuLong BaiLong BaiMobarakol IslamMobarakol IslamWan Jun NahWan Jun NahZhaoxi ZhangZhaoxi ZhangZhen ChenZhen Chen
Relationships as a list (45)

Publication timeline

Publication years for these 32 related works; 0 have no source publication date. This is a discovery sample, not a measure of research output or growth.

2021 · 1 work
2022 · 1 work
2023 · 2 works
2024 · 3 works
2025 · 1 work
2026 · 24 works

Semantic map

Model: all-minilm. Positions are a two-dimensional approximation of embedding similarity; proximity is not a citation or proof of agreement. Results come from the snapshot’s candidate window.

No current, compatible embeddings are available for related works in this snapshot. A semantic map appears after background embedding and snapshot generation.