Search arXivSearch

arXiv subjects

Xi Xiang

Publications and source records attributed to Xi Xiang.

3 recordsLinked to original sources

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

cs.CV

Acoustic Three-dimensional Chern Insulators with Arbitrary Chern Vectors

The Chern vector is a vectorial generalization of the scalar Chern number, being able to characterize the topological phase of three-dimensional (3D) Chern insulators. Such a vectorial generalization extends the applicability of Chern-type bulk-boundary correspondence from one-dimensional (1D) edge states to two-dimensional (2D) surface states, whose unique features, such as forming nontrivial torus knots or links in the surface Brillouin zone, have been demonstrated recently in 3D photonic crystals. However, since it is still unclear how to achieve an arbitrary Chern vector, so far the surface-state torus knots or links can emerge, not on the surface of a single crystal as in other 3D topological phases, but only along an internal domain wall between two crystals with perpendicular Chern vectors. Here, we extend the 3D Chern insulator phase to acoustic crystals for sound waves, and propose a scheme to construct an arbitrary Chern vector that allows the emergence of surface-state torus knots or links on the surface of a single crystal. These results provide a complete picture of bulk-boundary correspondence for Chern vectors, and may find use in novel applications in topological acoustics.

physics.app-ph

The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods

For an autonomous robotic system, monitoring surgeon actions and assisting the main surgeon during a procedure can be very challenging. The challenges come from the peculiar structure of the surgical scene, the greater similarity in appearance of actions performed via tools in a cavity compared to, say, human actions in unconstrained environments, as well as from the motion of the endoscopic camera. This paper presents ESAD, the first large-scale dataset designed to tackle the problem of surgeon action detection in endoscopic minimally invasive surgery. ESAD aims at contributing to increase the effectiveness and reliability of surgical assistant robots by realistically testing their awareness of the actions performed by a surgeon. The dataset provides bounding box annotation for 21 action classes on real endoscopic video frames captured during prostatectomy, and was used as the basis of a recent MIDL 2020 challenge. We also present an analysis of the dataset conducted using the baseline model which was released as part of the challenge, and a description of the top performing models submitted to the challenge together with the results they obtained. This study provides significant insight into what approaches can be effective and can be extended further. We believe that ESAD will serve in the future as a useful benchmark for all researchers active in surgeon action detection and assistive robotics at large.

cs.CV