Search arXiv⌕ Search

arXiv subjects

Mingxuan Xie

Publications and source records attributed to Mingxuan Xie.

3 recordsLinked to original sources

Seeing and Solving Are Not Enough for Vision-Language Models

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.

cs.CV↗

A Multi-Level Framework for Multi-Objective Hypergraph Partitioning: Combining Minimum Spanning Tree and Proximal Gradient

This paper proposes an efficient hypergraph partitioning framework based on a novel multi-objective non-convex constrained relaxation model. A modified accelerated proximal gradient algorithm is employed to generate diverse $k$-dimensional vertex features to avoid local optima and enhance partition quality. Two MST-based strategies are designed for different data scales: for small-scale data, the Prim algorithm constructs a minimum spanning tree followed by pruning and clustering; for large-scale data, a subset of representative nodes is selected to build a smaller MST, while the remaining nodes are assigned accordingly to reduce complexity. To further improve partitioning results, refinement strategies including greedy migration, swapping, and recursive MST-based clustering are introduced for partitions. Experimental results on public benchmark sets demonstrate that the proposed algorithm achieves reductions in cut size of approximately 2\%--5\% on average compared to KaHyPar in 2, 3, and 4-way partitioning, with improvements of up to 35\% on specific instances. Particularly on weighted vertex sets, our algorithm outperforms state-of-the-art partitioners including KaHyPar, hMetis, Mt-KaHyPar, and K-SpecPart, highlighting its superior partitioning quality and competitiveness. Furthermore, the proposed refinement strategy improves hMetis partitions by up to 16\%. A comprehensive evaluation based on virtual instance methodology and parameter sensitivity analysis validates the algorithm's competitiveness and characterizes its performance trade-offs.

cs.LG↗

Dynamic Graph Representation Learning for Passenger Behavior Prediction

Passenger behavior prediction aims to track passenger travel patterns through historical boarding and alighting data, enabling the analysis of urban station passenger flow and timely risk management. This is crucial for smart city development and public transportation planning. Existing research primarily relies on statistical methods and sequential models to learn from individual historical interactions, which ignores the correlations between passengers and stations. To address these issues, this paper proposes DyGPP, which leverages dynamic graphs to capture the intricate evolution of passenger behavior. First, we formalize passengers and stations as heterogeneous vertices in a dynamic graph, with connections between vertices representing interactions between passengers and stations. Then, we sample the historical interaction sequences for passengers and stations separately. We capture the temporal patterns from individual sequences and correlate the temporal behavior between the two sequences. Finally, we use an MLP-based encoder to learn the temporal patterns in the interactions and generate real-time representations of passengers and stations. Experiments on real-world datasets confirmed that DyGPP outperformed current models in the behavior prediction task, demonstrating the superiority of our model.

cs.LG↗