Search arXiv⌕ Search

arXiv subjects

Haoran Pang

Publications and source records attributed to Haoran Pang.

2 recordsLinked to original sources

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This independence assumption is often invalidated by the non-orthogonality of model weights and strong correlations between unit activations, potentially leading to performance degradation. To address this, we propose a Correlation-Aware Structured Pruning method. We formulate the pruning objective as a cardinality-constrained binary quadratic program that explicitly models cross-unit dependencies in the reconstruction error. Since this binary quadratic program is NP-hard and difficult to solve exactly, we develop a greedy interaction algorithm based on dependency-aware marginal costs to optimize unit selection. Furthermore, we incorporate a gradient-based strategy to achieve adaptive layer-wise sparsity allocation across the entire model. Extensive experiments on mainstream LLMs demonstrate that incorporating correlation information yields competitive accuracy-efficiency trade-offs compared to representative structured pruning baselines.

cs.CL↗

Antenna Placement Design for Interference Exploitation in Pinching-Antenna Systems

Pinching-antenna systems (PASs) have been proposed as a flexible antenna technology to fulfill the stringent requirements of high data rate and large-scale equipment deployment in future wireless networks. The principle of PA involves mapping a signal over dielectric waveguides for transmission. By adjusting the positions of pinching antennas (PAs) over the waveguides, with the aim of gain enhancement for line-of-sight links and the reduction of large-scale path loss. Symbol-level precoding (SLP) is a nonlinear precoding technique, which converts multi-user interference into constructive interference via beamforming design at symbol level. In this paper, we study the combination of SLP and PAS, leveraging the advantages of PAS to further enhance the ability of SLP to convert constructive interference. The transmit power minimization problem is formulated and solved for the multiple waveguides multiple PAs system by jointly beamforming and PAs' positions design under the SLP principle. The alternating optimization (AO) framework is applied to decouple the beamforming vector and the position coefficient of PA. For the given beamforming vectors, a new objective function is formulated with respect to the positions of the PAs. With the characteristics of the formulated objective function, the optimization problem for the position coefficients of PAs can be decomposed into multiple independent subproblems, each corresponding to a PA's position coefficient, and a projected gradient descent (PGD)-based method, constrained by the feasible movable region of each PA, is then developed to obtain the suboptimal position coefficients. The performance improvements achieved by the combination of PAS and SLP, as well as the effectiveness of the proposed algorithm are verified through the simulation results.

eess.SP↗