Search arXiv⌕ Search

arXiv subjects

Yanru Chen

Publications and source records attributed to Yanru Chen.

At least 19 recordsLinked to original sources

Cyclically Compatible Deformations of the Braid Arrangement

We prove a characteristic-polynomial shift formula for two-sided extensions of cyclically compatible deformations of the braid arrangement. For a nonnegative integer matrix $M=(m_{ij})$ with zero diagonal, let $\mathcal{A}_M$ be the arrangement \[ x_i-x_j=s,\qquad 1\le i<j\le n,\quad s\in[-m_{ij},m_{ji}]_{\mathbb{Z}}. \] Given $α,β\in\mathbb{N}^n$, define its two-sided extension $\mathcal{A}_M(α,β)$ by replacing this interval with \[ [-m_{ij}-α_i-β_j,\, m_{ji}+α_j+β_i]_{\mathbb{Z}}. \] Call $M$ cyclically compatible if all pairwise distinct $a,b,c$ with $1\le a,b,c\le n$ satisfy \[ m_{ac}\le m_{ab}+m_{bc}+1. \] Under this condition, for the reduced characteristic polynomial $\widetildeχ(\mathcal{A},t) =\dfrac{χ(\mathcal{A},t)}{t}$, we have \[ \widetildeχ(\mathcal{A}_M(α,β),t) = \widetildeχ(\mathcal{A}_M,t-|α|-|β|). \] The proof uses the finite-field method and a cyclic-gap enumeration formula. We also establish redistribution invariance, study weak-sum perturbations, and give applications to Shi, uniform interval, graphical, and Ferrers-type deformations.

math.CO↗

Modular Ideals and Factorizations of Affine Hyperplane Arrangements

Let $\mathcal{A}$ be an affine hyperplane arrangement, let $L(\mathcal{A})$ be its intersection semilattice, and let $χ_{\mathcal{A}}(t)$ be its characteristic polynomial. We introduce ideal decompositions of finite ranked meet-semilattices and prove that they induce purely combinatorial factorizations of characteristic polynomials. This framework recovers both Stanley's factorization associated with modular elements and Terao's factorization associated with nice partitions. For simple semimatroids, every join-closed ideal decomposition comes from a direct-sum decomposition. The two-factor case defines modular ideals, extending the role of modular elements from central to affine arrangements. Over an infinite field, every modular ideal of an arrangement intersection semilattice is realized by a suitable translated restriction of a subarrangement. We exhibit a nonempty Zariski-open set of valid translations, yielding essential realizations of the given modular ideal. We further prove that modular subarrangements correspond under coning to modular elements on the hyperplane at infinity and hence to M-ideals in the affine setting. Finally, every modular ideal gives an Orlik-Solomon graded vector-space decomposition, which becomes a graded-algebra decomposition when the two factors come from subarrangements.

math.CO↗

Kimi K2.5: Visual Agentic Intelligence

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.

cs.CL↗

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

cs.CL↗

Level Decompositions for Symmetric Deformations of the Braid Arrangement

Let $A\subseteq\mathbb R_{\ge0}$ be finite and nonempty, and let $\mathfrak{A}^A=(\mathcal{A}_1^A,\mathcal{A}_2^A,\ldots)$ be the associated sequence of symmetric deformations of the braid arrangement. Denote by $r_\ell(\mathcal{A}_n^A)$ the number of its level-$\ell$ regions and by $F_\ell(\mathfrak{A}^A,x)$ the corresponding exponential generating function. We prove $F_\ell(\mathfrak{A}^A,x)=\bigl(F_1(\mathfrak{A}^A,x)\bigr)^\ell$. As a consequence, the characteristic polynomial has the binomial-basis expansion $χ(\mathcal{A}_n^A,t)=\sum_{\ell=1}^{n}(-1)^{n-\ell}\,r_\ell(\mathcal{A}_n^A)\binom{t}{\ell}$. When $0\in A$ and $A^*=A\setminus\{0\}$ is nonempty, we refine a classical identity of Stanley level by level: $F_\ell(\mathfrak{A}^{A^*},x)=F_\ell(\mathfrak{A}^{A},1-e^{-x})$. Equivalently, the Catalan-type and semiorder-type level counts satisfy an unsigned Stirling convolution of the first kind. For the $m$-Catalan arrangement $\mathcal{A}_n^{[0,m]}$, we obtain $r_\ell(\mathcal{A}_n^{[0,m]})=n!\,\operatorname{Ran}_{m+1,m\ell}(n-\ell)$, where $\operatorname{Ran}_{p,r}(q)$ is a Raney number. This realizes Raney numbers as refined region counts and answers a question of Deshpande, Menon, and Sarkar. The proofs use labeled Dyck paths, interval orders, and exponential sequences of arrangements. We also realize the inverse Fu--Wang--Zhu bijection for $m$-Catalan regions by tableaux.

math.CO↗

Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage Devices

Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation. Prior OMS accelerators have largely been evaluated in isolation, making it difficult to understand system-level trade-offs across platforms. This paper presents the first workload-driven, cross-platform survey of accelerators for MS search by studying not only commodity platforms, but also emerging memory- and storage-centric architectures, including GPUs, near-storage FPGAs, DRAM near-memory processing, ReRAM/PCM in-memory processing, and 3D NAND/FeNAND in-storage processing, under consistent algorithmic and accuracy assumptions. Leveraging a binary hyperdimensional computing (HDC)-based OMS formulation that reduces similarity evaluation to lightweight bitwise primitives and tolerates device-level non-idealities, we enable a robust execution on memory-centric architectures despite device-level non-idealities and limited computing capability. Overall, this study identifies memory- and storage-centric architectures as a key architectural breakthrough for large-scale, high-speed search acceleration, delivering up to >100x speedup and >40,000x improvement in energy efficiency.

cs.AR↗

Characteristic Polynomials of Graph- and Digraph-Deleted Catalan Arrangements

We develop a finite-field stratification for characteristic polynomials of deletion subarrangements of the full $m$-Catalan arrangement. It reduces the count to cyclic placements of rigid blocks and yields falling-factorial expansions for deletions indexed by graphs, digraphs, and gain-labeled digraphs. The coefficients are graphical Stirling numbers for zero-layer deletions, directed matching numbers when the deleted layer $\ell$ satisfies $1\le \ell\le\lfloor m/2\rfloor$, directed path-cover numbers when $\lfloor m/2\rfloor<\ell\le m$, and admissible gain-labeled arc sets for multilayer deletions. For $\ell=m$, a complementary path-cover expansion yields factorization consequences. The method also gives formulas for directed Ish-type arrangements in terms of path-cycle covers and outdegrees.

math.CO↗

PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference

Large language models are increasingly deployed on edge devices with tight power and area budgets. While mixed-precision GEMM reduces arithmetic complexity, quantized inference is often dominated by dequantization and nonlinear operators. Lookup Table (LUT)-based method mitigates these costs by precomputing outputs and replacing repeated arithmetic with table lookups, but existing designs incur significant capacity and lookup-latency overheads. This paper presents PALUTE, a LUT-based Processing-In-Memory accelerator built on Monolithic 3D DRAM for efficient edge LLM inference. PALUTE enables in-DRAM LUT queries that exploit the vertical organization of M3D DRAM memory array tiles to achieve high parallelism with low area overhead. A near-memory LUT generator supports low-latency LUT generation for both GEMM and element-wise unary nonlinear operators, while a system-level tiering and scheduling strategy minimizes data movement across memory tiers. Evaluation using cycle-accurate simulation and RTL synthesis shows that PALUTE achieves 1,264 TPS end-to-end throughput at 0.16 W, improving energy efficiency by 12.8$\times$ over CHIME and 1.6$\times$ over FIGLUT, improving area efficiency by 2.0$\times$ over PIMPAL under W4A4 across Qwen3-4B models.

cs.AR↗

Convolution-type Identity for Characteristic Polynomials of Geometric Semilattices

We establish a convolution formula for the characteristic polynomial of a finite geometric semilattice $M$: \[ χ(M,st)=\sum_{X\in \underline{M}} s^{r-{\rm rk}_{\underline{M}}(X)}χ(\underline{M}^X,t)\,χ(M_{(X)},s), \] where $\underline{M}$ denotes the centralization of $M$, and $M_{(X)}$ denotes the localization at $X$. This generalizes a nice formula of Southerland, Southern, and Zhou, which is recovered at $s=1$. When specialized to hyperplane arrangements, the identity yields a new expansion closely related to Wang's convolution formula. We further provide a combinatorial interpretation of the convolution formula using the finite field method over $\mathbb{F}_{p^2}$ and $\mathbb{F}_p$.

math.CO↗

GEN-Graph: Heterogeneous PIM Accelerator for General Computational Patterns in Graph-based Dynamic Programming

While graph-based dynamic programming (DP) is a cornerstone of genomics and network analytics, its efficiency is hampered by fundamentally conflicting computational patterns. Matrix-centric DP drives regular, compute-bound network analytics, while topology-centric DP handles irregular, memory-bound genomic traversals. These two categories of DP have substantially different computation patterns and dataflows, which makes it difficult for a single homogeneous processing-in-memory (PIM) architecture to efficiently support both. This work presents GEN-Graph, a novel heterogeneous PIM chiplet that integrates two types of specialized compute tiles within a 2.5D package: Matrix-tile, a processing-using-memory (PUM) tile optimized for matrix-centric workloads, such as all-pairs shortest path (APSP); and traversal-tile, a processing-near-memory (PNM) tile optimized for traversal-centric DP workloads, such as DNA sequence alignment. Our hardware-software co-design employs recursive partitioning and reconfigurable windowed bit-parallel logic to ensure exact computation. Results show the matrix tile achieves 42.8x speedup and 392x energy efficiency over the NVIDIA H100 GPU for APSP. For sequence-to-graph alignment, the traversal tile sustains 2.56 million reads/s (short-reads) and 39.3 thousand reads/s (long-reads), outperforming state-of-the-art accelerators by up to 2.56x in throughput. GEN-Graph provides the first scalable, exact solution for general DP dataflows by matching hardware specialization to algorithmic structure.

cs.AR↗

Attention Residuals

Residual connections with PreNorm are standard in modern LLMs, yet they accumulate all layer outputs with fixed unit weights. This uniform aggregation causes uncontrolled hidden-state growth with depth, progressively diluting each layer's contribution. We propose Attention Residuals (AttnRes), which replaces this fixed accumulation with softmax attention over preceding layer outputs, allowing each layer to selectively aggregate earlier representations with learned, input-dependent weights. To address the memory and communication overhead of attending over all preceding layer outputs for large-scale model training, we introduce Block AttnRes, which partitions layers into blocks and attends over block-level representations, reducing the memory footprint while preserving most of the gains of full AttnRes. Combined with cache-based pipeline communication and a two-phase computation strategy, Block AttnRes becomes a practical drop-in replacement for standard residual connections with minimal overhead. Scaling law experiments confirm that the improvement is consistent across model sizes, and ablations validate the benefit of content-dependent depth-wise selection. We further integrate AttnRes into the Kimi Linear architecture (48B total / 3B activated parameters) and pre-train on 1.4T tokens, where AttnRes mitigates PreNorm dilution, yielding more uniform output magnitudes and gradient distribution across depth, and improves downstream performance across all evaluated tasks.

cs.CL↗

Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

Multi-object tracking (MOT) involves analyzing object trajectories and counting the number of objects in video sequences. However, 2D MOT faces challenges due to positional cost confusion arising from partial occlusion. To address this issue, we present the novel Occlusion-Aware SORT (OA-SORT) framework, a plug-and-play and training-free framework that includes the Occlusion-Aware Module (OAM), the Occlusion-Aware Offset (OAO), and the Bias-Aware Momentum (BAM). Specifically, OAM analyzes the occlusion status of objects, where a Gaussian Map (GM) is introduced to reduce background influence. In contrast, OAO and BAM leverage the OAM-described occlusion status to mitigate cost confusion and suppress estimation instability. Comprehensive evaluations on the DanceTrack, SportsMOT, and MOT17 datasets demonstrate the importance of occlusion handling in MOT. On the DanceTrack test set, OA-SORT achieves 63.1% and 64.2% in HOTA and IDF1, respectively. Furthermore, integrating the Occlusion-Aware framework into the four additional trackers improves HOTA and IDF1 by an average of 2.08% and 3.05%, demonstrating the reusability of the occlusion awareness.

cs.CV↗

Kimi K2: Open Agentic Intelligence

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

cs.LG↗

Level of Faces for Exponential Sequence of Arrangements

In this paper, we introduce the bivariate exponential generating function $F_l(x,y)$ for the number of level-$l$ faces of an exponential sequence of arrangements (ESA), and establish the formula $F_l(x,y)=\big(F_1(x,y)\big)^l$ with a combinatorial interpretation. Its specialization at $x=0$ recovers a result first obtained by Chen et al. [3,4] for certain classic ESAs and later generalized to all ESAs by Southerland et al. [8]. As a byproduct, we obtain that an alternating sum of the number of level-$l$ faces is invariant with respect to the choice of ESA, and is exactly the Stirling number of the second kind. We also extend the binomial-basis expansion theorem [3,4,14] and Stanley's formula on ESAs [9] from characteristic polynomials to Whitney polynomials.

math.CO↗

RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs

All-pairs shortest paths (APSP) remains a major bottleneck for large-scale graph analytics, as data movement with cubic complexity overwhelms the bandwidth of conventional memory hierarchies. In this work, we propose RAPID-Graph to address this challenge through a co-designed processing-in-memory (PIM) system that integrates algorithm, architecture, and device-level optimizations. At the algorithm level, we introduce a recursion-aware partitioner that enables an exact APSP computation by decomposing graphs into vertex tiles to reduce data dependency, such that both Floyd-Warshall and Min-Plus kernels execute fully in-place within digital PIM arrays. At the architecture and device levels, we design a 2.5D PIM stack integrating two phase-change memory compute dies, a logic die, and high-bandwidth scratchpad memory within a unified advanced package. An external non-volatile storage stack stores large APSP results persistently. The design achieves both tile-level and unit-level parallel processing to sustain high throughput. On the 2.45M-node OGBN-Products dataset, RAPID-Graph is 5.8x faster and 1,186x more energy efficient than state-of-the-art GPU clusters, while exceeding prior PIM accelerators by 8.3x in speed and 104x in efficiency. It further delivers up to 42.8x speedup and 392x energy savings over an NVIDIA H100 GPU.

cs.AR↗

CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference

The proliferation of large language models (LLMs) is accelerating the integration of multimodal assistants into edge devices, where inference is executed under stringent latency and energy constraints, often exacerbated by intermittent connectivity. These challenges become particularly acute in the context of multimodal LLMs (MLLMs), as high-dimensional visual inputs are transformed into extensive token sequences, thereby inflating the key-value (KV) cache and imposing substantial data movement overheads to the LLM backbone. To address these issues, we present CHIME, a chiplet-based heterogeneous near-memory acceleration for edge MLLMs inference. CHIME leverages the complementary strengths of integrated monolithic 3D (M3D) DRAM and RRAM chiplets: DRAM supplies low-latency bandwidth for attention, while RRAM offers dense, non-volatile storage for weights. This heterogeneous hardware is orchestrated by a co-designed mapping framework that executes fused kernels near data, minimizing cross-chiplet traffic to maximize effective bandwidth. On FastVLM (0.6B/1.7B) and MobileVLM (1.7B/3B), CHIME achieves up to 54x speedup and up to 246x better energy efficiency per inference as compared to the edge GPU NVIDIA Jetson Orin NX. It sustains 116.5-266.5 token/J compared to Jetson's 0.7-1.1 token/J. Furthermore, it delivers up to 69.2x higher throughput than the state-of-the-art PIM accelerator FACIL. Compared to the M3D DRAM-only design, CHIME's heterogeneous memory further improves energy efficiency by 7% and performance by 2.4x.

cs.AR↗

HDDB: Efficient In-Storage SQL Database Search Using Hyperdimensional Computing on Ferroelectric NAND Flash

Hyperdimensional Computing (HDC) encodes information and data into high-dimensional distributed vectors that can be manipulated using simple bitwise operations and similarity searches, offering parallelism, low-precision hardware friendliness, and strong robustness to noise. These properties are a natural fit for SQL database workloads dominated by predicate evaluation and scans, which demand low energy and low latency over large fact tables. Notably, HDC's noise-tolerance maps well onto emerging ferroelectric NAND (FeNAND) memories, which provide ultra-high density and in-storage compute capability but suffer from elevated raw bit-error rates. In this work, we propose HDDB, a hardware-software co-design that combines HDC with FeNAND multi-level cells (MLC) to perform in-storage SQL predicate evaluation and analytics with massive parallelism and minimal data movement. Particularly, we introduce novel HDC encoding techniques for standard SQL data tables and formulate predicate-based filtering and aggregation as highly efficient HDC operations that can happen in-storage. By exploiting the intrinsic redundancy of HDC, HDDB maintains correct predicate and decode outcomes under substantial device noise (up to 10% randomly corrupted TLC cells) without explicit error-correction overheads. Experiments on TPC-DS fact tables show that HDDB achieves up to 80.6x lower latency and 12,636x lower energy consumption compared to conventional CPU/GPU SQL database engines, suggesting that HDDB provides a practical substrate for noise-robust, memory-centric database processing.

cs.AR↗

Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

cs.CL↗