Search arXiv⌕ Search

arXiv · 2610.02385

Fine-Grained Analysis of SIMD-Based Hash Table Implementations

Abstract

In recent years, several variants of classical hash table schemes have been developed by engineers in order to take advantage of the processor's internal parallelism using SIMD instructions, which make it possible to operate on multiple bytes simultaneously. At a small additional memory cost, this enables a significant speedup, making it a data structure that is increasingly popular in practice, when very high performance is required. In this article, we provide a detailed theoretical analysis of the dynamics of such hash tables. From a methodological standpoint, we use and adapt a technique developed by Wormald in the 1990s to study dynamic graphs. This approach, which can be adapted to many variants, enables us to accurately estimate the quantities of interest by capturing the dynamics of the data structure through systems of differential equations. Although complex, we provide an explicit description of the solutions of these systems, which can furthermore be efficiently approximated numerically. Our main results are stated with high probability, which is significantly more precise than average-case analyses, and they match experimental results remarkably well, even for hash tables of moderate size.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cyril Nicaud, Pablo Rotondo. 2026-10-01. Fine-Grained Analysis of SIMD-Based Hash Table Implementations. https://arxiv.org/abs/2610.02385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Fair Diversity Maximization via Local Search

Diversity maximization is a fundamental optimization problem with applications in machine learning, data summarization, information retrieval, and recommendation systems. In many such applications, the data are partitioned into groups, and the selected subset must satisfy prescribed group quotas. We study Fair Diversity Maximization: given a set of points in a metric space partitioned into $m$ groups, the goal is to select exactly $k_i$ points from each group $i$ while maximizing the minimum pairwise distance among the selected points. The best previously known approximation guarantee is $m+1$, which grows linearly with the number of groups. We show that this dependence on $m$ is not fundamental. We present a new local-search framework that yields a $4$-approximation for any constant number of groups, with no restrictions on the metric space or on the size of the selected set. To the best of our knowledge, this is the first constant-factor approximation whose guarantee is independent of the number of groups in this general setting. Our framework maintains all group quotas exactly while progressively eliminating violations of the diversity objective. We further develop a specialized algorithm for two groups that achieves a $2$-approximation, improving the previous best factor of $3$. This factor is optimal: unless $\mathrm{P}=\mathrm{NP}$, no polynomial-time algorithm can achieve an approximation factor strictly better than $2$, even for the unconstrained case.

cs.DS↗

Methods for Counting and Enumerating Set Partitions

Set partitions are arrangements of distinct objects into groups. After a brief review of the subject, we consider the task of counting and enumerating set partitions. The number of set partitions, known as Bell number, is a rapidly increasing number and does not have an explicit formula. We study approximate expressions for the Bell number given in the literature. We find that an asymptotic formula of Moser and Wyman gives a surprisingly accurate approximation to the Bell number even for small set sizes. Furthermore, a simple expression due to Berend and Tasssa can be conveniently used to approximate the Bell number for small set sizes. Next, we consider enumeration of set partitions. The problem of listing all set partitions arises in a variety of settings, in particular in combinatorial optimization tasks. Algorithms for enumerating all set partitions are reviewed. The focus is on non-recursive algorithms without Gray code constructions. We compare the classic algorithm of Hutchinson with three more modern ones. Empirically, it is found that all of them scale exponentially with the set size. While the exact compiler and optimization settings do matter, it can be concluded that the algorithm of Djokic et al. is the fastest one, thus it is recommended for practical use.

cs.DS↗

Equalizing Closeness Centralities via Edge Additions

Graph modification problems with the goal of optimizing some measure of a given node's network position have a rich history in the algorithms literature. Less commonly explored are modification problems with the goal of equalizing positions, though this class of problems is well-motivated from the perspective of equalizing social capital, i.e., algorithmic fairness. In this work, we study how to add edges to make the closeness centralities of a given pair of nodes more equal. We formalize several versions of this problem: Closeness Ratio Improvement, which aims to maximize the ratio of closeness centralities between two specified nodes, and Closeness Gap Minimization, which aims to minimize the absolute difference of centralities. For the former, we present a quasilinear-time $\frac{6}{11}$-approximation, complemented by a bicriteria inapproximability bound. In contrast to this positive result, we show that Closeness Gap Minimization admits no multiplicative approximation, unless P=NP. We also establish NP-hardness for All-Pairs Closeness Ratio Improvement, which aims to maximize the minimum ratio of closeness centralities across all node pairs.

cs.DS↗