Search arXiv⌕ Search

arXiv · 2609.36513

Local Search for Fair Max-Min Diversification

Abstract

Given $n$ points in a metric space, Max-Min diversification asks for a subset of $k$ points maximizing the minimum pairwise distance between the selected points. This is arguably the most fundamental notion of diversity with applications across a wide range of domains. We consider this problem under partition constraints, previously studied as Fair Max-Min Diversification (FMMD). Here, each point has a color in $[m]$, and a feasible solution must contain exactly $k_i$ points of color $i$, where $k_1,\ldots,k_m$ are prescribed parameters satisfying $\sum_i k_i=k$. We give the first constant factor approximation for the problem using local search, that runs in time $f(m)\cdot \operatorname{poly}(n)$, in which all constraints are satisfied exactly. All previously known algorithms either provided an $\widetilde Θ(m)$ approximation factor, had running times exponential in the solution size $k$, or satisfied the fairness constraints only approximately or in expectation. We further generalize our result to the problem where each point may belong to an arbitrary subset of colors. Given lower and upper bounds $\ell_i$ and $u_i$ for every color $i$, the goal is to find $k$ points whose color counts satisfy all these bounds while maximizing their diversity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sepideh Mahabadi, Shyam Narayanan, Varun Sivashankar. 2026-09-29. Local Search for Fair Max-Min Diversification. https://arxiv.org/abs/2609.36513

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Auction-Based Algorithms for Matroid Intersection: Near-Linear Query Complexity and Constant-Pass Semi-Streaming

In this paper, we develop a new auction-based framework for matroid intersection and use it to obtain improved approximation algorithms in several computational settings. Our framework is inspired by Fleiner's generalized stable matching algorithm and extends the semi-streaming auction algorithm for bipartite matching due to Assadi, Liu, and Tarjan. Using this framework, for any $\varepsilon > 0$, we present a simple $(1-\varepsilon)$-approximation algorithm in the rank-oracle model whose query complexity matches that of the current fastest algorithm. Furthermore, by extending this result, we obtain the first $(1-\varepsilon)$-approximation algorithm for the weighted problem that requires only a near-linear number of rank-oracle queries, achieving the best known rank-oracle query complexity for the problem. We also obtain a $(1-\varepsilon)$-approximation semi-streaming algorithm for matroid intersection in the multi-pass streaming model, where the elements of the ground set arrive sequentially. It is the first algorithm achieving this approximation ratio using a constant number of passes and nearly linear space in the ranks of the matroids. When viewed in the standard offline setting, the same algorithm yields the first deterministic $(1-\varepsilon)$-approximation algorithm for matroid intersection that requires only a near-linear number of independence-oracle queries.

cs.DS↗

Faster network motif discovery by counting isomorphic subtrees

We develop a new algorithm for counting the number of subgraphs of a network isomorphic to a given query graph (#SubgraphIsomorphism), motivated by network motif search. High-degree vertices (hubs), common in real-world networks, contribute to a combinatorial explosion in the number of subgraphs, making existing motif search algorithms intractable for motif sizes greater than $\approx 8$ on a wide variety of networks of interest. Our procedure leverages the $k$-core decomposition and a novel subtree-counting technique to quickly scan the periphery of a network. These two innovations allow our algorithm to significantly speed up its predecessors in practice, especially as most real-world networks have a relatively large periphery. We prove that #RootedSubtreeIsomorphism, a key subroutine in our algorithm, is #P-complete via a reduction from counting bipartite matchings. We provide analytic upper bounds on our algorithm's execution time, and evaluate its performance on 11 real-world networks of varying topologies.

cs.DS↗

Optimal VC Dimension of Contrastive Learning with Margin

Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $ϕ:[n]\rightarrow \mathbb{R}^{d}$, if $\|ϕ(i)-ϕ(k)\|_2>(1+α)\cdot\|ϕ(i)-ϕ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.

cs.DS↗