Search arXivSearch

arXiv · 2410.11803

Exact Hyper-Rectangular Clustering via Adaptive Subset Selection

Abstract

We study the hyper-rectangular clustering problem (HRCP), where the goal is to partition a set of points into a fixed number of axis-aligned clusters while minimizing the total span. Existing exact approaches, based on mathematical optimization formulations, are limited to relatively small instances due to their strong dependence on the number of data points. We propose an incremental exact algorithm that exploits a key structural property of the problem: optimal cluster boundaries are determined by a subset of points, while interior points do not affect the objective value. The algorithm iteratively solves HRCP on carefully selected subsets of points and expands them only when necessary. A simple optimality condition ensures that, once a solution covers the entire dataset, it is optimal for the original problem. We introduce sampling strategies designed to identify points likely to lie on cluster boundaries, guiding the incremental process toward informative subsets. Computational experiments show that the proposed approach significantly improves scalability, solving instances with up to 10,000 points and substantially outperforming monolithic formulations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Diego Delle Donne, Javier Marenco, Eduardo Moreno. 2026-07-31. Exact Hyper-Rectangular Clustering via Adaptive Subset Selection. https://arxiv.org/abs/2410.11803

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Factorisability of Low Dimensional Non-Negative Integer Matrices

We consider the problem of determining if a given two-dimensional nonnegative integer matrix $M$ is the product of two such matrices, excluding trivial units. A matrix $M$ with no such factorisation is called prime and therefore belongs to the minimal (infinite rank) generator of $2 \times 2$ matrices over the natural numbers, otherwise it is called composite. We also consider the problem of finding a (non-unique) factorisation of a composite matrix. Our results have applications in computational group theory and the theory of codes, where such matrices are called incidence matrices. We analyse the complexity of primality and finding a factorisation for a composite matrix, providing a first efficient algorithm.

cs.DM

The parameterised complexity of generalised temporal domination on temporal graphs with modular structure

Inspired by the static problem $(α,β)$-Dominating Set, we propose a general temporal domination problem, called $(α,β)$-Temporal Dominating Set ($(α,β)$-TDS). We show that this problem encompasses Temporal Dominating Set, and additionally provides first temporal extensions of problems such as $k$-Dominating Set and $α$-Dominating Set. In this paper, we study the parameterised complexity of $(α,β)$-TDS with respect to temporal neighbourhood diversity (TND), temporal modular-width (TMW), and temporal cliquewidth (TCW). We obtain fixed parameter tractability results for all values of $α$ and $β$ with respect to TND; W[1]-hardness with respect to TMW and TCW whenever $β$ is in the problem input, or whenever $α\in (0,1)$ and $β$ is a fixed constant; and para-NP-hardness with respect to TCW when $α= 0$ and $β= 1$, or $α= 1$ and $β= 0$.

cs.DM