Search arXiv⌕ Search

arXiv · 2610.02387

SOLO: Certified-Recall Metric Similarity Search with Scan-Only Sampled Inverted Lists

Abstract

We present SOLO, an index for approximate nearest-neighbor search in general metric spaces whose serving path contains no ranking heuristic of any kind: a query is routed to the $k_s$ nearest points of a random sample of the database, and every object in the touched posting lists is evaluated with the true distance. Because nothing must outrank anything, recall equals a coverage probability computable from the stored index: one ground-truth pass over a query sample certifies every operating point at once, without serving any of them -- a recall certificate, and for a navigable graph no analogous object exists at any price. The whole index is one recursive rule -- sample the collection, post each object to its $b$ nearest sample points, split any list that outgrows a bound, always scan the leaves -- and its operating surface obeys an equal-work law, recall $\approx f(b \cdot k_s)$, whose level is a one-scalar signature of the dataset. The same scan-only structure gives a serving floor no graph architecture reaches once the router is itself indexed by the same rule: Deep-100M served at recall 0.9977 from 1 GB of resident memory (enforced cap, 10.7 bytes per object) and at 0.9964 from 256 MB, Deep-1B at recall 0.9925 from 512 MB (and from 96 MB at depth 3), inserts that are one search, and deletes that are exact. Throughput is competitive where the hardware allows it -- up to $1.8\times$ a tuned HNSW at $10^8$ on a two-socket 32-core server, with operating points to the right of where that graph saturates -- and the tables report it against HNSW, DiskANN, GRAFT, NAPP, misi, and SPANN's assignment rule on the same hardware and ground truth.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Édgar Chávez. 2026-10-01. SOLO: Certified-Recall Metric Similarity Search with Scan-Only Sampled Inverted Lists. https://arxiv.org/abs/2610.02387

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grain Theory: Type-Level Granularity Correctness in Data Pipelines

Data pipelines fail silently when a transformation changes the grain of the data, the level of detail each element represents, without the designer intending it; fan traps and chasm traps are the familiar symptoms. Dimensional modeling has used grain informally since the 1990s. We define grain by irreducibility and isomorphism alone, with no appeal to functional dependencies, and compute it compositionally over products, sums and the inductive and coinductive fixed points, so it is defined on arrays, event streams and the other nested and recursive types where dependency theory has no referent, and the grain of a type's entity is its entity key, where collections of different grain integrate. Grain is more than a key: one structure under different declared grains denotes different things. The grain ordering is a partial order on types, axiomatized by Armstrong-style axioms whose completeness transfers on product types. For the relational algebra we prove inference rules, chief among them an equi-join theorem computing the exact, minimal grain. Grain projection commutes with transformation and grain lifts compose, and the algorithm CalcG decides pipeline grain-correctness at design time, from the schema alone, in time linear in the pipeline. The errors then follow from a single comparison: a grain the calculus computes, set against a grain the design intends, stands in exactly one relation to it in the grain ordering, and that position names the failure. Fan traps, dropped grain fields, ungrounded multi-version reads, wrong-grain aggregations and behavioral-class violations are the instances of this classification; chasm traps, a data-instance failure the theory does not model, are localized but not decided. The theory is mechanized in Lean 4.

cs.DB↗

Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents

However, building reliable data agents remains difficult because real enterprise data is fragmented across many heterogeneous database systems, with duplicated and inconsistent data, and key information often buried in unstructured text, requiring agents to go beyond just writing SQL or data science scripts to answer questions. Existing benchmarks tackle only individual pieces of this end-to-end workflow (e.g., text-to-SQL over a single database) and are increasingly saturated and contaminated. We present a new benchmark for LLM agents, the Data Agent Benchmark (DAB), grounded in a study of enterprises building production data agents across six industries. DAB comprises 104 queries across 17 datasets and 4 database management systems. On DAB, the best state-of-the-art agent achieves only 57% pass@1. We analyze agent failure modes and distill takeaways for future data-agent development. Our benchmark and experiment code are published at github.com/ucbepic/DataAgentBench.

cs.DB↗

Query Performance Tuning with Optimal Exploration of Optimizer Cost Model Parameter Space

Modern query optimizers use analytical cost models to estimate the cost of a given query plan. Such cost models are typically functions of a set of "cost units" that specify unit CPU cost when processing a row or unit IO cost when accessing a disk page. These cost units are traditionally viewed as platform-dependent constants, that is, they require a one-shot calibration when a database is deployed on a hardware/software platform, but are fixed afterward regardless of the query being optimized for. Some very recent work has taken a different perspective by viewing these cost units as tunable parameters that we call "cost model parameters (CMPs)" in this paper. However, so far there is no approach that offers any optimality guarantee for the tuning results. We present a new approach to systematically explore the query plan space spanned by the CMPs and find the best plan in terms of execution time. Compared to alternative exploration approaches that use random search (RS) or Bayesian optimization (BO), our new approach is deterministic and, therefore, avoids the undesirable instability that is inevitable when applying RS or BO. Moreover, it is guaranteed to find all candidate plans in the query plan space without suffering from the overhead of an exhaustive enumeration. We also present a set of optimization techniques to reduce the overall evaluation time spent on executing the candidate plans found, a factor that is often overlooked by previous work but is critical from a practical point of view. Experimental evaluation on top of PostgreSQL and Microsoft SQL Server demonstrates the efficacy of tuning the CMPs, which can find query plans that are orders of magnitude faster in execution time than the ones found by RS or BO.

cs.DB↗