Sequential Sampling for Binary Classification: Two LLMs are (Almost) All You Need
We study how to orchestrate multiple large language models (LLMs) for reliable binary classification when models differ in accuracy, monetary cost, and response time. A decision-maker adaptively chooses which LLM to query, updates its belief after each response, and stops once sufficient confidence has been reached. Our results reveal a sharp contrast between exact and near-optimal orchestration. In the worst case, an exactly optimal policy may need to query every available LLM. Yet under high reliability requirements, a simple policy using at most two LLMs is asymptotically optimal. The policy identifies one specialist for each possible outcome and adaptively switches between them according to the accumulated evidence. We derive a universal lower bound on the achievable cost and show that this two-LLM policy approaches the bound at the best possible rate. Two LLMs are not merely sufficient; they can be essential. When models have complementary strengths, relying on the best single LLM can be arbitrarily more costly than using two specialists. How the models are queried also matters: committing to a fixed query plan can waste substantial resources because it cannot stop early or redirect queries as evidence accumulates. In a weak-signal regime, we analytically show that the best static design costs four times as much as the optimal sequential policy. Numerical experiments using 18 LLMs on a public benchmark show that these benefits remain substantial at practically relevant confidence levels. The main message is simple: near-optimal LLM orchestration need not require coordinating a large portfolio of models. Identifying two complementary specialists and switching between them as evidence evolves can capture most of the value of adaptive inference.