An Inspectable LLM Council for Multi-Model Research Answer Aggregation
Long-form research questions often require a single answer that combines details, reconciles conflicting estimates, and retains qualifications spread across several large language model (LLM) outputs. We present an inspectable LLM Council that separates synthesis into two stages: an analyst converts three independent candidate answers into a structured state, and a writer composes the final response from that state. We evaluate the workflow under closed-book conditions on 100 tasks from the Deep Research Accuracy, Completeness, and Objectivity (DRACO) benchmark, retaining 600 final answers, 100 analyst states, and 720 whole-answer judgments across the main and supplementary experiments. Under the primary rubric judge, Council scores 70.98, exceeding GPT and Gemini by 6.18 and 8.34 points; its 0.47-point difference from Claude has a 95% paired interval of $[-0.69,1.61]$. In the full supplementary comparison, Council scores 1.87 points above direct fusion (95% paired interval $[0.33,3.48]$) and covers 51.1% of positive rubric criteria covered by exactly one candidate, compared with 43.7% for direct fusion. On an exploratory 68-task subset with matching returned model labels, direct fusion scores higher, so the full comparison cannot isolate the analyst state's effect. Traced cases show how estimates, provenance qualifications, and selection decisions pass through the saved state into final answers.