arXiv · 2609.39958
Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking
Abstract
Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating directly from a short prompt. Every judge scores the system higher on all seventeen development deliverables. Margins against direct Opus generation from a short prompt range from -4.7 to +0.8 points. Judges agree on broad progress across development rounds but agree less on final-deck rankings than on pooled scores. Repeated grading also shifts scores on unchanged decks, making small improvements difficult to distinguish from judge variability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ludovic Gibert, Matis Despujols, Andre-Louis Rochet. 2026-09-30. Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking. https://arxiv.org/abs/2609.39958
Cite the original work for its findings. Save a collection to share your selection of sources.