arXiv · 2609.30571
HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
Abstract
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar. 2026-09-24. HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases. https://arxiv.org/abs/2609.30571
Cite the original work for its findings. Save a collection to share your selection of sources.