Search arXivSearch

arXiv · 2509.21493

Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation

Abstract

We propose Sci2Pol-Bench and Sci2Pol-Corpus, the first benchmark and training dataset for evaluating and fine-tuning large language models (LLMs) on policy brief generation from a scientific paper. We build Sci2Pol-Bench on a five-stage taxonomy to mirror the human writing process: (i) Autocompletion, (ii) Understanding, (iii) Summarization, (iv) Generation, and (v) Verification. It features 18 tasks in multiple-choice and open-ended formats. Specifically, for the Generation stage, we show that BERTScore and ROUGE scores fail to capture the quality of brief writing, and introduce a new LLM-based evaluation metric aligned with expert judgement. Using this benchmark, we evaluate 13 leading open-source and commercial LLMs to uncover key limitations. To improve LLM performance on brief writing, we curate the Sci2Pol-Corpus for fine-tuning. We start by linking each cited scientific paper to its corresponding policy document, drawn from 5.6 million policy records. This produces 140,000 candidate pairs. We then employ an LLM-as-a-judge to filter high-quality examples, followed by in-context polishing using three expert-written samples as references. This process yields a final set of 639 new pairs. Finally, we fine-tune three models on Sci2Pol-Corpus: LLaMA-3.18B, Gemma-12B, and Gemma-27B. Fine-tuning leads to consistent performance improvements across Sci2Pol-Bench. Notably, after fine-tuning, Gemma-27B surpasses the much larger GPT-4o and DeepSeek-V3 (671B). These demonstrate the effectiveness of our corpus in bridging the gap between science and policy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Weimin Wu, Alexander C. Furnas, Eddie Yang, Gefei Liu, Akhil Pandey Akella, Xuefeng Song, Dashun Wang, Han Liu. 2026-02-19. Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation. https://arxiv.org/abs/2509.21493

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Dimension Bridging for 3D RANS with Neural Network Accelerated Gaussian Functional Regression

In many computational science and engineering problems, repeatedly solving fully resolved physics-based models to design for a quantity of interest (QoI) can quickly become intractable, requiring the use of low-fidelity models to predict the same QoI but introduce errors where some features are neglected or are otherwise inaccurately resolved. We use Gaussian Functional Regression (GFR) to learn a correction to a 2D Reynolds-Averaged Navier-Stokes (RANS) model to predict the aerodynamic coefficients from a 3D RANS model. This model pair has a disparity in the governing physics from the reduced dimensionality, a previously unexplored application for GFR. Empirically, our results show that with a proper choice of low-dimensional (LD) model, the proposed kernel allows for the use of fewer high-dimensional (HD) evaluations to regress a response surface to the same level of accuracy as standard stationary kernels. Moreover, the new kernel provides more informative uncertainty quantification, which we show is advantageous when used to drive an adaptive sampling algorithm. Finally, we propose a novel neural network accelerated kernel, which we show offers predictions in good agreement while speeding up evaluations by millions of times in wall clock measurements, bringing the computational budget within the real-time regime.

cs.CE

SabreAgent: Language Models at Design Time for Lost-Sales Inventory Control

SabreAgent uses a language model at design time to construct two components for lost-sales inventory control: a product-specific seasonal prior and a validation-selected capped base-stock policy family. During operation, statistical forecasting and inventory optimization use these frozen artifacts to determine orders, with zero language-model calls. We evaluate the approach on the $1{,}320$ instances of InventoryBench. Under the benchmark's cost assumptions, the operations-research core draws on a zero-lead-time optimality result and a projected-inventory rule for positive deterministic lead times. The latter computes replenishment shortfalls by propagating inventory using sales along simulated demand paths. The seasonal prior adds forecast variants alongside the original forecaster, and the selected policy family handles stochastic lead times with order destruction. SabreAgent scores $0.6311$, compared with $0.5380$ for the strongest published baseline, and ranks first in all six benchmark cells. Ablations attribute most of the gain to the OR core. In the paired analysis, the seasonal component adds $1.79\%$ across the three real-data cells, and the search component adds $2.3\%$ across the two stochastic-lead-time cells. These results demonstrate how model-generated priors and policy structure can improve an OR controller through design-time use.

cs.CE

Hierarchical Multi-Task Learning with Liquidity-Aware Signals for Stock Forecasting

Stock price forecasting is a long-standing challenge in computational finance, driven by the inherent randomness of markets and complex temporal patterns. While recent deep-learning models have raised forecasting accuracy by jointly modeling inter-stock and temporal price dynamics, they conflate inter-stock relationships with intra-stock temporal dependencies and focus solely on the univariate objective of price movement. To address these limitations, we propose LiMT, a Hierarchical Multi-Task Learning framework that integrates liquidity-aware signals for stock price forecasting. LiMT employs a Market Regime Encoder (MRE) module that first extracts contemporaneous cross-stock dependencies, then models each stock's temporal dynamics, yielding a unified latent state. Building on this latent state, we introduce a Liquidity-Driven Learning (LDL) module, a mixture-of-experts architecture that features cross-task gating mechanisms to jointly predict price movement, volatility, and trading volume. We further design an Adaptive Portfolio Optimization (APO) mechanism that converts multi-task forecasts into executable portfolio weights under transaction-cost and liquidity constraints. Extensive experiments on the CSI300 and CSI500 benchmarks show that LiMT performs best among strong neural and tree-based baselines across the reported metrics. In realistic CSI300 backtests, APO improves annualized return from 3.99% to 10.01% and Sharpe ratio from 1.22 to 1.86 over equal weighting, showing that the multi-task forecasts translate into deployable portfolio gains.

cs.CE