Search arXiv⌕ Search

arXiv subjects

Lingxiao Guan

Publications and source records attributed to Lingxiao Guan.

3 recordsLinked to original sources

What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents

Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We introduce compositional controllability to address these questions. A comparison window covers one stage, several stages, or the whole agent. Our central result bounds the gap between observed and controlled score differences using only nuisance outside the window. This yields an admissibility test applied before scores are inspected. Inadmissible comparisons are refused. For admissible pairs, an ordering is certified only when the score gap exceeds the combined sampling and nuisance radii; otherwise, it remains undecided. These decisions give each system a rank interval. We introduce BioLitBench, a benchmark of 2,042 biomedical articles represented as structured claim graphs. Among seven published pipelines, a conventional statistical analysis declares a winner in 14 of 21 pairwise comparisons. Yet the top-ranked system alone received the target review's bibliography. To isolate pipeline performance, our test requires matched inputs and a fixed backbone model. It refuses 11 of the 21 comparisons, including every comparison involving the top-ranked system. Seven of the 14 conventional conclusions fall within these refused pairs. The same comparison windows support stage-level training. We train SCRIBE on Qwen3.8-27B using rewards measured at each stage's exit. Under matched evidence, SCRIBE achieves a certified rank interval of [1,2], with certified advantages over all evaluated published pipelines and the evaluated Claude and OpenAI agents. Under same pool, SCRIBE matches the strongest published retriever and is certified above three published pipelines.

cs.LG↗

Biomedical Question Answering via Multi-Level Summarization on a Local Knowledge Graph

In Question Answering (QA), Retrieval Augmented Generation (RAG) has revolutionized performance in various domains. However, how to effectively capture multi-document relationships, particularly critical for biomedical tasks, remains an open question. In this work, we propose a novel method that utilizes propositional claims to construct a local knowledge graph from retrieved documents. Summaries are then derived via layerwise summarization from the knowledge graph to contextualize a small language model to perform QA. We achieved comparable or superior performance with our method over RAG baselines on several biomedical QA benchmarks. We also evaluated each individual step of our methodology over a targeted set of metrics, demonstrating its effectiveness.

cs.CL↗

Frequency dependence of near-surface oceanic kinetic energy from drifter observations and global high-resolution models

The geographical variability, frequency content, and vertical structure of near-surface oceanic kinetic energy (KE) are important for air-sea interaction, marine ecosystems, operational oceanography, pollutant tracking, and interpreting remotely sensed velocity measurements. Here, KE in high-resolution global simulations (HYbrid Coordinate Ocean Model; HYCOM, and Massachusetts Institute of Technology general circulation model; MITgcm), at the sea surface (0 m) and 15 m, are respectively compared with KE from undrogued and drogued surface drifters. Global maps and zonal averages are computed for low-frequency ($<$ 0.5 cpd), near-inertial, diurnal, and semi-diurnal bands. Both models exhibit low-frequency equatorial KE that is low relative to drifter values. HYCOM near-inertial KE is higher than in MITgcm, and closer to drifter values, probably due to more frequently updated atmospheric forcing. HYCOM semi-diurnal KE is lower than in MITgcm, and closer to drifter values, likely due to inclusion of a parameterized topographic internal wave drag. A concurrent tidal harmonic analysis in the diurnal band demonstrates that much of the diurnal flow is non-tidal. We compute a simple proxy of near-surface vertical structure, the ratio of 0 m KE to 0 m KE plus 15 m KE in model outputs, and undrogued KE to undrogued KE plus drogued KE in drifter observations. Over most latitudes and frequency bands, model ratios track the drifter ratios to within error bars. Values of this ratio demonstrate significant vertical structure in all frequency bands except the semidiurnal band. Latitudinal dependence in the ratio is greatest in diurnal and low-frequency bands.

physics.ao-ph↗