Search arXiv⌕ Search

arXiv · 2609.30749

ORCA: Evaluating LLMs on Data Science Code Translation

Abstract

Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which comprises 1,600 carefully curated grounding-level tasks across 3 representative domains: Data Querying, Data Manipulation, and Deep Learning; and ORCA-PROJECT, which contains 200 translation tasks over complete data science projects across 7 data science task types. Each task is accompanied by annotated reference translations and test cases for validating functional equivalence. We further incorporate a multi-stage quality verification process that thoroughly verifies task correctness and test case robustness. Experimental results demonstrate challenges in DSCT, with even frontier LLMs showing limited performance. Specifically, Claude-Opus-4.6 achieves a success rate of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT, indicating considerable room for improvement in DSCT. We also observe a clear directional preference in DSCT, where translation is consistently easier when the source code expresses the task through more explicit, fine-grained operations. Motivated by this, we propose an intent-augmented method, in which the model first infers source-code intent and then uses it as additional context for translation, achieving average absolute success-rate gains of 4.80% and 5.33% on ORCA-MAIN and ORCA-PROJECT, respectively.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng. 2026-09-25. ORCA: Evaluating LLMs on Data Science Code Translation. https://arxiv.org/abs/2609.30749

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution

This note is a substantially revised version of the author's earlier preprint (version~1, 2023), which proposed a ``Categorical Graph Model for Conflict Resolution'' (C-GMCR). Several of the central definitions and claims of version~1 are incorrect, and this version replaces them. We first observe that the states and one-step moves of a graph model do not form a category; the correct construction is the free category on a quiver whose arrows are moves labelled by the decision maker (DM) who controls them. Controller labels are functorial: they define a functor into the free monoid on the set of DMs, and the legal move sequences underlying coalition reachability are exactly the paths whose label words contain no immediate repetition. Preferences, in contrast, are not functorial. We show that a map representing a DM's preference in a preordered set exists only if the preference is transitive, and that even then it extends to a functor on the path category if and only if every move, regardless of its controller, is weakly improving for the focal DM; in that case general metarationality, symmetric metarationality and sequential stability all collapse to Nash stability for the focal DM. Preferences therefore enter the categorical picture as a selection of a subquiver of improvements, not as a functor. We restate the four standard stability concepts in path language, work them out on the Prisoner's Dilemma, and take a first step towards comparing graph models via induced embeddings: Nash stability is reflected by induced embeddings, whereas general metarationality, symmetric metarationality and sequential stability are neither preserved nor reflected. A section lists each correction to version~1 explicitly.

cs.AI↗

From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark

Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs. Our code is available at https://github.com/you68681/kaar.

cs.AI↗

GLOVE: Global Verifier for LLM Memory-Environment Realignment

Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through internal model cognition, such as reflection, for editing memory entries. However, these assumptions often break down in practical environments with dynamic drifts. We propose the Global Verifier (GLOVE), a framework that introduces a new design dimension for LLM memory systems by establishing a relative notion of truth. Through active probing to detect inconsistencies between retrieved memories and fresh observations, GLOVE enables memory-environment realignment by verifying and updating memory without access to ground-truth supervision or strong reliance on model introspection. We evaluate GLOVE on diverse benchmarks spanning web navigation, planning, and control, augmented with controlled environmental drifts that introduce non-stationarity beyond the original benchmark settings. Our results show that GLOVE substantially improves agent success rates, suggesting a robust pathway to cognitive agents capable of self-evolving.

cs.AI↗