Data Science Code Translation (DSCT) is the conversion of code between data‑science libraries while preserving functional equivalence, enabling interoperability across ecosystems. Large Language Models (LLMs) have made notable strides in Data Science Code Generation (DSCG), yet their DSCT capabilities remain under‑examined. We introduce the ORCA benchmark with two complementary settings. ORCA‑MAIN comprises 1,600 carefully curated grounding‑level tasks spanning Data Querying, Data Manipulation, and Deep Learning. ORCA‑PROJECT contains 200 translation tasks covering seven data‑science task types. Each task includes annotated reference translations and test cases for functional validation, and a multi‑stage quality verification process ensures task correctness and test robustness. Experiments reveal that DSCT is still challenging; even state‑of‑the‑art LLMs perform modestly. Claude‑Opus‑4.6 achieves a 56.92% success rate on ORCA‑MAIN and 33.67% on ORCA‑PROJECT, indicating ample room for improvement. We also observe a clear directional bias: translation is easier when the source code expresses the task with explicit, fine‑grained operations. Motivated by this, we propose an intent‑augmented approach that first infers the source‑code intent and then supplies it as additional context for translation, yielding average absolute success‑rate gains of 4.80% on ORCA‑MAIN and 5.33% on ORCA‑PROJECT.
Review: ORCA offers a systematic evaluation framework for DSCT, exposing current LLM limitations in cross‑library code migration and demonstrating that intent‑aware prompting can meaningfully boost performance, paving the way for future research.