Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. This study systematically investigates the challenges posed by culturally loaded translation in LLM-based MT systems.
We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges:
- Task challenges: Frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content.
- Human evaluation challenges: Evaluator backgrounds lead to substantial disagreement in translation judgments.
- Automatic evaluation challenges: Widely used metrics fail to reliably assess translation quality for this task.
These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.
Blogger's Review: This paper delves into the complexities of culturally loaded translation, highlighting the limitations of LLMs in this domain and prompting a reflection on the inadequacies of machine translation technology in handling cultural nuances. Future research should focus more on understanding cultural contexts to enhance translation quality.