NeFut Logo NeFut
中 Admin Login

[CS.AI] Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Multimodal Machine Translation (MMT) incorporates images related to the source text to resolve ambiguities and improve translation quality.

Existing multimodal large language models often ignore visual cues after fusion, making the enhancement of visual sensitivity an active research problem.

We introduce a training strategy called Metric‑based Loss Weighting (MLW). The key idea is to increase the loss weight for target tokens that benefit from the accompanying image, thereby forcing the model to rely more on visual context. Tokens are identified using the Point‑wise Cross‑mutual Information (PCXMI) metric, defined as: $$PCXMI(t)=\log\frac{P(y_t\mid x, v)}{P(y_t\mid x)}$$ where $x$ is the source text, $v$ the image features, and $y_t$ the $t$‑th target token.

We also propose a Congruency‑based PCXMI that measures the agreement between visual and textual contexts by comparing the output distributions with and without visual information. Experiments show that either PCXMI or Congruency‑based PCXMI alone improves visual grounding, while their combination yields the best performance.

We fine‑tune three pretrained multimodal LLMs on image‑guided machine translation for English‑Chinese, English‑German, and English‑French directions, and evaluate on the CoMMuTE contrastive dataset. MLW outperforms standard fine‑tuning by up to more than 7 percentage points in accuracy, while preserving strong overall translation quality.

Review

Original Source: https://arxiv.org/abs/2609.31169

[h] Back to Home