NeFut Logo NeFut
中 Admin Login

[CS.AI] Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Existing work usually evaluates large language models (LLMs) on multilingual kinship understanding with multiple‑choice tests, treating the task as a recognition problem. This study instead prompts five open‑weight LLMs to generate kinship terms in three non‑Western languages (Hindi, Tamil, Korean) across two communicative tasks, and compares them with a matched four‑option selection baseline.\ \ On identical relation‑language cells, GPT‑OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in only 36.00% of the corresponding generation attempts; Llama‑3.370B shows a similar pattern (77.92% vs. 24.24%). Because the four‑option condition displays candidate terms and does not require script production, the gap is interpreted as an evaluation format issue rather than direct evidence that lexical knowledge is missing.\ \ When prompts explicitly specify the third language (L3), accuracy varies sharply: GLM‑5.1 reaches 72.29% while Llama‑3.370B drops to 24.24%. The paternal‑lineage advantage is language‑specific—large in Hindi, weak or reversed in Korean—while Tamil shared‑term pairs serve as a control for measurement variation.\ \ These findings demonstrate that generating culturally specific kinship terms remains difficult even when the relationship is explicitly stated, motivating the inclusion of generation‑based evaluation alongside multiple‑choice testing.\ \ Review

Original Source: https://arxiv.org/abs/2609.26942

[h] Back to Home