Audio large language models let users interact via speech. When an input recording is severely degraded, the model may misinterpret the query and answer based on an incorrect transcription. This paper studies model‑conditional transcription reliability: can an Audio LLM recognize when its own transcription is unreliable?
We first prompt the model to assess the reliability of its own transcription and find that it is a poor judge, predicting reliability in most cases. Existing signals—speech‑quality predictors, generation uncertainty, and transcript‑conditioned WER estimation—provide limited help for detecting transcription failures. In contrast, we discover that reliability is strongly encoded in the audio‑encoder representations.
Leveraging this, we build a lightweight reliability predictor that operates on frozen audio‑encoder embeddings and classifies reliability before generation. The predictor can request clarification when a voice query is deemed unreliable, while allowing reliable queries to proceed unchanged. Experiments show macro‑F1 scores of 81.10% in‑domain and 78.09% cross‑domain, outperforming the strongest baselines by 10.33 and 11.93 points respectively.
We also demonstrate that reliability labels transfer across Audio LLM families, and that transfer performance correlates with the alignment of model‑specific reliability boundaries.
Review