NeFut Logo NeFut
中 Admin Login

[CS.AI] Pretrained ASR Pseudo‑labeling for Noisy Police Audio

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Pre‑trained automatic speech recognition (ASR) systems perform poorly on noisy Broadcast Police Communication (BPC), hampering analysis of police decision‑making. Pseudo‑labeling offers an unsupervised way to improve ASR without costly human transcriptions, yet its effectiveness in highly noisy domains remains unclear.

We systematically evaluate two foundation ASR models—OpenAI Whisper and Qwen3‑ASR—on BPC corpora from Baltimore and Chicago. Internal confidence metrics (log‑probabilities and STAR scores) fail to separate high‑quality from low‑quality pseudo‑labels. To address this, we introduce an external LLM‑as‑a‑judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts.

The LLM‑judging filter is more aggressive than internal metrics and substantially reduces the word error rate (WER) of the pseudo‑labeled training sets for both corpora, although a sizable gap remains compared to an oracle filter. We also propose a cross‑model pseudo‑labeling scheme where one model is fine‑tuned on pseudo‑labels generated by the other, and then the roles are swapped. This bidirectional fine‑tuning shows promise as a future direction for pseudo‑labeling research.

Review: The work demonstrates the practical value of LLM‑based filtering in extremely noisy domains and introduces a novel cross‑model pseudo‑labeling strategy, offering a viable path for adapting ASR to challenging audio environments.

Original Source: https://arxiv.org/abs/2609.30469

[h] Back to Home