Solution:
Root Cause: Permissive Decoder Probability Thresholds during Low-Confidence Inferences
Whisper and similar ASR models evaluate internal metrics during inference:
no_speech_prob (the probability that a segment contains no speech) and
logprob (the average log probability of the generated tokens). When
no_speech_threshold is configured too low (e.g., below 0.6) or when
logprob_threshold is set too permissive (e.g.,
-1.0), the decoder fails to abort text generation on low-confidence segments, allowing hallucinated tokens to pass through as valid transcriptions.
# Diagnostic Verification:
1. Inspect the JSON/dictionary payload of segment outputs during silent or noisy intervals.
2. Check if segments with
no_speech_prob > 0.5 and
avg_logprob < -1.0 are still being emitted by the decoder engine.
# Step-by-Step Fix:
1. Tune Decoder Threshold Flags in Python / CLI:
Adjust no_speech_threshold and logprob_threshold parameters to force early rejection of low-confidence tokens: python
segments, info = model.transcribe(
"input_audio.wav",
no_speech_threshold=0.6,
logprob_threshold=-0.8,
compression_ratio_threshold=2.4
)
2. Tune Whisper CLI Arguments:
Pass strict threshold flags when using the standard CLI: bash
whisper input.wav --no_speech_threshold 0.6 --logprob_threshold -0.8 --compression_ratio_threshold 2.4
3. Implement Manual Filter for Low Logprob Segments:
Filter output arrays programmatically: python
valid_segments = [s for s in segments if s.no_speech_prob < 0.6 and s.avg_logprob > -0.8]
# Prevention & Long-Term Monitoring:
Monitor average avg_logprob scores across transcription datasets to establish baseline confidence curves for specific audio domains.