Hugging Face finds top speech-to-text models reproduce benchmark transcripts even when the audio contradicts them — type0 | type0