Voice models can speak well but listen poorly
Most voice AI is graded on how accurately it turns speech into text, which ignores everything a transcript throws away: tone, emotion, hesitation, who is speaking, and the noise in the room. Hume AI built Real World VoiceEQ to test exactly that layer. The benchmark covers more than 40 voice models across four capability groups, speech recognition, text-to-speech, speech-to-speech, and speech understanding, scored on over 60 metrics. What sets it apart is the grading itself: over a million human ratings collected across different accents, demographics, and acoustic conditions, including 785,000 ratings for text-to-speech and 48,000 for speech-to-speech.
Two findings stand out. First, there is no all-rounder. No single model lands in the top five across every capability group, so a system that sounds great may be mediocre at understanding, and the reverse. Second, and more pointed, models are much stronger at speaking than at listening. They generate fluent, natural-sounding speech, yet they struggle to pick up the paralinguistic cues, the tone shift or the pause, that carry meaning a transcript never records. Hume also reports that automated speech-language models agree poorly with human raters on these subjective judgments, which is the argument for gathering human ratings at this scale in the first place.
Why it matters
If you build a voice agent, the number you probably track is transcription accuracy, and VoiceEQ says that misses the half of the problem where products actually feel broken. A model that cannot tell a frustrated customer from a calm one will misfire no matter how clean its speech sounds, so test the listening side, not just the talking side, before you ship.