
Summary
- Hume AI CEO Andrew Ettinger unveiled the company's internal voice AI evaluation framework, called CLEAR, on the MTS podcast.
- Appropriate responses and comprehension come first, he said, while sounding natural and human-like is only the third factor.
- The fourth and final factor is whether that quality holds up consistently across hundreds or thousands of interactions.
Hume AI CEO Andrew Ettinger has pushed back on what's often treated as the single most important benchmark in voice AI. According to him, how human a voice sounds is only the third most important thing when evaluating a voice AI system. In a video posted on X, Ettinger introduced CLEAR, the framework Hume AI uses internally to score voice AI systems, and laid out that argument.
What gets checked before voice quality
Back in August, Hume AI worked with Hugging Face to measure and publish findings on benchmark memorization in open-source speech recognition (ASR) models. Given the company's track record of building tools to score and validate voice AI systems, CLEAR looks like a natural extension of that same work.
Here's the order Ettinger laid out for CLEAR. The first thing it checks is whether the conversation actually holds together — can the system respond appropriately to what the other person says? If it fails here, he said, nothing else in the framework matters. Second comes listening comprehension. A system can speak beautifully, but if it can't accurately understand what the other person is saying, the conversation falls apart anyway.
Only after clearing those first two bars does CLEAR check the third factor: does it sound human? Ettinger explained that even when the content of a response and the system's comprehension are both solid, an unnatural-sounding voice makes users reluctant to keep trusting it. The fourth and final factor is whether all of this — context and timing included — stays consistent across hundreds or thousands of interactions. His reasoning: whether you're deploying the system as a cable company support agent or something a parent trusts enough to use for a medical consultation, doing well once or twice isn't enough.

"Getting the words and comprehension right isn't enough if it doesn't sound right"
"Even if those two things are right, if it doesn't sound natural, people won't want to keep trusting it," Ettinger said. That's a notable statement coming from the head of a company operating in a market where most marketing has leaned almost entirely on "how human does it sound" as speech synthesis technology has advanced. Yet here's the CEO of one such company ranking that exact quality third.

Why the order is flipped
As voice AI pushes deeper into real-world use cases like call centers and medical consultations, the implication is that keeping a conversation coherent and on-context matters more, and needs to be validated first, than whether the voice sounds pleasant. However natural a voice sounds, if it answers a question incorrectly or forgets something said earlier, users lose trust in the entire system after just one failure. On the flip side, once responses and comprehension are solidly in place, naturalness becomes the thing that holds onto that trust over time — that's the picture the CLEAR framework paints.
Editor's take
Nearly every piece of voice AI marketing leads with "indistinguishable from a human voice." So it's telling that the CEO of a company that has spent time building evaluation tools for this exact market ranks that quality third. It's a signal that there's a gap between what the industry is selling and what's actually needed. Any team that's deployed voice AI has probably run into the same thing — the demo voice sounds fantastic, but once it's live, users bail after two or three exchanges. The cause is rarely the tone of the voice. It's usually a missed context or a wrong answer.
For teams in Korea looking to adopt voice AI for support or agent use cases, this ordering doubles as a solid checklist. Instead of putting voice naturalness first when choosing a vendor, start by pulling a few hundred real conversation logs and checking whether the system responded appropriately and understood correctly. Naturalness comes after that, and consistency last.
Expect other voice AI companies to publish their own evaluation frameworks in the coming months, each pushing the message that "here's how we validate our system." In a market where a single benchmark score rarely settles who's better, the evaluation methodology itself becomes a marketing asset.





Comments