Back to News Feed
TechCrunch AI32d agoMarina Temkin

Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human

While artificial intelligence agents have become increasingly proficient at navigating complex customer support inquiries, a glaring hurdle remains: most users can instantly distinguish between a human representative and a machine. Smallest.ai, a startup established in late 2024, is betting that the next major breakthrough in voice technology won't come from simply scaling up large language models (LLMs), but from developing specialized, compact models engineered specifically for fluid human conversation.

The company’s primary objective is ambitious: to render interactions with AI agents indistinguishable from human-to-human dialogue. To achieve this, Smallest.ai is crafting a lightweight voice model that mimics the human cognitive process—listening, internalizing, and responding simultaneously.

A New Approach to Conversational Latency

To accelerate this mission, Smallest.ai has successfully secured $13 million in a Series A funding round. The investment was led by Seligman Ventures, with additional backing from Sierra Ventures and 3one4 Capital, bringing the startup’s total capital raised to over $21 million.

According to founder and CEO Sudarshan Kamath, the current architecture of most LLMs is fundamentally ill-suited for real-time voice.

"The way an LLM works is you give it an entire prompt, and then it starts thinking. If you think about how we are talking, I’m not giving you like a large clipping of my audio, and then you start thinking."

In a text-based interface, a slight delay is negligible. In a live voice call, however, even a fraction of a second of latency feels jarring and unnatural. Smallest.ai’s model acts as a real-time intelligence layer, enabling seamless conversations with virtually zero response lag.

The Two-Model Strategy

Kamath envisions a future where AI agents operate on a dual-model framework:

  • The Voice Model: A specialized, compact engine designed for real-time, low-latency interaction.
  • The Foundational LLM: A robust, "offline" model triggered only when the system encounters complex queries outside its immediate knowledge base.

In this scenario, the agent might briefly place a user on hold to "research" a difficult issue, mirroring the behavior of a human customer service representative. Unlike massive foundational models, Smallest.ai’s proprietary tech is optimized for voice-specific challenges, including:

  • Handling diverse global accents.
  • Supporting dozens of languages.
  • Maintaining performance in noisy, real-world environments.

Market Positioning and Competition

The startup is already gaining traction, with a client roster that includes industry players like RingCentral and Truecaller. Kamath views any company operating in the customer support space—including emerging firms like Sierra and Decagon—as a potential partner. When asked why these well-funded support platforms wouldn't simply build their own voice models, Kamath noted that focusing on voice engineering would be a significant distraction from their core business objectives.

Smallest.ai enters a competitive landscape populated by heavyweights like ElevenLabs and Cartesia, as well as regional innovators like Sarvam. However, while many competitors diversify their efforts into audio dubbing or podcast generation, Smallest.ai remains laser-focused on its singular goal.

"We want our models to break the Turing test. You should speak to our model and not know it’s AI or human. That’s the sole focus of the company," Kamath stated.

By prioritizing the nuances of human speech over raw computational scale, Smallest.ai is positioning itself to define the next generation of enterprise voice interaction.

#model