Why Smallest.ai Is Betting Small Models Can Fix Voice AI

Smallest.ai has raised $13 million in Series A funding to build faster, more natural voice agents for enterprise customers. Its approach centers on small, specialized voice models that can listen, think, and speak in real time, while handing harder questions to larger models when needed.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

This is mostly a routine funding and product story about faster enterprise voice agents, with only mild concerns around humanlike automation.

Why Smallest.ai Is Betting Small Models Can Fix Voice AI

Voice AI is getting better at handling customer support, but the experience still often gives itself away. Small pauses, awkward turn-taking, accent problems, and trouble with noisy calls can make a machine feel unmistakably unlike a person.

Smallest.ai, a startup founded in late 2024, is trying to solve that problem from a different direction. Instead of treating faster large language models as the main answer, the company is building smaller, specialized models designed specifically for real-time human conversation.

A New Funding Round For Faster Voice Agents

Smallest.ai has raised $13 million in a Series A round led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital. The round brings the startup’s total funding to over $21 million.

The company’s goal is direct: make speaking with an AI agent feel indistinguishable from speaking with a human. That is a high bar, especially in customer support, where people expect quick answers, natural interruptions, and responses that match the rhythm of a live call.

The startup’s founder and CEO, Sudarshan Kamath, argues that the next step for voice agents is not simply a bigger model. Smallest.ai is developing a small voice model that is meant to work more like a person in conversation: listening, thinking, and speaking at the same time.

“While I’m speaking to you, you’re already thinking, and you might interrupt me if I talk for too long,” Sudarshan Kamath (pictured left), founder and CEO of Smallest.ai, told TechCrunch, adding that this is exactly how the startup’s model is designed to work.

Why Latency Feels Worse In Speech

In a text chat, a delay can feel normal. A user types a question, the system processes the prompt, and an answer appears. In a phone conversation, even a short pause can feel strange because people do not usually wait for a full block of speech before beginning to understand it.

Kamath described the difference by contrasting conversation with the way a large language model typically works. A large model receives a full prompt and then starts producing a response. That flow can work in text, but it does not match the flow of spoken interaction.

“The way an LLM works is you give it an entire prompt, and then it starts thinking,” Kamath said.

For Smallest.ai, that gap is the central product problem. A convincing voice agent needs to respond with minimal lag, follow the user’s speech as it happens, and avoid the mechanical feeling that comes from waiting too long before answering.

The company says its model acts as a real-time intelligence layer for natural customer conversations on specific topics. The key promise is not just that the system can talk, but that it can keep up with the pace and interruptions of an actual conversation.

A Two-Model Future For Customer Support

Smallest.ai’s approach does not remove large foundational models from the workflow. Instead, it assigns different jobs to different types of models.

The small voice model handles the real-time exchange when a customer is asking about topics inside its limited knowledge base. If the conversation moves outside that scope, the system can hand the question to a large foundational model and briefly put the customer on hold to “research” the issue.

Kamath believes this pattern could become common across AI agents: one small voice model for the live conversation, and an “offline” LLM that is called when the problem becomes more complex.

That division is important because it frames voice AI as more than speech layered on top of a chatbot. In this view, the real-time model is responsible for the human-feeling part of the interaction, while the larger model is reserved for heavier reasoning or knowledge work when needed.

Where Smallest.ai Is Focusing

Smallest.ai is concentrating on voice-specific challenges instead of trying to build a broad foundational model. The company is working on nuances that matter in live calls, including diverse accents, dozens of languages, and noisy environments.

Its existing customers include companies in the voice space, including RingCentral and Truecaller. Kamath also said any customer support company, including newer ones like Sierra and Decagon, could be a potential customer for the startup.

That customer base helps explain the company’s positioning. Many customer support companies need strong voice capabilities, but Kamath argues that building a dedicated voice model may pull them away from their core business.

“Extremely good at doing voice is a distraction from their core business.”

Smallest.ai is competing with ElevenLabs, Cartesia, and regional players like Sarvam. But the company is drawing a narrower boundary around its product than some competitors. While other voice AI companies also apply the technology to areas such as audio dubbing and podcasting, Smallest.ai is focused strictly on real-time conversational voice agents for enterprise customers.

The Bigger Bet Behind Small Voice Models

The bet behind Smallest.ai is that enterprise voice AI will be judged less by raw model size and more by conversational realism. A support agent that understands a customer but responds too slowly can still feel broken. A system that speaks well but struggles with accents, background noise, or interruptions can still remind users that they are dealing with software.

That makes real-time performance a product feature, not just a technical detail. For customer support, the model needs to make the call feel fluid while still knowing when to pause and route harder questions to a larger system.

Smallest.ai’s stated ambition is to make the machine disappear from the experience. As Kamath put it, “We want our models to break the Turing test.”

If the company’s model works as intended, the test for voice AI may become simple: whether a customer notices the difference at all.