Short version: Something has been accepted into the Cartesia Startups Grant Program. We're thrilled — and more to the point, it lands exactly where our users have been pushing us.
Voice has quietly become one of the most requested things people build on Something. Not chatbots with a speaker icon bolted on — actual phone-call-shaped software: intake lines, practice partners, companions, screening calls. The multilingual wellness companion we wrote up recently was built in an afternoon and speaks six languages, and it's not an outlier.
Voice is also where the technical bar is highest. Nobody notices a chat interface that takes a moment to respond. Everybody notices a voice that pauses too long before answering, because that pause is the difference between a conversation and a transaction.
Why Cartesia
Cartesia builds real-time speech models on a different architecture from most of the field — designed for live interaction rather than adapted to it after the fact. Their stack handles both sides of a conversation: turning text into speech, and turning speech back into text as it streams.
Two things about that matter for the kind of agents people build here.
Latency you don't hear. In a voice agent, the time spent generating speech is time the caller spends sitting in silence wondering whether the line dropped. Models built for streaming from the ground up close that gap in a way that adapted ones generally can't.
Languages, plural. This is the piece we care about most. Our own most-shared voice build was a wellness companion speaking Hindi, Bengali, Tamil, Telugu, Marathi and English — and the hardest part of multilingual voice has never been the logic. It's finding speech models that don't treat everything outside English as a degraded special case.
The gap between the language you're offered and the language you actually think in is where most voice products lose the people who need them most.
What it means going forward
Nothing you have to do. Voice agents built on Something keep working the way they do today — describe what you want, pick your languages, get a live URL.
What the grant buys is room to make that better without watching a meter while we experiment. The work it funds is aimed at four things: faster responses on live calls, so agents interrupt and reply more like a person does; broader language coverage with the same quality bar applied outside English; more voices, with finer control over tone and pacing; and enough headroom to test at real concurrency before a launch rather than after it.
That last one is the unglamorous one that matters most. Voice agents fail in a specific way — fine in testing, fine with five users, then somebody shares the link and a crowd calls at once.
Being straight with you
This is a grant, not an integration that shipped this morning. We're announcing the partnership and the work it funds — as pieces land in the product, we'll write them up here with the same treatment as everything else on this blog.
Thanks
To the team at Cartesia for backing an early product, and for a program genuinely built for the stage we're at rather than one that assumes you already have a Series A.
If you're building voice agents — in one language or six — Something has a free tier. And if you want to build on Cartesia directly, their docs are a good place to start.
Build a voice agent this afternoon
Describe it in a sentence, pick your languages, get a live URL. Free tier, no card required.
