Guides · 4 min read

322 Voices, 142 Languages, No Server: What It Takes to Fit That Much Speech in a Browser Tab

PrivateAI's text-to-speech tool covers 322+ voices across 142 languages entirely on-device. Here's why voice and language breadth — not raw speed — is the hard problem for browser-based TTS, and how that number stacks up against the open-source and paid alternatives.

322 Voices, 142 Languages, No Server: What It Takes to Fit That Much Speech in a Browser Tab

The Number That’s Easy to Skim Past

Most write-ups of browser-based AI tools reach for the same headline: nothing gets uploaded, there’s no server in the loop. PrivateAI fits that framing, but its text-to-speech tool carries a second number that’s worth stopping on: 322+ voices across 142 languages, generated entirely on the user’s own GPU, alongside a separate speech-to-text tool covering 99 languages. That’s not the spec sheet of a browser demo bolted onto a “no upload” pitch. It’s closer in scale to what a metered cloud API offers — running with no account, no server, and no bill.

Why Breadth Is the Hard Part

Speed is what most coverage of on-device AI focuses on, because it used to be the blocker: WebGPU only reached a majority of browsers in the last couple of years, and before that, running any real model client-side meant a slow, janky demo. But speed and breadth are different engineering problems. A single-language, single-voice model is a solved problem on consumer hardware. Covering 142 languages with hundreds of distinct voices means either bundling many separate models — one per language or language family — or building one model general enough to hold that much phonetic and prosodic variation without ballooning past what a browser tab can reasonably download and hold in memory. Neither path is free, and it’s the reason most local, browser-deployable TTS projects ship a fraction of that language count.

What the Open-Source Landscape Actually Ships

The two engines most commonly cited for running speech synthesis client-side illustrate the gap. Kokoro, an 82-million-parameter model released under Apache 2.0 that reached the top of the TTS Arena leaderboard, ships 54 voices across 9 language variants — English, Japanese, Mandarin, Spanish, French, Hindi, Italian, Portuguese, and Korean — with its weights cached in browser storage and synthesis running on WebGPU or WebAssembly. Piper, an established local TTS engine also used in projects like Home Assistant, covers more ground — 100+ voices across 30+ languages — running through WebAssembly rather than needing a GPU at all. Both are real, actively maintained projects, and both stop well short of 142 languages. That’s the comparison that makes PrivateAI’s number worth noting: it isn’t just “TTS that happens to run locally,” it’s local TTS with language coverage closer to a general-purpose cloud voice API than to the open-source engines built for the same browser-only constraint.

What the Paid Alternative Costs

The other side of that comparison is what this kind of coverage normally costs when a vendor runs it in the cloud instead. ElevenLabs, the voice-AI platform most often benchmarked for breadth and quality, charges $0.05 per 1,000 characters on its Flash/Turbo models and $0.10 per 1,000 on its Multilingual v2/v3 models, with its largest voice library running to roughly 3,000 voices across its paid tiers. Azure’s Text-to-Speech service prices its neural voices at $16 per million characters, rising to $22 per million for its newer HD tier, after a 500,000-character monthly free allowance. Both are reasonably priced as cloud APIs go — but both are metered, because every character synthesized costs the vendor real GPU time somewhere. A tool that puts a comparable language count on a user’s own GPU sidesteps that meter by construction, not by discounting it.

The Trade a Browser Tab Actually Makes

None of this means on-device TTS is a strictly better architecture — it’s a different one, suited to a specific job. The reason narrow, well-scoped tools like text-to-speech, OCR, and background removal are what run well in a browser tab, rather than an open-ended reasoning model, is that they’re small enough to fit the memory and compute a consumer GPU actually has. Voice and language breadth pushes directly against that constraint, which is exactly why it’s the harder number to move — easier to make a model fast than to make it fluent in 142 languages without it outgrowing the device it has to run on.

Why the Number Matters Beyond the Spec Sheet

For a company whose other five products build tools for people working across languages by necessity — manufacturing teams reading drawings in Japanese, meeting notes captured across markets, legacy systems built for a different encoding era — a browser tool that reads text aloud in 142 languages without a server or a subscription is a small but concrete example of the same bet: that on-device processing can match, not just approximate, what a metered cloud service offers.


Curious what text-to-speech sounds like with no server, no account, and 142 languages to choose from? Try PrivateAI free — or reach out at [email protected].

See Our Work

From MinuteAI to AgentKits — explore the products and projects we've shipped.

View Portfolio

Related Articles