just remember: despite they would have you believe they are are a united front... they are better thought of as a "loose federation of warring tribes".
I’m reducing into silliness I’m sure, I don’t have any breadth or depth of knowledge here.
Edit: wonder why Apple is ostensibly different. MS seems similar, and don’t know enough about AWS—maybe I’ve seen complaints about them being disjointed, but not as much as Google & Microsoft.
My personal account, 3.6 Flash Lite and 3.5 Thinking.
Meanwhile, I can go hog wild and drain my bank account on GCP. I don’t though, because my family has to eat.
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: https://deepmind.google/models/gemma/gemma-4/
[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
I found all these on this subreddit: https://www.reddit.com/r/TextToSpeech/
I've also see comments like "Microsoft Edge's read-aloud feature is amazing for TTS" but I haven't tried it myself.
Personally I use Kokoro with a python front-end on my Macbook, I linked to it in this comment: https://news.ycombinator.com/item?id=49818923 it outputs MP3, so I just copy those to my phone and listen as audiobooks.
How long is this stored? What could go wrong? :P
Having a voice under 1Mo is crazy, even if it sounds robotic.
Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?
And there's no pricing listed anywhere.
I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.
If you're after quality then Qwen3 TTS is a very good model esp. if you take some effort to craft a voice file. It is slow, so isn't practical for real-time voices (like assistants). It can also occasionally switch to a different voice to the one provided, so you may want to break up the text being processed.
I've not yet tried other recent/recentish models.
If you are after performance then two options from older models are:
1. flite with a HTS (Hidden Markov Model) voice like cmu_us_rms (male) or cmu_us_slt (female);
2. espeak/espeak-ng with an MBROLA (an Overlapped Add model) voice (mb-us1, mb-de5-en, etc.).
Alternatively, you could try using Qwen3 TTS or over voice changing model with the CMU Arctic (http://www.festvox.org/cmu_arctic/) voice data which includes audio for the rms and slt voices among others.
If you're feeling adventurous you could also try fine tuning one of the TTS models on that data to create a custom voice, though the data is likely to be in the training data for the voices, so using an audio sample may be sufficient depending on the TTS model.
https://tools.simonwillison.net/gemini-tts-playground#compos...
I listened to some of the voices. The male voices are believable while all the female voices sound the same and artificial. For some reason, it also reminds me of the voices in Toy Story movies.
Bias in the training data?
Price per hour:
- 3.8 Flash TTS, standard: $0.81
- 3.8 Flash TTS, batch: $0.41
- 3.8 Flash‑Lite TTS, standard: $0.54
- 3.8 Flash‑Lite TTS, batch: $0.27
They all sound like Americans putting in their best fake British accent.
$0.50 per hour pricing could last a long time with back and forth conversation use.
Nice work. It’s awesome to have these capabilities so close at hand and so trivially easy to integrate with.
All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters.
They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.
I can see they want to create an ecosystem, but I see no focus in any one area.
On top of that I would imagine the research side of Google might disk over things from one modality being useful to another - something like an audio optimising or memory optimising for text to speech could maybe also be useful for translation or world models. Etc etc
The original is the best: https://youtu.be/qxWwEPeUuAg
[deleted]
Also HN: Fuck those voice actors and their careers.
[dead]
[dead]
Also famously, the million types of context menus in Windows.
Google and Microsoft are in the disjointed camp. Apple is not. AWS, I do not know.
Googly way to put it! Nice. Might steal that :)
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
[deleted]
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.
[dead]
[dead]
https://abc.xyz/investor/board-and-governance/google-code-of...
ctrl/cmd+f "evil"
I don't know why this argument is brought up all the time anyway, it literally means nothing. They can name themselves "Don't Be Evil Inc" and continue to do evil stuff cause evil isn't an objective measure. If squeezing juice out of puppies made money, any business can just say it's "not evil."
And if you think they're evil why would you trust them to follow their own guideline of not doing evil? An evil corp would be more likely to just hide behind that phrase, not quietly remove it as some subtle hint that they want to be openly and proudly evil all of a sudden.
I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?
MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?
Good idea, not something I've considered yet. Wouldn't take much to add it since there's already a feature for selecting a quotation and assigning it an intonation. Same infrastructure could be reused to select arbitrary text and assign stage directions.
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
[dead]
Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS
I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/
I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!
My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...
Please let me know if you have any questions or feedback.
It was especially nice during a bike trip along the Rhine, I listened to a lot of the history of the industrial area and its cities.
Different models have different strengths. If you throw an entire ebook at a model you're going to get a different result than if you craft a perfect 10 second sentence with a model that supports voice direction and emotion tags, so you should try a bunch depending on your use case.
https://aistudio.google.com/generate-speech?model=gemini-3.8...
Full session here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Um, what?
Imagine someone showing you that they've trained their dog to hold a paintbrush and paint. There would be no contradiction between "this is incredible" and "these paintings suck".
[deleted]
They sounded exactly like that person... but it's like they're angry at me.
Great toys but not production ready.
Official pricing can be seen here: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-fla...
Basically spread many AI's everywhere, get people develop/user their ecosystem, and lock them in eventually.
But I find their models' intelligence lacking still.
I don't ever use their built-in gemini features (i have paid gmail) because they don't work well.
e.g. I ask gemini to format my Google docs per Google's material design spec with spacing, etc. It does a real bad job. Many times it does it line by line, and when I finaly get it do it for the whole doc, it does it sloppy, and extremely slow (takes 5 minutes for 10 page doc)
And yeah integration with Google Workspace is still surprisingly not as good as you would expect.
At best, good, but not great.
Onomatopoeia? Sure it is there, and some fillers (or whatever you call those little sounds). But moans?
I would not want that in my product.
so one of my interests is reducing the payload size for video games
the vast majority of the image sizes have been audio recordings, and its been that way in different qualities for the last two decades. this is still the case as more varied and comprehensive audio is pursued by studios at unfathomable expense and still failing to cross a bar of realism
good voice models are just a few gigabytes in comparison and can supplant all of that, and be run locally at this point. Future ubiquitous hardware configurations in consumer devices will make inference dedicated and computationally cheaper and faster
although AAA studios are hamstrung and will be deeply unpopular if they stopped booking voice actors
everyone else who would have never had the capital for voice actors will just use this and have richer experiences until they themselves are AAA studios from the market buying their rich experiences
this will vastly supplant the assumed and uninspired “tricking humans” use case from that video. once it crosses a threshold of ease, the applications will expand
No que por los dos?
Don’t worry, we will be able to destroy careers and scam people at the same time. We don’t have to pick and choose!
but let’s not pretend solo developers were ever going to have a voice actor suite and hire that talent en masse
the transactions were never going to happen
and now the outcome will be better than the studios that are making those transactions
1. Vibe code a local recording dashboard with mic selection, record/replay, and reading prompts. I ended up with about 12 minutes of recordings which was like a 150 or something clips.
2. Review the transcripts, trim excess silence, normalize levels, reduce background hiss. Used whisper to help find flubbed word substitutions (happens).
3. Fine-tune Qwen3-TTS 1.7B on my RTX 3090. This took some debugging because the trainer/runtime combination had misaligned loss targets and training/inference mismatches.
4. Vibe code listening dashboards to compare checkpoints and learning rates until I had something that seemed reasonable.
It was honestly pretty vibe coding friendly.
[deleted]
But I think you’re dancing on the grave of an entire industry, and every person that’s going to lose their life savings due to this.
the market doesn’t want 15 year lead times and overly expensive and delayed games that don't experiment on anything
every friction plaguing the industry is solved by distributed indie developers being able to make richer experiences faster and cheaper