New practice tool unlocked: Talk synths?

I’m deep into vocaloid, and use a lot of singing voice synthesizers. That was my gateway to learning Japanese. However, I randomly recalled today that talk synthesizers are a thing?! They’re fancier, more customizable versions of TTS software that have a lot of different characters. I know VoiceVOX is free (but is a bit difficult/arcane to get running?) and there are numerous paid ones like A.I.Voice, Cevio/Voisona Talk, Voicepeak, etc. that have pretty big names in the Vocaloid/broader vsynth space

Right now I only have Koharu Rikka for Voicepeak (she was given away free to people who own her Synthesizer V bank) but I’m looking to get more, preferably on the cheap via second-hand activation trading. I think I can hit up a friend to see if they have Frimomen (a voice given away free with basically every Voicepeak purchase), but it would be nice to have a large library.

Talk synths are fun to listen to, and are also good to practice sentence writing in Japanese. It also forces me to use my Japanese IME because it won’t parse romaji correctly.

Thought I would make a post here in case anyone else has had the idea, or wants to practice for themself!

I was just looking at TTS stuff a couple hours ago. Free stuff is a lot better than it used to be when I checked this out years ago. Orca TTS app is a pretty common default on many Linux distributions, but the old standard voices were pretty robotic. The newer voices are starting to actually be decent. It also supports some other synthesizers that are starting to be pretty natural like Piper. I came across a project that cloned Rocky’s voice from Project Hail Mary and was running a Raspberry Pi, amaze amaze.

While looking into those I came across AI model synthesizers like KittenTTS, Kokoro TTS, F5-TTS, and IndexTTS2. These newer AI models can clone voices from as little as 4 seconds of audio, translate languages and keep the voice in the output language, and F5 and Index have emotion settings that actually work.. They’re pretty impressive. I’ve seen AI translation like this years ago, but nothing that could be run locally on comparably small models (seeing stuff like this years ago was one of the things that slowed my Japanese studying actually lol.. until I finally decided I want to actually know it).

I’ve known the vocaloid stuff existed (I watch some vtubers.. and there is Miku..), but haven’t gone looking into how easy they were to get into. I’ll check some of your mentions out

There’s a lot in the way of AI voice cloning (arguably too much :skull: ) but the ones I mentioned are all commercial (or free for Voicevox) and are made with consenting voice providers getting paid for their time. And because it’s commercial (or long-term free) they run pretty well and work locally. A.I.Voice 2 and Voicepeak (Edit: also Voisona and the most recent Cevio banks) are also using AI models of their voice providers, which greatly enhances clarity, but it’s ethical/consensual.

Fun fact, despite being called A.I.Voice, A.I.Voice1 does not actually use AI at all. It’s just made by A.I.Talk, but it uses concatenative synthesis (stringing pre-recorded sounds together, like Utau/Vocaloid 1-5)

Most of the vocaloid things being paid makes sense given their usage. I’ve seen videos related to Miku stuff about what goes into making some of those voices and it’s a huge amount of small sound recordings to make it sound real. There are way more sounds in languages than I had initially thought before looking into them and pitch accent related stuff.

I just want something to create audio to listen to a bit of text stuff while doing other things. It doesn’t have to be amazing and I’m not planning to get into really doing anything else with it, so free options are good enough. I could definitely see these being usable for listening practice, though