Elektrine lite

← Feed

@delta@mk.absturztau.be

Post #2756069

2026-05-12 02:53 UTC

being able to mess with parameters to set a variety of different voices is so good and every TTS engine from now on needs this enough AI TTS engines that try too hard to be 'realistic' and 'natural' and are just 1 model = 1 voice and if you want a different voice you download another model that is also just 1 voice give me TTS models with dozens of parameters to make whatever variety of voices i want

Replies (2)

  • @delta@mk.absturztau.be 2026-05-12 04:05

    i can sort of imagine how creating a similar voice engine might work, but i am 0% qualified to even consider implementing such a thing in my current state but i can imagine doing it somewhat modularly, and taking some notes from how good ol utau worked have a base sound sample source (could be generated, or using a bank of prerecorded phoneme samples) a sample transformation engine (like utau) and combine that with some kind of system that takes input text and transforms it into the list of phonemes with pitch/intonation data for the sample transformation engine to use

    Open ##2756070

  • @piku@blahaj.zone 2026-05-12 02:59

    @delta@mk.absturztau.be https://github.com/dylanpdx/talkmodachi edit: oop nvm u said living the dream

    Open ##2756071