Then I ask ChatGPT and it recommended GPT-SoVITS to me, after I looked up the price of GPU servers, I decided to run the program on my local MacBook ,then use it to deal with the cloning voice generation requests from my website server, the general steps are like:
- I submitted my article
- Deepseek receives the article and generate a Japanese summary in 50 words and sends back to the server
- The server sent the text to my local MacBook and MacBook generate a .m4a file . Then the server gets the audio and bounds the voice file with summary
About the voice training, I extracted a 4-min audio file containing only Maki’s voice from the original anime and then use it to train the model.
What I want to suggest you who want to do the same is that you’d better not leave long silent sections in the audio you extracted since it will lead to ur final output voice containing meaningless silent parts or lots of mistakes in dealing with the sentence segmentation, and, sometimes, the voice will even be short of half part of the original text. That’s why I had to train my model twice.
That’s all!
So happy to hear maki’s voice
とても楽しいだね
おやすみなさい、地球人