About this video
In this video, I experiment with voice cloning and try to figure out if it’s possible to:
— create a digital deputy — generate HR responses in my voice — and in general, understand where the boundary of ethics lies — I give an example with my re-voicing of World of Warcraft
The code from the stream is in the pinned post.
Let’s analyze:
— the concept of digital portraits — multi-vector models and RAG — Qwen3-TTS vs Coqui — LoRA, Flash Attention, and accelerating generation — what is cheaper: cache or loading the model into RAM — practical limitations and MVP
This is not a tutorial, but a live engineering experiment from a stream — with fallacies, memes, and strange results (Starcoder, hello).
Inside:
• short vs. long sample for cloning • text generation for voice • experiment with an HR response • reflections on ethics • ideas for application
The video is based on a stream — without cutting out context and with explanations along the way.
If you: — want to gain a deeper understanding of programming — feel that “code is not everything” — or are simply curious about how everything is connected
Join the streams, subscribe to the channel and blog on Telegram.
00:00:00 – Compilation of funny moments 00:01:00 – Problem statement 00:01:27 – How to re-dub World Of Warcraft into Russian using AI? 00:03:10 – Concept of digital portraits and substitutions, multi-vector models, and RAG 00:12:45 – Formulating problems, roadmap, and MVP 00:15:00 – Attributes of a digital portrait 00:15:59 – We analyze Voice Memos and create a short sample for cloning 00:17:23 – I repeat the problem statement and project goals 00:20:42 – About Qwen3-TTS in relation to coqui-ai/TTS 00:21:55 – We launch and prepare the Python code 00:23:59 – We listen to voice cloning based on a short sample 00:24:59 – About Flash Attention, LoRA, Mac Studio M1 Max, and accelerating voice cloning 00:26:45 – We record a sample based on the text of the book “The Hobbit” 00:29:28 – We record a generated piece for pattern matching and comparison 00:30:15 – We clone the voice for a long sample 00:33:20 – We listen to voice cloning based on a long sample 00:34:47 – We analyze the results 00:36:40 – We clarify the project goals 00:38:00 – How it can work with video 00:38:45 – Meme about digital substitutes 00:39:00 – About the applicability of RCTF, tonality, Jung, and Tolkien in the approach to generating text for voice cloning 00:42:20 – Random asks “what’s going on?” 00:43:00 – Experiment to generate a text response from HR for a technical question 00:44:15 – We apologize to the neural network 00:44:20 – About DDD hatred 00:45:51 – Starcoder went crazy and generated a pasta about a balcony, but about elixir with profanity 00:47:10 – We generate an audio clone for the HR’s answer about elixir 00:48:30 – Sometimes, fully loading the model into RAM for each generation is cheaper than caching 00:49:20 – We listen to voice cloning for the HR’s answer about elixir 00:50:00 – We listen to all voice clones 00:50:45 – Demand for cloning for World Of Warcraft 00:51:30 – Ideas for applying voice cloning 00:52:48 – Why am I even talking about this and streaming? 00:53:30 – About the ethical issues of voice cloning 00:54:00 – Applicability of the project
Where can I be found?
▶ Twitch (streams): https://www.twitch.tv/marat_zimnurov
▶ Telegram — about programming and digests: https://t.me/digitable_blog
▶ A post about color theory that I mention in the video: https://t.me/digitable_blog/30
▶ Projects and services: https://digitable.ru/
▶ GitHub: https://github.com/the-homeless-god
All other contact details can be found in the channel header.
You can support the channel by subscribing here and on Telegram https://t.me/digitable_blog or on Twitch via the donation button. If there is a reason and demand, I will create a Boosty page in the future, but for now, I will take it easy.
Watch on Rutube: https://rutube.ru/video/42489a46a3416720a21d02c883508f6c/