readme/notes

This commit is contained in:
Brandon Thomas
2022-02-28 05:46:37 -05:00
parent ba326bfcfd
commit b50205847c
3 changed files with 52 additions and 1 deletions
+8
View File
@@ -1,3 +1,4 @@
*.pyc
*.swo
*.swp
*secret*sh
@@ -13,3 +14,10 @@ production_database_secret_query.sh
secrets.toml
secrets.yaml
target/
# Models for testing
future-todo/ControllableTalkNet/models
tts/hifigan/hifisr
# Other garbage to prune/remove
tts/hifigan/LJSpeech-1.1/
+35 -1
View File
@@ -1,6 +1,40 @@
tts-core
========
The core TTS models that power FakeYou.
I'm in the progress of pulling these out of the service monorepo so we can iterate faster.
TODO
----
Core cleanup
- [ ] Cleanup Tacotron code
- [ ] Integrate WaveGlow and HifiGan more cleanly
- [ ] Define the HTTP server interface
- [ ] Define the model checking interface
- [ ] Clean up in-memory LRU caching
- [ ] There are two copies of Nemo. Why? Clean this up.
- [ ] Deployability
- [ ] Clean up requirements.txt and upgrade packages
- [ ] Clean up Dockerfile
Make the models sound better
- [ ] Arpabet
- [ ] G2P
- [ ] Arpabet custom dictionary
- [ ] User-supplied Arpabet strings in braces
- [ ] Word break down 888 -> eight hundred eighty eight
- [ ] Support I18N
- [ ] Accent character support
Introduce new models
- [ ] TalkNet
- [ ] GlowTTS + MelGan (vocoder)
- [ ] Make these work under the same venv
Lots more I haven't covered yet...
+9
View File
@@ -0,0 +1,9 @@
TalkNet
=======
TalkNet is not currently in a usable state. Ideally we can have both Tacotron and
Talknet models operate at the same time from the same uniform codebase.
The reason we want them to work together is so that we can leverage the same GPU
resources and use LRU in-memory caching.
This copy of TalkNet was taken from a public MIT-licensed repo and notebook (TODO: link)