mirror of
https://github.com/storytold/storyteller-ml.git
synced 2026-10-09 00:09:55 +00:00
Update ReadMe
This commit is contained in:
+11
-12
@@ -35,38 +35,37 @@ python fakeyou_infer.py [OPTIONS]
|
||||
|
||||
### Options
|
||||
|
||||
- `--use_pretrained_gpt`: Use pretrained GPT model. If enabled, the script will use the default pretrained models for GPT and SoVITS and enable the ZeroShot capabilities.
|
||||
- `--use_pretrained_gpt`: Use pretrained GPT model. If enabled, the script will use the default pretrained models for GPT and SoVITS.
|
||||
- `--use_pretrained_sovits`: Use pretrained SoVITS model. If enabled, the script will use the default pretrained models for GPT and SoVITS.
|
||||
- `--gpt_model`: Path to GPT model checkpoint. Required if `--use_pretrained_gpt` is not enabled.
|
||||
- `--sovits_model`: Path to SoVITS model checkpoint. Required if `--use_pretrained_gpt` is not enabled.
|
||||
- `--ref_audio`: Path to reference wav file. Required unless `--ref_free` is enabled.
|
||||
- `--ref_text`: Path to reference text file. Required unless `--ref_free` is enabled.
|
||||
- `--sovits_model`: Path to SoVITS model checkpoint. Required if `--use_pretrained_sovits` is not enabled.
|
||||
- `--ref_audio`: Path to reference wav file. Required.
|
||||
- `--ref_text`: Path to reference text file (optional). If left blank `""` or not passed, it will disable the reference text and use the reference audio.
|
||||
- `--target_language`: Language of the target text. Choices: `chinese`, `english`, `japanese`, `chinese+english`, `japanese+english`, `automatic`.
|
||||
- `--ref_language`: Language of the reference text. Choices: `chinese`, `english`, `japanese`, `chinese+english`, `japanese+english`, `automatic`. Required unless `--ref_free` is enabled.
|
||||
- `--ref_language`: Language of the reference text. Choices: `chinese`, `english`, `japanese`, `chinese+english`, `japanese+english`, `automatic`.
|
||||
- `--target_text`: Path to the target text file. Required.
|
||||
- `--how_to_cut`: How to cut the text for synthesis. Default: `No slice`. Options: `No slice`, `Slice once every 4 sentences`, `Cut per 50 characters`, `Slice by Chinese punct`, `Slice by English punct`, `Slice by every punct`.
|
||||
- `--top_k`: Top K sampling. Default: 20.
|
||||
- `--top_p`: Top P sampling. Default: 0.6.
|
||||
- `--temperature`: Sampling temperature. Default: 0.6.
|
||||
- `--ref_free`: Perform reference-free TTS.
|
||||
- `--output_path`: Path to save the output wav file. Default: `output`.
|
||||
|
||||
### Example Commands
|
||||
|
||||
|
||||
#### Using Custom Models with Reference Audio
|
||||
|
||||
```sh
|
||||
python fakeyou_infer.py --gpt_model GPT_weights/your_gpt_model.ckpt --sovits_model SoVITS_weights/your_sovits_model.pth --ref_audio data/voice/sponge.wav --ref_text data/voice/ref.txt --target_text data/voice/input.txt --target_language english --ref_language english --how_to_cut Slice once every 4 sentences --output_path output/
|
||||
python fakeyou_infer.py --gpt_model GPT_weights/your_gpt_model.ckpt --sovits_model SoVITS_weights/your_sovits_model.pth --ref_audio data/voice/sponge.wav --ref_text data/voice/ref.txt --target_text data/voice/input.txt --target_language english --ref_language english --how_to_cut "Slice once every 4 sentences" --output_path output/
|
||||
```
|
||||
|
||||
#### Using Custom Models Without Reference Audio
|
||||
#### Using Custom Models Without Reference Text
|
||||
|
||||
```sh
|
||||
python fakeyou_infer.py --gpt_model GPT_weights/your_gpt_model.ckpt --sovits_model SoVITS_weights/your_sovits_model.pth --target_text data/voice/input.txt --target_language english --ref_free --how_to_cut "Slice once every 4 sentences" --output_path output/
|
||||
python fakeyou_infer.py --gpt_model GPT_weights/your_gpt_model.ckpt --sovits_model SoVITS_weights/your_sovits_model.pth --ref_audio data/voice/sponge.wav --target_text data/voice/input.txt --target_language english --ref_language english --how_to_cut "Slice once every 4 sentences" --output_path output/
|
||||
```
|
||||
|
||||
#### Using Pretrained Models for using ZeroShot capabilities
|
||||
#### Using Pretrained Models
|
||||
|
||||
```sh
|
||||
python fakeyou_infer.py --use_pretrained_gpt --ref_audio data/voice/sponge.wav --ref_text data/voice/ref.txt --target_text data/voice/input.txt --target_language english --how_to_cut Slice once every 4 sentences --output_path output/
|
||||
python fakeyou_infer.py --use_pretrained_gpt --use_pretrained_sovits --ref_audio data/voice/sponge.wav --ref_text data/voice/ref.txt --target_text data/voice/input.txt --target_language english --ref_language english --how_to_cut "Slice once every 4 sentences" --output_path output/
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user