mirror of
https://github.com/wassname/TTS.git
synced 2026-09-09 11:16:00 +08:00
🔥 XTTS implementation
This commit is contained in:
@@ -53,6 +53,7 @@
|
||||
models/overflow.md
|
||||
models/tortoise.md
|
||||
models/bark.md
|
||||
models/xtts.md
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Bark 🐶
|
||||
# 🐶 Bark
|
||||
|
||||
Bark is a multi-lingual TTS model created by [Suno-AI](https://www.suno.ai/). It can generate conversational speech as well as music and sound effects.
|
||||
It is architecturally very similar to Google's [AudioLM](https://arxiv.org/abs/2209.03143). For more information, please refer to the [Suno-AI's repo](https://github.com/suno-ai/bark).
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Tortoise 🐢
|
||||
# 🐢 Tortoise
|
||||
Tortoise is a very expressive TTS system with impressive voice cloning capabilities. It is based on an GPT like autogressive acoustic model that converts input
|
||||
text to discritized acouistic tokens, a diffusion model that converts these tokens to melspeectrogram frames and a Univnet vocoder to convert the spectrograms to
|
||||
the final audio signal. The important downside is that Tortoise is very slow compared to the parallel TTS models like VITS.
|
||||
|
||||
@@ -0,0 +1,108 @@
|
||||
# ⓍTTS
|
||||
ⓍTTS is a super cool Text-to-Speech model that lets you clone voices in different languages by using just a quick 3-second audio clip. Built on the 🐢Tortoise,
|
||||
ⓍTTS has important model changes that make cross-language voice cloning and multi-lingual speech generation super easy.
|
||||
There is no need for an excessive amount of training data that spans countless hours.
|
||||
|
||||
This is the same model that powers [Coqui Studio](https://coqui.ai/), and [Coqui API](https://docs.coqui.ai/docs), however we apply
|
||||
a few tricks to make it faster and support streaming inference.
|
||||
|
||||
### Features
|
||||
- Voice cloning with just a 3-second audio clip.
|
||||
- Cross-language voice cloning.
|
||||
- Multi-lingual speech generation.
|
||||
- 24khz sampling rate.
|
||||
|
||||
### Code
|
||||
Current implementation only supports inference.
|
||||
|
||||
### Languages
|
||||
As of now, XTTS-v1 supports 13 languages: English, Spanish, French, German, Italian, Portuguese,
|
||||
Polish, Turkish, Russian, Dutch, Czech, Arabic, and Chinese (Simplified).
|
||||
|
||||
Stay tuned as we continue to add support for more languages. If you have any language requests, please feel free to reach out.
|
||||
|
||||
### License
|
||||
This model is licensed under [Coqui Public Model License](https://coqui.ai/cpml).
|
||||
|
||||
### Contact
|
||||
Come and join in our 🐸Community. We're active on [Discord](https://discord.gg/fBC58unbKE) and [Twitter](https://twitter.com/coqui_ai).
|
||||
You can also mail us at info@coqui.ai.
|
||||
|
||||
Using 🐸TTS API:
|
||||
|
||||
```python
|
||||
from TTS.api import TTS
|
||||
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v1", gpu=True)
|
||||
|
||||
# generate speech by cloning a voice using default settings
|
||||
tts.tts_to_file(text="It took me quite a long time to develop a voice, and now that I have it I'm not going to be silent.",
|
||||
file_path="output.wav",
|
||||
speaker_wav="/path/to/target/speaker.wav",
|
||||
language="en")
|
||||
|
||||
# generate speech by cloning a voice using custom settings
|
||||
tts.tts_to_file(text="It took me quite a long time to develop a voice, and now that I have it I'm not going to be silent.",
|
||||
file_path="output.wav",
|
||||
speaker_wav="/path/to/target/speaker.wav",
|
||||
language="en",
|
||||
decoder_iterations=30)
|
||||
```
|
||||
|
||||
Using 🐸TTS Command line:
|
||||
|
||||
```console
|
||||
tts --model_name tts_models/multilingual/multi-dataset/xtts_v1 \
|
||||
--text "Bugün okula gitmek istemiyorum." \
|
||||
--speaker_wav /path/to/target/speaker.wav \
|
||||
--language_idx tr \
|
||||
--use_cuda true
|
||||
```
|
||||
|
||||
Using model directly:
|
||||
|
||||
```python
|
||||
from TTS.tts.configs.xtts_config import XttsConfig
|
||||
from TTS.tts.models.xtts import Xtts
|
||||
|
||||
config = XttsConfig()
|
||||
config.load_json("/path/to/xtts/config.json")
|
||||
model = Xtts.init_from_config(config)
|
||||
model.load_checkpoint(config, checkpoint_dir="/path/to/xtts/", eval=True)
|
||||
model.cuda()
|
||||
|
||||
outputs = model.synthesize(
|
||||
"It took me quite a long time to develop a voice and now that I have it I am not going to be silent.",
|
||||
config,
|
||||
speaker_wav="/data/TTS-public/_refclips/3.wav",
|
||||
gpt_cond_len=3,
|
||||
language="en",
|
||||
)
|
||||
```
|
||||
|
||||
|
||||
## Important resources & papers
|
||||
- VallE: https://arxiv.org/abs/2301.02111
|
||||
- Tortoise Repo: https://github.com/neonbjb/tortoise-tts
|
||||
- Faster implementation: https://github.com/152334H/tortoise-tts-fast
|
||||
- Univnet: https://arxiv.org/abs/2106.07889
|
||||
- Latent Diffusion:https://arxiv.org/abs/2112.10752
|
||||
- DALL-E: https://arxiv.org/abs/2102.12092
|
||||
|
||||
|
||||
## XttsConfig
|
||||
```{eval-rst}
|
||||
.. autoclass:: TTS.tts.configs.xtts_config.XttsConfig
|
||||
:members:
|
||||
```
|
||||
|
||||
## XttsArgs
|
||||
```{eval-rst}
|
||||
.. autoclass:: TTS.tts.models.xtts.XttsArgs
|
||||
:members:
|
||||
```
|
||||
|
||||
## XTTS Model
|
||||
```{eval-rst}
|
||||
.. autoclass:: TTS.tts.models.xtts.XTTS
|
||||
:members:
|
||||
```
|
||||
Reference in New Issue
Block a user