A PyTorch implementation of Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis.

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

Tacotron (with Dynamic Convolution Attention)

A PyTorch implementation of Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis. Audio samples can be found here.

^{Fig 1:Tacotron (with Dynamic Convolution Attention).}

^{Fig 2:Example Mel-spectrogram and attention plot.}

Quick Start

Ensure you have Python 3.6 and PyTorch 1.7 or greater installed. Then install this package with:

pip install tacotron

Example Usage

import torch
import soundfile as sf
from univoc import Vocoder
from tacotron import load_cmudict, text_to_id, Tacotron

# download pretrained weights for the vocoder (and optionally move to GPU)
vocoder = Vocoder.from_pretrained(
    "https://github.com/bshall/UniversalVocoding/releases/download/v0.2/univoc-ljspeech-7mtpaq.pt"
).cuda()

# download pretrained weights for tacotron (and optionally move to GPU)
tacotron = Tacotron.from_pretrained(
    "https://github.com/bshall/Tacotron/releases/download/v0.1/tacotron-ljspeech-yspjx3.pt"
).cuda()

# load cmudict and add pronunciation of PyTorch
cmudict = load_cmudict()
cmudict["PYTORCH"] = "P AY1 T AO2 R CH"

text = "A PyTorch implementation of Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis."

# convert text to phone ids
text = torch.LongTensor(text_to_id(text, cmudict)).unsqueeze(0).cuda()

# synthesize audio
with torch.no_grad():
    mel, _ = tacotron.generate(text)
    wav, sr = vocoder.generate(mel.transpose(1, 2))

# save output
sf.write("location_relative_attention.wav", wav, sr)

Train from Scatch

Clone the repo:

git clone https://github.com/bshall/Tacotron
cd ./Tacotron

Install requirements:

pip install -r requirements.txt

Download and extract the LJ-Speech dataset:

wget https://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2
tar -xvjf LJSpeech-1.1.tar.bz2

Download the train split here and extract it in the root directory of the repo.
Extract Mel spectrograms and preprocess audio:

python preprocess.py in_dir=path/to/LJSpeech-1.1 out_dir=datasets/LJSpeech-1.1

Train the model:

python train.py checkpoint_dir=ljspeech dataset_dir=datasets/LJSpeech-1.1 text_dir=path/to/LJSpeech-1.1/metadata.csv

Pretrained Models

Pretrained weights for the LJSpeech model are available here.

Notable Differences from the Paper

Trained using a batch size of 64 on a single GPU (using automatic mixed precision).
Used a gradient clipping threshold of 0.05 as it seems to stabilize the alignment with the smaller batch size.
Used a different learning rate schedule (again to deal with smaller batch size).
Used 80-bin (instead of 128 bin) log-Mel spectrograms.

Acknowlegements

Project details

These details have not been verified by PyPI

Project links

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.1.1

Nov 24, 2020

This version

0.1.0

Nov 13, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tacotron-0.1.0.tar.gz (907.6 kB view hashes)

Uploaded Nov 13, 2020 Source

Built Distribution

tacotron-0.1.0-py3-none-any.whl (910.1 kB view hashes)

Uploaded Nov 13, 2020 Python 3

Hashes for tacotron-0.1.0.tar.gz

Hashes for tacotron-0.1.0.tar.gz
Algorithm	Hash digest
SHA256	`a570334088a2635006e40847e989dc446ab6b766c12b9726115086fb73382d64`
MD5	`09eda31b2d7edef9e96ea58f0adcd1d8`
BLAKE2b-256	`6c4ef5da038e55b6fc0b3aff40c4aba96abc7981f8d91839a3ed84cd5aeebb6d`

Hashes for tacotron-0.1.0-py3-none-any.whl

Hashes for tacotron-0.1.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`18d36ce143a831b4cd79598eabb0e0470a569b1171b0ed48113f4b58eb123484`
MD5	`d3f21f4d717df5db1866442b181ca261`
BLAKE2b-256	`87b524e3e9b6dd87f175e3e97ccf971c50da0b8039ea2edf2cdf448c54d6e2dd`