Spectral voice conversion for text-to-speech synthesis

A new voice conversion algorithm that modifies a source speaker's speech to sound as if produced by a target speaker is presented. It is applied to a residual-excited LPC text-to-speech diphone synthesizer. Spectral parameters are mapped using a locally linear transformation based on Gaussian mixture models whose parameters are trained by joint density estimation. The LPC residuals are adjusted to match the target speakers average pitch. To study effects of the amount of training on performance, data sets of varying sizes are created by automatically selecting subsets of all available diphones by a vector quantization method. In an objective evaluation, the proposed method is found to perform more reliably for small training sets than a previous approach. In perceptual tests, it was shown that nearly optimal spectral conversion performance was achieved, even with a small amount of training data. However, speech quality improved with increases in the training set size.

Voice transformationusing PSOLA techniqueVoice transformation using PSOLA techniqueTransformation offormants for voice…Transformation of formants for voice conversion using artificial neural networksStatistical methods forvoice quality…Statistical methods for voice quality transformationInterpolation propertiesof linear prediction…Interpolation properties of linear prediction parametric representationsHarmonic plus noisemodels for speech…Harmonic plus noise models for speech, combined with statistical methods, for speech and speaker modificationLocal models andGaussian mixture models…Local models and Gaussian mixture models for statistical data processingVoice conversion bycodebook mapping of lin…Voice conversion by codebook mapping of line spectral frequencies and excitation spectrumVoice conversionalgorithm based on…Voice conversion algorithm based on Gaussian mixture model with dynamic frequency warping of STRAIGHT spectrumVoice conversion withsmoothed GMM and MAP…Voice conversion with smoothed GMM and MAP adaptationHigh Quality VoiceConversion through…High Quality Voice Conversion through Phoneme-Based Linear Mapping Functions with STRAIGHT for MandarinSpeaking-aid systemsusing GMM-based voice…Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speechParametric VoiceConversion Based on…Parametric Voice Conversion Based on Bilinear Frequency Warping Plus Amplitude ScalingVulnerability of speakerverification systems…Vulnerability of speaker verification systems against voice conversion spoofing attacks: The case of telephone speechA study on spoofingattack in…A study on spoofing attack in state-of-the-art speaker verification: the telephone speech caseComparing ANN and GMM ina voice conversion…Comparing ANN and GMM in a voice conversion frameworkVoice Conversion UsingRNN Pre-Trained by…Voice Conversion Using RNN Pre-Trained by Recurrent Temporal Restricted Boltzmann MachinesVoice conversion usingGeneral Regression…Voice conversion using General Regression Neural NetworkVoice conversion usingdeep neural networks…Voice conversion using deep neural networks with speaker-independent pre-trainingIntroduction to VoicePresentation Attack…Introduction to Voice Presentation Attack Detection and Recent AdvancesSpectral voiceconversion for…Spectral voice conversion for text-to-speech synthesisEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.