Papers
arxiv:2607.25351

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

Published on Sep 7
Authors:

Abstract

Some text-to-speech systems ship a synthesis model and preset style vectors but withhold the reference encoder that turns a recording into a style vector, so a user cannot obtain a style for a new voice. We recover that vector without the encoder by inverting the released pipeline with gradient descent: all weights stay frozen and only the style vector is optimized, against time-pooled WavLM statistics of one recording of the target. The objective discards the time axis, so no transcript is needed. Two constraints from the release make the result a drop-in style: every row of the timbre style stays on the unit sphere where the presets lie, and the small duration style, which a time-pooled loss cannot see, is fitted to the recording's speaking rate. On SupertonicTTS, over 147 speakers and 100 held-out sentences each, ECAPA-TDNN similarity rises from 0.129 to 0.419, every recovered style starts from a preset and ends closer to its target than that preset was, and the pooled word error rate stays below that of the presets. A verifier at its equal-error point accepts 52% of the recovered voices as the target speaker, against 1% of the presets.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.25351
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.25351 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.25351 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.25351 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.