и посмотреть медиа
Speech Technology - речь
и посмотреть медиа
Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.
Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.
AI bot for text-to-speech and voice cloning in Telegram. Create audio using neural network quickly and easily.
🗣️ Speech Technology — канал для любителей голосовых технологий. Узнайте о новейших системах распознавания речи, синтезе голоса и ИИ-ассистентах. От Siri до продвинутых нейросетей.
🔬 Здесь публикуются статьи о разработках, исследованиях и применении в повседневной жизни. Обзоры софта, аппаратных решений и тенденции рынка.
🎤 Будьте в курсе прорывов в области речевых технологий. Полезно для разработчиков, лингвистов и всех интересующихся ИИ.
Topic of today - autoregression vs diffusion in TTS First of all the presentation from Meta. Claims diffusion/autoregression decision could be dynamic Ссылка скрыта Second paper from today on the similar topic Ссылка скрыта DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong Shim Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, since the model receives the full input text before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that together defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3x faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates the hallucinations observed in AR generation.
Открыть канал и посмотреть медиаOnly registered users can share their opinion.
Be the first to share your impression of this resource!
channel with varietyITContent: news, reviews, tips for professionals and technology lovers.
News, statistics and interesting moments from aviation of Uzbekistan.