Logo
TGCATALOG
Catálogo Selecciones Blog
Speech Technology - речь

Speech Technology - речь

Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.

Sin valoraciones
25
10.09.2026
25
10.09.2026
Sin valoraciones
25
10.09.2026
Redirección segura vía bot
О канале

Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.

Подписчиков 1,718
Тематика Tecnología
Язык Español
Ссылка t.me/speechtech

También te recomendamos

Kekaton AI  Voz en off y clon de voz
Kekaton AI Voz en off y clon de voz
Bot

bot de IA para clonar texto a voz y voz en Telegram. Cree audio utilizando la red neuronal de forma rápida y fácil.

Descripción

🗣️ Speech Technology — канал для любителей голосовых технологий. Узнайте о новейших системах распознавания речи, синтезе голоса и ИИ-ассистентах. От Siri до продвинутых нейросетей.

🔬 Здесь публикуются статьи о разработках, исследованиях и применении в повседневной жизни. Обзоры софта, аппаратных решений и тенденции рынка.

🎤 Будьте в курсе прорывов в области речевых технологий. Полезно для разработчиков, лингвистов и всех интересующихся ИИ.

Últimas publicaciones

Speech Technology - речь
Speech Technology - речь
🔒Открыть пост
и посмотреть медиа
Topic of today - autoregression vs diffusion in TTS First of all the presentation from Meta. Claims diffusion/autoregression decision could be dynamic Ссылка скрыта Second paper from today on the similar topic Ссылка скрыта DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong Shim Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, since the model receives the full input text before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that together defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3x faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates the hallucinations observed in AR generation.
Speech Technology - речь
Evolución de suscriptores
+0.2% últimos 30 días
Actuales
1,718
Hace un mes
1,715
Crecimiento medio
+0 / día
Actualizado
hace 4 horas

Reseñas de canal Speech Technology - речь

Inicia sesión para dejar una reseña

Solo los usuarios registrados pueden compartir su opinión.

Aún no hay reseñas

¡Sé el primero en compartir tu experiencia con este recurso!

Recursos similares

Catálogo de motocicletas y mercancías de China con artículos relevantes en stock.

Canal

Canal con noticias y materiales para usuarios de la gazstation de aplicaciones.

Canal

Descargas de vídeos y música de Instagram, TikTok, YouTube.

Bot
Cambiar a tema claro
Inicio Catálogo Selecciones Blog Entrar