и посмотреть медиа
Speech Technology - речь
и посмотреть медиа
Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.
Канал посвящён технологиям речи: распознавание, синтез и голосовые помощники. Актуальные новости и обзоры.
AI бот для озвучки текста и клонирования голоса в Telegram. Создавай аудио с помощью нейросети быстро и просто.
🗣️ Speech Technology — канал для любителей голосовых технологий. Узнайте о новейших системах распознавания речи, синтезе голоса и ИИ-ассистентах. От Siri до продвинутых нейросетей.
🔬 Здесь публикуются статьи о разработках, исследованиях и применении в повседневной жизни. Обзоры софта, аппаратных решений и тенденции рынка.
🎤 Будьте в курсе прорывов в области речевых технологий. Полезно для разработчиков, лингвистов и всех интересующихся ИИ.
You can download really huge datasets these days Ссылка скрыта LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training. - 1.3B video URLs from CommonCrawl - 80M downloaded videos - 10M video hours - 55M captioned clips - 300M frame-caption pairs This repository contains 1.7 million audio clips taken from BVD-V-55M and sampled for uniqueness of the source video, so that the subset maximises source diversity rather than clip count. Each clip comes with a caption, its language, and the timestamps locating it in the source video. Ссылка скрыта
Открыть канал и посмотреть медиаСсылка скрыта English/Chinese only but really good quality. 1st place on ArtificalAnalysis leaderboard
Открыть канал и посмотреть медиаKytai released pockettts training code Ссылка скрыта
Открыть канал и посмотреть медиаTopic of today - autoregression vs diffusion in TTS First of all the presentation from Meta. Claims diffusion/autoregression decision could be dynamic Ссылка скрыта Second paper from today on the similar topic Ссылка скрыта DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong Shim Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, since the model receives the full input text before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that together defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3x faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates the hallucinations observed in AR generation.
Открыть канал и посмотреть медиаDeepmind also releases something Ссылка скрыта PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.
Открыть канал и посмотреть медиаPiotr is a reincarnation of Lennart Ссылка скрыта
Открыть канал и посмотреть медиаТолько зарегистрированные пользователи могут делиться своим мнением.
Станьте первым, кто поделится своим впечатлением об этом ресурсе!
Канал опытных трейдеров-охотников за трендами в Binance Futures. 100% качественные сигналы для прибыльной торговли.
Сигналы USDⓈ-M фьючерсов для Binance: 83-93% точности, 2-3 бесплатных в неделю, с инструкциями и поддержкой.
Сигналы пампов Binance: timely оповещения для трейдеров крипторынка.