图像 · 视频 4 个事件
本分类 RSS ↗昨天
-
有限来源图像 · 视频
Launch HN:Speko(YC S26)——语音 AI 领域的 OpenRouter译
Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public benchmarked options, and…
Hacker News Frontpage 8月10日 · 周一
-
已确认开源Agent · 编程
用 NVIDIA Magpie TTS 构建低延迟多语言语音 Agent:开放权重、部署全可控译
Every voice interaction has a latency budget. By the time a user hears your application respond, you've already spent precious milliseconds capturing audio, transcribing speech, running an LLM, retri…
Hugging Face Blog 8月3日 · 周一
-
已确认图像 · 视频
我们如何在六个月内构建响应式语音 AI 实时系统译
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
OpenAI News 7月23日 · 周四
-
已确认图像 · 视频
Nunchaku 4-bit 扩散推理登陆 Diffusers译
Large diffusion transformers can create stunning images (or even videos, audio snippets, and now text), but loading a modern text-to-image model in BF16 precision often requires 20-30 GB of VRAM, whi…
Hugging Face Blog