2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Awesome AI Agents › Audio & Voice Assistants
Tortoise is a sophisticated text-to-speech (TTS) system designed with a strong emphasis on multi-voice capabilities and highly realistic prosody and intonation. The project provides all the necessary code to run Tortoise TTS in inference mode, enabling users to generate speech from text with a focus on quality and naturalness. The system leverages both an autoregressive decoder and a diffusion decoder, which traditionally have low sampling rates, but recent improvements have significantly enhanced its speed and reduced latency, making it more practical for real-time applications. The repository includes detailed installation instructions for various platforms, including Windows, macOS (Apple Silicon), and Docker environments, ensuring accessibility for users with different setups. It requires an NVIDIA GPU for local installations to achieve optimal performance. The project also offers a live demo hosted on Hugging Face Spaces, allowing users to experience the TTS capabilities without local setup. Tortoise supports multiple usage modes, including single phrase speech synthesis, streaming via socket server, and reading large text files with options for fast or standard processing. It also provides a programmatic API for integration into other applications, with support for advanced features like DeepSpeed, key-value caching, and half-precision floating-point operations to optimize performance. The project is inspired by Mojave desert flora and fauna, with the name "Tortoise" reflecting its initial slower speed, which has since been improved. It acknowledges contributions from the broader AI and machine learning community, including Hugging Face, and draws inspiration from notable research papers in the field of generative models and vocoders. Overall, Tortoise is a high-quality, multi-voice TTS system that balances realism and performance, making it suitable for developers and researchers interested in advanced speech synthesis technologies.
https://github.com/neonbjb/tortoise-tts
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
LiveKit Agents is an open-source framework for building real-time voice AI agents with integrated speech-to-text, large language models, text-to-speech, and telephony capabilities.
01 is an open-source voice interface platform that enables natural language voice control across desktop, mobile, and ESP32 devices, offering extensive customization and hardware support.
A reusable Android library in Kotlin that adds a full voice conversation pipeline to an app, chaining microphone capture, voice activity detection, speech-to-text, an Anthropic Claude response and text-to-speech.
ElevenLabs UI is a component library built on shadcn/ui that provides customizable React components to help developers build multimodal agent and audio applications faster.
ElatoAI enables real-time AI speech interaction on Arduino ESP32 devices using OpenAI, Gemini, and Eleven Labs AI models with secure websocket communication and a web app for control.
Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.
joinly is a self-hosted MCP server that lets an AI agent join Google Meet, Zoom or Teams calls in a browser, listen through speech-to-text and reply by voice or chat in real time.