2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Awesome AI Agents › Audio & Voice Assistants
The LiveKit Agents framework is a powerful open-source platform designed for building real-time voice AI agents capable of seeing, hearing, and speaking. It provides a comprehensive ecosystem that integrates speech-to-text (STT), large language models (LLM), text-to-speech (TTS), and real-time APIs to create flexible and customizable voice AI applications. The framework supports server-side agentic applications and offers built-in job scheduling and task distribution through dispatch APIs, enabling efficient connection of end users to agents. LiveKit Agents also includes extensive WebRTC client support via LiveKit's open-source SDKs, compatible with nearly all major platforms, and seamless telephony integration allowing agents to make and receive phone calls. The platform facilitates data exchange with clients using RPCs and other data APIs, ensuring smooth interaction between agents and users. Being fully open-source, users can run the entire stack on their own servers, including the widely used LiveKit media server. The framework is designed with core concepts such as Agents (LLM-based applications with defined instructions), AgentSessions (containers managing interactions with end users), and entrypoints (starting points for interactive sessions). It offers practical usage examples, including simple voice agents and multi-agent handoff scenarios, demonstrating how to build and deploy voice AI agents with various functionalities. The framework supports plugins for popular model providers and provides detailed documentation and guides for installation and usage. LiveKit Agents is actively developed and encourages community contributions, making it a versatile and evolving tool for creating advanced voice AI solutions.
https://github.com/livekit/agents
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Tortoise is a high-quality multi-voice text-to-speech system focused on realistic prosody and intonation, offering various usage modes and advanced performance optimizations.
01 is an open-source voice interface platform that enables natural language voice control across desktop, mobile, and ESP32 devices, offering extensive customization and hardware support.
A reusable Android library in Kotlin that adds a full voice conversation pipeline to an app, chaining microphone capture, voice activity detection, speech-to-text, an Anthropic Claude response and text-to-speech.
ElevenLabs UI is a component library built on shadcn/ui that provides customizable React components to help developers build multimodal agent and audio applications faster.
ElatoAI enables real-time AI speech interaction on Arduino ESP32 devices using OpenAI, Gemini, and Eleven Labs AI models with secure websocket communication and a web app for control.
Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.
joinly is a self-hosted MCP server that lets an AI agent join Google Meet, Zoom or Teams calls in a browser, listen through speech-to-text and reply by voice or chat in real time.