2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Awesome AI Agents › Audio & Voice Assistants
ElevenLabs UI is a specialized component library and custom registry designed to accelerate the development of multimodal agents and audio applications. Built on top of the shadcn/ui framework, it offers a collection of pre-built, customizable React components tailored specifically for agentic and audio-centric use cases. These components include orbs, waveforms, voice agents, audio players, and more, providing developers with ready-to-use building blocks that simplify the creation of sophisticated audio and agent applications. The library integrates seamlessly with Next.js projects and requires Node.js 18 or later, shadcn/ui initialization, and Tailwind CSS configuration. Installation is streamlined through a CLI tool that allows developers to add either all components at once or select individual components as needed. This CLI can be used directly via npx or through the shadcn/ui CLI, making it flexible and easy to incorporate into existing workflows. ElevenLabs UI not only enhances productivity by reducing the time needed to build complex UI elements but also ensures consistency and quality through its curated set of components. The project is open for contributions, encouraging developers to fork the repository, make improvements, and submit pull requests. Licensed under the MIT license, ElevenLabs UI is engineered by ElevenLabs and is well-documented with examples and a contributing guide to support community involvement. Overall, ElevenLabs UI is a valuable resource for developers looking to build advanced audio and agent applications efficiently, leveraging a robust set of UI components that integrate smoothly with modern web development technologies.
https://github.com/elevenlabs/ui
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Tortoise is a high-quality multi-voice text-to-speech system focused on realistic prosody and intonation, offering various usage modes and advanced performance optimizations.
LiveKit Agents is an open-source framework for building real-time voice AI agents with integrated speech-to-text, large language models, text-to-speech, and telephony capabilities.
01 is an open-source voice interface platform that enables natural language voice control across desktop, mobile, and ESP32 devices, offering extensive customization and hardware support.
A reusable Android library in Kotlin that adds a full voice conversation pipeline to an app, chaining microphone capture, voice activity detection, speech-to-text, an Anthropic Claude response and text-to-speech.
ElatoAI enables real-time AI speech interaction on Arduino ESP32 devices using OpenAI, Gemini, and Eleven Labs AI models with secure websocket communication and a web app for control.
Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.
joinly is a self-hosted MCP server that lets an AI agent join Google Meet, Zoom or Teams calls in a browser, listen through speech-to-text and reply by voice or chat in real time.