2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Awesome AI Agents › Audio & Voice Assistants
ElatoAI is an innovative project that enables real-time AI speech interaction using advanced AI models and hardware integration. It leverages OpenAI Realtime API, Gemini Live API, and Eleven Labs AI Agents to provide seamless, uninterrupted conversations lasting over 15 minutes globally. The system is designed to work on the Arduino ESP32 platform, utilizing Secure WebSockets (WSS) and Deno edge functions to facilitate communication between the AI models and the hardware device. This makes it ideal for AI toys, AI companions, and other AI-enabled devices. The project includes a DIY hardware design for the ESP32 device, which can be controlled via a web application built with Next.js and React. The app allows users to select from various AI characters, engage in real-time conversations, and create personalized AI personas. The architecture consists of three main components: a frontend client hosted on Vercel, an edge server running Deno functions to manage websocket connections and API calls, and the ESP32 IoT client that handles audio transmission and reception. ElatoAI supports multiple AI speech models and offers features such as real-time speech-to-speech conversion, voice modulation, and character customization. The project provides detailed setup instructions, including running a local Supabase backend, configuring environment variables, and flashing the firmware onto the ESP32 device. It also offers options for using a hosted edge server or running a local server for development purposes. With its open-source MIT license, ElatoAI encourages community contributions and experimentation. It is well-documented with tutorials, demo videos, and a supportive Discord community. The project aims to make AI voice interaction accessible and practical for various applications, from interactive toys to personal AI assistants, showcasing the potential of combining cutting-edge AI with embedded systems.
https://github.com/akdeb/ElatoAI
ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.
Tortoise is a high-quality multi-voice text-to-speech system focused on realistic prosody and intonation, offering various usage modes and advanced performance optimizations.
LiveKit Agents is an open-source framework for building real-time voice AI agents with integrated speech-to-text, large language models, text-to-speech, and telephony capabilities.
01 is an open-source voice interface platform that enables natural language voice control across desktop, mobile, and ESP32 devices, offering extensive customization and hardware support.
A reusable Android library in Kotlin that adds a full voice conversation pipeline to an app, chaining microphone capture, voice activity detection, speech-to-text, an Anthropic Claude response and text-to-speech.
ElevenLabs UI is a component library built on shadcn/ui that provides customizable React components to help developers build multimodal agent and audio applications faster.
Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.
joinly is a self-hosted MCP server that lets an AI agent join Google Meet, Zoom or Teams calls in a browser, listen through speech-to-text and reply by voice or chat in real time.