Awesome AI AgentsAudio & Voice Assistants

openinterpreter/01

⭐ 5156 Python added to this list on 2025-07-30 repository created 2024-01-11

The project "01" by OpenInterpreter is an open-source voice interface platform designed for desktop, mobile, and ESP32 chip devices. Inspired by futuristic concepts like the Rabbit R1 and the Star Trek computer, it aims to provide a natural language voice interface that allows users to interact with their devices through spoken commands. The platform is powered by Open Interpreter and supports a variety of capabilities including executing code, browsing the web, managing files, and controlling third-party software. It is designed to be versatile, with server options tailored for different hardware capabilities: a Light Server optimized for low-power devices such as ESP32, and a Livekit Server for devices with higher processing power. The project offers clients for Android, iOS, ESP32, and desktop environments, making it accessible across multiple device types. Users can also build their own hardware devices or explore other hardware options provided by the project. Customization is a key feature, allowing users to modify behavior, language models, system messages, and more through editable profiles. The project is experimental and under rapid development, with warnings about safety and the lack of basic safeguards, advising users to avoid running it on devices with sensitive information or access to paid services until a stable release is available. The community aspect is emphasized with contributions welcomed, a Discord community for support, and comprehensive documentation including installation guides, configuration instructions, safety considerations, and API references. Overall, "01" is a cutting-edge, open-source voice interface platform that aims to bring intelligent voice control to a wide range of devices, combining advanced AI capabilities with hardware flexibility and user customization.

https://github.com/openinterpreter/01

aiandroid-appbrowse-webcontrol-softwarecustomizationdesktopdesktop-clientesp32esp32-clientexecute-codeexperimentalhardwareios-applight-serverlivekit-servermanage-filesmobilenatural-languageopen-interpreteropen-sourcesafetyvoice-controlvoice-interface

Also in Audio & Voice Assistants

2noise/ChatTTS

ChatTTS is a generative text-to-speech model optimized for natural and expressive dialogue-based speech synthesis, supporting multi-speaker and fine-grained prosody control for conversational AI applications.

neonbjb/tortoise-tts

Tortoise is a high-quality multi-voice text-to-speech system focused on realistic prosody and intonation, offering various usage modes and advanced performance optimizations.

livekit/agents

LiveKit Agents is an open-source framework for building real-time voice AI agents with integrated speech-to-text, large language models, text-to-speech, and telephony capabilities.

ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK

A reusable Android library in Kotlin that adds a full voice conversation pipeline to an app, chaining microphone capture, voice activity detection, speech-to-text, an Anthropic Claude response and text-to-speech.

elevenlabs/ui

ElevenLabs UI is a component library built on shadcn/ui that provides customizable React components to help developers build multimodal agent and audio applications faster.

akdeb/ElatoAI

ElatoAI enables real-time AI speech interaction on Arduino ESP32 devices using OpenAI, Gemini, and Eleven Labs AI models with secure websocket communication and a web app for control.

bolna-ai/bolna

Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.

joinly-ai/joinly

joinly is a self-hosted MCP server that lets an AI agent join Google Meet, Zoom or Teams calls in a browser, listen through speech-to-text and reply by voice or chat in real time.