Awesome AI Agents › Weekly › 2025-05-02
2025-05-02
392 projects added
Agent Communication & Protocols
LiveKit Python SDKs provide real-time video, audio, and data communication capabilities for Python applications, enabling easy integration with LiveKit Cloud or self-hosted servers for building interactive and streaming applications.
SPADE is a Python-based multi-agent systems platform using XMPP for communication, enabling the development of intelligent agents that interact with other agents and humans in real-time.
http2-wrapper is a Node.js package that enables HTTP/2 support using the familiar HTTP/1 API, allowing seamless integration and improved performance without rewriting existing code.
Emcee is a command-line tool that provides a Model Context Protocol server for OpenAPI-based web applications, enabling AI assistants like Claude Desktop to connect with external APIs and services dynamically through standardized tool integration.
Agency is a minimal and scalable Python framework for building flexible agentic systems using the Actor model, supporting concurrency, networked communication, and detailed observability.
A project demonstrating a working pattern for SSE-based MCP clients and servers enabling decoupled, cloud-native communication between MCP servers and clients.
A comprehensive directory and resource hub for Google's Agent2Agent (A2A) protocol enabling secure communication and collaboration between AI agents.
A comprehensive repository offering practical guides, utilities, and server implementations for exploring and developing with the Model Context Protocol (MCP), an open standard for connecting AI applications to external data sources and tools.
Agent Integration & Deployment Tools
Bespoke Automata is a platform for creating, testing, and deploying complex AI agents locally with a graphical interface and API endpoints.
GeniA is an open-source AI platform that acts as a virtual engineering team member, automating and assisting with a wide range of software engineering tasks in production environments through seamless Slack integration and extensible tool support.
FusionInventory Agent is a generic management agent that automates IT asset management, network discovery, inventory, and deployment tasks, integrating with GLPI for centralized control.
Agent-FLAN is a project that designs data and methods for effective fine-tuning of large language models to significantly improve their agent capabilities and reduce hallucinations.
Eidolon is an open-source AI Agent Server and SDK that enables developers to build, deploy, and customize modular agent-based services with seamless communication and enterprise readiness.
NGINX Agent is a tool for remotely managing, configuring, and monitoring NGINX instances, providing real-time metrics and event collection to enhance server performance and administration.
inspectIT Ocelot is a zero-configuration Java agent for dynamically collecting application performance, tracing, and behavior data, integrating seamlessly with popular open-source monitoring tools.
AWS Generative AI CDK Constructs is an open-source library providing well-architected, reusable infrastructure patterns to simplify building and deploying generative AI solutions on AWS using the AWS Cloud Development Kit.
CrewAI Quickstart is a collection of notebooks, code templates, and guides designed to help developers build and experiment with agentic workflows using CrewAI's tools and integrations.
GLPI Agent is a comprehensive management agent that performs inventory, network discovery, and deployment tasks, integrating seamlessly with GLPI servers for efficient IT asset management.
TapeAgents is a flexible framework that uses a structured replayable log to support all stages of Large Language Model Agent development, including building, debugging, serving, and optimizing agents.
WorkflowAI is an open-source platform that enables teams to rapidly build, test, and deploy AI features with support for multiple AI models, flexible deployment options, and advanced observability and cost tracking.
BondAI is a comprehensive platform for building and deploying highly capable single and multi-agent AI systems with extensive integrations and easy-to-use interfaces.
Adaline Gateway is a fully local, production-grade Super SDK that provides a unified and powerful interface for calling over 200 large language models with advanced features like batching, retries, caching, and custom plugins.
BeeAI is an open platform that enables discovery, execution, and composition of AI agents from any framework, supporting multi-agent workflows and seamless integration across languages.
Surfkit is a versatile toolkit for building, sharing, running, and managing AI agents that operate on devices locally or in the cloud.
AgentStore is a scalable platform for dynamically integrating heterogeneous agents to automate operating system tasks independently or collaboratively, providing a specialized generalist computer assistant.
The Jenkins EC2 Plugin enables Jenkins to dynamically provision and manage Amazon EC2 instances as build agents, optimizing build infrastructure scalability and cost efficiency.
The Haystack Cookbook is a collection of example notebooks demonstrating the use of the Haystack AI framework for various search and retrieval tasks, aimed at helping users learn and implement Haystack effectively.
Vespper is an open-source AI-powered on-call developer that integrates with observability and incident management tools to provide real-time root cause analysis and insights within Slack, helping engineers resolve incidents faster and more efficiently.
Tiger is an open-source, community-driven tool ecosystem that enables AI agents to control and interact with computer systems and online resources through integrations with popular AI frameworks like crewAI, LangChain, and AutoGen.
PerfMon Server Agent is a Java-based tool for remotely accessing and monitoring a wide range of system performance metrics on servers using a simple plain-text protocol.
Terrarium is a secure, low-latency Python sandbox environment for executing untrusted code using Pyodide, designed for easy deployment on cloud platforms like Google Cloud Run.
Open Assistant API is an open-source, self-hosted AI assistant API framework that supports multi-model LLM integration, RAG engines, custom functions, and tool extensions with seamless OpenAI compatibility.
Graphsignal is an AI optimization platform with a Python tracer that helps developers monitor, analyze, and optimize the performance and resource usage of machine learning models and applications.
OCSInventory-ocsreports is a web console for OCS Inventory NG that enables centralized asset management, network discovery, and software deployment across large IT infrastructures.
Stride AI Agents is an open-source repository providing tools and knowledge to build, implement, and master autonomous AI systems for diverse applications across industries.
Proficient AI provides TypeScript/JavaScript SDKs and tools to quickly add and manage conversational AI agents powered by large language models in applications.
django-ai-assistant integrates AI Assistants with Django to build intelligent applications using large language models and advanced AI features like AI Tool Calling and Retrieval-Augmented Generation.
Agentica is an open-source TypeScript framework that simplifies building reliable AI agents with Large Language Models by providing schema-driven function calling, automatic error correction, and seamless OpenAPI integration.
JS Agent is a composable and extensible framework for building AI agents and applications using JavaScript and TypeScript, supporting multiple language models and providing robust tools for agent development.
FireAct is a comprehensive framework for fine-tuning language agents, providing tools, data, and models to facilitate advanced training and evaluation of language models like Llama and GPT.
E2B Infrastructure is an open-source project providing the foundational infrastructure for AI code interpreting, enabling deployment and management of AI agents in cloud environments using Terraform.
Palico AI is an integrated framework that enables developers to build, improve, and productionize Large Language Model applications with flexibility, advanced features, and extensive integrations.
Embodied-agents is a modular and extensible toolkit that integrates state-of-the-art multi-modal transformer models into robotics stacks, enabling seamless local and remote AI inference for language, motion, and sensory tasks in robotic systems.
Rigging is a lightweight and flexible framework that simplifies the use of large language models in production by providing structured prompt definitions, broad model support, and advanced features like async batching and integrated tracing.
StackImpact Go Profiler is a production-grade performance profiling tool for Go applications that provides continuous CPU, memory, latency, and error monitoring with detailed insights and minimal overhead.
AgentLabs is an open-source universal frontend platform that enables developers to deploy and manage AI agents with real-time streaming, authentication, analytics, and payment features, supporting both cloud and self-hosted environments.
Connery SDK is an open-source SDK and CLI that streamlines the development of AI plugins and actions by providing a standardized REST API and tools for easy plugin server creation.
Apache SkyWalking Go is a Golang auto-instrumentation agent that provides native tracing, metrics, and logging capabilities for Golang projects, enabling effective observability and performance monitoring in distributed systems.
GoRelic is a deprecated Go runtime monitoring agent that collects detailed performance metrics and sends them to New Relic for application monitoring and analysis.
ArcadeAI/arcade-ai is a developer platform with a Python SDK and CLI that enables building, deploying, and managing tools for AI agents to interact with the world efficiently.
MCP Link automates the conversion of any OpenAPI V3 API into a fully compatible MCP Server, enabling seamless integration of RESTful APIs into AI Agent ecosystems without code modification.
The Steamship Python SDK enables developers to build, deploy, and interact with AI agents, packages, and plugins on the Steamship platform, providing tools for scalable AI application development and integration.
Pinpoint C Agent is a cross-platform tool that enables monitoring of PHP and Python applications by integrating them with the Pinpoint APM system using C/C++ APIs and auto-injection techniques.
A TypeScript toolkit for interacting with Bitcoin and Stacks blockchains, offering wallet management, smart contract interactions, and blockchain operations.
Morpheus is a decentralized network that powers smart agents by integrating AI and blockchain technology to enable natural language interactions with Ethereum smart contracts and wallets.
Sagentic.ai Agent Framework is a unified platform for building, running, and scaling autonomous agents with easy project setup, development tools, and community support.
Awesome Devins is a curated list of open-source AI agents inspired by Devin, aimed at assisting software engineering and development tasks with community collaboration and cloud runtime support.
NPi is an open-source platform providing Tool-use APIs that enable AI agents to interact with and operate various software tools and applications through natural language programming interfaces.
Modus is an open-source, serverless framework that enables developers to build scalable, AI-powered applications and agentic systems using WebAssembly with support for Go and AssemblyScript.
Claude Artifacts React is an open-source tool that enables quick and easy deployment of React code snippets generated by Claude Artifacts to Vercel or Cloudflare Pages for prototyping and sharing.
Instrukt is a terminal-based integrated AI environment that enables users to create, instruct, and manage modular AI agents securely within sandboxed containers, featuring document indexing, tool attachment, and a rich terminal interface for efficient AI workflows.
Tambo AI is a React library that simplifies building AI-powered assistants and agents with features like thread management, state persistence, and AI-generated components for creating generative and personalized user experiences.
Characterfile is a simple and easy-to-use file format and toolset for generating and transmitting character data from social media and chat exports, compatible with AI agents like Eliza.
Jenkinsci/docker-agent provides Docker images for Jenkins agents, including base and inbound agents, facilitating containerized Jenkins automation with support for multiple platforms and JDK versions.
Sematext Docker Agent is a deprecated tool for collecting host and container metrics, logs, and events, with users directed to updated Sematext monitoring solutions for Docker, Kubernetes, and infrastructure.
EdgeChains.js is a full-stack Generative AI library that simplifies deployment, prompt management, and scalable execution of GenAI applications using declarative Jsonnet configurations and automatic parallelism.
A deprecated library that enabled streaming for OpenAI Assistants API using Astra Assistants, now replaced by the astra-assistants package.
Chipper is a modular and lightweight AI interface that enables users to build and customize Retrieval-Augmented Generation (RAG) pipelines with local and cloud models, featuring document processing, web scraping, and secure API proxy capabilities in a fully containerized environment.
any-agent is a Python library that provides a unified interface to multiple agent frameworks, enabling easy switching and evaluation of AI agents through a single API.
MCP Containers offers containerized versions of hundreds of Model Context Protocol (MCP) servers, enabling easy, secure, and up-to-date deployment of these servers via Docker containers.
npcpy is a versatile Python-based AI toolkit that provides an agent-based framework and multiple command-line interfaces for integrating, testing, and deploying AI models and agents in daily workflows.
The Microsoft AI Agents Hackathon 2025 is a free, virtual three-week event that educates and empowers developers to build innovative AI agents using leading frameworks and SDKs, supported by expert sessions and community collaboration.
TEN Framework is a real-time, distributed, cloud-edge collaborative multimodal AI agent framework supporting multiple programming languages for building high-performance AI applications.
ReCall is a reinforcement learning framework that trains large language models to reason and use arbitrary tools agentically, enabling advanced tool-based reasoning and general-purpose AI agents.
ACI.dev is an open-source platform that connects AI agents to over 600 tool integrations with multi-tenant authentication, granular permissions, and access via direct function calls or a unified MCP server, enabling production-ready AI agent development without infrastructure complexity.
Daytona is a secure and elastic infrastructure platform designed for safely running AI-generated code with high performance and scalability.
mcp-use is an open-source unified client library that enables seamless integration of any LangChain-supported LLM with MCP servers to build custom agents with tool access.
EVM MCP Server is a Model Context Protocol server providing AI agents with unified blockchain services across 30+ EVM-compatible networks, supporting token transfers, smart contract interactions, and ENS name resolution.
Audio & Voice Assistants
An AI-powered smart speaker system that uses speech recognition, text-to-speech, and OpenAI's language models to enable voice-driven conversations and web search capabilities on PC/Mac and Raspberry Pi platforms.
Voice Lab is a comprehensive open-source framework for testing, evaluating, and optimizing voice agents powered by large language models, enabling systematic performance analysis, cost optimization, and prompt refinement through customizable metrics and scenarios.
MixedVoices is an analytics and evaluation platform for voice agents that provides tools to track, visualize, and optimize conversational AI performance through call flow analysis, machine learning metrics, and call quality assessment.
Bolna is an open-source end-to-end platform for building voice-first conversational AI agents by orchestrating telephony, speech recognition, language models, and speech synthesis technologies.
Autonomous Research & Content Generation
Agency is a Go library that enables developers to build and explore generative AI applications and autonomous agents using Large Language Models through a clean, idiomatic, and efficient approach.
Bedrock Engineer is an autonomous software development agent application using Amazon Bedrock that offers customizable AI agents for file operations, web search, code analysis, and more, providing an interactive and rich development assistant experience.
Advanced_RAG is a project offering practical Python notebooks to explore and implement advanced Retrieval-Augmented Generation techniques using Langchain, OpenAI GPTs, and META LLAMA3 for enhancing large language models with contextual knowledge.
An advanced deep research assistant using a multi-agent iterative approach with the OpenAI Agents SDK to generate detailed research reports compatible with various LLMs and tools.
A curated repository compiling outstanding projects, tools, applications, benchmarks, and research related to autonomous AI agents based on GPT and large language models, serving as a comprehensive resource for the AI community.
GPT-Instagram is an autonomous multi-agent AI application that analyzes a user's Instagram history to generate personalized and potentially viral Instagram post recommendations reflecting their personality.
AgentSquare is a research project providing an automatic search platform for designing modular large language model agents, supporting multiple tasks and encouraging community contributions.
xLAM is a family of Large Action Models by Salesforce AI Research designed to enhance AI agent systems through standardized multi-turn trajectory data and advanced function-calling capabilities.
Qwen2-Boundless is a fine-tuned language model based on Qwen2-1.5B-Instruct, designed to handle sensitive and controversial topics while supporting general question answering, primarily optimized for Chinese language tasks.
Stackwise is an open-source collection of AI applications fostering community collaboration and innovation in AI development.
Contoso Creative Writer is a multi-agent AI solution using Azure OpenAI and related Azure services to help users write well-researched, product-specific articles based on user topics and instructions.
Graph Chain-of-Thought (Graph-CoT) is a framework that enhances Large Language Models by enabling them to perform step-by-step reasoning on structured textual graphs, improving their grounding and accuracy in answering complex questions.
GraphAgent is an agentic graph language assistant that integrates language models with graph models to handle complex structured and unstructured data for predictive and generative tasks.
CHRONOS is a novel retrieval-based approach for news timeline summarization that uses iterative self-questioning to generate efficient and scalable chronological summaries, supported by a new large open-domain dataset called Open-TLS.
Ask.py is a Python program that implements a search-extract-summarize flow using AI models to perform web searches, extract relevant information, and generate summarized answers with customizable options and multiple LLM endpoint support.
INSIGHT is an autonomous AI system that automates medical research by managing tasks through AI agents, integrating biomedical APIs, and indexing results for comprehensive analysis and discovery.
Rankify is a comprehensive Python toolkit for unified retrieval, re-ranking, and retrieval-augmented generation, integrating numerous benchmark datasets and state-of-the-art models for advanced research and development in information retrieval.
ChadGPT is a comprehensive project focused on understanding, self-hosting, and applying large language models using open-source tools and exploring future AI and Linux system integrations.
Agent OS is an experimental framework and runtime designed to build sophisticated, long-running, and self-coding autonomous AI agents that can write and execute their own code to interact with the world, featuring a Git-inspired data storage system and a demo conversational agent called Jetpack.
AMEX-codebase is an enhanced toolkit built on LLaMA2-Accessory for fine-tuning and inference of large language models with data preparation utilities and experimental demonstrations.
The LLM Prompt Library is a comprehensive collection of experimental prompt templates and tools designed for multiple major large language models, supporting advanced prompt engineering and optimization across various domains.
WriteHERE is an open-source AI writing framework that uses recursive planning, heterogeneous task integration, and dynamic adaptation to enhance long-form writing for fiction and technical reports.
Syzygy-of-Thoughts is a project implementing a novel reasoning framework that enhances large language model performance on complex tasks by integrating algebraic methods into the Chain of Thought approach.
Sidekick is a native macOS app that enables offline interaction with a local large language model to retrieve and analyze information from files, folders, and websites on the user's Mac, offering advanced features like function calling, memory, and image generation.
Browser & Web Automation
WebVLN is a novel research project that extends vision-and-language navigation to websites, enabling AI agents to navigate web environments using natural language instructions and web-specific content like HTML for improved performance.
GPT-4V Web Agent is an AI-powered tool that can visually perceive, control, and navigate web browsers to automate web scraping, data extraction, and interactive browsing tasks.
Berkeley-NLP/Agent-Eval-Refine is a framework for autonomous evaluation and refinement of digital agents that interact with web and mobile environments, featuring state-of-the-art models and tools for improving agent performance.
RCI Agent for MiniWoB++ is a Python-based project that uses pre-trained language models to solve computer tasks in the MiniWoB++ benchmark by leveraging a novel RCI prompting scheme for improved task execution guided by natural language.
WebShop is a simulated e-commerce environment designed to train and evaluate grounded language agents for complex real-world web interaction tasks involving product search and purchase based on natural language instructions.
AgentLab is an open-source framework for developing, testing, and benchmarking scalable and reproducible web agents across diverse tasks using BrowserGym benchmarks and large-scale parallel experiments.
WebDreamer is a project that uses large language models as world models of the internet to enable model-based planning for web agents, improving their efficiency and performance on web interaction tasks.
A curated list of tools, frameworks, and resources for building AI web agents that autonomously browse and interact with the web.
A curated and continuously updated collection of public projects, papers, codebases, datasets, and tools focused on Large Language Model-based Web Agents and Tools for various personal and professional applications.
Self-MAP is a project that advances multi-turn instruction following for conversational web agents using fine-tuned models and the MT-Mind2Web dataset, as presented in ACL 2024.
Open-CUAK is an open-source platform for teaching, hiring, and managing reliable and scalable automation agents, primarily focused on browser automation, offering a free alternative to OpenAI Operator with privacy and flexibility.
WebMarker is a tool that visually marks interactive elements on web pages to enhance their use with vision-language models for improved web automation and interaction.
Banana-lyzer is an open-source AI agent evaluation framework for web tasks that uses Playwright to run structured and reproducible tests on static snapshots of websites, supporting diverse datasets and multiple evaluation criteria.
BrowserGPT enables users to control their web browser using natural language commands by integrating GPT-4 with Playwright for seamless AI-driven browser automation.
HyperAgent is an AI-enhanced browser automation tool that uses large language models to enable natural language commands for scalable, customizable web automation and data extraction.
Browser Automation
WebCanvas is an open-source framework for building, training, and evaluating web agents in live, dynamic web environments using key-node based evaluation and modular components.
surf.new is an open-source playground for testing and visualizing AI-powered web agents interacting with the web in real-time, built with modern web technologies and supporting multiple AI providers.
Dendrite is an open-source framework for building intelligent web AI agents that can authenticate, interact with, and extract data from any website in a human-like manner.
Notte is an open-source full stack framework that creates intelligent web browsing agents using a perception layer to enable fast, reliable, and cost-effective interactions with websites through large language models.
A Chrome extension that enhances ChatGPT by providing live previews, syntax highlighting, and interaction capabilities for code snippets within the chat interface.
WebLINX is a benchmark and toolkit for building and evaluating conversational web navigation agents through multi-turn dialogues on real-world websites.
AgentLLM is a proof-of-concept project that enables autonomous agents powered by open-source large language models to run efficiently and privately within web browsers using GPU acceleration.
Fuji-Web is an AI-powered browser extension that autonomously navigates websites and executes user tasks through a sidepanel interface, simplifying online task automation with a single command.
wge is a research codebase for training sample-efficient reinforcement learning agents to perform web tasks specified by natural language instructions using workflow-guided exploration.
Chatbots
Tiledesk Server is an open-source API platform for building advanced multi-channel live chat and chatbot customer support solutions with human-in-the-loop capabilities.
Loyal Elephie is a memory-enabled AI chatbot optimized for local large language models, featuring controllable memory, hybrid search, secure access, and multi-language support.
llm-movieagent is a project that implements an intelligent agent using a semantic layer to enable a large language model to interact with a Neo4j graph database for movie information retrieval and personalized recommendations.
Insights-bot is a chatbot that uses OpenAI GPT models to provide summaries and insights for information flows on Telegram, Slack, and Discord.
Notate is a cross-platform desktop chat application that enhances AI conversations with multi-model support, document analysis, local deployment, and privacy-focused features.
Dialogflow Web Integration v2 is an unofficial, feature-rich progressive web app that enables embedding and interacting with Dialogflow V2 conversational agents on websites, supporting voice, theming, and extensive browser compatibility.
Tiledesk-dashboard is an open-source web application dashboard for managing Tiledesk, a multi-channel live chat platform with AI-powered chatbots and human-in-the-loop capabilities for customer support and conversational app development.
The Data Questionnaire Agent Chatbot is a reverse chatbot that asks users questions about data integration practices and provides advice based on a knowledge base.
An autonomous HR chatbot prototype that uses ChatGPT, LangChain, and Pinecone to answer HR queries through integrated tools and a Streamlit front-end.
A hands-on 75-minute workshop repository for building a conversational AI agent using Azure AI Agent Service to analyze sales data and assist business decisions.
Code Automation and Testing
Mutahunter is an open-source, language-agnostic mutation testing tool that uses large language models to improve test suite effectiveness and software quality.
Devlooper is a program synthesis agent that autonomously generates and iteratively fixes code by running tests in a sandboxed environment until all tests pass.
ChatUITest is a project focused on automatically generating GUI test scripts and includes related tools and datasets for simulating and studying UI defects to improve automated UI testing.
Kodus AI is an open-source AI-powered tool that automates personalized, context-aware code reviews to help development teams catch bugs, enforce best practices, and maintain high-quality codebases efficiently.
Shippie is an AI-powered extensible code review agent that integrates into CI/CD pipelines to automatically detect code issues and improve software delivery speed.
Code Generation & Refactoring
BigCodeBench is a benchmark platform designed to evaluate and rank large language models on complex and practical code generation tasks, advancing AI capabilities towards artificial general intelligence in programming.
MyCoder is a command-line AI-powered coding assistant that supports modular tools, parallel task execution, GitHub integration, and interactive corrections to enhance developer productivity.
aider.el is an Emacs package that integrates the Aider AI pair programming tool to provide seamless AI-assisted coding, refactoring, and testing within Emacs.
Shire is an AI-driven coding agent language that automates programming tasks and integrates with various development tools to create customized AI copilots for enhanced software development workflows.
The English Compiler is a proof-of-concept AI tool that converts English markdown specifications directly into functional code, demonstrating a future where natural language can replace traditional programming.
EvalGPT is a code interpreter framework that leverages large language models to automate code writing, execution, and error handling for user-defined tasks, inspired by Google's Borg system.
Instructa AI Prompts is an open-source repository providing curated AI prompts and rules to enhance AI-assisted coding workflows across multiple popular AI coding tools.
Codai is an AI-powered code assistant that provides intelligent code suggestions, refactoring, and reviews through a session-based CLI by deeply understanding the full context of a project.
wcgw is a shell and coding agent that integrates with Claude and ChatGPT desktop apps to enable AI-driven coding, building, and command execution on local machines.
Web3GPT is an AI-powered platform that streamlines smart contract development and deployment across multiple EVM-compatible blockchains using natural language and specialized AI agents.
Amazon Q Developer CLI enhances terminal productivity on macOS and Linux by providing IDE-style autocomplete, natural language chat, contextual awareness, and AI-driven automation for popular command-line tools.
Friday is an AI-powered developer assistant that uses GPT-4 to help users quickly generate complete Node.js applications through customizable prompts.
Cody is an AI coding assistant that enables interactive natural language queries on your codebase with real-time updates and customizable file monitoring.
Llama-github is an open-source Python library that enables LLM chatbots and AI agents to retrieve and generate context from GitHub public projects for solving complex coding tasks efficiently.
Clean Coder AI is an intelligent 2-in-1 AI system that acts as both a Scrum Master and Developer, automating project planning, management, and coding tasks with advanced AI agents integrated with Todoist.
Kwaak is an open-source tool that uses autonomous AI agents to help developers reduce technical debt by running multiple agents in parallel to improve code quality, write tests, update documentation, and create pull requests, all from a chat-like terminal interface.
Awesome Vibe Coding is a curated list of resources and tools that support the innovative practice of vibe coding, where developers collaborate with AI to write and deploy code more intuitively and efficiently.
A Python SDK that provides a programmatic interface to interact with intelligent code generation agents from Codegen, enabling AI-assisted software development through natural language prompts.
Micro Agent is a minimalistic autonomous agent powered by OpenAI GPT-4 that autonomously writes and tests Python software based on a given purpose, serving as a research tool towards AGI and AI-driven programming.
bumpgen is an AI-powered tool that automates upgrading npm packages and fixes breaking changes in TypeScript and TSX projects.
DynaSaur is a dynamic large language model agent framework that uses Python code generation to extend and adapt its actions beyond predefined sets, achieving top performance on the GAIA benchmark.
Autoview is an AI-powered tool that automatically generates TypeScript frontend UI components from schemas or Swagger/OpenAPI documents, streamlining frontend development and integration.
Cognitive Architecture Frameworks
KnowAgent is a framework that enhances large language model-based agents' planning abilities by augmenting them with an extensive action knowledge base and iterative self-learning from generated action trajectories.
Stately Expert is a framework for building intelligent AI agents powered by state machines and enhanced with observations, feedback, and flexible decision-making policies, integrating multiple AI model providers via the Vercel AI SDK.
OpenNARS is an open-source general-purpose AI reasoning system based on Non-Axiomatic Reasoning System (NARS) designed to process tasks and learn adaptively under conditions of insufficient knowledge and resources.
Saplings is a library that enables building AI agents with advanced reasoning capabilities using tree search algorithms like Monte Carlo Tree Search, A*, and greedy best-first search for improved decision-making and task performance.
ENVISIONS is a neural-symbolic self-training framework that enhances large language models through environment-guided interactive evolution to improve reasoning and problem-solving abilities.
CLIN is a continually learning language agent designed for rapid task adaptation and generalization within the ScienceWorld environment using advanced language models like GPT-4.
A-MEM is an advanced agentic memory system for LLM agents that dynamically organizes, links, and evolves memories using intelligent indexing and semantic search to enhance agent performance on complex tasks.
Lumos is an open-source project that develops modular language agents with unified data formats and competitive performance, leveraging LLAMA-2 and GPT-4 annotations for complex interactive tasks.
SwiftSage V2 is a generative agent system that combines fast and slow thinking processes using large language models and in-context reinforcement learning to solve complex reasoning tasks through iterative feedback and code execution.
Cybergod is an autonomous AGI program designed to perform any task to earn money independently, featuring advanced terminal interaction, reinforcement learning, and an integrated digital economy system.
Mentals AI is a tool for creating and operating AI agents using markdown files, featuring recursive loops, memory, and tool integration without traditional programming.
Dynalang is an advanced agent that leverages diverse language types to predict the future and solve tasks using a multimodal world model, demonstrated across multiple embodied AI environments.
Agent-Driver is an autonomous driving system that uses large language models as cognitive agents to integrate human-like reasoning and common sense into driving, significantly outperforming traditional methods on the nuScenes benchmark.
AgentDock is an open-source framework for building sophisticated AI agents with configurable determinism, enabling reliable and creative AI applications through a node-based architecture and multi-stage workflows.
Courses and Tutorials
A comprehensive collection of Jupyter notebooks supporting the "Building LLMs for Production" book, covering foundational concepts, advanced techniques, and practical applications of large language models in production environments.
Claude-React-Jumpstart is a beginner-friendly step-by-step tutorial for setting up and running Claude-generated React code locally using Vite, Tailwind CSS, and Shadcn UI components.
PhiloAgents Course is an open-source educational project that teaches how to build an AI-powered game simulation engine impersonating historical philosophers, combining philosophy with modern AI technology through hands-on modules covering agent development, RAG systems, system architecture, and LLMOps best practices.
An AI-powered tool that analyzes GitHub repositories and generates beginner-friendly tutorials with visualizations to simplify understanding complex codebases.
Data Integration and Specialized Solutions
LAYRA is a visual-first Retrieval-Augmented Generation system that preserves document layout and visual elements, enabling advanced document understanding through pure visual embeddings and a scalable web-based platform.
Wren Engine is a semantic engine for MCP clients and AI agents that provides precise, context-aware access to enterprise business data with governance and interoperability across modern data stacks.
Data Processing & ETL Agents
Jaiqu is an AI-driven tool that automatically transforms JSON data from any schema into any desired schema using GPT-4 generated jq queries for seamless and repeatable data mapping.
json-repair is a Go-compatible open-source tool designed to automatically detect and repair JSON anomalies commonly produced by Large Language Models, providing a reliable and zero-dependency solution for fixing malformed JSON data.
Instructor Cloud is a cloud-based AI service that extracts models from text in real-time using GPT-4 and FastAPI.
FuzzTypes is a Pydantic extension that provides autocorrecting annotation types for enhanced data validation and normalization, enabling automatic correction and structured data composition in Python applications.
Streamline Analyst is an AI-powered open-source data analysis agent that automates data cleaning, preprocessing, model selection, and visualization to simplify and accelerate the entire data analysis workflow.
AI Filesystem (aifs) is a local semantic search tool that indexes and embeds various file types in a folder to enable fast and efficient semantic search over local directories.
Swiftide is a Rust library for building fast, streaming indexing, querying, and agentic applications using large language models with a modular and extendable API.
Desktop Automation
ScreenAgent is a computer control agent driven by a Visual Language Large Model that interacts with real computer screens through mouse and keyboard operations to complete multi-step tasks.
AssistGUI is an LLM-based desktop GUI assistant that automates operations in productivity software by controlling mouse and keyboard actions to enhance user efficiency.
AskUI Vision Agent is a multi-platform automation framework that enables secure and flexible native device automation using AI-driven UI commands and agentic instructions, designed for enterprise deployment.
InfiGUIAgent is a multimodal generalist GUI agent that uses advanced reasoning and reflection capabilities to enhance task automation on computing devices through a two-stage supervised fine-tuning approach.
GUI-Thinker is an AI-driven framework that enables adaptive and self-reflective control of PC graphical user interfaces, surpassing state-of-the-art desktop GUI agents on the WorldGUI Benchmark.
Documentation & Testing Assistants
anythingllm-docs is a comprehensive and well-structured documentation repository for the AnythingLLM platform by Mintplex Labs, facilitating user and developer engagement through detailed guides and active maintenance.
AUITestAgent is the first automatic natural language-driven GUI testing tool for mobile apps that fully automates GUI interaction and function verification based on natural language test requirements.
Kaizen is an AI-powered automation suite that streamlines software development by automating code reviews, test generation, documentation, and issue detection to boost developer productivity and innovation.
Educational & Learning Agents
Groundhog is an AI coding assistant designed as a teaching tool to help users understand and leverage the inner workings of coding agents like Cursor through detailed code explanations and a modern Rust-based CLI.
LearningX is a Python project offering comprehensive examples and tutorials on classical and deep reinforcement learning as well as basic and advanced machine learning techniques.
Financial & Trading Systems
FinGen is an experimental AI-powered finance agent that uses GPT-4-turbo, LangChain agents, and the Polygon finance API to provide interactive financial analysis through a generative user interface.
gym-fx is a Forex trading simulator environment for OpenAI Gym that enables testing and optimization of custom trading agents using historical market data and configurable trading parameters.
LangChainBitcoin is a suite of AI tools that enable LangChain agents to interact directly with Bitcoin and the Lightning Network, including managing Bitcoin balances and traversing APIs requiring Lightning-based authentication.
Okcash is a decentralized, community-driven cryptocurrency project that has evolved into a multichain digital cash ecosystem supporting seamless transactions and DeFi integration across multiple blockchain networks.
Gaming & Simulation Agents
MineStudio is a comprehensive package that provides tools and APIs for developing, training, and deploying AI agents in Minecraft, featuring simulation, data handling, model training, inference, and benchmarking capabilities.
Poke-env is a Python interface for creating and training reinforcement learning and rule-based bots to battle on Pokemon Showdown, providing a comprehensive environment for AI-driven Pokemon battles.
OpenSteer is an open-source C++ library and interactive demo for creating and prototyping steering behaviors for autonomous characters in games and animation.
JARVIS-1 is an open-world multimodal language model-based agent designed to perform complex multi-task planning and embodied control in the Minecraft environment, leveraging memory-augmented models for enhanced performance.
Gigax is a runtime environment for AI-powered NPCs that perform complex actions in real-time using fine-tuned large language models running locally on your hardware.
CustomNavMesh is an enhanced Unity navigation system that enables agents to avoid other non-moving agents by dynamically switching between agent and obstacle states, improving pathfinding realism and agent interaction.
ChatSim is an editable scene simulation platform for autonomous driving that leverages LLM-agent collaboration and advanced rendering techniques to create realistic and manipulable driving scenarios.
Odyssey is a framework that empowers Minecraft agents with advanced open-world skills using a fine-tuned LLaMA-3 model and a comprehensive skill library to enable autonomous exploration and complex task planning.
Suspicion-Agent is an AI project implementing a theory of mind aware GPT-4 agent to play imperfect information games, enhancing strategic reasoning and decision-making.
Simulator Controller is a modular and extendable sim racing application featuring AI-powered assistants that provide real-time race management, strategy, and driving coaching to create a realistic and immersive racing experience.
IDE Integrations
DevChat is an open-source platform that enables developers to create customized AI-driven workflows using natural language commands directly within their IDEs, enhancing productivity and automating development tasks.
Blinky is an open-source AI-powered debugging agent integrated into VSCode that helps developers identify and fix backend code errors using large language models and advanced debugging techniques.
MCP LLMS-TXT Documentation Server is an open-source MCP server that exposes llms.txt files to IDEs and applications, providing controlled, auditable access to LLM documentation for enhanced development workflows.
GPT Runner is a versatile tool that enables AI-powered conversations with project files and manages AI presets to significantly boost development efficiency through CLI, web, and VSCode interfaces.
Image Processing & Analysis Agents
AgentLego is an open-source library that enhances large language model agents with a rich set of multimodal tool APIs for visual, speech, and image processing capabilities, supporting easy integration and remote tool access.
Phi-3-MLX is an AI framework optimized for Apple Silicon that integrates vision and language models for advanced multimodal AI tasks including text generation, visual question answering, and code execution.
Overeasy is a framework for orchestrating zero-shot computer vision models to build custom end-to-end pipelines for tasks like bounding box detection, classification, and segmentation without requiring large annotated datasets.
IoT & Smart Home Agents
Protofy is an open-source AI-driven machine automation platform that uses a large language model to control smart and industrial devices through a natural language autopilot system.
Low-Code/No-Code Platforms
Craftgen.ai is an open-source, no-code AI platform that enables users to build, customize, and deploy dynamic AI workflows and agents using a scalable graph architecture and comprehensive model support.
Medical & Healthcare
A curated repository of research papers and resources reviewing the transformative applications and challenges of Large Language Models in medicine and healthcare.
MedAgents is a Multi-disciplinary Collaboration framework leveraging large language models to facilitate zero-shot medical reasoning through expert gathering, iterative analysis, and consensus-driven decision making.
TxAgent is an AI agent that leverages multi-step reasoning and a comprehensive toolbox of biomedical tools to provide personalized and precise therapeutic treatment recommendations based on real-time biomedical knowledge and clinical guidelines.
Multi-Agent Medical Assistant is an AI-powered chatbot system that integrates multi-agent orchestration, advanced retrieval-augmented generation, and medical image analysis to assist healthcare professionals, researchers, and patients with medical diagnosis and healthcare research.
Multi-Agent Collaboration Systems
AI Flow is an advanced AI agentic framework designed to create dynamic, evolving digital AI agents with human-like interactions and autonomous capabilities on the BNB Chain.
AutoAct is an automatic agent learning framework that enables efficient and autonomous question answering through self-planning and division-of-labor among specialized sub-agents without relying on large-scale annotated data or closed-source models.
AgentKit is a TypeScript framework for building multi-agent AI networks with deterministic routing, shared state management, and rich tooling via MCP.
Daydreams is a cross-chain generative agent framework that enables intelligent agents to execute complex tasks across multiple blockchain networks and APIs with advanced reasoning and persistent memory capabilities.
NLSOM is a multi-agent AI system where diverse agents communicate in natural language to collaboratively solve tasks through a Mindstorm process, enhanced by a reward mechanism to optimize performance.
Sotopia is an open-ended, scalable social learning environment designed to evaluate and facilitate social intelligence in language agents through multi-agent interactions.
Devyan is an AI-powered software development assistant that orchestrates a team of GPT-based agents to collaboratively design, implement, test, and review programming solutions based on user input.
Westworldjs is a JavaScript framework designed for creating and managing multi-agent simulations, enabling complex interactions among autonomous agents in a simulated environment.
Anda is a Rust-based AI agent framework integrating ICP blockchain and Trusted Execution Environments to create a secure, autonomous, and composable network of AI agents aimed at building a super AGI system.
Fabrice AI is a lightweight, functional, and composable framework for building collaborative AI agents that solve complex tasks through workflows and task delegation.
LangGraph for Java is a library for building stateful, multi-agent applications with large language models, designed to work with langchain4j and offering advanced features like async support, checkpoints, and graph visualization.
Clippinator is an AI-driven programming assistant that autonomously plans, writes, debugs, and tests software projects using a multi-agent system based on GPT-4, enhancing developer productivity through automation and human collaboration.
A curated list and resource hub of top AI agents showcasing various autonomous and specialized AI systems for diverse applications.
OpenAGI is an open-source framework designed to enable the development of autonomous, human-like AI agents with long-term memory capabilities, applicable across various domains such as education, finance, and healthcare.
L2MAC is a groundbreaking LLM-based multi-agent framework that functions as a general-purpose stored-program automatic computer to generate extensive, unbounded outputs for complex tasks like large codebases and books, overcoming traditional LLM context window limitations.
CortexON is an open-source multi-agent AI system designed to automate complex everyday tasks and business processes through dynamic collaboration of specialized agents.
Open Multi-Agent Canvas is an open-source multi-agent chat interface that enables managing multiple agents in one dynamic conversation and supports MCP servers for advanced research and task management.
QRev is an open-source AI-first alternative to Salesforce that uses AI agents to automate and scale sales organizations with a modern, customizable CRM platform.
Agents-Flex is a Java-based LLM application framework similar to LangChain, providing tools for prompt management, function calling, memory, embedding, and multi-agent chains to build sophisticated language model applications.
AI Agent Roadmap is a curated resource that explores and lists the latest AI agent frameworks and vector databases, providing developers and researchers with a comprehensive guide to autonomous AI systems and tools.
The AEA Framework by Fetch.ai enables the creation of Autonomous Economic Agents that act independently to generate economic value for their owners in digital environments.
Evolving Agents Toolkit (EAT) is a Python-based framework for building autonomous, adaptive multi-agent ecosystems that orchestrate, evolve, and manage AI agents and tools to achieve complex, goal-driven workflows.
Jason is an open-source interpreter for an extended version of AgentSpeak, enabling the development of customizable multi-agent systems based on BDI agent-oriented logic programming.
EDSL is a domain-specific language for designing, conducting, and analyzing AI-powered surveys and experiments with multiple AI agents and large language models, facilitating computational social science and market research.
LAMBDA is an open-source multi-agent data analysis system that uses large language models to perform code-free, natural language-driven data analysis with iterative code generation and debugging, automatic report generation, and Jupyter Notebook export capabilities.
Westworld is an experimental Python library for multi-agent simulation and optimization, inspired by Unity ML Agents, supporting various environments and agent interactions with features for visualization and reinforcement learning integration.
PilottAI is a Python framework for building scalable, fault-tolerant, and intelligent multi-agent systems with advanced orchestration and LLM integration.
This project implements a multiagent debate framework to improve factuality and reasoning in language models, demonstrated across various tasks like math, biographies, and MMLU benchmarks.
VacAIgent is an AI-powered trip planning application that uses autonomous CrewAI agents and a Streamlit interface to create personalized vacation itineraries based on user preferences.
Mahilo is a flexible multi-agent framework that enables real-time, human-supervised interaction and context sharing among intelligent agents for enhanced collaboration and productivity.
SwarmZero SDK is a Python library for creating and managing AI agents and collaborative swarms using multiple large language model providers.
L3AGI is an open-source framework that enables AI agents to collaborate as effectively as human teams, providing tools for creating autonomous assistants, integrating diverse data sources, and managing multi-agent systems through a user-friendly interface and APIs.
Plexe is a machine learning framework that enables users to build and train models using natural language prompts and an AI-powered multi-agent system.
Autono is a highly robust autonomous agent framework based on the ReAct paradigm that enhances adaptive decision-making and multi-agent collaboration for solving complex tasks efficiently.
MetaGPT is a multi-agent AI framework that simulates a software company by assigning GPT-based roles to collaboratively automate software development from natural language requirements.
L3AGI is an open-source framework that enables AI assistants to collaborate as effectively as human teams, providing tools for autonomous agents, data integration, and team management.
Orchestration Frameworks
Jido is an autonomous agent framework for Elixir that enables building distributed, adaptive agent systems with composable workflows, real-time monitoring, and robust testing tools.
Agentsflow is a web-based drag-and-drop platform for creating, connecting, and running autonomous AI agents using the autogen framework, designed for both beginners and advanced users.
Enact is a Python framework for building, tracking, and improving generative software that integrates with machine learning models and APIs, offering features like recursive execution tracking, rewind/replay, and human-in-the-loop input sampling.
TrustGraph is an Autonomous Knowledge Operations Platform that automates and manages scalable, reliable AI agent operations using RAG pipelines, unified LLM access, and enterprise-grade infrastructure with full observability.
LionAGI is a comprehensive Intelligence Operating System framework that enables structured, multi-model AI orchestration with advanced reasoning, tool integrations, and real-time observability for building sophisticated AI workflows.
Flows AI is a lightweight, type-safe AI workflow orchestrator built on Vercel AI SDK that enables developers to create and execute complex AI workflows by orchestrating multiple AI agents with flexible input and output contracts.
Rill Flow is a high-performance, scalable distributed workflow orchestration engine designed for managing complex distributed workloads and integrating large language model services efficiently.
aiFlows is a modular and collaborative AI workflow framework that enables concurrent execution and peer-to-peer distributed collaboration through reusable and customizable computational building blocks called Flows.
A curated and comprehensive repository of references and resources on Azure OpenAI services and Large Language Models, covering topics from RAG and agent frameworks to prompting, fine-tuning, and LLM applications.
FastAgency is an open-source framework that enables developers to quickly and efficiently deploy multi-agent AI workflows from prototype to production with unified interfaces, API integration, and scalable distributed system support.
VoltAgent is an open-source TypeScript framework for building and orchestrating AI agents powered by Large Language Models, enabling developers to create scalable, customizable, and maintainable AI applications with modular components and visual monitoring.
RAG and Business Analytics
Super-Rag is a high-performance RAG pipeline offering summarization, retrieval, reranking, and code interpretation for AI applications through a simple and production-ready REST API.
Research Lists and Survey Projects
A comprehensive and organized repository of academic papers focused on GUI agents, covering datasets, benchmarks, models, frameworks, and multimodal vision-language foundation models to support research and development in GUI agent technologies.
A curated repository collecting typical Retrieval-Augmented Generation (RAG) papers, systems, surveys, and benchmarks from 2022 to 2024 to support research and development in the RAG field.
A comprehensive and regularly updated repository surveying research papers on Large Language Model-based agents applied to various software engineering tasks and agent architectures.
XLang Paper Reading is a curated collection of research papers focused on building and evaluating language model agents through executable language grounding, enabling natural language instructions to be transformed into executable actions in real-world environments.
A comprehensive curated repository of over 500 research papers and repositories covering a wide range of topics related to Large Language Models and their applications.
A comprehensive and regularly updated collection of ICLR papers and open-source projects focusing on large language models and NLP research from 2021 to 2025.
This repository hosts a comprehensive survey paper on OS Agents, which are MLLM-based agents that operate within operating system environments to automate tasks on computers, phones, and browsers, providing insights into foundational models, frameworks, evaluation, and safety aspects.
A comprehensive survey repository compiling research papers and insights on tool learning with large language models, focusing on benefits, implementation workflows, benchmarks, and future directions.
LLM-Agent-Benchmark-List is a curated and continuously updated collection of benchmarks for evaluating the performance and capabilities of large language models and agent-powered AI systems across various domains.
A comprehensive repository and survey paper focused on the latest advances and research in long chain-of-thought reasoning for large language models, providing resources, taxonomy, and future directions to enhance AI reasoning capabilities.
A knowledge base application designed to centralize and accelerate research on computer control agents by aggregating projects and benchmarks information.
A comprehensive and up-to-date research collection of papers on Large Language Model (LLM) agents, covering methodologies, applications, challenges, and key categories such as construction, collaboration, evolution, tools, security, benchmarks, and real-world applications.
A curated repository compiling academic papers and resources on the evaluation, alignment, and application of Large Language Models in Social Science, with a focus on psychology and intrinsic human values.
CL4R1T4S is a project that provides transparency by collecting and sharing the hidden system prompts and guidelines used by major AI models and agents to promote trust and understanding of AI behavior.
Research Platforms & Simulators
Synapse is an innovative agent that uses trajectory-as-exemplar prompting with memory to achieve state-of-the-art performance in computer control tasks on MiniWoB++ and Mind2Web benchmarks.
Spider2-V is a NeurIPS 2024 research project providing a virtual machine environment and dataset to evaluate multimodal agents' ability to automate data science and engineering workflows using state-of-the-art vision-language models.
DigiRL is a research project providing code and resources for training autonomous reinforcement learning agents to control Android devices in real-world environments using novel training algorithms and multiple training modes.
Security & Privacy Agents
EIA is a research project demonstrating a novel environmental injection attack that manipulates web agents to leak private user information by injecting malicious elements into benign websites' HTML, evading traditional malware detection and challenging current defense mechanisms.
JDBG is a powerful Java dynamic reverse engineering and debugging tool that operates at runtime, enabling deep inspection and manipulation of Java applications through bytecode and object analysis.
Agentic Radar is a security scanner tool designed to analyze and assess the security and operational aspects of agentic workflows using large language models, providing detailed reports and visualizations to identify vulnerabilities and improve transparency.
PentAGI is a fully autonomous AI-driven penetration testing system that integrates professional security tools, multi-agent AI collaboration, and scalable microservices architecture to deliver comprehensive and automated security assessments.
This project provides official code and data for studying and evaluating the adversarial robustness of multimodal large language model agents through various adversarial attacks and detailed evaluation methods.
Rogue is an intelligent automated web vulnerability scanner that uses Large Language Models to mimic human penetration testing for discovering and validating web application security weaknesses.
xnumon is a macOS monitoring agent that logs system activities to detect malware and intrusions, providing detailed event tracking and context information similar to Windows sysmon.
Aetherius AI Assistant is a private, locally-operated multi-agent AI framework with realistic long-term memory and versatile capabilities, designed for user control and privacy using open-source LLMs and advanced memory retrieval techniques.
Hades is an open-source Host-Based Intrusion Detection System leveraging eBPF and netlink technologies to monitor and analyze system events for detecting potential intrusions.
LiveRecall is an open-source tool that captures and encrypts screen snapshots, enabling users to recall them through natural language queries using semantic search technology.
PopupAttack is a research codebase demonstrating how adversarial pop-ups can successfully attack vision-language autonomous agents, causing them to misclick and fail tasks in environments like OSWorld and VisualWebArena.
Cloud Guardian is an AI-based Cybersecurity Assistant trained on expert interviews to provide deep knowledge and practical tools for cloud security in public and hybrid cloud environments.
mind-sdk-deepseek-rust is a Native Rust SDK by Mind Network that enables AI-driven predictions encrypted with Fully Homomorphic Encryption and submitted on-chain for decentralized model consensus.
An example API that uses a Stacks smart contract to verify access to resources and returns HTTP 402 responses for unpaid access attempts.
AdvWeb is a research project providing tools and methodologies for controllable black-box adversarial attacks on vision-language model-powered web agents to study their security and robustness.
jrasp-agent is a Java Runtime Application Self Protection system that enhances JVM application security by modifying bytecode to detect and block vulnerabilities in real time with minimal performance impact.
Shermie-Proxy is a versatile proxy packet capture tool that supports multiple protocols including HTTP, HTTPS, WebSocket, TCP, and Socks5, enabling real-time interception and modification of network data.
IvyCheck Python SDK enables developers to integrate hallucination, PII, and prompt injection detection checks into their Python applications for enhanced content verification and security.
Awesome Mind Network is a curated collection of open-source resources, SDKs, and tools by Mind Network designed to empower developers and researchers building privacy-preserving technologies, Agentic AI, and decentralized infrastructure.
Invariant Guardrails is a rule-based security layer that monitors and controls AI agent behavior to ensure safe and robust operation of LLM and MCP-powered applications.
Simulation & Benchmarking Environments
simple_rl is a simple and reproducible Python framework for experimenting with reinforcement learning algorithms and Markov Decision Processes.
VisualAgentBench (VAB) is a comprehensive benchmark for evaluating and developing large multimodal models as visual foundation agents across diverse embodied, GUI, and visual design tasks.
MMInA is a benchmark and environment for evaluating the long-chain reasoning abilities of multimodal internet agents across diverse tasks and domains.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
MobileAgentBench is an automated benchmarking framework for evaluating mobile large language model agents using common mobile applications on Android emulators or devices.
AgentStudio is a comprehensive toolkit providing environments, tools, and benchmarks for developing and evaluating general virtual agents that interact with computer software.
Open-operator-evals is an open-source benchmarking framework that evaluates the performance of web operators and agents using multiple metrics and a reproducible dataset to provide transparent and statistically sound comparisons.
AndroidWorld is a comprehensive environment and benchmark for autonomous agents to interact with and control Android devices through a live emulator, featuring diverse tasks and integration with web-based benchmarks.
CRAB is a Python-centric framework for building and benchmarking multimodal embodied language model agents across diverse cross-platform environments with a novel benchmarking suite and unified interface.
VisualWebArena is a benchmark for evaluating autonomous multimodal language agents on complex and realistic web-based visual tasks.
B-MoCA is a benchmarking testbed for evaluating mobile device control agents across diverse Android virtual device configurations using tools like Appium and Android Debug Bridge.
TravelPlanner is a benchmark for evaluating language agents in real-world travel planning tasks involving complex tool use and multiple constraints.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
WebWalker is a benchmarking project by Alibaba-NLP that evaluates large language models' capabilities in web traversal tasks using the WebWalkerQA dataset and a multi-agent framework for effective memory management.
MobileSafetyBench is a testbed for evaluating the safety and helpfulness of autonomous agents controlling mobile devices using Android emulators and automated interaction tools.
AgentStudio is a comprehensive toolkit offering environments, tools, and benchmarks to develop and evaluate general virtual agents capable of interacting with diverse computer software through GUI and API actions.
Task Automation & Workflow Orchestration
Giselle is an open-source AI app builder that enables seamless human-AI collaboration through agentic workflows, available as both a cloud service and self-hosted platform.
Lecca.io is a versatile AI platform for configuring, customizing, and automating Large Language Models with powerful tools and workflows to build intelligent AI agents.
BabyCommandAGI is a Python-based project that combines Command Line Interface (CLI) and Large Language Models (LLM) to automate and manage tasks dynamically through continuous interaction and execution.
Cognify is an AI tool that automatically optimizes generative AI agents and workflows to improve quality, reduce latency, and lower costs using hierarchical workflow-level autotuning.
InterfaceAgent is a versatile framework for creating intelligent system and interface agents that manage and automate tasks across mobile and desktop applications using advanced AI models and robust error handling.
Lemon Agent is a Plan-Validate-Solve agent designed for accurate and reliable workflow automation with extensive tool integrations and user-configurable workflows.
DockerShrink is an AI-powered command-line tool that optimizes and reduces the size of Docker images for Node.js applications by applying advanced techniques like multi-stage builds and lightweight base images.
Godmode is an AI-driven platform based on Auto-GPT that has rapidly gained one million users in three months, offering advanced autonomous AI capabilities through godmode.space.
maclaunch is a command-line tool for managing and controlling macOS startup items, allowing users to list, enable, or disable various system and user services without modifying configuration files.
AIBTC AI Agent Crew is a Python-based project that uses AI agents powered by CrewAI and Langchain to automate tasks related to Bitcoin and the Stacks blockchain, featuring a Streamlit UI and blockchain integration tools.
MobA is a novel mobile task automation system using a two-level agent architecture powered by multimodal large language models to efficiently manipulate mobile phones and execute complex tasks.
Agent Pilot is a versatile workflow automation platform that enables users to create, organize, and execute complex AI-driven workflows with customizable interfaces and multi-member collaboration.
A curated list of workflow automation software, tools, articles, and resources aimed at improving productivity and efficiency by automating repetitive tasks for individuals and teams.
AgentScript is an open-source SDK that enables building AI agents by generating and executing code plans in a safe runtime, allowing advanced workflow control, state management, and human interaction.
Agent Workflow Memory (AWM) is a system that induces, integrates, and utilizes workflows via agent memory to improve task-solving efficiency in both offline and online settings, achieving state-of-the-art results on platforms like WebArena and Mind2Web.
MoLing is a dependency-free, locally deployed office AI assistant MCP server that enables system and browser interactions for file management, command execution, and automation across multiple operating systems.
SheetCopilot is a framework that uses Large Language Models to assist users in manipulating spreadsheets through natural language commands, improving productivity and accessibility in software tasks.
AutoDroid is a research system that enables large language models to perform intelligent task automation on smartphones by interacting with app interfaces using screenshots and UI hierarchies.
Pandora FMS is an open-source, scalable monitoring solution that integrates network, server, application, and infrastructure monitoring with customizable alerts, reporting, and multitenant support.
BeeBot is an Autonomous AI Assistant designed to perform a wide range of practical tasks autonomously using intelligent tool selection and LLM integration, currently on hold pending advancements in AI capabilities.
PromethAI is an open-source Python-based framework that provides autonomous AI agents to help users navigate decision-making, set personalized goals, and execute tasks efficiently.
PC Agent is a framework that enables autonomous digital agents to mimic human cognition for controlling computers and completing complex tasks through human-computer interaction data collection, cognition completion, and a multi-agent system.
Open Agent Studio is an open-source cross-platform desktop application that enables advanced AI-driven Agentic Process Automation as a robust alternative to traditional RPA tools like UIpath, leveraging GPT models and semantic understanding for resilient automation workflows.
Quantalogic is a comprehensive AI agent framework that leverages large language models to enable dynamic problem-solving, structured workflows, and conversational AI with tool integrations for developers and enterprises.
ComfyBench is a benchmarking framework that evaluates LLM-based agents' ability to autonomously design collaborative AI systems by generating and executing workflows in ComfyUI.
ClickClickClick is a framework that enables autonomous control of Android devices and computers using any large language model (LLM), supporting multiple interfaces including web, CLI, and API for task automation.
Yourgoal is a Swift port of BabyAGI, an AI-powered task management system that autonomously creates, prioritizes, and executes tasks using OpenAI and Pinecone APIs.
Maige is an open-source AI-powered tool that automates GitHub issue and pull request management using natural language workflows to streamline repository maintenance.
Netbox-agent automates the creation and updating of hardware and network inventory in Netbox by using standard system tools to gather and synchronize infrastructure data.
ByteChef is an open-source, low-code platform for API integration and workflow automation that enables organizations and SaaS products to build, extend, and automate workflows across multiple applications and services with ease.
Training Datasets
GUI-World is a dataset and benchmark for evaluating and advancing multimodal large language models in understanding and interacting with dynamic graphical user interfaces, complemented by GUI-Vid, a fine-tuned VideoLLM for GUI video analysis.
MultiUI is a large-scale dataset and codebase for training and evaluating models on text-rich visual understanding of webpage user interfaces, featuring 7.3 million samples, pre-trained models, and comprehensive benchmark evaluations.
GUI Odyssey is a comprehensive dataset for training and evaluating cross-app navigation agents on mobile devices, featuring 7,735 episodes across multiple devices, apps, and navigation tasks.
UI Interaction
Auto-GUI is a multimodal AI agent framework that predicts user interface actions using a novel chain-of-action technique, enabling direct interaction with interfaces without environment parsing or application-specific APIs.
AGUVIS is a unified pure vision-based framework for autonomous GUI agents that operate across multiple platforms, leveraging a novel two-stage training pipeline and inner monologue for enhanced planning and reasoning.
This project introduces the Tree-of-Lens agent for layout-aware screen reading of graphical user interfaces based on user-indicated points, enhancing GUI understanding and navigation verification.
CoAT is a novel framework and benchmark dataset for enhancing GUI agents' decision-making in Android applications using Chain-of-Action-Thought modeling.
Mobile Next MCP is a platform-agnostic Model Context Protocol server enabling scalable mobile automation and interaction with iOS and Android devices, simulators, and emulators through accessibility snapshots and coordinate-based controls.
UGround is a universal visual grounding project for GUI agents enabling AI to navigate and interact with digital interfaces as humans do, achieving state-of-the-art results on multiple benchmarks.
OS-Atlas is a foundation action model for generalist GUI agents that enables precise visual grounding and interaction with UI elements in screenshots using advanced vision-language models.
Aria-UI is an open-source, fast, and context-aware action grounding system for GUI instructions, achieving state-of-the-art performance in dynamic agent tasks like AndroidWorld and OSWorld.
GUICourse is a project providing datasets, code, and models to train and evaluate versatile GUI agents by enhancing vision language models' abilities to understand and interact with graphical user interfaces.
VideoGUI is a multi-modal benchmark for GUI automation derived from instructional videos, focusing on visual-centric software and hierarchical task annotations to advance intelligent GUI interaction systems.
GPT-4V in Wonderland leverages large multimodal models for zero-shot navigation of smartphone GUIs, enabling AI agents to perform complex tasks like shopping on apps without prior training.
LlamaTouch is a scalable and faithful testbed for evaluating mobile UI automation agents by comparing execution traces with annotated essential states in real-world mobile environments.
Mobile-Env is a universal platform for training and evaluating agents that interact with mobile GUIs, supporting flexible task definitions and multiple observation and action modalities for Android apps.
A curated and continuously updated collection of research papers, models, tools, and datasets focused on UI agents that interact with various user interfaces across platforms like web, mobile, and OS environments.
GUI Action Narrator is a project that provides a dataset and framework for captioning GUI videos by detecting cursor actions and extracting keyframes to generate detailed narrations of GUI activities.
OS-Genesis is a project that automates the synthesis of high-quality and diverse GUI agent trajectory data using reverse task synthesis and a trajectory reward model for effective end-to-end training of GUI agents.
prompt2ui is a Next.js-based web application that converts textual prompts into interactive user interfaces using AI integration, designed for fun and experimentation.
Agent UI is a modern, customizable chat interface for AI agents featuring real-time streaming, tool call visualization, reasoning steps, and multi-modal content support, built with Next.js and Tailwind CSS.
Seq2act is a research project that maps natural language instructions to mobile UI action sequences using machine learning models for phrase extraction and grounding, enhancing mobile interaction automation.
SeeClick is a state-of-the-art project providing models, data, and code for advanced visual GUI agents with a focus on GUI grounding, featuring the ScreenSpot benchmark and superior performance on multi-platform GUI element prediction.
ScreenSpot-Pro is a comprehensive project providing tools, datasets, and benchmarks for GUI grounding tailored to professional high-resolution computer use, advancing AI-driven human-computer interaction.
Video Processing Agents
StreamRAG is a GPT-powered video search and streaming agent that enables real-time video search, summarization, and publishing of searchable video collections on the ChatGPT store.
Virtual Assistants
Whiz is a terminal copilot that uses OpenAI's language models to convert natural language requests into executable terminal commands, enhancing command-line productivity and ease of use.
Astra Assistants API is an open-source drop-in replacement for the OpenAI Assistants API v2, supporting multiple LLM providers, streaming, persistent threads, and flexible deployment options.
Rodel Agent is a Windows desktop application that integrates chat, text-to-speech, image generation, and machine translation using mainstream AI services to provide a comprehensive AI experience.
Suna is an open-source generalist AI assistant that automates real-world tasks through natural conversation, combining browser automation, file management, web crawling, and API integrations in a secure, modular architecture.
Newer issue
Older issue