xlang-ai/OSWorld
OSWorld is a benchmarking platform for evaluating multimodal AI agents performing open-ended tasks in real computer environments using virtual machines and various virtualization technologies.
Awesome AI Agents › Sensor Fusion Agents
AppWorld is a sophisticated and high-fidelity execution environment designed to simulate a controllable world of applications and people for benchmarking interactive coding agents. It features nine day-to-day applications accessible through 457 APIs, populated with digital activities of approximately 100 simulated individuals living in a virtual world. This environment enables the creation and evaluation of natural, diverse, and challenging autonomous agent tasks that require rich and interactive coding capabilities. The project provides a comprehensive benchmark for testing the performance and capabilities of interactive coding agents in a realistic and complex setting. The repository includes extensive resources such as a website, task explorer, API explorer, leaderboard, videos, blog posts, and a research paper published as the ACL'24 Best Resource Paper. Users can install the AppWorld package in a Python 3.11+ environment and access a command-line interface (CLI) that supports various commands for installation, data download, task exploration, environment serving, interactive playground, agent running, evaluation, and leaderboard management. AppWorld supports a minimal working ReAct agent example and offers detailed guides for building base databases, tasks, and customizing agent implementations. The environment emphasizes code execution safety and agent development restrictions to ensure secure and controlled interactions. It also provides options for serving the environment and APIs with or without Docker, facilitating flexible deployment. The project is licensed under Apache 2.0 and welcomes contributions from the community. It aims to advance research in interactive coding agents by providing a realistic and controllable testbed that mimics real-world app usage and human interactions, enabling the development and benchmarking of autonomous agents capable of complex coding and API interactions.
https://github.com/StonyBrookNLP/appworld
OSWorld is a benchmarking platform for evaluating multimodal AI agents performing open-ended tasks in real computer environments using virtual machines and various virtualization technologies.
Open IBM framework and benchmark for building, orchestrating and evaluating domain-specific AI agents in industrial asset operations, with MCP servers over sensors, failure modes, time series and work orders.
WebArena is a self-hostable web environment designed for building and evaluating autonomous agents capable of realistic web navigation and interaction tasks.
AndroidEnv is a Python library by DeepMind that provides a Reinforcement Learning platform by simulating Android devices, enabling agents to interact with real-world Android applications through touchscreen gestures for diverse RL tasks.
AndroidWorld is a comprehensive environment and benchmark for autonomous agents to interact with and control Android devices through a live emulator, featuring diverse tasks and integration with web-based benchmarks.
An Extended Kalman Filter implementation that fuses GPS, IMU, and encoder sensor data to accurately estimate the pose of a ground robot in a navigation frame.
TravelPlanner is a benchmark for evaluating language agents in real-world travel planning tasks involving complex tool use and multiple constraints.
Open-operator-evals is an open-source benchmarking framework that evaluates the performance of web operators and agents using multiple metrics and a reproducible dataset to provide transparent and statistically sound comparisons.