IBM/AssetOpsBench
Open IBM framework and benchmark for building, orchestrating and evaluating domain-specific AI agents in industrial asset operations, with MCP servers over sensors, failure modes, time series and work orders.
Awesome AI Agents › Sensor Fusion Agents
OSWorld is a comprehensive benchmarking platform designed for evaluating multimodal agents tasked with open-ended activities within real computer environments. Presented at NeurIPS 2024, this project provides a unique environment where AI agents can interact with virtual machines running real operating systems such as Ubuntu and Windows, enabling the testing of their capabilities in practical, real-world scenarios. The platform supports various virtualization technologies including VMware, VirtualBox, and Docker with KVM support, making it versatile for different hardware and server setups. OSWorld offers a rich set of tools and interfaces for setting up virtual environments, running tasks, and evaluating agent performance through screenshots, command execution, and video recordings. It includes baseline agents, such as those using GPT-4V, to facilitate benchmarking and comparison. The project is well-documented with a detailed README, installation guides, and example scripts to help users quickly get started. It also provides a data viewer, evaluation examples, and a Discord community for support and collaboration. OSWorld aims to push the boundaries of AI agent research by providing a realistic and challenging testbed for multimodal interaction, task execution, and autonomous decision-making in computer environments. This makes it a valuable resource for researchers and developers working on AI agents, reinforcement learning, and human-computer interaction in complex digital settings.
https://github.com/xlang-ai/OSWorld
Open IBM framework and benchmark for building, orchestrating and evaluating domain-specific AI agents in industrial asset operations, with MCP servers over sensors, failure modes, time series and work orders.
WebArena is a self-hostable web environment designed for building and evaluating autonomous agents capable of realistic web navigation and interaction tasks.
AndroidEnv is a Python library by DeepMind that provides a Reinforcement Learning platform by simulating Android devices, enabling agents to interact with real-world Android applications through touchscreen gestures for diverse RL tasks.
AndroidWorld is a comprehensive environment and benchmark for autonomous agents to interact with and control Android devices through a live emulator, featuring diverse tasks and integration with web-based benchmarks.
An Extended Kalman Filter implementation that fuses GPS, IMU, and encoder sensor data to accurately estimate the pose of a ground robot in a navigation frame.
TravelPlanner is a benchmark for evaluating language agents in real-world travel planning tasks involving complex tool use and multiple constraints.
AppWorld is a high-fidelity execution environment simulating a world of apps and people to benchmark and evaluate interactive coding agents through diverse and challenging autonomous tasks.
Open-operator-evals is an open-source benchmarking framework that evaluates the performance of web operators and agents using multiple metrics and a reproducible dataset to provide transparent and statistically sound comparisons.