google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
TheAgentCompany is a comprehensive benchmarking platform designed to evaluate the performance of large language model (LLM) agents on real-world professional tasks within a simulated software company environment. The project addresses the growing interest in how AI agents can assist or autonomously perform work-related tasks by interacting with digital environments similar to human digital workers. These agents perform tasks such as browsing the web, writing code, running programs, and communicating with coworkers, providing a realistic and practical assessment of their capabilities. The benchmark includes a diverse set of tasks representing various professional roles including software engineers, product managers, data scientists, human resources, financial staff, and administrators. The tasks cover a wide range of data types and challenges such as coding, conversational interactions, mathematical reasoning, image processing, and text comprehension. This diversity ensures a thorough evaluation of the agents' abilities across different domains and task complexities. The platform is designed for ease of use and extensibility. It supports quick setup of servers either locally or on the cloud using Docker and Docker Compose, with pre-configured services like GitLab, Plane, ownCloud, and RocketChat to simulate a real company environment. Tasks are containerized as Docker images, each containing scripts for initialization, evaluation, and instructions for the agents. The benchmark can be run using the OpenHands platform or manually, allowing flexibility in how evaluations are conducted. Evaluation methods include deterministic and LLM-based evaluators, with a comprehensive scoring system that assesses both final results and intermediate subcheckpoints. The system supports multiple agent interactions and provides a leaderboard to track progress. The project is open-source under the MIT license and encourages contributions from the community. Overall, TheAgentCompany offers a robust and extensible framework for benchmarking AI agents on consequential real-world tasks, providing valuable insights for industry adoption and economic policy regarding AI's impact on the labor market.
https://github.com/TheAgentCompany/TheAgentCompany
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.
VisualWebArena is a benchmark for evaluating autonomous multimodal language agents on complex and realistic web-based visual tasks.