google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
AgentStudio is a comprehensive toolkit designed to facilitate the development, evaluation, and benchmarking of general virtual agents capable of interacting with any computer software. It provides a trinity of environments, tools, and benchmarks aimed at creating robust, general, and open-ended virtual agents. The environment is lightweight and interactive, featuring highly generic observation and action spaces such as video observations and GUI/API actions, which significantly expand the task space for agents to operate in real-world settings. The toolkit includes tools for creating online benchmark tasks, annotating GUI elements, and labeling actions in videos, enabling detailed and structured data generation for training and evaluation. AgentStudio offers a suite of online benchmark tasks that cover both GUI interactions and function calling, complete with auto-evaluation and language feedback to assess agent performance effectively. Additionally, AgentStudio provides three benchmark datasets—GroundUI, IDMBench, and CriticBench—that focus on fundamental agent abilities including GUI grounding, learning from videos, and success detection. These datasets help in gaining deeper insights into agent capabilities beyond overall performance metrics. The project supports a wide range of tasks spanning API usages like terminal commands and Gmail interactions, as well as GUI software such as VS Code. It also includes tools for benchmark task creation and validation, step-level GUI element annotation, and trajectory-level video-action recording and refinement. AgentStudio is well-documented and actively maintained, with a focus on open-source collaboration and community contributions. It is licensed under AGPL v3 and encourages contributions to improve its capabilities. The project is supported by a detailed paper available on arXiv and a dedicated project page for further resources and updates. Keywords extracted include environments, tools, benchmarks, virtual agents, computer software interaction, generic observation spaces, GUI actions, API actions, online benchmark tasks, auto-evaluation, language feedback, benchmark datasets, GUI grounding, video learning, success detection, task creation, annotation tools, video-action labeling, open-source, and agent capabilities.
https://github.com/SkyworkAI/agent-studio
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.