google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
VisualWebArena is a benchmark designed to evaluate multimodal autonomous language agents through a diverse set of complex web-based visual tasks. It extends the execution-based evaluation framework introduced by WebArena, focusing on realistic and varied challenges that test the capabilities of autonomous multimodal agents. The project includes a comprehensive suite of tasks that simulate real-world web environments, such as classifieds, shopping, Reddit, Wikipedia, and homepage navigation, among others. VisualWebArena supports end-to-end evaluation by providing standalone environments, configuration files, and scripts to facilitate testing and training of agents. It integrates with advanced language models like GPT-3.5, GPT-4V, and Gemini, enabling researchers to benchmark their agents' performance on visual and interactive web tasks. The project also offers pre-installed Amazon Machine Images for easy setup, human and GPT-4V + SoM agent trajectories for analysis, and demo scripts to run agents on arbitrary web pages. VisualWebArena emphasizes reproducibility and extensibility, allowing users to customize environments and evaluation protocols. It supports multimodal observations, including accessibility trees and image-based inputs, to enhance agent perception and interaction. The repository includes detailed installation instructions, environment setup guides, and example commands to run evaluations and demos. VisualWebArena is a valuable resource for advancing research in autonomous multimodal agents, providing a realistic benchmark that bridges language understanding, vision, and web interaction in a unified framework.
https://github.com/web-arena-x/visualwebarena
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.