google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.
VisualWebArena is a benchmark for evaluating autonomous multimodal language agents on complex and realistic web-based visual tasks.
CRAB is a Python-centric framework for building and benchmarking multimodal embodied language model agents across diverse cross-platform environments with a novel benchmarking suite and unified interface.
OpenRCA is a Microsoft benchmark that measures how well large language models perform root cause analysis over software telemetry, shipping datasets, an evaluation harness and an RCA-agent baseline.
A community-driven repository providing comprehensive benchmark scores, pricing, and detailed data for large language models, accessible via an interactive dashboard.
simple_rl is a simple and reproducible Python framework for experimenting with reinforcement learning algorithms and Markov Decision Processes.
VisualAgentBench (VAB) is a comprehensive benchmark for evaluating and developing large multimodal models as visual foundation agents across diverse embodied, GUI, and visual design tasks.
AgentStudio is a comprehensive toolkit offering environments, tools, and benchmarks to develop and evaluate general virtual agents capable of interacting with diverse computer software through GUI and API actions.
MMInA is a benchmark and environment for evaluating the long-chain reasoning abilities of multimodal internet agents across diverse tasks and domains.
MobileAgentBench is an automated benchmarking framework for evaluating mobile large language model agents using common mobile applications on Android emulators or devices.
MobileSafetyBench is a testbed for evaluating the safety and helpfulness of autonomous agents controlling mobile devices using Android emulators and automated interaction tools.
B-MoCA is a benchmarking testbed for evaluating mobile device control agents across diverse Android virtual device configurations using tools like Appium and Android Debug Bridge.
AgentStudio is a comprehensive toolkit providing environments, tools, and benchmarks for developing and evaluating general virtual agents that interact with computer software.
FedMABench is an open-source benchmark platform for federated training and evaluation of mobile agents on decentralized heterogeneous user data, featuring diverse datasets, federated algorithms, and advanced base models.