Awesome AI Agents › Multimodal Model Benchmarks

Multimodal Model Benchmarks

21 projects

google-deepmind/acme

Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.

⭐ 4061 Python

THUDM/AgentBench

AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.

⭐ 3730 Python

microsoft/WindowsAgentArena

Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.

⭐ 896 Python

WooooDyy/AgentGym

AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.

⭐ 844 Python added 2025-03-09

Ayanami0730/deep_research_bench

Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.

⭐ 829 Python

TheAgentCompany/TheAgentCompany

TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.

⭐ 779 Python added 2025-03-09

eleurent/rl-agents

rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.

⭐ 727 Python added 2025-04-19

LehengTHU/Agent4Rec

Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.

⭐ 503 Python added 2025-04-19

web-arena-x/visualwebarena

VisualWebArena is a benchmark for evaluating autonomous multimodal language agents on complex and realistic web-based visual tasks.

⭐ 486 Python added 2025-03-09

camel-ai/crab

CRAB is a Python-centric framework for building and benchmarking multimodal embodied language model agents across diverse cross-platform environments with a novel benchmarking suite and unified interface.

⭐ 426 Python added 2025-04-19

microsoft/OpenRCA

OpenRCA is a Microsoft benchmark that measures how well large language models perform root cause analysis over software telemetry, shipping datasets, an evaluation harness and an RCA-agent baseline.

⭐ 419 Python

JonathanChavezTamales/llm-leaderboard

A community-driven repository providing comprehensive benchmark scores, pricing, and detailed data for large language models, accessible via an interactive dashboard.

⭐ 356 JavaScript added 2025-06-16

david-abel/simple_rl

simple_rl is a simple and reproducible Python framework for experimenting with reinforcement learning algorithms and Markov Decision Processes.

⭐ 335 Python added 2025-03-09

THUDM/VisualAgentBench

VisualAgentBench (VAB) is a comprehensive benchmark for evaluating and developing large multimodal models as visual foundation agents across diverse embodied, GUI, and visual design tasks.

⭐ 276 Python added 2025-04-13

ltzheng/agent-studio

AgentStudio is a comprehensive toolkit offering environments, tools, and benchmarks to develop and evaluate general virtual agents capable of interacting with diverse computer software through GUI and API actions.

⭐ 232 Python added 2025-05-02

shulin16/MMInA

MMInA is a benchmark and environment for evaluating the long-chain reasoning abilities of multimodal internet agents across diverse tasks and domains.

⭐ 54 Python added 2025-04-19

MobileAgentBench/mobile-agent-bench

MobileAgentBench is an automated benchmarking framework for evaluating mobile large language model agents using common mobile applications on Android emulators or devices.

⭐ 37 Python added 2025-04-19

jylee425/mobilesafetybench

MobileSafetyBench is a testbed for evaluating the safety and helpfulness of autonomous agents controlling mobile devices using Android emulators and automated interaction tools.

⭐ 37 Jupyter Notebook added 2025-04-19

jylee425/b-moca

B-MoCA is a benchmarking testbed for evaluating mobile device control agents across diverse Android virtual device configurations using tools like Appium and Android Debug Bridge.

⭐ 33 Jupyter Notebook added 2025-04-19

SkyworkAI/agent-studio

AgentStudio is a comprehensive toolkit providing environments, tools, and benchmarks for developing and evaluating general virtual agents that interact with computer software.

⭐ 18 Python added 2025-04-19

wwh0411/FedMABench

FedMABench is an open-source benchmark platform for federated training and evaluation of mobile agents on decentralized heterogeneous user data, featuring diverse datasets, federated algorithms, and advanced base models.

⭐ 17 Python added 2025-06-16