google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
AgentGym is a comprehensive framework designed to facilitate the development, evaluation, and evolution of large language model (LLM)-based agents across a wide variety of interactive environments. The project is motivated by the goal of building generalist AI agents capable of handling diverse tasks and evolving themselves in different settings. AgentGym provides a unified platform that supports real-time feedback, concurrency, and scalability, making it easier for researchers and developers to explore and benchmark LLM-based agents. The framework includes 14 distinct environments spanning web navigation, text games, household tasks, digital games, embodied tasks, tool usage, and programming. Each environment is accessible via encapsulated HTTP services, allowing for modular and decoupled interaction. AgentGym also offers a high-quality trajectory dataset called AgentTraj and a benchmark suite named AgentEval, both available on Hugging Face. These resources enable standardized evaluation and training of agents. A novel method called AgentEvol is introduced within the project to explore the potential for agent self-evolution beyond previously seen data, demonstrating competitive performance with state-of-the-art models. The platform architecture separates environment servers from the core agent controller, which manages agent evaluation, data collection, and training. The project encourages community contributions to expand the range of environments and capabilities. Overall, AgentGym represents a significant step towards creating versatile, evolving AI agents by providing a rich ecosystem of tools, datasets, benchmarks, and environments for research and development in the field of LLM-based agents.
https://github.com/WooooDyy/AgentGym
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.
VisualWebArena is a benchmark for evaluating autonomous multimodal language agents on complex and realistic web-based visual tasks.