google-deepmind/acme
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
Awesome AI Agents › Multimodal Model Benchmarks
The project "llm-leaderboard" by JonathanChavezTamales is a comprehensive, community-driven repository that compiles extensive data and benchmark scores for large language models (LLMs). It serves as a centralized platform where users can compare and explore various language models through an interactive dashboard available at llm-stats.com. The repository includes detailed information on hundreds of LLMs, covering aspects such as model parameters, context window sizes, licensing details, capabilities, and more. It also provides data on provider pricing and performance metrics like throughput and latency, alongside standardized benchmark results. The project emphasizes data quality and accuracy, implementing a community review process for all changes and encouraging multiple source citations to ensure reliable information. Contributors are welcomed to update model data and provider information by following contribution guidelines and data formats specified in the repository. The leaderboard feature showcases a ranked list of language models based on various performance benchmarks, including GPQA, MMLU, MATH, HumanEval, and others, with detailed metrics such as input and output context sizes and release dates. The repository also fosters a community environment with a Discord channel for discussions and collaboration. Overall, this project is a valuable resource for researchers, developers, and enthusiasts interested in the performance and pricing of LLMs, facilitating informed decisions and comparisons in the rapidly evolving field of language models.
https://github.com/JonathanChavezTamales/llm-leaderboard
Acme is a flexible and scalable research framework providing modular reinforcement learning components and agents for developing and benchmarking RL algorithms.
AgentBench is a comprehensive benchmark platform designed to evaluate large language models as autonomous agents across diverse environments and tasks, facilitating research and development in LLM-based agent capabilities.
Windows Agent Arena is a scalable platform for testing and benchmarking multi-modal AI agents in a realistic Windows OS environment using Azure ML for large-scale deployment.
AgentGym is a versatile framework for developing, evaluating, and evolving large language model-based agents across diverse interactive environments with real-time feedback and scalability.
Benchmark and evaluation pipeline for deep research agents, scoring long-form generated reports on quality with the RACE framework and on citation trustworthiness with the FACT framework, plus a public leaderboard.
TheAgentCompany is a benchmarking platform that evaluates large language model agents on diverse real-world professional tasks within a simulated software company environment to assess their performance and impact on work-related activities.
rl-agents is a comprehensive collection of reinforcement learning and planning algorithms with tools for experiment management, benchmarking, and performance monitoring in gym-compatible environments.
Agent4Rec is a recommender system simulator using 1,000 LLM-powered generative agents initialized from MovieLens-1M to simulate realistic user interactions with personalized movie recommendations.