Awesome AI AgentsMulti-Agent Surveys

zhangxjohn/LLM-Agent-Benchmark-List

⭐ 170 added to this list on 2025-04-19 repository created 2024-01-29

The LLM-Agent-Benchmark-List project is a comprehensive and continuously updated resource that compiles a wide range of benchmarks for evaluating large language models (LLMs) and agent-powered systems. As LLMs become increasingly integral to artificial intelligence research and applications, this project addresses the critical need to systematically assess their performance across various dimensions. The repository organizes benchmarks into categories such as surveys, tool use, reasoning, knowledge, graph processing, video generation, code comprehension and generation, alignment, and agent capabilities. Each category includes references to recent academic papers and associated project pages, providing users with direct access to state-of-the-art evaluation methodologies and datasets. The project emphasizes the importance of evaluating not just raw performance but also the practical utility and reasoning abilities of LLMs, especially as they relate to advancing toward Artificial General Intelligence (AGI). By aggregating diverse benchmarks, the project serves as a valuable guide for researchers and developers seeking to understand the strengths and limitations of different LLMs and their applications in real-world scenarios. It also highlights ongoing research trends and emerging challenges in the field, such as tool manipulation, logical reasoning, social reasoning, knowledge integration, and code generation accuracy. The inclusion of agent-specific benchmarks further supports the evaluation of LLMs in interactive and autonomous roles, reflecting the growing interest in agent-based AI systems. Overall, the LLM-Agent-Benchmark-List is a vital resource for the AI community, fostering collaboration and informed development through its curated and accessible benchmark listings.

https://github.com/zhangxjohn/LLM-Agent-Benchmark-List

academic-papersagentagent-evaluationagialignmentartificial-intelligencebenchmarkcode-comprehensioncode-generationevaluationgraph-processingknowledge-integrationlarge-language-modelsllmnatural-language-processingreasoningresearch-surveysurveytool-usevideo-generation

Also in Multi-Agent Surveys

sindresorhus/awesome

A curated collection of awesome lists covering a wide range of technology topics and development resources.

Paitesanshi/LLM-Agent-Survey

A comprehensive survey on the construction, application, and evaluation of autonomous agents powered by large language models, providing a foundational resource for researchers and practitioners in the field.

AGI-Edgerunners/LLM-Agents-Papers

A comprehensive repository listing academic papers related to large language model (LLM) based agents, covering surveys, enhancement techniques, interactions, applications, automation, training, scaling, stability, and infrastructure.

taichengguo/LLM_MultiAgents_Survey_Papers

A comprehensive repository and survey of research papers on Large Language Model based Multi-Agent systems, covering frameworks, orchestration, problem solving, world simulation, and datasets.

LightChen233/Awesome-Long-Chain-of-Thought-Reasoning

A comprehensive repository and survey paper focused on the latest advances and research in long chain-of-thought reasoning for large language models, providing resources, taxonomy, and future directions to enhance AI reasoning capabilities.

yinizhilian/ICLR2025-Papers-with-Code

A comprehensive and regularly updated collection of ICLR papers and open-source projects focusing on large language models and NLP research from 2021 to 2025.

FudanSELab/Agent4SE-Paper-List

A comprehensive and regularly updated repository surveying research papers on Large Language Model-based agents applied to various software engineering tasks and agent architectures.

OS-Agent-Survey/OS-Agent-Survey

This repository hosts a comprehensive survey paper on OS Agents, which are MLLM-based agents that operate within operating system environments to automate tasks on computers, phones, and browsers, providing insights into foundational models, frameworks, evaluation, and safety aspects.