Awesome AI AgentsData Integration and Specialized Solutions

ruc-datalab/DeepAnalyze

⭐ 4617 Python added to this list on 2025-11-08 repository created 2025-10-11

DeepAnalyze is an innovative project that introduces the first agentic large language model (LLM) designed specifically for autonomous data science. It aims to automate the entire data science pipeline, enabling users to perform a wide range of data-centric tasks without human intervention. These tasks include data preparation, analysis, modeling, visualization, and report generation. DeepAnalyze supports open-ended data research by working with diverse data sources such as structured data (databases, CSV, Excel), semi-structured data (JSON, XML, YAML), and unstructured data (TXT, Markdown). The system ultimately produces analyst-grade research reports, making it a powerful tool for data scientists and analysts. The project is fully open-source, providing access to the model, code, training data, and demo, allowing users to deploy or extend their own data analysis assistants. DeepAnalyze has gained significant attention in the community, with over 1,000 GitHub stars and extensive social media engagement shortly after its release. It is developed by researchers from Renmin University of China and Tsinghua University. DeepAnalyze offers multiple user interfaces, including a WebUI and a JupyterUI, the latter integrating with Jupyter Lab to facilitate interaction through notebooks. The project supports deployment via vllm and provides detailed instructions for installation, setup, and usage. It encourages community contributions to improve the system and share use cases. Overall, DeepAnalyze represents a significant advancement in autonomous data science by leveraging agentic LLMs to automate complex data workflows, making data science more accessible and efficient.

https://github.com/ruc-datalab/DeepAnalyze

agentagenticagentic-aiagentic-llmaiai-scientistautonomous-data-sciencechatbotchatgptdatadata-analysisdata-centric-tasksdata-engineeringdata-modelingdata-preparationdata-researchdata-sciencedata-science-pipelinedata-visualizationdatabasegptjupyteruillamallmopen-sourceqwenrenmin-university-of-chinareport-generationresearch-reportssciencesemi-structured-datastructured-datatsinghua-universityunstructured-datavllmwebui

Also in Data Integration and Specialized Solutions

run-llama/llama_index

LlamaIndex is a leading data framework that enables building LLM-powered applications by providing tools for data ingestion, structuring, and advanced querying to augment large language models with private and external data.

pingcap/tidb

TiDB is an open-source, cloud-native, distributed SQL database offering ACID guarantees, horizontal scalability, high availability, HTAP capabilities, and MySQL compatibility.

getzep/graphiti

Graphiti is a framework for building and querying real-time, temporally-aware knowledge graphs designed to support AI agents in dynamic environments with efficient incremental updates and hybrid retrieval methods.

dolthub/dolt

Dolt is a SQL database with Git-like features, enabling full version control over data, including branching, merging, and diffing of data.

vectordotdev/vector

Vector is a high-performance, end-to-end observability data pipeline designed to collect, transform, and route logs, metrics, and traces.

cube-js/cube

Cube Core is an open-source semantic layer that enables AI, business intelligence, and embedded analytics by providing a flexible, API-driven platform supporting multiple SQL data sources and real-time analytics.

llmware-ai/llmware

llmware is a unified framework for building enterprise Retrieval-Augmented Generation (RAG) pipelines using small, specialized language models integrated with secure knowledge sources for efficient AI applications.

cocoindex-io/cocoindex

CoCoIndex is an ultra-performant data transformation framework for AI, specializing in incremental processing for data indexing and real-time semantic search with Python and Rust.