Awesome AI AgentsData Integration and Specialized Solutions

zilliztech/deep-searcher

⭐ 8262 Python added to this list on 2025-04-19 repository created 2025-02-08

DeepSearcher is an open-source project designed as a deep research alternative for reasoning and searching on private data. It integrates advanced large language models (LLMs) such as OpenAI's models, DeepSeek, Grok, Claude, Llama, and others with vector databases like Milvus and Zilliz Cloud to enable highly accurate search, evaluation, and reasoning capabilities based on private enterprise data. The project is tailored for enterprise knowledge management, intelligent question-answering systems, and information retrieval scenarios, ensuring data security while maximizing the use of internal data. DeepSearcher supports multiple embedding models and vector database management with data partitioning for efficient retrieval. It also offers document loading from local files and is developing web crawling capabilities. The system allows flexible configuration of LLMs and embedding providers, supporting a variety of models and APIs, including OpenAI, DeepSeek, SiliconFlow, TogetherAI, XAI Grok, Claude, Google Gemini, PPIO, and Ollama. The project provides a quick start guide for installation via pip or development mode using the uv tool, with examples for setting up environment variables and configuring different LLMs. DeepSearcher aims to provide enterprises with a powerful tool to leverage their private data for intelligent insights and comprehensive reporting through a combination of vector search and large language model reasoning.

https://github.com/zilliztech/deep-searcher

claudedata-securitydeep-researchdeepsearcherdeepseekdocument-loaderembedding-modelsenterprise-knowledge-managementgrokinformation-retrievalintelligent-q&alarge-language-modelsllamallmsmilvusopen-sourceopenaiprivate-datapythonvector-databasesweb-crawlingzilliz-cloud

Also in Data Integration and Specialized Solutions

run-llama/llama_index

LlamaIndex is a leading data framework that enables building LLM-powered applications by providing tools for data ingestion, structuring, and advanced querying to augment large language models with private and external data.

pingcap/tidb

TiDB is an open-source, cloud-native, distributed SQL database offering ACID guarantees, horizontal scalability, high availability, HTAP capabilities, and MySQL compatibility.

getzep/graphiti

Graphiti is a framework for building and querying real-time, temporally-aware knowledge graphs designed to support AI agents in dynamic environments with efficient incremental updates and hybrid retrieval methods.

dolthub/dolt

Dolt is a SQL database with Git-like features, enabling full version control over data, including branching, merging, and diffing of data.

vectordotdev/vector

Vector is a high-performance, end-to-end observability data pipeline designed to collect, transform, and route logs, metrics, and traces.

cube-js/cube

Cube Core is an open-source semantic layer that enables AI, business intelligence, and embedded analytics by providing a flexible, API-driven platform supporting multiple SQL data sources and real-time analytics.

llmware-ai/llmware

llmware is a unified framework for building enterprise Retrieval-Augmented Generation (RAG) pipelines using small, specialized language models integrated with secure knowledge sources for efficient AI applications.

cocoindex-io/cocoindex

CoCoIndex is an ultra-performant data transformation framework for AI, specializing in incremental processing for data indexing and real-time semantic search with Python and Rust.