Awesome AI AgentsData Integration and Specialized Solutions

Canner/wren-engine

⭐ 663 Java added to this list on 2025-04-09 repository created 2022-05-09

Wren Engine is a semantic engine designed specifically for Model Context Protocol (MCP) clients and AI agents, aiming to enhance enterprise-level data interactions by providing precise semantic understanding and governance over business data. It addresses the challenges enterprises face when dealing with complex structured data stored across cloud warehouses, relational databases, and secure filesystems. Unlike traditional systems that offer raw data access, Wren Engine adds a crucial semantic layer that interprets user intent, maps it accurately to the correct data, applies trusted calculations, and enforces user-based permissions and access controls. This ensures that AI-driven workflows operate with clarity and precision, especially in business contexts where terms like "active customer," "net revenue," or "churn rate" need exact definitions. The project is part of the MCP ecosystem, which is an open standard connecting large language models (LLMs) with various tools, databases, and enterprise systems. Wren Engine empowers AI agents by embedding semantic understanding directly into MCP clients, enabling them to interact with business data in a context-aware and governance-compliant manner. It supports interoperability with modern data stacks such as PostgreSQL, MySQL, and Snowflake, and is designed to be embeddable into any MCP client or AI agent workflow. Wren Engine consists of four main modules: ibis-server (a web server powered by FastAPI and Ibis), wren-core (the semantic core written in Rust and powered by Apache DataFusion), wren-core-py (Python bindings for the core), and mcp-server (the MCP server powered by the MCP Python SDK). The project is currently in beta, with active development and biweekly releases planned. The mission of Wren Engine is to fuel the next wave of AI agents by building a foundation for future MCP clients and enterprise data access, focusing on context-aware, composable systems that scale AI adoption across teams with better understanding and automation. The community is active, with support channels like Discord and GitHub for feedback and issue tracking.

https://github.com/Canner/wren-engine

agentagentic-aiaiai-agentsai-workflowsapache-datafusionbusiness-databusiness-intelligencecomposable-systemscontext-aware-aidatadata-accessdata-analysisdata-analyticsdata-governancedata-lakedata-modelingdata-securitydata-warehouseenterprise-datafastapihacktoberfestllmllm-integrationmcpmcp-servermodel-context-protocolmysqlpostgresqlpython-bindingrustsemanticsemantic-enginesemantic-layersemantic-understandingsnowflakesql

Also in Data Integration and Specialized Solutions

run-llama/llama_index

LlamaIndex is a leading data framework that enables building LLM-powered applications by providing tools for data ingestion, structuring, and advanced querying to augment large language models with private and external data.

pingcap/tidb

TiDB is an open-source, cloud-native, distributed SQL database offering ACID guarantees, horizontal scalability, high availability, HTAP capabilities, and MySQL compatibility.

getzep/graphiti

Graphiti is a framework for building and querying real-time, temporally-aware knowledge graphs designed to support AI agents in dynamic environments with efficient incremental updates and hybrid retrieval methods.

dolthub/dolt

Dolt is a SQL database with Git-like features, enabling full version control over data, including branching, merging, and diffing of data.

vectordotdev/vector

Vector is a high-performance, end-to-end observability data pipeline designed to collect, transform, and route logs, metrics, and traces.

cube-js/cube

Cube Core is an open-source semantic layer that enables AI, business intelligence, and embedded analytics by providing a flexible, API-driven platform supporting multiple SQL data sources and real-time analytics.

llmware-ai/llmware

llmware is a unified framework for building enterprise Retrieval-Augmented Generation (RAG) pipelines using small, specialized language models integrated with secure knowledge sources for efficient AI applications.

cocoindex-io/cocoindex

CoCoIndex is an ultra-performant data transformation framework for AI, specializing in incremental processing for data indexing and real-time semantic search with Python and Rust.