Awesome AI AgentsData Integration and Specialized Solutions

chainbase-labs/manuscript-core

⭐ 692 Java added to this list on 2025-04-19 repository created 2024-09-24

Manuscript is a revolutionary blockchain data streaming framework developed by Chainbase Labs. It enables seamless integration of on-chain and off-chain blockchain data into various target data storage systems for unrestricted querying and analysis. Manuscript acts as a protocol, framework, and toolkit that simplifies and unifies data access and processing methods across the Chainbase global blockchain data network. It supports programmability, interoperability, and monetization of blockchain data, allowing developers and users to customize data workflows, integrate multi-chain and off-chain data, and create a fair data value exchange ecosystem. Manuscript supports multiple programming languages including Golang, Rust, Python, Node.js, Java, C/C++, Zig, and WebAssembly, and offers diverse data access methods such as SQL, DataFrames, HTTPS, gRPC, FTP, WebDAV, and FUSE. It can handle various data formats like JSON, CSV, ORC, XML, XLSX, and BLOB, and supports data storage services including RPC, S3, IPFS, Azblob, HDFS, Google Drive, BigQuery, WebDAV, MySQL, and PostgreSQL. The framework includes a GUI and CLI client, supports local and distributed deployment with Docker and Kubernetes, and integrates with stream processing frameworks like Flink. Manuscript's vision is to enable data trade within the Chainbase ecosystem, providing a unified language and interface for accessing blockchain data across any service, format, or programming language. The project roadmap includes enhancements such as user-defined functions for blockchain data parsing, advanced data processing APIs, lightweight Kubernetes deployment, distributed edge node coordinators, and support for RPC and substream data formats. Overall, Manuscript empowers developers to build complex data analysis pipelines and cross-chain applications with ease, fostering innovation and growth in the blockchain data ecosystem.

https://github.com/chainbase-labs/manuscript-core

agentsaiapisazblobbigqueryblobblockchainchainbaseclicryptocsvdatadata-analysisdata-formatsdata-integrationdata-queryingdata-storagedata-streamingdata-tradedataframesdockeredge-computingflinkframeworkftpfusegoogle-drivegraphqlgrpcguihdfshttpsinteroperabilityipfsjsonkubernetesmonetizationmulti-chainmysqloff-chain-dataon-chain-dataorcpostgresqlprogrammabilityprotocolrpcs3sqlstreamingtoolkituser-defined-functionsweb3webdavxlsxxml

Also in Data Integration and Specialized Solutions

run-llama/llama_index

LlamaIndex is a leading data framework that enables building LLM-powered applications by providing tools for data ingestion, structuring, and advanced querying to augment large language models with private and external data.

pingcap/tidb

TiDB is an open-source, cloud-native, distributed SQL database offering ACID guarantees, horizontal scalability, high availability, HTAP capabilities, and MySQL compatibility.

getzep/graphiti

Graphiti is a framework for building and querying real-time, temporally-aware knowledge graphs designed to support AI agents in dynamic environments with efficient incremental updates and hybrid retrieval methods.

dolthub/dolt

Dolt is a SQL database with Git-like features, enabling full version control over data, including branching, merging, and diffing of data.

vectordotdev/vector

Vector is a high-performance, end-to-end observability data pipeline designed to collect, transform, and route logs, metrics, and traces.

cube-js/cube

Cube Core is an open-source semantic layer that enables AI, business intelligence, and embedded analytics by providing a flexible, API-driven platform supporting multiple SQL data sources and real-time analytics.

llmware-ai/llmware

llmware is a unified framework for building enterprise Retrieval-Augmented Generation (RAG) pipelines using small, specialized language models integrated with secure knowledge sources for efficient AI applications.

cocoindex-io/cocoindex

CoCoIndex is an ultra-performant data transformation framework for AI, specializing in incremental processing for data indexing and real-time semantic search with Python and Rust.