Awesome AI AgentsNL AI Frameworks

ScrapeGraphAI/Scrapegraph-ai

⭐ 30916 Python added to this list on 2025-04-19 repository created 2024-01-27

ScrapeGraphAI is a Python-based web scraping library that leverages large language models (LLMs) and graph logic to create intelligent scraping pipelines for extracting data from websites and local documents such as XML, HTML, JSON, and Markdown. The library simplifies the web scraping process by allowing users to specify the information they want to extract, and it automatically handles the scraping and data extraction tasks. It supports multiple scraping pipelines, including single-page scrapers like SmartScraperGraph, multi-page scrapers like SearchGraph, and specialized pipelines that generate Python scripts or audio files from scraped data. The library is designed to work with various LLMs, including OpenAI, Groq, Azure, Gemini, and local models via Ollama, providing flexibility in how the scraping intelligence is powered. ScrapeGraphAI offers a user-friendly API and SDKs in Python and Node.js, making it easy to integrate into different projects. The library also supports parallel processing for multi-page scraping tasks, enhancing efficiency. Installation is straightforward via pip, with additional setup for Playwright to fetch website content. The project is well-documented with extensive guides, examples, and multilingual support, and it encourages community contributions through its Discord server and GitHub repository. Overall, ScrapeGraphAI aims to transform websites into clean, organized data for AI agents and data analytics, providing a cost-effective and effortless data extraction solution.

https://github.com/ScrapeGraphAI/Scrapegraph-ai

aiai-powered-scrapingapiautomated-scraperautomationazuredata-analyticsdata-extractiongeminigpt-3gpt-4groqlarge-language-modelsllama3llmmachine-learningmulti-page-scrapingollamaopenaiplaywrightpythonscscrapingscraping-pipelinesscraping-pythonscrapingwebscriptcreatorgraphsdksearchgraphsmartscrapergraphspeechgraphweb-scrapingwebscraping

Also in NL AI Frameworks

infiniflow/ragflow

RAGFlow is an open-source Retrieval-Augmented Generation engine that leverages deep document understanding and Large Language Models to provide accurate, citation-backed question-answering from complex and diverse data sources.

pydantic/pydantic

Pydantic is a Python library for fast and extensible data validation using Python type hints, enabling developers to define and validate data models efficiently.

apache/doris

Apache Doris is a high-performance, real-time analytical database with a storage-compute integrated architecture, designed for sub-second query response and high concurrency in diverse data analysis scenarios.

business-science/ai-data-science-team

AI Data Science Team is a Python library featuring AI-powered agents that automate and accelerate common data science tasks, including data cleaning, feature engineering, machine learning, and exploratory data analysis, to improve productivity and efficiency.

ucbepic/docetl

DocETL is a system for creating and executing complex document processing and data transformation pipelines powered by large language models, featuring an interactive UI playground and a Python package for production use.

hitsz-ids/synthetic-data-generator

Synthetic Data Generator (SDG) is a comprehensive framework for generating high-quality, privacy-compliant synthetic structured tabular data using advanced statistical and LLM-based models, optimized for big data applications.

Mintplex-Labs/vector-admin

VectorAdmin is a universal and user-friendly tool suite for managing multiple vector databases, providing full control over vector data with multi-user support, embedding management, and cloud deployment capabilities.

scouter-project/scouter

Scouter is an open source Application Performance Management (APM) tool that monitors and manages the performance of software applications and system resources across various platforms and services.