Awesome AI AgentsWeb Automation Systems

platonai/PulsarRPA

⭐ 3 Kotlin added to this list on 2025-03-16 repository created 2018-03-12

PulsarRPA is a high-performance, distributed, and open-source Robotic Process Automation (RPA) framework designed for large-scale web automation and data extraction. It specializes in browser automation, web content understanding, and comprehensive data scraping from complex and dynamic websites. The framework supports scalable crawling with browser rendering and AJAX data extraction, making it suitable for spider-grade web scraping tasks. PulsarRPA integrates Large Language Models (LLM) to enable natural language web content analysis and intuitive content description, enhancing the automation capabilities with AI-powered features such as automatic field extraction and pattern recognition. The project offers a variety of user-friendly interfaces, including simple curl commands for beginners to chat about webpages or extract data, and advanced SQL-like query interfaces (X-SQL) for complex data extraction and web business intelligence. It also provides native API support for expert users to control browser actions, perform one-line scraping, and implement advanced RPA crawling workflows. PulsarRPA emphasizes developer convenience with features like one-line data extraction, simple API integration, and extended SQL for web data mining. Key features include advanced bot protection techniques such as IP rotation and privacy context management, parallel page rendering for high efficiency, and a block-resistant design to ensure reliable operation. The framework is cost-effective, capable of processing over 100,000 pages per day with minimal hardware requirements. It supports multiple storage options including local file systems, MongoDB, HBase, and Gora, and offers comprehensive monitoring with detailed logging and metrics. PulsarRPA is scalable with a fully distributed architecture suitable for enterprise-level deployments. It provides smart retry mechanisms, precise scheduling, and complete lifecycle management to ensure quality assurance. The project is well-documented with demo videos, advanced guides, and example codes, making it accessible to users of varying skill levels. Contact and support information is readily available, enhancing community engagement and assistance.

https://github.com/platonai/PulsarRPA

ai-agentsai-crawlerai-integrationai-rpaajax-data-extractionapi-integrationbot-protectionbrowser-automationcost-effectivecrawlerdata-extractiondistributed-systemgorahbasehigh-performanceip-rotationlarge-language-modelsllmloggingmongodbmonitoringnatural-language-processingparallel-processingpattern-recognitionprivacy-managementrobotic-process-automationrpascalable-architecturescraperspider-grade-crawlingsql-like-queryweb-automationweb-scrapingworkflow-automationx-sql

Also in Web Automation Systems

firecrawl/firecrawl

Firecrawl is an advanced web data API that crawls and scrapes entire websites to convert content into clean, LLM-ready markdown or structured data for AI applications.

nanobrowser/nanobrowser

Nanobrowser is an open-source Chrome extension that enables AI-powered web automation through a multi-agent system using user-configured LLM API keys, offering a privacy-focused and cost-effective alternative to commercial tools like OpenAI Operator.

steel-dev/steel-browser

Steel Browser is an open-source browser API that enables developers to build AI-powered web agents and automation tools with full browser control, session management, proxy support, and debugging features, simplifying web automation without infrastructure overhead.

fake-useragent/fake-useragent

fake-useragent is a Python package that provides an up-to-date and customizable user-agent faker using a real-world database for realistic user-agent strings.

ishan0102/vimGPT

vimGPT is a project that combines GPT-4V's vision capabilities with the Vimium keyboard navigation extension to enable AI-assisted web browsing through visual and keyboard interactions.

brightdata/brightdata-mcp

Bright Data MCP is a powerful Model Context Protocol server that enables AI agents and applications to access and extract real-time web data seamlessly, bypassing geo-restrictions and bot protections for enhanced web scraping and navigation.

spider-rs/spider

A concurrency-first web crawler and scraper written in Rust that streams pages as they arrive and renders JavaScript only when needed, with bindings for Node.js and Python and an MCP server for agents.

nottelabs/notte

Notte is an open-source full stack framework that creates intelligent web browsing agents using a perception layer to enable fast, reliable, and cost-effective interactions with websites through large language models.