Awesome AI AgentsImage Processing & Analysis Agents

roboflow/inference

⭐ 2445 Python added to this list on 2025-04-19 repository created 2023-07-31

Roboflow Inference is a powerful platform that transforms any computer or edge device into a command center for computer vision projects. It enables users to self-host fine-tuned models and access state-of-the-art foundation models such as Florence-2, CLIP, and SAM2. The platform supports building and deploying complex workflows that allow for object detection, classification, segmentation, and integration of large multimodal models. Users can combine machine learning with traditional computer vision techniques like OCR, barcode reading, QR code scanning, and template matching. The system offers comprehensive monitoring, recording, and analysis of predictions, along with management of cameras and video streams. Notifications can be sent when specific events occur, and the platform can connect with external systems and APIs. It is highly extensible, allowing users to add their own code and models to workflows, making it suitable for production-scale deployments. The platform provides a quickstart guide with Docker and NVIDIA Container Toolkit support for GPU acceleration, and includes a Jupyter notebook server for interactive exploration. Example workflows demonstrate practical applications such as detecting small objects, multi-model consensus, active learning, license plate reading, face blurring, and background removal. The API and SDK facilitate integration and automation, enabling users to run models and workflows on images and video streams programmatically. Tutorials and videos are available to help users build AI-powered applications like self-serve checkouts and smart parking systems. Overall, Roboflow Inference is a comprehensive, flexible, and scalable solution for deploying computer vision models and workflows on edge devices and computers, empowering developers to create intelligent visual applications with ease.

https://github.com/roboflow/inference

agentsai-applicationsapiapisbarcode-readingclassificationclipcomputer-visiondeploymentdockeredge-deviceextensibilityexternal-systemsfine-tuned-modelsflorence-2foundation-modelsgpu-accelerationinferenceinference-apiinference-serverinstance-segmentationjetsonjupyter-notebooklarge-multimodal-modelsmachine-learningmonitoringnotificationsobject-detectionocronnxproduction-deploymentpythonqr-codesam2sdksegmentationself-serve-checkoutsmart-parkingtemplate-matchingtensorrttraditional-computer-visiontutorialsvideo-streamsvitworkflowsyolo11yolov12yolov5yolov8

Also in Image Processing & Analysis Agents

img2threejs/img2threejs

Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.

apple/ml-ferret

Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.

THUDM/CogVLM

CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.

11cafe/jaaz

Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.

SamurAIGPT/Generative-Media-Skills

Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.

QIN2DIM/hcaptcha-challenger

hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.

mbzuai-oryx/groundingLMM

GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.

LYL1015/JarvisArt

Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.