Awesome AI Agents › Image Processing & Analysis Agents

Image Processing & Analysis Agents

15 projects

img2threejs/img2threejs

Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.

⭐ 15923 Python added 2026-08-10

apple/ml-ferret

Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.

⭐ 8663 Python added 2025-04-19

THUDM/CogVLM

CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.

⭐ 6740 Python added 2025-04-19

11cafe/jaaz

Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.

⭐ 6637 TypeScript added 2025-06-16

SamurAIGPT/Generative-Media-Skills

Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.

⭐ 3652 Shell added 2026-03-01

QIN2DIM/hcaptcha-challenger

hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.

⭐ 2493 Python added 2025-04-09

roboflow/inference

Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.

⭐ 2445 Python added 2025-04-19

mbzuai-oryx/groundingLMM

GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.

⭐ 968 Python added 2025-04-19

LYL1015/JarvisArt

Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.

⭐ 864 JavaScript

taco-group/4KAgent

NeurIPS 2025 multi-agent system that upscales any image to 4K resolution, pairing a vision-language perception agent that plans restoration with a restoration agent that executes, reflects and rolls back.

⭐ 823

VILA-Lab/FigMirror

Agent skill for Claude Code and Codex that copies the visual style of a reference paper figure onto your own data through a Drawer and Reviewer loop, producing an editable matplotlib script and a camera-ready PDF.

⭐ 511 Python

LYL1015/JarvisEvo

Self-evolving photo editing agent from Tencent Hunyuan and Xiamen University that drives Adobe Lightroom through a multimodal model trained with paired editor and evaluator optimisation.

⭐ 420 Python

InternLM/agentlego

AgentLego is an open-source library that enhances large language model agents with a rich set of multimodal tool APIs for visual, speech, and image processing capabilities, supporting easy integration and remote tool access.

⭐ 415 Python added 2025-03-09

overeasy-sh/overeasy

Overeasy is a framework for orchestrating zero-shot computer vision models to build custom end-to-end pipelines for tasks like bounding box detection, classification, and segmentation without requiring large annotated datasets.

⭐ 392 HTML added 2025-03-09

JosefAlbers/Phi-3-Vision-MLX

Phi-3-MLX is an AI framework optimized for Apple Silicon that integrates vision and language models for advanced multimodal AI tasks including text generation, visual question answering, and code execution.

⭐ 279 Jupyter Notebook added 2025-03-09