Awesome AI AgentsImage Processing & Analysis Agents

11cafe/jaaz

⭐ 6637 TypeScript added to this list on 2025-06-16 repository created 2025-06-01

Jaaz is an AI design agent that serves as a local and free alternative to Lovart, offering powerful capabilities for designing, editing, and generating images, posters, storyboards, and more. It features a creative canvas board that supports fast iterations and layout publishing, enabling users to unleash their creativity efficiently. The agent is powered by large language models (LLMs) that can intelligently write prompts and batch generate images or entire storyboards, making it a versatile tool for creative professionals and enthusiasts alike. Jaaz supports integration with Ollama for local LLM usage and ComfyUI for free local image generation, including models like Stable Diffusion and Flux Dev. Users can edit images conversationally using Flux Kontext, which allows for object removal, style transfer, specific element edits, and consistent character generation—all through chat interactions. The platform also offers an infinite canvas and storyboard feature to facilitate creative workflows. Future updates aim to include video generation and editing capabilities through tools like Wan2.1 and Kling, expanding the creative possibilities further. Jaaz supports multiple API providers such as Claude, OpenAI, Gemini, and Replicate, allowing users to leverage various image generation models like GPT-4O, Recraft, Flux, and Google Imagen. The project is designed for ease of use, requiring users to add LLM API keys or install Ollama for local models, and add image generation API keys like Replicate. It supports multiple operating systems with downloadable versions for macOS and Windows, and offers manual installation instructions for Linux or local builds. Jaaz is ideal for generating creative content such as storyboards, posters, and images with high-quality, realistic outputs and flexible editing options, making it a comprehensive AI-powered design assistant.

https://github.com/11cafe/jaaz

agentaiai-design-agentaiagentaiimageaiimagegeneratoraitoolaitoolsapi-integrationcharacter-generationchat-based-editingclaudecomfyuicreative-canvascreative-workflowsfluxflux-devflux-kontextgeminigoogle-imagengpt-4oimage-generationinfinite-canvasklingllmlocal-alternativelovartmanusobject-removalollamaopenaiposter-designrecraftreplicatestable-diffusionstoryboard-creationstyle-transfervideo-generationwan2.1

Also in Image Processing & Analysis Agents

img2threejs/img2threejs

Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.

apple/ml-ferret

Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.

THUDM/CogVLM

CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.

SamurAIGPT/Generative-Media-Skills

Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.

QIN2DIM/hcaptcha-challenger

hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.

roboflow/inference

Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.

mbzuai-oryx/groundingLMM

GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.

LYL1015/JarvisArt

Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.