Awesome AI AgentsImage Processing & Analysis Agents

THUDM/CogVLM

⭐ 6740 Python added to this list on 2025-04-19 repository created 2023-09-18

CogVLM and CogAgent are advanced open-source visual language models developed by THUDM, designed to push the boundaries of multimodal AI capabilities. CogVLM-17B integrates 10 billion visual parameters with 7 billion language parameters, enabling sophisticated image understanding and multi-turn dialogue at a resolution of 490x490 pixels. It achieves state-of-the-art results across 10 classic cross-modal benchmarks such as NoCaps, Flicker30k captioning, RefCOCO series, Visual7W, GQA, ScienceQA, VizWiz VQA, and TDIUC, demonstrating its robust performance in visual and language tasks. Building on CogVLM, CogAgent-18B enhances these capabilities with 11 billion visual parameters and 7 billion language parameters, supporting higher resolution image understanding at 1120x1120 pixels. CogAgent extends functionality to include GUI image agent capabilities, allowing it to interact with graphical user interfaces visually. It achieves state-of-the-art generalist performance on 9 cross-modal benchmarks including VQAv2, OK-VQ, TextVQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE, and excels in GUI operation datasets like AITW and Mind2Web. The project offers multiple deployment options including a web demo, CLI tools, and integration with Huggingface transformers, supporting efficient inference with quantization techniques to reduce GPU memory requirements. It also provides fine-tuning capabilities and supports multi-GPU model parallelism. The repository includes comprehensive documentation, example prompts, and a new web UI for user-friendly interaction. Recent updates highlight the release of CogVLM2, a next-generation model based on llama3-8b, which rivals or surpasses GPT-4V in many scenarios. The project is actively maintained with frequent updates, new datasets, and improved checkpoints enhancing robustness and performance. CogVLM and CogAgent represent cutting-edge tools for researchers and developers working on visual language understanding, multimodal AI, and interactive AI agents.

https://github.com/THUDM/CogVLM

ai-agentcogagentcogvlmcross-modal-benchmarkscross-modalityfine-tuninggui-agenthuggingface-integrationimage-understandinglanguage-modelllama3-8bmulti-gpumulti-modalmulti-turn-dialoguemultimodal-aipretrained-modelsquantizationstate-of-the-artvisual-expertvisual-language-modelvisual-language-modelsweb-demo

Also in Image Processing & Analysis Agents

img2threejs/img2threejs

Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.

apple/ml-ferret

Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.

11cafe/jaaz

Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.

SamurAIGPT/Generative-Media-Skills

Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.

QIN2DIM/hcaptcha-challenger

hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.

roboflow/inference

Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.

mbzuai-oryx/groundingLMM

GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.

LYL1015/JarvisArt

Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.