img2threejs/img2threejs
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Awesome AI Agents › Image Processing & Analysis Agents
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.
CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.
Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.
Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.
hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.
Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.
GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.
Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.
NeurIPS 2025 multi-agent system that upscales any image to 4K resolution, pairing a vision-language perception agent that plans restoration with a restoration agent that executes, reflects and rolls back.
Agent skill for Claude Code and Codex that copies the visual style of a reference paper figure onto your own data through a Drawer and Reviewer loop, producing an editable matplotlib script and a camera-ready PDF.
Self-evolving photo editing agent from Tencent Hunyuan and Xiamen University that drives Adobe Lightroom through a multimodal model trained with paired editor and evaluator optimisation.
AgentLego is an open-source library that enhances large language model agents with a rich set of multimodal tool APIs for visual, speech, and image processing capabilities, supporting easy integration and remote tool access.
Overeasy is a framework for orchestrating zero-shot computer vision models to build custom end-to-end pipelines for tasks like bounding box detection, classification, and segmentation without requiring large annotated datasets.
Phi-3-MLX is an AI framework optimized for Apple Silicon that integrates vision and language models for advanced multimodal AI tasks including text generation, visual question answering, and code execution.