img2threejs/img2threejs
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Awesome AI Agents › Image Processing & Analysis Agents
FigMirror is a research plotting tool from VILA-Lab that reproduces the visual style of an existing scientific figure using your own data. The user supplies a reference figure, usually a screenshot from a paper, plus the data to plot, and FigMirror returns an editable matplotlib script together with a rendered PDF that looks as if it belonged to the same paper family. It ships as a skill for Claude Code and for Codex, and also as a local web application that adds upload, preview, iteration history and refinement in the browser. An install script auto-detects which agent runtime is present and installs the skill plus, on the Codex path, two custom agents. The method is an agentic Drawer and Reviewer loop with separated roles. The Drawer writes and runs plotting code, the Reviewer inspects the rendered output against the reference at several zoom levels, draws bounding boxes on problem areas and passes annotated feedback into the next iteration, and the loop repeats until the style matches. The project describes an additional grounded measurement step that extracts concrete quantities from the reference rather than relying on visual impression alone, and an internal hybrid style scorer used for repeatable comparison between method variants. A companion gallery on Hugging Face offers 139 paper figures across 25 chart families as ready-made references for users who have no suitable reference at hand. The audience is researchers and students who spend repeated cycles hand-tuning matplotlib parameters to make figures presentable for a submission. The output stays as ordinary Python code, so it can be edited and re-run without the agent afterwards.
https://github.com/VILA-Lab/FigMirror
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.
CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.
Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.
Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.
hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.
Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.
GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.