img2threejs/img2threejs
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Awesome AI Agents › Image Processing & Analysis Agents
Ferret is an advanced multimodal large language model (MLLM) developed by Apple that specializes in referring and grounding tasks with fine granularity and open vocabulary. It is designed to accept any form of referring and grounding requests and respond effectively, making it a versatile tool for multimodal understanding and interaction. The project accompanies a research paper and includes several key contributions such as the Ferret model itself, which integrates a hybrid region representation and a spatial-aware visual sampler to enhance its ability to perform detailed and accurate referring and grounding. Additionally, the project provides the GRIT dataset, a large-scale, hierarchical, and robust dataset with approximately 1.1 million samples for instruction tuning, and Ferret-Bench, a multimodal evaluation benchmark that tests referring, grounding, semantics, knowledge, and reasoning capabilities. The repository offers comprehensive resources including code, checkpoints for 7B and 13B parameter models, training scripts, and evaluation tools. It supports training on multiple GPUs and provides detailed instructions for setup, including dependencies and environment configuration. The project also features a Gradio-based interactive demo that allows users to experience the model's capabilities in real-time by running a local server and model worker. Ferret is built on top of existing models like Vicuna and LLaVA, leveraging their strengths while extending functionality to multimodal tasks involving spatial and visual grounding. The project is intended for research use only, with licensing restrictions that limit commercial applications. It is a significant contribution to the field of multimodal AI, enabling more natural and precise interactions with visual and textual data through a unified model architecture.
https://github.com/apple/ml-ferret
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.
Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.
Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.
hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.
Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.
GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.
Research release of an intelligent photo retouching agent built on a multimodal LLM that plans edits from natural language and drives Adobe Lightroom through a dedicated agent protocol.