img2threejs/img2threejs
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Awesome AI Agents › Image Processing & Analysis Agents
Overeasy is a powerful and flexible framework designed to orchestrate zero-shot computer vision models, enabling users to create custom end-to-end pipelines for various image processing tasks without the need for large annotated datasets. The project focuses on leveraging pre-trained zero-shot models to perform tasks such as bounding box detection, classification, and soon segmentation, making it accessible for users to build sophisticated computer vision solutions efficiently. Overeasy introduces the concept of Agents, which are specialized tools that handle specific image processing functions, and Workflows, which allow users to define sequences of these Agents to process images in a structured and modular manner. The framework also supports Execution Graphs to manage and visualize the entire image processing pipeline, providing clear insights into the workflow's operation and intermediate results. Detections in Overeasy represent the outputs such as bounding boxes, segmentation masks, and classification labels, facilitating easy interpretation and further processing. Installation is straightforward via pip, and the project offers extensive documentation and examples, including a Colab notebook for users without local GPU resources. A notable example demonstrates a workflow to detect whether a person is wearing personal protective equipment (PPE) on a worksite by chaining multiple Agents like bounding box detection, non-maximum suppression, image splitting, classification, and result mapping. The project emphasizes ease of use, modularity, and the ability to combine multiple zero-shot models to create tailored vision pipelines. Overeasy is well-suited for developers and researchers looking to implement zero-shot vision tasks without the overhead of dataset collection and training, promoting rapid prototyping and deployment of vision applications. The project is open-source, MIT licensed, and actively maintained with support channels available for users. Keywords include zero-shot vision, computer vision, bounding box detection, classification, segmentation, workflows, agents, execution graphs, image processing, pre-trained models, and pipeline orchestration.
https://github.com/overeasy-sh/overeasy
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.
CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.
Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.
Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.
hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.
Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.
GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.