img2threejs/img2threejs
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Awesome AI Agents › Image Processing & Analysis Agents
AgentLego is an open-source library designed to enhance large language model (LLM) based agents by providing a rich set of versatile tool APIs. It extends the capabilities of LLM agents with multimodal functionalities such as visual perception, image generation and editing, speech processing, and visual-language reasoning. The library offers a flexible tool interface that allows users to easily add custom tools with various types of arguments and outputs, making it highly adaptable to different use cases. AgentLego supports seamless integration with popular LLM-based agent frameworks like LangChain, Transformers Agents, and Lagent, enabling developers to incorporate its tools into their existing workflows effortlessly. It also supports tool serving and remote access, which is particularly beneficial for tools that require heavy machine learning models or specialized environments such as GPUs and CUDA. The library includes a wide range of tools categorized into general abilities, speech-related tools, image-processing tools, and AI-generated content (AIGC) tools. General tools include a Python interpreter-based calculator and Google search. Speech tools cover text-to-speech and speech-to-text functionalities. Image-processing tools offer capabilities like image description, optical character recognition (OCR), visual question answering (VQA), human body pose estimation, face landmark detection, edge extraction, depth image generation, sketch scribble creation, object detection, and segmentation. AIGC tools focus on generating and editing images based on text prompts or other inputs, including text-to-image generation, image expansion, object removal and replacement, image stylization, and advanced generation techniques using ControlNet and ImageBind series. These tools enable complex image manipulations and multimodal content creation. AgentLego is easy to install via pip and provides detailed documentation and examples for quick starts. It is released under the Apache 2.0 license, with users advised to comply with the licenses of the underlying models used. Overall, AgentLego is a comprehensive toolkit that significantly enhances the functionality and versatility of LLM agents by integrating a broad spectrum of multimodal tools and capabilities.
https://github.com/InternLM/agentlego
Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.
Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.
CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.
Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.
Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.
hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.
Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.
GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.