Awesome AI AgentsImage Processing & Analysis Agents

JosefAlbers/Phi-3-Vision-MLX

⭐ 279 Jupyter Notebook added to this list on 2025-03-09 repository created 2024-05-27

Phi-3-MLX is an AI framework optimized for Apple Silicon Macs that integrates multimodal vision and language models, specifically the Phi-3.5-vision and Phi-3.5-mini-128K models, using the MLX framework. It supports a variety of AI tasks including advanced text generation, visual question answering, code execution, and more. The project emphasizes performance optimization for Apple Silicon hardware, with features like model and cache quantization to improve efficiency and reduce resource usage. It supports batched generation for handling multiple prompts simultaneously and offers a flexible agent system that can be customized for different AI workflows. The framework also includes LoRA fine-tuning capabilities, allowing users to adapt models to specific tasks or datasets. Additionally, it provides API integration for extended functionalities such as image generation and text-to-speech conversion. The project is designed to be user-friendly, with easy installation via pip and straightforward usage examples for both command line and Python scripting. Core functionalities include visual question answering, constrained beam decoding for structured text generation, multiple-choice question answering, and multi-turn conversational agents that can analyze images, generate and modify code, and interact with external APIs. Custom toolchains can be created for specialized workflows, enhancing the flexibility and applicability of the framework. The minimum hardware requirements are Apple Silicon Macs with at least 8GB RAM, with recommendations for 16GB RAM for optimal performance. Overall, Phi-3-MLX provides a powerful and efficient platform for running advanced vision and language AI models locally on Apple Silicon devices, catering to developers and researchers interested in multimodal AI applications.

https://github.com/JosefAlbers/Phi-3-Vision-MLX

agentai-frameworkapiapi-integrationapple-siliconbatched-generationbgptcache-quantizationcode-executionconstrained-beam-decodingcustom-toolchainsfine-tuningfinetuningflexible-agent-systemimage-generationlanguage-modelllmlocal-ailoralora-fine-tuninglstmmacmac-optimizationmacosmetalmlxmlx-frameworkmodel-quantizationmulti-agent-systemsmulti-turn-conversationmultimodalmultimodal-modelphi-3phi-3-5phi-3-miniphi-3-visionphi-3.5-mini-128kphi-3.5-visionretnettext-generationtext-to-speechvision-modelvisual-question-answeringvlm

Also in Image Processing & Analysis Agents

img2threejs/img2threejs

Agent skill for Claude Code and Codex that rebuilds the object in a reference image as procedural Three.js TypeScript code through a gated, vision-reviewed, token-efficient pipeline.

apple/ml-ferret

Ferret is an end-to-end multimodal large language model developed by Apple that excels in fine-grained referring and grounding tasks with open vocabulary, supported by a large-scale dataset and evaluation benchmark for research purposes.

THUDM/CogVLM

CogVLM and CogAgent are state-of-the-art open-source visual language models designed for advanced image understanding, multi-turn dialogue, and GUI agent capabilities, achieving top performance on multiple cross-modal benchmarks.

11cafe/jaaz

Jaaz is a local and free AI design agent that enables users to design, edit, and generate images, posters, and storyboards with advanced AI-powered tools and a creative canvas for fast iterations.

SamurAIGPT/Generative-Media-Skills

Collection of agent skills and workflow recipes giving Claude Code, Cursor, Gemini CLI and OpenCode access to image, video and audio generation models through muapi-cli and an MCP server.

QIN2DIM/hcaptcha-challenger

hCaptcha Challenger is an open-source project that uses multimodal large language models and advanced machine learning techniques to automate solving hCaptcha challenges without relying on third-party services or scripts.

roboflow/inference

Roboflow Inference is a platform that turns any computer or edge device into a command center for deploying and managing computer vision models and workflows, enabling advanced AI-powered visual applications.

mbzuai-oryx/groundingLMM

GLaMM is a groundbreaking multimodal AI model that generates natural language responses integrated with object segmentation masks, enabling advanced visual grounding and conversational tasks.