wandb/openui
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Awesome AI Agents › GUI Action Mapping
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Mobile Next MCP is a platform-agnostic Model Context Protocol server enabling scalable mobile automation and interaction with iOS and Android devices, simulators, and emulators through accessibility snapshots and coordinate-based controls.
GPT-4V-Act is an AI agent that uses GPT-4V(ision) to interact with web user interfaces through mouse and keyboard inputs, enabling enhanced accessibility and automation.
egjs is a modular collection of JavaScript components designed to simplify and accelerate the development of customizable web applications with a focus on ease of use and performance.
SeeClick is a state-of-the-art project providing models, data, and code for advanced visual GUI agents with a focus on GUI grounding, featuring the ScreenSpot benchmark and superior performance on multi-platform GUI element prediction.
OS-Atlas is a foundation action model for generalist GUI agents that enables precise visual grounding and interaction with UI elements in screenshots using advanced vision-language models.
Aria-UI is an open-source, fast, and context-aware action grounding system for GUI instructions, achieving state-of-the-art performance in dynamic agent tasks like AndroidWorld and OSWorld.
A curated and continuously updated collection of research papers, models, tools, and datasets focused on UI agents that interact with various user interfaces across platforms like web, mobile, and OS environments.
Auto-GUI is a multimodal AI agent framework that predicts user interface actions using a novel chain-of-action technique, enabling direct interaction with interfaces without environment parsing or application-specific APIs.
prompt2ui is a Next.js-based web application that converts textual prompts into interactive user interfaces using AI integration, designed for fun and experimentation.
OS-Genesis is a project that automates the synthesis of high-quality and diverse GUI agent trajectory data using reverse task synthesis and a trajectory reward model for effective end-to-end training of GUI agents.
GUICourse is a project providing datasets, code, and models to train and evaluate versatile GUI agents by enhancing vision language models' abilities to understand and interact with graphical user interfaces.
GPT-4V in Wonderland leverages large multimodal models for zero-shot navigation of smartphone GUIs, enabling AI agents to perform complex tasks like shopping on apps without prior training.
CoAT is a novel framework and benchmark dataset for enhancing GUI agents' decision-making in Android applications using Chain-of-Action-Thought modeling.
LlamaTouch is a scalable and faithful testbed for evaluating mobile UI automation agents by comparing execution traces with annotated essential states in real-world mobile environments.
Mobile-Env is a universal platform for training and evaluating agents that interact with mobile GUIs, supporting flexible task definitions and multiple observation and action modalities for Android apps.
Seq2act is a research project that maps natural language instructions to mobile UI action sequences using machine learning models for phrase extraction and grounding, enhancing mobile interaction automation.
GUI Action Narrator is a project that provides a dataset and framework for captioning GUI videos by detecting cursor actions and extracting keyframes to generate detailed narrations of GUI activities.