wandb/openui
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Awesome AI Agents › GUI Action Mapping
The GUI Action Narrator project introduces a dataset called Act2Cap and a framework named GUI Narrator designed for GUI video captioning. This project focuses on interpreting high-resolution screenshots and keyframe extraction in GUI actions by utilizing cursor detection. The Act2Cap dataset consists of sequences of 10-frame GUI screenshots that depict atomic actions, with annotations describing the cursor's actions such as left click, right click, double click, typing, or dragging. The framework enhances the understanding of GUI actions by generating narrations based on these screenshots, which are supported by visual prompts and cropped images derived from cursor detection. The project provides models for cursor detection and narration, along with a test benchmark available on Hugging Face. Users can download the dataset and checkpoints for cursor detection and keyframe extraction, and run inference code to generate visual prompts and cropped images. The repository includes instructions for installing required packages and running the model. This project is valuable for research and development in GUI video captioning, human-computer interaction, and automated GUI action understanding, offering tools and data to analyze and narrate GUI activities effectively.
https://github.com/showlab/GUI-Narrator
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Mobile Next MCP is a platform-agnostic Model Context Protocol server enabling scalable mobile automation and interaction with iOS and Android devices, simulators, and emulators through accessibility snapshots and coordinate-based controls.
GPT-4V-Act is an AI agent that uses GPT-4V(ision) to interact with web user interfaces through mouse and keyboard inputs, enabling enhanced accessibility and automation.
egjs is a modular collection of JavaScript components designed to simplify and accelerate the development of customizable web applications with a focus on ease of use and performance.
SeeClick is a state-of-the-art project providing models, data, and code for advanced visual GUI agents with a focus on GUI grounding, featuring the ScreenSpot benchmark and superior performance on multi-platform GUI element prediction.
OS-Atlas is a foundation action model for generalist GUI agents that enables precise visual grounding and interaction with UI elements in screenshots using advanced vision-language models.
Aria-UI is an open-source, fast, and context-aware action grounding system for GUI instructions, achieving state-of-the-art performance in dynamic agent tasks like AndroidWorld and OSWorld.
A curated and continuously updated collection of research papers, models, tools, and datasets focused on UI agents that interact with various user interfaces across platforms like web, mobile, and OS environments.