wandb/openui
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Awesome AI Agents › GUI Action Mapping
OS-Genesis is a project that focuses on automating the construction of GUI agent trajectories through a novel approach called reverse task synthesis. The project aims to synthesize high-quality and diverse GUI agent trajectory data without the need for human supervision or predefined tasks. This is achieved by leveraging an interaction-driven pipeline that incorporates a trajectory reward model, enabling effective end-to-end training of GUI agents. The repository contains code and data supporting the research paper titled "OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis." The project provides resources for training and evaluating GUI agents, particularly in mobile and web environments. It includes models trained on datasets for AndroidControl, AndroidWorld, and web interfaces, with links to Hugging Face for accessing these models and training data. The project also offers raw collected triples of state-action-state transitions, along with screenshots and JSON data, allowing users to reproduce the reverse task synthesis process without re-collecting data from emulators. OS-Genesis supports evaluation on the AndroidControl benchmark with provided scripts for inference and evaluation. The project integrates with other tools like InternVL2 and Qwen2-VL for training purposes. Overall, OS-Genesis represents a significant advancement in automating GUI interaction data generation, facilitating the development of intelligent GUI agents capable of understanding and interacting with graphical user interfaces in a diverse and scalable manner.
https://github.com/OS-Copilot/OS-Genesis
OpenUI is an open-source tool that enables users to describe user interfaces using natural language and see them rendered live, leveraging large language models to generate UI components across multiple frontend frameworks.
Mobile Next MCP is a platform-agnostic Model Context Protocol server enabling scalable mobile automation and interaction with iOS and Android devices, simulators, and emulators through accessibility snapshots and coordinate-based controls.
GPT-4V-Act is an AI agent that uses GPT-4V(ision) to interact with web user interfaces through mouse and keyboard inputs, enabling enhanced accessibility and automation.
egjs is a modular collection of JavaScript components designed to simplify and accelerate the development of customizable web applications with a focus on ease of use and performance.
SeeClick is a state-of-the-art project providing models, data, and code for advanced visual GUI agents with a focus on GUI grounding, featuring the ScreenSpot benchmark and superior performance on multi-platform GUI element prediction.
OS-Atlas is a foundation action model for generalist GUI agents that enables precise visual grounding and interaction with UI elements in screenshots using advanced vision-language models.
Aria-UI is an open-source, fast, and context-aware action grounding system for GUI instructions, achieving state-of-the-art performance in dynamic agent tasks like AndroidWorld and OSWorld.
A curated and continuously updated collection of research papers, models, tools, and datasets focused on UI agents that interact with various user interfaces across platforms like web, mobile, and OS environments.