{"id":85063,"url":"https://github.com/trycua/acu","name":"acu","description":"A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.","projects_count":162,"last_synced_at":"2026-08-15T23:00:20.432Z","repository":{"id":271162165,"uuid":"912449106","full_name":"trycua/acu","owner":"trycua","description":"A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.","archived":false,"fork":false,"pushed_at":"2025-09-26T04:09:48.000Z","size":141,"stargazers_count":1725,"open_issues_count":12,"forks_count":133,"subscribers_count":15,"default_branch":"main","last_synced_at":"2026-08-09T18:56:54.507Z","etag":null,"topics":["ai","ai-research","awesome","computer","computer-use","gui-agent","ui-agent"],"latest_commit_sha":null,"homepage":"https://x.com/i/communities/1874549355442802764","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/trycua.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":".github/FUNDING.yml","license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null},"funding":{"github":"trycua","patreon":null,"open_collective":null,"ko_fi":null,"tidelift":null,"community_bridge":null,"liberapay":null,"issuehunt":null,"lfx_crowdfunding":null,"polar":null,"buy_me_a_coffee":null,"thanks_dev":null,"custom":null}},"created_at":"2025-01-05T15:56:34.000Z","updated_at":"2026-08-09T09:16:22.000Z","dependencies_parsed_at":"2025-02-16T23:23:28.364Z","dependency_job_id":"b3fffad3-e288-484a-9d83-61026b4ee532","html_url":"https://github.com/trycua/acu","commit_stats":null,"previous_names":["francedot/acu","trycua/acu"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/trycua/acu","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/trycua%2Facu","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/trycua%2Facu/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/trycua%2Facu/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/trycua%2Facu/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/trycua","download_url":"https://codeload.github.com/trycua/acu/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/trycua%2Facu/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36698916,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-06T04:43:03.162Z","status":"online","status_checked_at":"2026-08-15T02:00:05.847Z","response_time":94,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2025-02-23T00:01:38.897Z","updated_at":"2026-08-15T23:00:20.432Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Papers","Projects","Articles","Contributing"],"sub_categories":["Frameworks \u0026 Models","Dataset","Benchmark","Automation","Environment \u0026 Sandbox","UI Grounding","Surveys","Safety"],"readme":"\u003cdiv align=\"center\"\u003e\n\u003ch1\u003e\n  \u003cdiv class=\"image-wrapper\" style=\"display: inline-block;\"\u003e\n    \u003cimg src=\"img/logo.png\" alt=\"logo\" height=\"100\" style=\"display: block; margin: auto;\"\u003e\n  \u003c/div\u003e\n  \n  [![X Community](https://img.shields.io/badge/Community-black?logo=x\u0026style=flat-square)](https://x.com/i/communities/1874549355442802764)\n  \n  ACU - Awesome Agents for Computer Use\n\u003c/h1\u003e\n\u003c/div\u003e\n\n\u003e An AI Agent for Computer Use is an autonomous program that can **reason** about tasks, **plan** sequences of actions, and **act** within the domain of a computer or mobile device in the form of clicks, keystrokes, other computer events, command-line operations and internal/external API calls. These agents combine perception, decision-making, and control capabilities to interact with digital interfaces and accomplish user-specified goals independently.\n\nA curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.\n\n## Table of Contents\n\n- [ACU - Awesome Agents for Computer Use](#acu---awesome-agents-for-computer-use)\n  - [Table of Contents](#table-of-contents)\n  - [Articles](#articles)\n  - [Papers](#papers)\n    - [Surveys](#surveys)\n    - [Frameworks \u0026 Models](#frameworks--models)\n    - [UI Grounding](#ui-grounding)\n    - [Dataset](#dataset)\n    - [Benchmark](#benchmark)\n    - [Safety](#safety)\n  - [Projects](#projects)\n    - [Open Source](#open-source)\n      - [Frameworks \u0026 Models](#frameworks--models-1)\n      - [Environment \u0026 Sandbox](#environment--sandbox)\n      - [Automation](#automation)\n    - [Commercial](#commercial)\n      - [Frameworks \u0026 Models](#frameworks--models-2)\n  - [Contributing](#contributing)\n\n## Articles\n- [Anthropic | Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku](https://www.anthropic.com/news/3-5-models-and-computer-use)\n- [Bill Gates | AI is about to completely change how you use computers](https://www.gatesnotes.com/AI-agents)\n- [Ethan Mollick | When you give a Claude a mouse](https://www.oneusefulthing.org/p/when-you-give-a-claude-a-mouse)\n- [OpenAI | Introducing Operator: A research preview of an agent that can use its own browser to perform tasks for you](https://openai.com/index/introducing-operator)\n\n## Papers\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eSurveys\u003c/b\u003e\u003c/summary\u003e\n\n### Surveys\n\n- [AI Agents for Computer Use: A Review of Instruction-based Computer Control, GUI Automation, and Operator Assistants](https://arxiv.org/abs/2501.16150) (Jan. 2025)\n  - Comprehensive review establishing taxonomy of computer control agents (CCAs) from environment, interaction, and agent perspectives, analyzing 86 CCAs and 33 datasets\n\n- [GUI Agents: A Survey](https://arxiv.org/abs/2412.13501) (Dec. 2024)\n  - General survey of GUI agents\n\n- [Large Language Model-Brained GUI Agents: A Survey](https://arxiv.org/abs/2411.18279) (Nov. 2024)\n  - Focus on LLM-based approaches\n  - [Website](https://vyokky.github.io/LLM-Brained-GUI-Agents-Survey/)\n\n- [GUI Agents with Foundation Models: A Comprehensive Survey](https://arxiv.org/abs/2411.04890) (Nov. 2024)\n  - Comprehensive overview of foundation model-based GUI agents\n  \n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eFrameworks \u0026 Models\u003c/b\u003e\u003c/summary\u003e\n\n### Frameworks \u0026 Models\n\n- [Reinforcement Learning for Long-Horizon Interactive LLM Agents](https://arxiv.org/abs/2502.01600) (Feb. 2025)\n  - Novel RL approach (LOOP) for training IDAs directly in target environments\n  - 32B parameter agent outperforms OpenAI o1 by 9 percentage points on AppWorld\n\n- [Large Action Models: From Inception to Implementation](https://arxiv.org/abs/2412.10047) (Dec. 2024)\n  - Comprehensive framework for developing LAMs that can perform real-world actions beyond language generation\n  - Details key stages including data collection, model training, environment integration, grounding and evaluation\n\n- [Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation](https://openreview.net/forum?id=jR6YMxVG9i) (Dec. 2024)\n  - Novel reward-guided navigation approach\n\n- [SpiritSight Agent: Advanced GUI Agent with One Look](https://openreview.net/forum?id=jY2ow7jRdZ) (Dec. 2024)\n  - Single-shot GUI interaction approach\n\n- [AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs](https://openreview.net/forum?id=wl4c9jvcyY) (Dec. 2024)\n  - Novel approach for automatic GUI functionality annotation\n\n- [Simulate Before Act: Model-Based Planning for Web Agents](https://openreview.net/forum?id=JDa5RiTIC7) (Dec. 2024)\n  - Novel model-based planning approach using LLM world models\n\n- [Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet Agents](https://arxiv.org/abs/2412.13194) (Dec. 2024)\n  - Novel autonomous skill discovery framework for web agents\n  - [Code](https://yanqval.github.io/PAE/)\n\n- [Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents](https://openreview.net/forum?id=3Gzz7ZQLiz) (Dec. 2024)\n  - Novel framework for contextualizing web pages to enhance LLM agent decision making\n\n- [Digi-Q: Transforming VLMs to Device-Control Agents via Value-Based Offline RL](https://openreview.net/forum?id=CjfQssZtAb) (Dec. 2024)\n  - Novel value-based offline RL approach for training VLM device-control agents\n\n- [Magentic-One](https://www.microsoft.com/en-us/research/uploads/prod/2024/11/MagenticOne.pdf) (Nov. 2024)\n  - Multi-agent system with orchestrator-led coordination\n  - Strong performance on GAIA, WebArena, and AssistantBench\n\n- [Agent Workflow Memory](https://arxiv.org/abs/2409.07429) (Sep. 2024)\n  - Novel workflow memory framework for agents\n  - [Code](https://github.com/zorazrw/agent-workflow-memory)\n\n- [The Impact of Element Ordering on LM Agent Performance](https://arxiv.org/abs/2409.12089) (Sep. 2024)\n  - Novel study on element ordering's impact on agent performance\n  - [Code](https://github.com/waynchi/gui-agent)\n\n- [Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents](https://arxiv.org/abs/2408.07199) (Aug. 2024)\n  - Novel reasoning and learning framework\n  - [Website](https://www.multion.ai/blog/introducing-agent-q-research-breakthrough-for-the-next-generation-of-ai-agents-with-planning-and-self-healing-capabilities)\n\n- [OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models](https://aclanthology.org/2024.acl-demos.8/) (Aug. 2024)\n  - Open platform for web-based agent deployment\n  - [Code](https://github.com/boxworld18/OpenWebAgent)\n\n- [Agent-e: From autonomous web navigation to foundational design principles in agentic systems](https://arxiv.org/abs/2407.13032) (Jul. 2024)\n  - Hierarchical architecture with flexible DOM distillation\n  - Novel denoising method for web navigation\n\n- [Apple Intelligence Foundation Language Models](https://arxiv.org/pdf/2407.21075) (Jul. 2024)\n  - Vision-Language Model with Private Cloud Compute\n  - Novel foundation model architecture\n\n- [Tree search for language model agents](https://arxiv.org/abs/2407.01476) (Jul. 2024)\n  - Multi-step reasoning and planning with best-first tree search\n  - Novel approach for LLM-based agents\n\n- [DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning](https://arxiv.org/abs/2406.11896) (Jun. 2024)\n  - Novel reinforcement learning approach\n  - [Code](https://github.com/DigiRL-agent/digirl)\n\n- [Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration](https://arxiv.org/abs/2406.01014) (Jun. 2024)\n  - Multi-agent collaboration for mobile device operation\n  - [Code](https://github.com/X-PLUG/MobileAgent)\n\n- [Octopus Series: On-device Language Models for Computer Control](https://arxiv.org/abs/2404.01549) (Apr. 2024)\n  - v4: Graph of language models with functional tokens integration (Apr. 2024)\n  - v3: Sub-billion parameter multimodal model for edge devices (Apr. 2024)\n  - v2: Super agent for Android and iOS (Apr. 2024)\n  - v1: Function calling of software APIs (Apr. 2024)\n  - [Website](https://www.nexa4ai.com/octopus-v3)\n  - [Code](https://github.com/NexaAI/octopus-v4)\n\n- [AutoWebGLM: Bootstrap and reinforce a large language model-based web navigating agent](https://arxiv.org/abs/2404.03648) (Apr. 2024)\n  - Novel approach for real-world web navigation and bilingual benchmark\n  - [Code](https://github.com/THUDM/WebGLM)\n\n- [Cradle: Empowering Foundation Agents towards General Computer Control](https://arxiv.org/abs/2403.03186) (Mar. 2024)\n  - Focus on general computer control using Red Dead Redemption II as a case study\n  - [Code](https://github.com/BAAI-Agents/Cradle)\n\n- [Android in the Zoo: Chain-of-Action-Thought for GUI Agents](https://arxiv.org/abs/2403.02713) (Mar. 2024)\n  - Novel Chain-of-Action-Thought framework for Android interaction\n  - [Code](https://github.com/IMNearth/CoAT)\n\n- [ScreenAgent: A Computer Control Agent Driven by Visual Language Large Model](https://arxiv.org/abs/2402.07945) (Feb. 2024)\n  - Vision-language model for computer control\n  - [Code](https://github.com/niuzaisheng/ScreenAgent)\n\n- [OS-Copilot: Towards Generalist Computer Agents with Self-Improvement](https://arxiv.org/abs/2402.07456) (Feb. 2024)\n  - Vision-Language Model for PC interaction\n  - [Code](https://github.com/OS-Copilot/OS-Copilot)\n\n- [UFO: A UI-Focused Agent for Windows OS Interaction](https://arxiv.org/abs/2402.07939) (Feb. 2024)\n  - Specialized for Windows OS interaction\n  - [Code](https://github.com/microsoft/UFO)\n\n- [CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation](https://arxiv.org/abs/2402.11941) (Feb. 2024)\n  - Novel comprehensive environment perception (CEP) approach for exhaustive GUI perception\n  - Introduces conditional action prediction (CAP) for reliable action response\n\n- [Intention-inInteraction (IN3): Tell Me More!](https://arxiv.org/abs/2402.09205) (Feb. 2024)\n  - Novel benchmark for evaluating user intention understanding in agent designs\n  - Introduces model experts for robust user-agent interaction\n\n- [Dual-view visual contextualization for web navigation](https://arxiv.org/abs/2402.04476) (Feb. 2024)\n  - Novel approach for automatic web navigation with language instructions\n  - Key: HTML elements, visual contextualization\n\n- [ScreenAI: A Vision-Language Model for UI and Infographics Understanding](https://arxiv.org/abs/2402.04615) (Feb. 2024)\n  - Specialized for mobile UI and infographics understanding\n  - Novel approach for visual interface comprehension\n\n- [GPT-4V(ision) is a Generalist Web Agent, if Grounded](https://arxiv.org/abs/2401.01614) (Jan. 2024)\n  - Demonstrates GPT-4V capabilities for web interaction\n  - [Code](https://github.com/OSU-NLP-Group/SeeAct)\n\n- [Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception](https://arxiv.org/abs/2401.16158) (Jan. 2024)\n  - Visual perception for mobile device interaction\n  - [Code](https://github.com/X-PLUG/MobileAgent)\n\n- [WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models](https://arxiv.org/abs/2401.13919) (Jan. 2024)\n  - End-to-end approach for web interaction\n  - [Code](https://github.com/MinorJerry/WebVoyager)\n\n- [CogAgent: A Visual Language Model for GUI Agents](https://arxiv.org/abs/2312.08914) (Dec. 2023)\n  - Works across PC and Android platforms\n  - [Code](https://github.com/THUDM/CogVLM)\n\n- [AppAgent: Multimodal Agents as Smartphone Users](https://arxiv.org/abs/2312.13771) (Dec. 2023)\n  - Focused on smartphone interaction\n  - [Code](https://github.com/mnotgod96/AppAgent)\n\n- [LASER: LLM Agent with State-Space Exploration for Web Navigation](https://arxiv.org/abs/2309.08172) (Sep. 2023)\n  - Novel approach to web navigation\n  - [Code](https://github.com/Mayer123/LASER)\n\n- [AndroidEnv: A Reinforcement Learning Platform for Android](https://arxiv.org/abs/2105.13231) (May 2021)\n  - Reinforcement learning platform for Android interaction\n  - [Code](https://github.com/google-deepmind/android_env)\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eUI Grounding\u003c/b\u003e\u003c/summary\u003e\n\n### UI Grounding\n\n- [OmniParser for Pure Vision Based GUI Agent](https://arxiv.org/pdf/2408.00203) (Aug. 2024)\n  - Novel vision-based screen parsing method for UI screenshots\n  - Combines finetuned interactable icon detection and functional description models\n  - [Code](https://github.com/microsoft/OmniParser)\n\n- [Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs](https://arxiv.org/abs/2404.05719) (Apr. 2024)\n  - Mobile UI understanding\n  - [Code](https://github.com/apple/ml-ferret)\n\n- [SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents](https://arxiv.org/abs/2401.10935) (Jan. 2024)\n  - Advanced visual grounding techniques\n  - [Code](https://github.com/njucckevin/SeeClick)\n\n- [Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms](https://arxiv.org/abs/2410.18967) (Oct. 2024)\n  - Multimodal LLM for universal UI understanding across diverse platforms\n  - Introduces adaptive gridding for high-resolution perception\n  - Preprint\n\n- [Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents](https://arxiv.org/abs/2410.05243) (Oct. 2024)\n  - Universal approach to GUI interaction\n  - [Code](https://github.com/OSU-NLP-Group/UGround)\n\n- [OS-ATLAS: Foundation Action Model for Generalist GUI Agents](https://arxiv.org/abs/2410.23218) (Oct. 2024)\n  - Comprehensive action modeling\n  - [Code](https://github.com/OS-Copilot/OS-Atlas)\n\n- [UI-Pro: A Hidden Recipe for Building Vision-Language Models for GUI Grounding](https://openreview.net/forum?id=5wmAfwDBoi) (Dec. 2024)\n  - Novel framework for building VLMs with strong UI element grounding capabilities\n\n- [Grounding Multimodal Large Language Model in GUI World](https://openreview.net/forum?id=M9iky9Ruhx) (Dec. 2024)\n  - Novel GUI grounding framework with automated data collection engine and lightweight grounding module\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eDataset\u003c/b\u003e\u003c/summary\u003e\n\n### Dataset\n\n- [Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents](https://arxiv.org/pdf/2502.11357) (Feb. 2025)\n  - Scalable multi-agent pipeline that leverages exploration for diverse web agent trajectory synthesis.\n\n- [OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis](https://arxiv.org/abs/2412.19723) (Dec. 2024)\n  - Novel interaction-driven approach for automated GUI trajectory synthesis\n  - Introduces reverse task synthesis and trajectory reward model\n  - [Code](https://github.com/OS-Copilot/OS-Genesis)\n\n- [AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials](https://arxiv.org/abs/2412.09605) (Dec. 2024)\n  - Web tutorial-based trajectory synthesis\n\n- [ICAL: Continual Learning of Multimodal Agents by Transforming Trajectories into Actionable Insights](https://arxiv.org/abs/2406.14596) (Jun. 2024)\n  - Novel approach to continual learning from trajectories\n\n- [Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale](https://arxiv.org/abs/2409.15637) (Sep. 2024)\n  - Scalable demonstration generation\n\n- [UiPad: UI Parsing and Accessibility Dataset](https://huggingface.co/datasets/MacPaw/uipad) (Sep. 2024)\n  - MacOS desktop UI dataset with accessibility trees and evaluation questions\n\n- [Multi-Turn Mind2Web: On the Multi-turn Instruction Following](https://arxiv.org/pdf/2402.15057) (Feb. 2024)\n  - Multi-turn instruction dataset for web agents\n  - [Code](https://github.com/magicgh/self-map)\n\n- [CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation](https://aclanthology.org/2024.findings-acl.928/) (Aug. 2024)\n  - Chinese benchmark for agent evaluation\n  - [Code](https://github.com/tjunlp-lab/CToolEval)\n\n- [AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks](https://arxiv.org/abs/2407.15711) (Jul. 2024)\n  - Benchmark for realistic and time-consuming web tasks\n  - [Code](https://assistantbench.github.io)\n\n- [Mind2Web: Towards a Generalist Agent for the Web](https://arxiv.org/abs/2306.06070) (Jun. 2023)\n  - Large-scale web interaction dataset\n  - [Code](https://github.com/OSU-NLP-Group/Mind2Web)\n\n- [Android in the Wild: A Large-Scale Dataset for Android Device Control](https://arxiv.org/abs/2307.10088) (Jul. 2023)\n  - Large-scale dataset for Android interaction\n  - Real-world device control scenarios\n\n- [WebShop: Towards Scalable Real-World Web Interaction](https://arxiv.org/abs/2207.01206) (Jul. 2022)\n  - Dataset for grounded language agents in web interaction\n  - [Code](https://github.com/princeton-nlp/WebShop)\n\n- [Rico: A Mobile App Dataset for Building Data-Driven Design Applications](https://dl.acm.org/doi/10.1145/3126594.3126651) (Oct. 2017)\n  - Mobile app UI dataset\n  - Design-focused data collection\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eBenchmark\u003c/b\u003e\u003c/summary\u003e\n\n### Benchmark\n\n- [A3: Android Agent Arena for Mobile GUI Agents](https://arxiv.org/abs/2501.01149) (Jan. 2025)\n  - Novel evaluation platform with 201 tasks across 21 widely used third-party apps\n  - [Website](https://yuxiangchai.github.io/Android-Agent-Arena/)\n  - [Code](https://github.com/AndroidArenaAgent/AndroidArena)\n\n- [OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments](https://arxiv.org/abs/2404.07972) (Apr. 2024)\n  - Comprehensive evaluation framework\n  - [Code](https://github.com/xlang-ai/OSWorld)\n\n- [AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents](https://arxiv.org/abs/2405.14573) (May. 2024)\n  - Android-focused evaluation\n  - [Code](https://github.com/google-research/android_world)\n\n- [Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?](https://arxiv.org/abs/2407.10956) (Jul. 2024)\n  - Evaluation in data science workflows\n  - [Code](https://github.com/xlang-ai/Spider2-V)\n\n- [AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents](https://arxiv.org/abs/2407.18901) (Jul. 2024)\n  - Comprehensive benchmark with 750 natural tasks across 9 day-to-day apps and 457 APIs\n  - GPT-4o achieves only ~49% on normal tasks and ~30% on challenge tasks\n  - [Code](https://github.com/stonybrooknlp/appworld/)\n\n- [τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains](https://arxiv.org/abs/2406.12045) (Jun. 2024)\n  - Novel benchmark for evaluating agent-user interaction and policy compliance\n  - State-of-the-art agents achieve \u003c50% success rate and \u003c25% consistency (pass^8)\n  - [Code](https://github.com/sierra-research/tau-bench)\n\n- [MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents](https://arxiv.org/abs/2406.08184) (Jun. 2024)\n  - Mobile agent evaluation\n  - [Code](https://github.com/MobileAgentBench/mobile-agent-bench)\n\n- [VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks](https://arxiv.org/abs/2401.13649) (Jan. 2024)\n  - Web-focused evaluation\n  - [Code](https://github.com/web-arena-x/visualwebarena)\n\n- [Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale](https://arxiv.org/abs/2409.08264) (Sep. 2024)\n  - Windows OS-focused evaluation framework\n  - [Code](https://github.com/microsoft/WindowsAgentArena)\n  - [Website](https://microsoft.github.io/WindowsAgentArena/)\n\n- [Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction](https://arxiv.org/abs/2305.08144) (May. 2023)\n  - Mobile-focused evaluation framework\n  - [Code](https://github.com/X-LANCE/Mobile-Env)\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eSafety\u003c/b\u003e\u003c/summary\u003e\n\n### Safety\n\n- [Attacking Vision-Language Computer Agents via Pop-ups](https://arxiv.org/abs/2411.02391) (Nov. 2024)\n  - Security analysis of computer agents\n  - [Code](https://github.com/SALT-NLP/PopupAttack)\n\n- [EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage](https://arxiv.org/abs/2409.11295) (Sep. 2024)\n  - Privacy and security analysis\n\n- [GuardAgent: Safeguard LLM Agent by a Guard Agent via Knowledge-Enabled Reasoning](https://arxiv.org/abs/2406.09187) (Jun. 2024)\n  - Safety mechanisms for agents\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n## Projects\n\n### Open Source\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eFrameworks \u0026 Models\u003c/b\u003e\u003c/summary\u003e\n\n### Frameworks \u0026 Models\n\n- [AutoGen](https://github.com/microsoft/autogen)\n  - Framework for building AI agent systems.\n  - It simplifies the creation of event-driven, distributed, scalable, and resilient agentic applications.\n\n- [Auto-GPT](https://github.com/Significant-Gravitas/Auto-GPT)  \n  - Autonomous GPT-4 agent  \n  - Task automation focus\n\n- [Browser Use](https://github.com/browser-use/browser-use)  \n  - Make websites accessible for AI agents with vision + HTML extraction  \n  - Supports multi-tab management and custom actions with LangChain integration  \n\n- [Claude Computer Use Demo](https://github.com/PallavAg/claude-computer-use-macos)  \n  - MacOS implementation  \n  - Claude integration  \n\n- [Claude Minecraft Use](https://github.com/ObservedObserver/claude-minecraft-use)  \n  - Game automation  \n  - Specialized use case  \n\n- [Computer Use OOTB](https://github.com/showlab/computer_use_ootb)  \n  - Ready-to-use implementation  \n  - Comprehensive toolset  \n\n- [Cua](https://github.com/trycua)\n  - Computer Use Interface \u0026 Agent\n\n- [Cybergod](https://github.com/james4ever0/agi_computer_control)  \n  - Advanced computer control  \n\n- [Grunty](https://github.com/suitedaces/computer-agent)  \n  - Computer control agent  \n  - Task automation focus\n\n- [Inferable](https://github.com/inferablehq/inferable)  \n  - Distributed agent builder platform  \n  - Build tools with existing code  \n\n- [LaVague](https://github.com/lavague-ai/LaVague)  \n  - AI web agent framework  \n  - Modular architecture  \n\n- [Mac Computer Use](https://github.com/deedy/mac_computer_use)  \n  - MacOS-specific tools  \n  - Anthropic integration  \n\n- [NatBot](https://github.com/nat/natbot)  \n  - Browser automation  \n  - GPT-4 Vision integration\n \n- [Notte Browser Using Agent](https://github.com/nottelabs/notte)  \n  - Full-stack web AI agents framework (agents, automations, cloud browser sessions)\n  - Notte turns websites into structured, navigable maps described in natural language\n\n- [OpenAdapt](https://github.com/OpenAdaptAI/OpenAdapt)  \n  - AI-First Process Automation  \n  - Multimodal model integration  \n\n- [OpenInterface](https://github.com/AmberSahdev/Open-Interface/)  \n  - Open-source UI interaction framework  \n  - Cross-platform support  \n\n- [OpenInterpreter](https://github.com/OpenInterpreter/open-interpreter)  \n  - General-purpose computer control framework  \n  - Python-based, extensible architecture  \n\n- [Open Source Computer Use by E2B](https://github.com/e2b-dev/secure-computer-use/tree/os-computer-use)  \n  - Open-source implementation of computer control capabilities  \n  - Secure sandboxed environment for AI agents  \n\n- [Self-Operating Computer](https://github.com/OthersideAI/self-operating-computer) (Nov. 2023) \n  - The first Computer Use framework created\n  - Computer control framework  \n  - Vision-based automation\n\n- [Skyvern](https://github.com/skyvern-ai/skyvern)  \n  - AI web agent framework\n  - Automate browser-based workflows with LLMs using vision and HTML extraction\n\n- [Surfkit](https://github.com/agentsea/surfkit)  \n  - Device operation toolkit  \n  - Extensible agent framework  \n\n- [WebMarker](https://github.com/reidbarber/webmarker)  \n  - Web page annotation tool  \n  - Vision-language model support\n\n- [Upsonic](https://github.com/upsonic/upsonic)  \n  - Reliable agent framework that support MCP \n  - Integrated Browser Use and Computer Use\n \n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eUI Grounding\u003c/b\u003e\u003c/summary\u003e\n\n### UI Grounding\n\n- [AskUI/PTA-1](https://huggingface.co/AskUI/PTA-1)\n  - A small vision language model for computer \u0026 phone automation, based on Florence-2.\n  - With only 270M parameters it outperforms much larger models in GUI text and element localization. \n\n- [Microsoft/OmniParser](https://huggingface.co/microsoft/OmniParser)\n  - A general screen parsing tool, which interprets/converts UI screenshot to structured format, to improve existing LLM based UI agent\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eEnvironment \u0026 Sandbox\u003c/b\u003e\u003c/summary\u003e\n\n### Environment \u0026 Sandbox\n\n- [Cua](https://github.com/trycua)\n  - macOS/Linux Sandbox on Apple Silicon\n\n- [dockur/windows](https://github.com/dockur/windows)\n  - Windows inside a Docker container\n\n- [E2B Desktop Sandbox](https://github.com/e2b-dev/desktop)\n  - Secure desktop environment\n  - Agent testing platform\n\n- [qemus/qemu-docker](https://github.com/qemus/qemu-docker)\n  - Docker container for running virtual machines using QEMU\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eAutomation\u003c/b\u003e\u003c/summary\u003e\n\n### Automation\n- [nut.js](https://github.com/nut-tree/nut.js)\n  - Native UI automation\n  - JavaScript/TypeScript implementation\n\n- [PyAutoGUI](https://github.com/asweigart/pyautogui)\n  - Cross-platform GUI automation\n  - Python-based control library\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n### Commercial\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cb\u003eFrameworks \u0026 Models\u003c/b\u003e\u003c/summary\u003e\n\n### Frameworks \u0026 Models\n\n- [Anthropic Claude Computer Use](https://www.anthropic.com/news/3-5-models-and-computer-use)\n  - Commercial computer control capability\n  - Integrated with Claude 3.5 models\n\n- [Multion](https://www.multion.ai)\n  - AI agents that can fully complete tasks in any web environment.\n\n- [Runner H](https://www.hcompany.ai/)\n  - Advanced AI agent for real-world applications.\n  - Scores 67% on WebVoyager\n\n\u003cbr/\u003e\n\u003c/details\u003e\n\n## Contributing\n\nWe welcome and encourage contributions from the community! Here's how you can help:\n\n- **Add new resources**: Found a relevant paper, project, or tool? Submit a PR to add it\n- **Fix errors**: Help us correct any mistakes in existing entries\n- **Improve organization**: Suggest better ways to structure the information\n- **Update content**: Keep entries up-to-date with latest developments\n\nTo contribute:\n1. Fork the repository\n2. Create a new branch for your changes\n3. Submit a pull request with a clear description of your additions/changes\n4. Post in the [X Community](https://x.com/i/communities/1874549355442802764) to let everyone know about the new resource\n\nFor an example of how to format your contribution, please refer to [this PR](https://github.com/francedot/acu/pull/1).\n\n\u003cbr/\u003e\n\n*Thank you for helping spread knowledge about AI agents for computer use!*\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/trycua%2Facu/projects"}