
Yu WU, Yi YANG. A platform-aware survey of execution, verification, and risk of graphical user interface agents[J]. Journal of Zhejiang University Science C, 2026, 27(7): 1-26.
@article{title="A platform-aware survey of execution, verification, and risk of graphical user interface agents",
author="Yu WU, Yi YANG",
journal="Journal of Zhejiang University Science C",
volume="27",
number="7",
pages="1-26",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0103"
}
%0 Journal Article
%T A platform-aware survey of execution, verification, and risk of graphical user interface agents
%A Yu WU
%A Yi YANG
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 7
%P 1-26
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0103
TY - JOUR
T1 - A platform-aware survey of execution, verification, and risk of graphical user interface agents
A1 - Yu WU
A1 - Yi YANG
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 7
SP - 1
EP - 26
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0103
Abstract: Graphical user interface (GUI) agents are widely used in general-purpose digital automation. Large language models, vision-language models, and multimodal foundation models can now interpret screenshots, follow natural-language instructions, and execute grounded actions in websites, mobile applications, and desktop operating systems. However, despite their different execution conditions, these settings are commonly grouped under one label. This survey reviews representative frameworks, models, datasets, and benchmarks of GUI agents through a common technical stack: observation, grounding, planning, memory, execution, and verification. We advocate the treatment of web, mobile, and desktop/operating-system agents as distinct operational regimes with different observabilities, action semantics, hidden-state dependences, execution costs, and side-effect risks. We also describe a shift from next-action prediction to dependable execution with increased focus on outcome verification, execution efficiency, and trustworthiness. The survey concludes with future directions in belief-state tracking, adaptive sensing, hybrid GUI–tool execution, and human oversight for reliable computer use.
[1]Abhyankar R, Qi Q, Zhang YY, 2025. OSWorld-Human: benchmarking the efficiency of computer-use agents.
[2]Agashe S, Han JZ, Gan SY, et al., 2025. Agent S: an open agentic framework that uses computers like a human. Proc 13th Int Conf on Learning Representations, p.1-23.
[3]Amazon Web Services, 2025. Build agents to automate production UI workflows with Amazon Nova Act (GA). https://aws.amazon.com/about-aws/whats-new/2025/12/build-automate-production-ui-workflows-nova-act [Accessed on Apr. 1, 2026].
[4]Anthropic Claude 3.7 Team, 2025. Claude 3.7 Sonnet and Claude Code. https://www.anthropic.com/news/claude-3-7-sonnet [Accessed on Apr. 1, 2026].
[5]Anthropic Claude 4 Team, 2025. Introducing Claude 4. https://www.anthropic.com/news/claude-4 [Accessed on Apr. 1, 2026].
[6]Anthropic Claude Sonnet 4.6 Team, 2026. Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/claude-sonnet-4-6 [Accessed on Apr. 1, 2026].
[7]Anthropic Computer Use Team, 2024. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use [Accessed on Apr. 1, 2026].
[8]Anthropic Documentation Team, 2026. Models overview. https://docs.anthropic.com/en/docs/about-claude/models/overview [Accessed on Apr. 1, 2026].
[9]Anthropic Research Team, 2024. Developing a computer use model. https://www.anthropic.com/news/developing-computer-use [Accessed on Apr. 1, 2026].
[10]Anthropic Safety Team, 2026. Claude Sonnet 4.6 system card. https://www.anthropic.com/claude-sonnet-4-6-system-card [Accessed on Apr. 1, 2026].
[11]Anupam S, Brown D, Li S, et al., 2025. BrowserArena: evaluating LLM agents on real-world web navigation tasks. Workshop on Multi-Turn Interactions in Large Language Models at NeurIPS, p.1-19.
[12]Azam R, Abuelsaad T, Vempaty A, et al., 2024. Multimodal auto validation for self-refinement in web agents. Proc 38th Conf on Neural Information Processing Systems, Workshop on Open-World Agents, p.1-10.
[13]Baechler G, Sunkara S, Wang M, et al., 2024. ScreenAI: a vision-language model for UI and infographics understanding. Proc 33rd Int Joint Conf on Artificial Intelligence, p.3058-3068.
[14]Bai H, Zhou YF, Cemri M, et al., 2024. DigiRL: training in-the-wild device-control agents with autonomous reinforcement learning. Proc 38th Int Conf on Neural Information Processing Systems, Article 397.
[15]Bai S, Chen KQ, Liu XJ, et al., 2025a. Qwen2.5-VL technical report.
[16]Bai S, Cai YX, Chen RZ, et al., 2025b. Qwen3-VL technical report.
[17]Bai TT, Bai YF, Bao YP, et al., 2026. Kimi K2.5: visual agentic intelligence.
[18]Berkovitch O, Caduri S, Kahlon N, et al., 2025. Identifying user goals from UI trajectories. Proc ACM on Web Conf, p.2381-2390.
[19]Boisvert L, Thakkar M, Gasse M, et al., 2024. WorkArena++: towards compositional planning and reasoning-based common knowledge work tasks. Proc 38th Conf on Neural Information Processing Systems, Article 195.
[20]Bommasani R, Hudson DA, Adeli E, et al., 2021. On the opportunities and risks of foundation models. https://arxiv.org/abs/2108.07258v1
[21]Bonatti R, Zhao D, Bonacci F, et al., 2025. Windows Agent Arena: evaluating multi-modal OS agents at scale. Proc 42nd Int Conf on Machine Learning, p.4874-4910.
[22]Brown TB, Mann B, Ryder N, et al., 2020. Language models are few-shot learners. Proc 34th Int Conf on Neural Information Processing Systems, Article 159.
[23]Burns A, Arsan D, Agrawal S, et al., 2022. A dataset for interactive vision-language navigation with unknown command feasibility. Proc 17th European Conf on Computer Vision, p.312-328.
[24]ByteDance Seed, 2026. Seed1.8 model card: towards generalized real-world agency.
[25]Cao Y, Wang YY, Bu P, et al., 2025. AndroidLens: long-latency evaluation with nested sub-targets for Android GUI agents.
[26]Chai YX, Huang SY, Niu YZ, et al., 2025. AMEX: Android multi-annotation expo dataset for mobile GUI agents. Findings of the Association for Computational Linguistics: ACL, p.2138-2156.
[27]Chen DP, Huang Y, Wu SY, et al., 2025. GUI-World: a video benchmark and dataset for multimodal GUI-oriented understanding. Proc 13th Int Conf on Learning Representations, p.1-53.
[28]Chen Q, Pitawela D, Zhao CY, et al., 2024. WebVLN: vision-and-language navigation on websites. Proc 38th AAAI Conf on Artificial Intelligence, p.1165-1173.
[29]Chen WT, Cui JB, Hu JY, et al., 2025. GUICourse: from general vision language model to versatile GUI agent. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.21936-21959.
[30]Chen XT, Zhao ZH, Chen L, et al., 2021. WebSRC: a dataset for web-based structural reading comprehension. Proc Conf on Empirical Methods in Natural Language Processing, p.4173-4185.
[31]Chen XT, Li HC, Liang JQ, et al., 2024. EDGE: enhanced grounded GUI understanding with enriched multi-granularity synthetic data.
[32]Chen YR, Hu XY, Yin KT, et al., 2025. Evaluating the robustness of multimodal agents against active environmental injection attacks. Proc 33rd ACM Int Conf on Multimedia, p.11648-11656.
[33]Cheng KZ, Sun QS, Chu YG, et al., 2024. SeeClick: harnessing GUI grounding for advanced visual GUI agents. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.9313-9332.
[34]Cheng PZ, Wu Z, Wu ZR, et al., 2025. OS-Kairos: adaptive interaction for MLLM-powered GUI agents. Findings of the Association for Computational Linguistics: ACL, p.6701-6725.
[35]de Chezelles TLS, Gasse M, Drouin A, et al., 2025. The BrowserGym ecosystem for web agent research. https://arxiv.org/abs/2412.05467
[36]Deka B, Huang ZF, Franzen C, et al., 2017. Rico: a mobile app dataset for building data-driven design applications. Proc 30th Annual ACM Symp on User Interface Software and Technology, p.845-854.
[37]Deng X, Gu Y, Zheng BY, et al., 2023. Mind2Web: towards a generalist agent for the web. Proc 37th Int Conf on Neural Information Processing Systems, Article 1220.
[38]Dihan ML, Hashem T, Ali ME, et al., 2025. WebOperator: action-aware tree search for autonomous agents in web environment. https://arxiv.org/abs/2512.12692
[39]Dong LZ, Zhou ZQ, Yang SB, et al., 2025. Say one thing, do another? Diagnosing reasoning–execution gaps in VLM-powered mobile-use agents. https://arxiv.org/abs/2510.02204
[40]Drouin A, Gasse M, Caccia M, et al., 2024. WorkArena: how capable are web agents at solving common knowledge work tasks? Proc 41st Int Conf on Machine Learning, p.11642-11662.
[41]Du YQ, Konyushkova K, Denil M, et al., 2023. Vision-language models as success detectors. Proc 2nd Conf on Lifelong Learning Agents, p.120-136.
[42]Duan PT, Cheng CY, Li G, et al., 2024. UICrit: enhancing automated design evaluation with a UI critique dataset. Proc 37th Annual ACM Symp on User Interface Software and Technology, Article 46.
[43]El Hattami A, Thakkar M, Chapados N, et al., 2025. WebArena Verified. Proc 39th Conf on Neural Information Processing Systems Workshop, p.1-28.
[44]Fan Y, Ding L, Kuo CC, et al., 2024. Read anywhere pointed: layout-aware GUI screen reading with tree-of-lens grounding. Proc Conf on Empirical Methods in Natural Language Processing, p.9503-9522.
[45]Feng SD, Chen CY, 2026. How smart is your GUI agent? A framework for the future of software interaction.
[46]Furuta H, Matsuo Y, Faust A, et al., 2024. Exposing limitations of language model agents in sequential-task compositions on the web. https://arxiv.org/abs/2311.18751
[47]Gao DF, Hu SY, Bai ZC, et al., 2024a. AssistEditor: multi-agent collaboration for GUI workflow automation in video creation. Proc 32nd ACM Int Conf on Multimedia, p.11255-11257.
[48]Gao DF, Ji L, Bai ZC, et al., 2024b. AssistGUI: task-oriented PC graphical user interface automation. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.13289-13298.
[49]Gao LX, Zhang L, Wang SH, et al., 2024. MobileViews: a large-scale mobile GUI dataset.
[50]Google DeepMind Computer Use Team, 2025. Introducing the Gemini 2.5 computer use model. https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-computer-use-model [Accessed on Apr. 1, 2026].
[51]Google Gemini 2.0 Team, 2024. Introducing Gemini 2.0: our new AI model for the agentic era. https://blog.google/innovation-and-ai/models-and-research/google-deepmind/google-gemini-ai-update-december-2024 [Accessed on Apr. 1, 2026].
[52]Google Gemini 2.5 Team, 2025. Gemini 2.5: our most intelligent AI model. https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025 [Accessed on Apr. 1, 2026].
[53]Google Gemini 3 Team, 2025. A new era of intelligence with Gemini 3. https://blog.google/products-and-platforms/products/gemini/gemini-3 [Accessed on Apr. 1, 2026].
[54]Google Gemini 3.1 Flash-Lite Team, 2026. Gemini 3.1 flash-lite: built for intelligence at scale. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-lite [Accessed on Apr. 1, 2026].
[55]Google Gemini 3.1 Pro Team, 2026. Gemini 3.1 Pro: a smarter model for your most complex tasks. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro [Accessed on Apr. 1, 2026].
[56]Gou BY, Wang RH, Zheng BY, et al., 2025. Navigating the digital world as humans do: universal visual grounding for GUI agents. Proc 13th Int Conf on Learning Representations, p.1-33.
[57]Gu ZX, Zeng ZW, Xu ZY, et al., 2025. UI-Venus technical report: building high-performance UI agents with RFT.
[58]Guo D, Wu FM, Zhu FD, et al., 2025. Seed1.5-VL technical report.
[59]Guo DY, Yang DJ, Zhang HW, et al., 2025. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633-638.
[60]Gur I, Nachum O, Miao YJ, et al., 2023. Understanding HTML with large language models. Findings of the Association for Computational Linguistics: EMNLP, p.2803-2821.
[61]He HL, Yao WL, Ma KX, et al., 2024. WebVoyager: building an end-to-end web agent with large multimodal models. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.6864-6890.
[62]Hong WY, Wang WH, Lv QS, et al., 2024. CogAgent: a visual language model for GUI agents. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.14281-14290.
[63]Hong WY, Yu WM, Gu XT, et al. 2025. GLM-4.1V-Thinking: towards versatile multimodal reasoning with scalable reinforcement learning. https://arxiv.org/abs/2507.01006v1
[64]Hsiao YC, Zubach F, Baechler G, et al., 2025. ScreenQA: large-scale question–answer pairs over mobile app screenshots. Proc Conf of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, p.9427-9452.
[65]Hu XY, Xiong T, Yi B, et al., 2025. OS agents: a survey on MLLM-based agents for computer, phone and browser use. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.7436-7465.
[66]Huang J, Zeng ZX, Han WK, et al., 2025. ScaleTrack: scaling and back-tracking automated GUI agents.
[67]Huang KH, Qiu HY, Dai YT, et al., 2025. GUI-KV: efficient GUI agents via KV cache with spatio-temporal awareness. https://arxiv.org/abs/2510.00536
[68]Hui Z, Li YH, Zhao D, et al., 2025. WinSpot: GUI grounding benchmark with multimodal large language models. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.1086-1096.
[69]Jang LK, Li YH, Zhao D, et al., 2025. VideoWebArena: evaluating long context multimodal agents with video understanding web tasks. Proc 13th Int Conf on Learning Representations, p.1-25.
[70]Jia HR, Liao JT, Zhang X, et al., 2026. OSWorld-MCP: benchmarking MCP tool invocation in computer-use agents. Proc 14th Int Conf on Learning Representations, p.1-35.
[71]Jin YL, Li Z, Zhang CW, et al., 2024. Shopping MMLU: a massive multi-task online shopping benchmark for large language models. Proc 38th Conf on Neural Information Processing Systems, p.1-28.
[72]Kapoor R, Butala YP, Russak M, et al., 2024. OmniACT: a dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web. Proc 18th European Conf on Computer Vision, p.161-178.
[73]Koh JY, Lo R, Jang L, et al., 2024. VisualWebArena: evaluating multimodal agents on realistic visual web tasks. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.881-905.
[74]Kong QY, Zhang X, Yang ZY, et al., 2025. MobileWorld: benchmarking autonomous mobile agents in agent–user interactive, and MCP-augmented environments.
[75]Kumar P, Lau E, Vijayakumar S, et al., 2024. Refusal-trained LLMs are easily jailbroken as browser agents.
[76]Lai HY, Liu X, Iong IL, et al., 2024. AutoWebGLM: a large language model-based web navigating agent. Proc 30th ACM SIGKDD Conf on Knowledge Discovery and Data Mining, p.5295-5306.
[77]Lai HY, Gao JJ, Liu X, et al., 2025. AndroidGen: building an Android language agent under data scarcity. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.2727-2749.
[78]Lee J, Hahm D, Choi JS, et al., 2026. MobileSafetyBench: evaluating safety of autonomous agents in mobile device control. Proc 40th AAAI Conf on Artificial Intelligence, p.37565-37573.
[79]Lei B, Xu N, Payani A, et al., 2025. GUI-SPOTLIGHT: adaptive iterative focus refinement for enhanced GUI visual grounding. https://arxiv.org/abs/2510.04039
[80]Leung HF, Xi XY, Zuo F, 2025. AndroidControl-Curated: revealing the true potential of GUI agents through benchmark purification.
[81]Levy I, Wiesel B, Marreed S, et al., 2026. ST-WebAgentBench: a benchmark for evaluating safety and trustworthiness in web agents. Proc 14th Int Conf on Learning Representations, p.1-43.
[82]Li E, Waldo J, 2024. WebSuite: systematically evaluating why web agents fail.
[83]Li KX, Meng ZY, Lin HZ, et al., 2025. ScreenSpot-Pro: GUI grounding for professional high-resolution computer use. Proc 33rd ACM International Conference on Multimedia, p.8778-8786.
[84]Li SF, Kallidromitis K, Gokul A, et al., 2025. MobileWorldBench: towards semantic world modeling for mobile agents.
[85]Li SL, Bu XY, Wang WJ, et al., 2025. MM-BrowseComp: a comprehensive benchmark for multimodal browsing agents. https://arxiv.org/abs/2508.13186
[86]Li T, Li G, Zheng JJ, et al., 2024. MUG: interactive multimodal grounding on user interfaces. Findings of the Association for Computational Linguistics: EACL, p.231-251.
[87]Li TJJ, Popowski L, Mitchell TM, et al., 2021. Screen2Vec: semantic embedding of GUI screens and GUI components. Proc CHI Conf on Human Factors in Computing Systems, Article 578.
[88]Li WZ, Lin JB, Jiang ZS, et al., 2025. Chain-of-Agents: end-to-end agent foundation models via multi-agent distillation and agentic RL.
[89]Li YD, Zhang C, Jiang WJ, et al., 2024. AppAgent v2: advanced agent for flexible mobile interactions.
[90]Li ZH, You KE, Zhang HT, et al., 2025. Ferret-UI 2: mastering universal user interface understanding across platforms. Proc 13th Int Conf on Learning Representations, p.1-24.
[91]Liao ZY, Mo LB, Xu CJ, et al., 2025. EIA: environmental injection attack on generalist web agents for privacy leakage. Proc 13th Int Conf on Learning Representations, p.1-32.
[92]Liao ZY, Jones J, Jiang LX, et al., 2026. RedTeamCUA: realistic adversarial testing of computer-use agents in hybrid web-OS environments. Proc 14th Int Conf on Learning Representations, p.1-46.
[93]Lin HJ, Tan XY, Qin YL, et al., 2025. CUARewardBench: a benchmark for evaluating reward models on computer-using agent.
[94]Lin KQ, Li LJ, Gao DF, et al., 2024. VideoGUI: a benchmark for GUI automation from instructional videos. Proc 38th Conf on Advances in Neural Information Processing Systems, p.1-32.
[95]Lin KQ, Li LJ, Gao DF, et al., 2025. ShowUI: one vision-language-action model for GUI visual agent. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.19498-19508.
[96]Liu EZ, Guu K, Pasupat P, et al., 2018. Reinforcement learning on web interfaces using workflow-guided exploration. Proc 6th Int Conf on Learning Representations, p.1-15.
[97]Liu GY, Zhao PX, Liu L, et al., 2025. LearnAct: few-shot mobile GUI agent with a unified demonstration benchmark.
[98]Liu GY, Zhao PX, Liang YZ, et al., 2026. MemGUI-Bench: benchmarking memory of mobile GUI agents in dynamic environments. https://arxiv.org/abs/2602.06075
[99]Liu JP, Ou TY, Song YF, et al., 2025. Harnessing webpage UIs for text-rich visual understanding. Proc 13th Int Conf on Learning Representations, p.1-35.
[100]Liu SY, Liu MH, Zhou HC, et al., 2025. VeriWeb: verifiable long-chain GUI dataset. https://arxiv.org/abs/2508.04026v1
[101]Liu X, Yu H, Zhang HC, et al., 2024a. AgentBench: evaluating LLMs as agents. Proc 12th Int Conf on Learning Representations, p.1-58.
[102]Liu X, Qin B, Liang DZ, et al., 2024b. AutoGLM: autonomous foundation agents for GUIs.
[103]Liu YH, Li PX, Wei ZS, et al., 2026. InfiGUIAgent: a multimodal generalist GUI agent with native reasoning and reflection. Proc 19th Conf of the European Chapter of the Association for Computational Linguistics, p.1035-1051.
[104]Liu ZY, Xie JJ, Ding ZC, et al., 2025. ScaleCUA: scaling open-source computer use agents with cross-platform data.
[105]Lu JT, Zhang ZY, Yang FK, et al., 2024. Turn every application into an agent: towards efficient human–agent–computer interaction with API-first LLM-based agents. https://arxiv.org/abs/2409.17140v1
[106]Lu QF, Shao WQ, Liu ZT, et al., 2025. GUIOdyssey: a comprehensive dataset for cross-app GUI navigation on mobile devices. Proc IEEE/CVF Int Conf on Computer Vision, p.22404-22414.
[107]Lu YD, Yang JW, Shen YL, et al., 2024. OmniParser for pure vision based GUI agent.
[108]Lù XH, Kasner Z, Reddy S, 2024. WebLINX: real-world website navigation with multi-turn dialogue. Proc 41st Int Conf on Machine Learning, p.33007-33056.
[109]Luo TG, Logeswaran L, Johnson J, et al., 2025. Visual test-time scaling for GUI agent grounding. Proc IEEE/CVF Int Conf on Computer Vision, p.19989-19998.
[110]Ma XB, Wang YT, Yao Y, et al., 2025. Caution for the environment: multimodal LLM agents are susceptible to environmental distractions. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.22324-22339.
[111]Mialon G, Fourrier C, Swift C, et al., 2024. GAIA: a benchmark for general AI assistants. Proc 12th Int Conf on Learning Representations, p.1-25.
[112]Mozannar H, Bansal G, Tan C, et al., 2025. Magentic-UI: towards human-in-the-loop agentic systems.
[113]Mu J, Zhang CY, Ni CM, et al., 2025. GUI-360°: a comprehensive dataset and benchmark for computer-using agents.
[114]Murty S, Zhu H, Bahdanau D, et al., 2024. NNetNav: unsupervised learning of browser agents through environment interaction in the wild.
[115]Nakano R, Hilton J, Balaji S, et al., 2021. WebGPT: browser-assisted question-answering with human feedback.
[116]Nayak S, Jian XR, Lin KQ, et al., 2025. UI-Vision: a desktop-centric GUI benchmark for visual perception and interaction. Proc 42nd Int Conf on Machine Learning, p.45817-45851.
[117]Nguyen A, 2024. Improved GUI grounding via iterative narrowing.
[118]Nguyen D, Chen J, Wang Y, et al., 2025. GUI agents: a survey. Findings of the Association for Computational Linguistics: ACL, p.22522-22538.
[119]Nong SQ, Zhu JL, Wu R, et al., 2024. MobileFlow: a multimodal LLM for mobile GUI agent.
[120]OpenAI ChatGPT Agent Team, 2025. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent [Accessed on Apr. 1, 2026].
[121]OpenAI CUA Team, 2025. Computer-using agent. https://openai.com/index/computer-using-agent [Accessed on Apr. 1, 2026].
[122]OpenAI Developer Platform Team, 2025. Introducing GPT-5 for developers. https://openai.com/index/introducing-gpt-5-for-developers [Accessed on Apr. 1, 2026].
[123]OpenAI GPT-5 Team, 2025. Introducing GPT-5. https://openai.com/index/introducing-gpt-5 [Accessed on Apr. 1, 2026].
[124]OpenAI GPT-5.4 Mini and Nano Team, 2026. Introducing GPT-5.4 mini and nano. https://openai.com/index/introducing-gpt-5-4-mini-and-nano [Accessed on Apr. 1, 2026].
[125]OpenAI GPT-5.4 Team, 2026. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4 [Accessed on Apr. 1, 2026].
[126]OpenAI o3 Operator Team, 2025. Addendum to OpenAI o3 and o4-mini system card: OpenAI o3 Operator. https://openai.com/index/o3-o4-mini-system-card-addendum-operator-o3 [Accessed on Apr. 1, 2026].
[127]OpenAI Operator System Card Team, 2025. Operator system card. https://openai.com/index/operator-system-card [Accessed on Apr. 1, 2026].
[128]OpenAI Operator Team, 2025. Introducing operator. https://openai.com/index/introducing-operator [Accessed on Apr. 1, 2026].
[129]OpenGVLab, 2025. WebArena-Lite-v2 benchmark evaluation guide. https://github.com/OpenGVLab/ScaleCUA/tree/main/evaluation/WebArenaLiteV2 [Accessed on Apr. 1, 2026].
[130]Ou TY, Xu FF, Madaan A, et al., 2024. Synatra: turning indirect knowledge into direct demonstrations for digital agents at scale. Proc 38th Int Conf on Advances in Neural Information Processing Systems, Article 2908.
[131]Pahuja V, Lu YD, Rosset C, et al., 2025. Explorer: scaling exploration-driven web trajectory synthesis for multimodal web agents. Findings of the Association for Computational Linguistics: ACL, p.6300-6323.
[132]Pan LH, Wang BW, Yu C, et al., 2023. AutoTask: executing arbitrary voice commands by exploring and learning from mobile GUI.
[133]Pan YC, Kong DH, Zhou SD, et al., 2024. WebCanvas: benchmarking web agents in online environments. Proc ICML Workshop on Agentic Markets, p.1-26.
[134]Park J, Tang P, Das S, et al., 2025. R-VLM: region-aware vision language model for precise GUI grounding. Findings of the Association for Computational Linguistics: ACL, p.9669-9685.
[135]Park JS, O’Brien JC, Cai CJ, et al., 2023. Generative agents: interactive simulacra of human behavior. Proc 36th Annual ACM Symp on User Interface Software and Technology, Article 2.
[136]Pawlowski P, Zawistowski K, Lapacz W, et al., 2025. TinyClick: single-turn agent for empowering GUI automation. Proc 26th Annual Conf of the Int Speech Communication Association, p.3035-3039.
[137]Qin YJ, Ye YN, Fang JJ, et al., 2025. UI-TARS: pioneering automated GUI interaction with native agents.
[138]Qwen3 Team, 2025. Qwen3: think deeper, act faster. https://qwenlm.github.io/blog/qwen3 [Accessed on Apr. 1, 2026].
[139]Qwen3-Coder Team, 2025. Qwen3-coder: agentic coding in the world. https://qwenlm.github.io/blog/qwen3-coder [Accessed on Apr. 1, 2026].
[140]Qwen3Guard Team, 2025. Qwen3Guard: real-time safety for your token stream. https://qwenlm.github.io/blog/qwen3guard [Accessed on Apr. 1, 2026].
[141]Rawles C, Li A, Rodriguez D, et al., 2023. Android in the Wild: a large-scale dataset for Android device control. Proc 37th Conf on Neural Information Processing Systems, Datasets and Benchmarks Track, p.1-21.
[142]Rawles C, Clinckemaillie S, Chang YF, et al., 2025. AndroidWorld: a dynamic benchmarking environment for autonomous agents. Proc 13th Int Conf on Learning Representations, p.1-36.
[143]Rodriguez JA, Jian XR, Panigrahi SS, et al., 2025. BigDocs: an open and permissively-licensed dataset for training multimodal models on document and code tasks. Proc 13th Int Conf on Learning Representations, p.1-57.
[144]Sager PJ, Meyer B, Yan P, et al., 2026. A comprehensive survey of agents for computer use: foundations, challenges, and future directions. J Artif Intell Res, 85:34.
[145]Schick T, Dwivedi-Yu J, Dessi R, et al., 2023. Toolformer: language models can teach themselves to use tools. Proc 37th Int Conf on Advances in Neural Information Processing Systems, Article 2997.
[146]Shaw P, Joshi M, Cohan J, et al., 2023. From pixels to UI actions: learning to follow instructions via graphical user interfaces. Proc 37th Int Conf on Neural Information Processing Systems, Article 1490.
[147]Shen HW, Liu C, Li GL, et al., 2024. Falcon-UI: understanding GUI before following user instructions.
[148]Shen JH, Jain A, Xiao ZD, et al., 2024. ScribeAgent: towards specialized web agents using production-scale workflow data.
[149]Shen JH, Bai H, Zhang LJ, et al., 2025. Thinking vs. doing: agents that reason by scaling test-time interaction.
[150]Shi TL, Karpathy A, Fan LX, et al., 2017. World of Bits: an open-domain platform for web-based agents. Proc 34th Int Conf on Machine Learning, p.3135-3144.
[151]Shi YC, Yu WH, Huang JY, et al., 2025. Towards trustworthy GUI agents: a survey.
[152]Shi ZR, Mei K, Jin MY, et al., 2025. From commands to prompts: LLM-based semantic file system for AIOS. Proc 13th Int Conf on Learning Representations, p.1-24.
[153]Shinn N, Cassano F, Gopinath A, et al., 2023. Reflexion: language agents with verbal reinforcement learning. Proc 37th Int Conf on Neural Information Processing Systems, Article 377.
[154]Song YP, Bian YH, Tang YT, et al., 2024. VisionTasker: mobile task automation using vision based UI understanding and LLM task planning. Proc 37th Annual ACM Symp on User Interface Software and Technology, Article 49.
[155]Styles O, Miller S, Cerda-Mardini P, et al., 2024. WorkBench: a benchmark dataset for agents in a realistic workplace setting. Proc 1st Conf on Language Modeling, p.1-39.
[156]Steel.dev, 2026. WebArena leaderboard. https://leaderboard.steel.dev/leaderboards/webarena/ [Accessed on Apr. 1, 2026].
[157]Su Y, Awadallah AH, Khabsa M, et al., 2017. Building natural language interfaces to web APIs. Proc 26th ACM Int Conf on Information and Knowledge Management, p.177-186.
[158]Sun LT, Chen XY, Chen L, et al., 2022. META-GUI: towards multi-modal conversational agents on mobile GUI. Proc Conf on Empirical Methods in Natural Language Processing, p.6699-6712.
[159]Sun QS, Liu ZMZ, Ma C, et al., 2026. ScienceBoard: evaluating multimodal autonomous agents in realistic scientific workflows. Proc 14th Int Conf on Learning Representations, p.1-38.
[160]Toyama D, Hamel P, Gergely A, et al., 2021. AndroidEnv: a reinforcement learning platform for Android.
[161]Tur AD, Meade N, Lù XH, et al., 2025. SafeArena: evaluating the safety of autonomous web agents. Proc 42nd Int Conf on Machine Learning, p.60404-60441.
[162]Vaswani A, Shazeer N, Parmar N, et al., 2017. Attention is all you need. Proc 31st Int Conf on Neural Information Processing Systems, p.6000-6010.
[163]Venkatesh SG, Talukdar P, Narayanan S, 2022. UGIF: UI grounded instruction following.
[164]Vu MD, Wang H, Chen JS, et al., 2024. GPTVoiceTasker: advancing multi-step mobile task efficiency through dynamic interface exploration and learning. Proc 37th Annual ACM Symp on User Interface Software and Technology, Article 48.
[165]Wang B, Li G, Zhou X, et al., 2021. Screen2Words: automatic mobile UI summarization with multimodal learning. Proc 34th Annual ACM Symp on User Interface Software and Technology, p.498-510.
[166]Wang HM, Zou HY, Song HT, et al., 2025. UI-TARS-2 technical report: advancing GUI agent with multi-turn reinforcement learning.
[167]Wang JY, Xu HY, Ye JB, et al., 2024. Mobile-Agent: autonomous multi-modal mobile device agent with visual perception.
[168]Wang K, Xia TY, Gu ZX, et al., 2024. E-ANT: a large-scale dataset for efficient automatic GUI navigation.
[169]Wang LY, Deng YY, Zha YW, et al., 2024. MobileAgentBench: an efficient and user-friendly benchmark for mobile LLM agents.
[170]Wang P, Tao RH, Chen QG, et al., 2025. X-WebAgentBench: a multilingual interactive web benchmark for evaluating global agentic system. Findings of the Association for Computational Linguistics: ACL, p.19320-19335.
[171]Wang S, Liu WW, Chen JX, et al., 2024. GUI agents with foundation models: a comprehensive survey.
[172]Wang XL, Bloch J, Shao ZD, et al., 2025. WebInject: prompt injection attack to web agents. Proc Conf on Empirical Methods in Natural Language Processing, p.2010-2030.
[173]Wang YQ, Zhang HJ, Tian JQ, et al., 2025. Ponder & Press: advancing visual GUI agent towards general computer control. Findings of the Association for Computational Linguistics: ACL, p.1461-1473.
[174]Wen H, Wang HM, Liu JX, et al., 2023. DroidBot-GPT: GPT-powered UI automation for Android.
[175]Wen H, Tian SZ, Pavlov B, et al., 2025. AutoDroid-V2: boosting SLM-based GUI agents via code generation. Proc 23rd Annual Int Conf on Mobile Systems, Applications and Services, p.223-235.
[176]Wu CH, Shah RR, Koh JY, et al., 2025. Dissecting adversarial robustness of multimodal LM agents. Proc 13th Int Conf on Learning Representations, p.1-22.
[177]Wu FZ, Wu ST, Cao YL, et al., 2024. WIPI: a new web threat for LLM-driven web agents.
[178]Wu QZ, Xu WK, Liu W, et al., 2024. MobileVLM: a vision-language model for better intra- and inter-UI understanding. Findings of the Association for Computational Linguistics: EMNLP, p.10231-10251.
[179]Wu QZ, Gao PZ, Liu W, et al., 2025. BacktrackAgent: enhancing GUI agent with error detection and backtracking mechanism. Proc Conf on Empirical Methods in Natural Language Processing, p.4250-4272.
[180]Wu ZY, Wu ZY, Xu FZ, et al., 2025. OS-Atlas: a foundation action model for generalist GUI agents. Proc 13th Int Conf on Learning Representations, p.1-19.
[181]Xie B, Shao R, Chen GW, et al., 2025. GUI-Explorer: autonomous exploration and mining of transition-aware knowledge for GUI agent. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.5650-5667.
[182]Xie JL, Chen ZH, Zhang RF, et al., 2025. Large multimodal agents: a survey. Vis Intell, 3(1):24.
[183]Xie TB, Zhou F, Cheng ZJ, et al., 2024a. OpenAgents: an open platform for language agents in the wild. Proc 1st Conf on Language Modeling, p.1-36.
[184]Xie TB, Zhang DY, Chen JX, et al., 2024b. OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environments. Proc 38th Int Conf on Neural Information Processing Systems, Article 1650.
[185]Xie TB, Deng JQ, Li XC, et al., 2025. Scaling computer-use grounding via user interface decomposition and synthesis. Proc 39th Conf on Neural Information Processing Systems, p.1-58.
[186]XLANG Lab, 2025. OSWorld-Verified. https://xlang.ai/blog/osworld-verified [Accessed on Apr. 1, 2026].
[187]Xu CJ, Kang MT, Zhang JW, et al., 2025. AdvAgent: controllable blackbox red-teaming on web agents. Proc 42nd Int Conf on Machine Learning, p.69318-69330.
[188]Xu HY, Zhang X, Liu HW, et al., 2026. Mobile-Agent-v3.5: multi-platform fundamental GUI agents.
[189]Xu QT, Hong FL, Li B, et al., 2023. On the tool manipulation capability of open-source large language models.
[190]Xu TQ, Chen LY, Wu DJ, et al., 2025. CRAB: cross-environment agent benchmark for multimodal language model agents. Findings of the Association for Computational Linguistics: ACL, p.21607-21647.
[191]Xu YF, Liu X, Sun XQ, et al., 2025a. AndroidLab: training and systematic benchmarking of Android autonomous agents. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.2144-2166.
[192]Xu YF, Liu X, Liu XH, et al., 2025b. MobileRL: online agentic reinforcement learning for mobile GUI agents. https://arxiv.org/abs/2509.18119
[193]Xu YH, Lu DJ, Shen ZN, et al., 2025a. AgentTrek: agent trajectory synthesis via guiding replay with web tutorials. Proc 13th Int Conf on Learning Representations, p.1-22.
[194]Xu YH, Wang ZK, Wang JL, et al., 2025b. AGUVIS: unified pure vision agents for autonomous GUI interaction. Proc 42nd Int Conf on Machine Learning, p.69772-69805.
[195]X-WebArena Leaderboard Maintainers, 2026. X-WebArena-Leaderboard: public WebArena results spreadsheet. https://docs.google.com/spreadsheets/d/1M801lEpBbKSNwP-vDBkC_pF7LdyGU1f_ufZb_NWNBZQ [Accessed on Apr. 1, 2026].
[196]Yang JW, Zhang H, Li F, et al., 2023. Set-of-Mark prompting unleashes extraordinary visual grounding in GPT-4V.
[197]Yang K, Liu Y, Chaudhary S, et al., 2025. AgentOccam: a simple yet strong baseline for LLM-based web agents. Proc 13th Int Conf on Learning Representations, p.1-33.
[198]Yang P, Ci H, Shou MZ, 2025. macOSWorld: a multilingual interactive benchmark for GUI agents.
[199]Yang Y, Li DX, Dai YT, et al., 2026. GTA1: GUI test-time scaling agent. Proc 14th Int Conf on Learning Representations, p.1-29.
[200]Yang YH, Wang Y, Li DX, et al., 2025. Aria-UI: visual grounding for GUI instructions. Findings of the Association for Computational Linguistics: ACL, p.22418-22433.
[201]Yang YL, Yang XS, Li SD, et al., 2024. Security matrix for multimodal agents on mobile devices: a systematic and proof of concept study. https://arxiv.org/abs/2407.09295v1
[202]Yang Z, Dou ZY, Feng D, et al., 2026. Ferret-UI Lite: lessons from building small on-device GUI agents. https://machinelearning.apple.com/research/ferret-ui [Accessed on Apr. 1, 2026].
[203]Yao SY, Chen H, Yang J, et al., 2022. WebShop: towards scalable real-world web interaction with grounded language agents. Proc 36th Int Conf on Advances in Neural Information Processing Systems, Article 1508.
[204]Yao SY, Zhao J, Yu D, et al., 2023. ReAct: synergizing reasoning and acting in language models. Proc 11th Int Conf on Learning Representations, p.1-33.
[205]Ye JB, Zhang X, Xu HY, et al., 2025. Mobile-Agent-v3: fundamental agents for GUI automation.
[206]Yin XR, Luo X, Wu H, et al., 2025. Unlocking smarter device control: foresighted planning with a world model-driven code execution approach. Findings of the Association for Computational Linguistics: EMNLP, p.3982-4005.
[207]Ying KN, Meng FQ, Wang J, et al., 2024. MMT-Bench: a comprehensive multimodal benchmark for evaluating large vision-language models towards multitask AGI. Proc 41st Int Conf on Machine Learning, p.57116-57198.
[208]Yu S, Li G, Shi WY, et al., 2025. PolySkill: learning generalizable skills through polymorphic abstraction.
[209]Zhang BL, Xiong SY, Sui DB, et al., 2024. RealWeb: a benchmark for universal instruction following in realistic web services navigation. Proc IEEE Int Conf on Web Services, p.342-351.
[210]Zhang C, Yang Z, Liu JX, et al., 2025. AppAgent: multimodal agents as smartphone users. Proc CHI Conf on Human Factors in Computing Systems, Article 70.
[211]Zhang CY, He SL, Qian JX, et al., 2025a. Large language model-brained GUI agents: a survey. https://arxiv.org/abs/2411.18279
[212]Zhang CY, Li LQ, He SL, et al., 2025b. UFO: a UI-Focused Agent for Windows OS interaction. Proc Conf of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, p.597-622.
[213]Zhang DY, Chen L, Yu K, 2023. Mobile-Env: a universal platform for training and evaluation of mobile interaction. https://arxiv.org/abs/2305.08144v1
[214]Zhang DY, Zhang ST, Yang ZY, et al., 2025. ProgRM: build better GUI agents with progress rewards.
[215]Zhang JW, Yu YQ, Liao MH, et al., 2025. UI-Hawk: unleashing the screen stream understanding for mobile GUI agents. Proc Conf on Empirical Methods in Natural Language Processing, p.18217-18236.
[216]Zhang JY, Zhao C, Zhao YH, et al., 2024. MobileExperts: a dynamic tool-enabled agent team in mobile devices.
[217]Zhang MS, Xu ZQ, Zhu JL, et al., 2025. Phi-Ground Tech Report: advancing perception in GUI grounding. Technical Report No. MSR-TR-2025-63, Microsoft Research, Redmond, WA.
[218]Zhang SQ, Zhang ZS, Chen KH, et al., 2024. Dynamic planning for LLM-based graphical user interface automation. Findings of the Association for Computational Linguistics: EMNLP, p.1304-1320.
[219]Zhang YZ, Yu T, Yang DY, 2025. Attacking vision-language computer agents via pop-ups. Proc 63rd Annual Meeting of the Association for Computational Linguistics, p.8387-8401.
[220]Zhang Z, Lu YX, Fu YK, et al., 2025. AgentCPM-GUI: building mobile-use agents with reinforcement fine-tuning. Proc Conf on Empirical Methods in Natural Language Processing: System Demonstrations, p.155-180.
[221]Zhang ZS, Zhang A, 2024. You Only Look at Screens: multimodal chain-of-action agents. Findings of the Association for Computational Linguistics: ACL, p.3132-3149.
[222]Zhang ZY, Liu XY, Zhang XY, et al., 2025. UI-Evol: automatic knowledge evolving for computer use agents.
[223]Zhao HH, Gao DF, Shou MZ, 2025. WorldGUI: an interactive benchmark for desktop GUI automation from any starting point.
[224]Zheng BY, Gou BY, Kil J, et al., 2024. GPT-4V(ision) is a generalist web agent, if grounded. Proc 41st Int Conf on Machine Learning, p.61349-61385.
[225]Zhou HZ, Zhang X, Tong PR, et al., 2025. MAI-UI technical report: real-world centric foundation GUI agents.
[226]Zhou SY, Xu FF, Zhu H, et al., 2024. WebArena: a realistic web environment for building autonomous agents. Proc 12th Int Conf on Learning Representations, p.1-22.
[227]Zhu ZC, Tang H, Li YS, et al., 2024. MobA: multifaceted memory-enhanced adaptive planning for efficient mobile task automation.
CLC number: TP391.4
On-line Access: 2026-08-12
Received: 2026-04-13
Revision Accepted: 2026-06-06
Crosschecked: 2026-08-12
Cited: 0
Clicked: 22
Open peer comments: Debate/Discuss/Question/Opinion
<1>