METAL LAB

모바일·웹·데스크톱을 한 모델로 다루는 GUI 에이전트, UI-Venus-2가 실제 화면 조작 성능을 끌어올렸다

arXiv:2609.000282026-09-02

UI-Venus-2 Technical Report

모바일·웹·데스크톱을 한 모델로 다루는 GUI 에이전트, UI-Venus-2가 실제 화면 조작 성능을 끌어올렸다

UI-Venus-2는 화면을 보고 클릭·타이핑·스크롤 같은 동작을 수행해 작업을 자동화하는 GUI 에이전트로, 모바일 앱 170여 개와 데스크톱 OS까지 다루는 범위를 넓혔다. 딥리서치 기반으로 작업 지시문을 자동 생성하고, 궤적 단위와 샘플 단위 검증기로 강화학습 신호의 신뢰도를 높였다. 그 결과 UI-Venus-2-27B와 9B 모델을 여러 GUI 벤치마크에서 강력한 비교 모델들과 나란히 평가했다.

METAL LAB 해설 도표

UI-Venus-2 3단계 학습 파이프라인

증거 상태측정 결과와 예정된 검증이 함께 있음

  1. 중간학습모바일·웹·OS 환경에서 수집한 대규모 궤적으로 GUI 상호작용 지식을 먼저 주입한다.
  2. 도메인별 오프라인 강화학습Grounding, CAPTCHA, Mobile, Web, Computer 각 도메인을 개별적으로 스텝 단위 강화학습으로 최적화한다.
  3. 다중 교사 온폴리시 지식증류(MOPD)도메인별로 특화된 모델들을 하나의 UI-Venus-2 모델로 통합하며, 실행 동작 부분에 학습 신호를 집중시킨다.
  4. 궤적·샘플 검증시각적 키포인트와 다중 모델 투표로 궤적 단위, 행동 단위 검증을 수행해 신뢰할 수 있는 강화학습 보상을 공급한다.
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 모바일, 웹, 데스크톱 세 환경에서 화면을 보고 판단하고 실행하는 하나의 통합 반응-행동 루프로 동작하도록 설계됐다.
  2. 환경(170개 이상 다국어 모바일 앱과 데스크톱 OS), 작업(딥리서치 기반 지시문 생성), 검증(궤적 단위·샘플 단위 평가) 세 가지를 함께 확장하는 방식을 택했다.
  3. 학습은 대규모 궤적 기반 중간학습, 도메인별 오프라인 강화학습, 다중 교사 온폴리시 지식증류(MOPD) 세 단계로 진행됐다.
  4. MOPD 단계에서는 추론 문장과 실행 동작을 다르게 취급해, 실제 환경을 바꾸는 동작 부분에 더 집중적으로 학습 신호를 주는 구조를 도입했다.
  5. 로그인·회원가입 등에서 막히는 CAPTCHA를 처리하는 능력과, 위험할 수 있는 동작을 통제하는 안전 장치도 함께 통합했다.
Figure 1: Performances of UI-Venus-2 on GUI-agent benchmarks. Each panel compares UI-Venus-2-27B and UI-Venus-2-9B with some selected strong baselines. We favor standalone end-to-end systems evaluated on the closest available task subset and step budget; source-reported action scaffolds may still differ. MobileWorld uses GUI-only success rate on 117 tasks with 50 steps, WebVoyager uses the refreshed 595-task split, Odysseys uses average rubric score over 200 tasks, VenusBench-CAPTCHA uses micro Pass@1 over all 219 examples, and VenusBench-GD uses English-instruction micro-average accuracy. “*” denotes the results are reproduced by us.
Figure 1: Performances of UI-Venus-2 on GUI-agent benchmarks. Each panel compares UI-Venus-2-27B and UI-Venus-2-9B with some selected strong baselines. We favor standalone end-to-end systems evaluated on the closest available task subset and step budget; source-reported action scaffolds may still differ. MobileWorld uses GUI-only success rate on 117 tasks with 50 steps, WebVoyager uses the refreshed 595-task split, Odysseys uses average rubric score over 200 tasks, VenusBench-CAPTCHA uses micro Pass@1 over all 219 examples, and VenusBench-GD uses English-instruction micro-average accuracy. “*” denotes the results are reproduced by us.
Table 1: Performance comparison on various mobile GUI benchmarks. VenusBench-Mobile reports success rate on its 149-task primary pool. For MobileWorld, we report GUI-only success rate on 117 tasks under the 50-step setting; values in parentheses, when available, use 100 steps. MemGUI reports Main Results pass@1. “*” denotes the baseline results evaluated or reproduced by us.
ModelsMobileGymVenusBench-MobileAndroidWorldMobileWorldKnowUBenchMemGUI
General VLMs
Qwen3.5-9B (Qwen Team, 2026a)9.0*15.3*57.818.0(18.0)*33.36.2*
Qwen3.6-27B (Qwen Team, 2026b)24.6*28.0*70.336.8(41.9)*-25.7*
Claude-Opus-4.6 (Anthropic, 2026a)-36.5*-44.5--
Kimi-K2.6 (Moonshot AI, 2026a)38.7*31.2*-55.6-39.1
Kimi-K3 (Moonshot AI, 2026b)---74.4--
Seed-2.0-Pro (Seed, 2026)52.020.1*-63.251.665.6*
Seed-2.1-Pro (ByteDance Seed, 2026)---73.2--
GPT-5.6-Sol (OpenAI, 2026)---70.1--
GUI-specific Models
UI-Venus-1.5-8B (Team et al., 2026c)18.4*16.173.722.2*26.03.9*
UI-Venus-1.5-30B-A3B (Team et al., 2026c)21.5*21.577.617.1-10.9*
GUI-Owl-1.5-32B-Instruct (Xu et al., 2026)20.3*-69.843.9-10.9
MAI-UI-8B (Zhou et al., 2025b)21.5*12.770.727.526.017.2*
Qwen-UI-Agent-27B (Zhou et al., 2026)---82.1(85.5)--
Ours
UI-Venus-2-9B52.746.580.265.8(75.2)56.562.6
UI-Venus-2-27B60.548.784.076.1(82.9)59.770.3
Figure 2: System Overview of UI-Venus-2. The figure illustrates the task generation and trajectory collection process of UI-Venus-2. Diverse tasks are constructed to form a multi-domain task pool, and interaction trajectories are collected across mobile, browser, and computer environments, covering a broad range of real-world applications, websites, and desktop software.
Figure 2: System Overview of UI-Venus-2. The figure illustrates the task generation and trajectory collection process of UI-Venus-2. Diverse tasks are constructed to form a multi-domain task pool, and interaction trajectories are collected across mobile, browser, and computer environments, covering a broad range of real-world applications, websites, and desktop software.
Table 2: Performance comparison on computer-use agent benchmarks: OSWorld-Verified and DeskCraft (left), and OSWorld 2.0 (right) under the official 150-step budget with 108 tasks. Reported OSWorld-Verified baselines use the 361-task setting in their cited source and may use model-specific action scaffolds. For DeskCraft, we report an author-evaluated aggregate over the 538-task union of the Standard and Interactive splits, which differs from the benchmark’s official split-level reporting. OSWorld 2.0 results report the official Binary Accuracy and Partial Score metrics; baselines are taken from the official leaderboard, possibly with model-specific tool settings, and the reasoning-effort setting is labeled in parentheses for models with multiple official entries. “*” indicates baseline results evaluated by us.
ModelsOSWorld-VerifiedDeskCraft
General VLMs
Claude-Opus-4.8 (Zhou et al., 2026)83.4-
Qwen3.5-9B (Qwen Team, 2026a)41.814.6∗
Qwen3.6-27B (Qwen Team, 2026b)62.028.7∗
Kimi-K2.6 (Moonshot AI, 2026a)73.141.4∗
Seed-2.0-Pro (Seed, 2026)62.340.0∗
Seed-2.1-Pro (Zhou et al., 2026)78.8-
GPT-5.5 (Zhou et al., 2026)78.7-
GUI-specific Models
GUI-Owl-1.5-32B-Instruct (Xu et al., 2026)56.5-
Qwen-UI-Agent-27B (Zhou et al., 2026)79.5-
Ours
UI-Venus-2-9B70.848.0
UI-Venus-2-27B80.555.5
Figure 3: The Three-Stage Pipeline of UI-Venus-2. Following the overall training recipe of UI-Venus-1.5, UI-Venus-2 starts with large-scale trajectory-based mid-training to inject GUI interaction knowledge. The resulting model is then optimized independently for each domain using step-level Offline-RL, covering Grounding, CAPTCHA, Mobile, Web, and Computer tasks. Finally, the domain-specialized models are consolidated into the final UI-Venus-2 model via multi-teacher on-policy distillation.
Figure 3: The Three-Stage Pipeline of UI-Venus-2. Following the overall training recipe of UI-Venus-1.5, UI-Venus-2 starts with large-scale trajectory-based mid-training to inject GUI interaction knowledge. The resulting model is then optimized independently for each domain using step-level Offline-RL, covering Grounding, CAPTCHA, Mobile, Web, and Computer tasks. Finally, the domain-specialized models are consolidated into the final UI-Venus-2 model via multi-teacher on-policy distillation.
Table 3: Performance comparison on four live-web benchmarks: WebVoyager, Online-Mind2Web, REAL, and Odysseys. The Fara1.5 and GPT-5 (SoM) WebVoyager entries use the refreshed 595-task, 100-step robust protocol and are averaged over three runs; live-site states may vary by evaluation date. For Odysseys, we report both the averaged rubric score (Avg.) and the perfect rubric score (Perfect). "*" indicates our reproduced results.
ModelsWebVoyagerOnline-Mind2WebREALOdysseys
Avg.Perfect
General VLMs
Qwen3.5-9B (Qwen Team, 2026a)46.9∗27.3∗18.2∗42.6∗13.5∗
Qwen3.5-4B (Qwen Team, 2026a)42.910.7
Qwen3.6-27B (Qwen Team, 2026b)84.3∗55.3∗27.3∗39.5∗18.5∗
OpenAI Operator (OpenAI, 2025)87.061.3
GPT-5 (SoM) (Awadallah et al., 2026)90.6
GPT-5.4 (Singh et al., 2025)55.433.5
Seed-2.0-Pro (Seed, 2026)85.1∗68.5∗74.4∗60.2∗30.1∗
GLM-5V-Turbo (Hong et al., 2026)88.5
Claude-Opus-4.6 (Anthropic, 2026a)88.068.944.5
Claude-Sonnet-4.6 (Anthropic, 2026b)49.831.0
Kimi-K2.6 (Moonshot AI, 2026a)76.8∗74.4∗
GUI-specific Models
UI-TARS-1.5 (Seed, 2025b)84.875.8
UI-Venus-1.5-30B-A3B (Team et al., 2026c)76.038.0∗
GUI-Owl-1.5-32B-Thinking (Xu et al., 2026)82.144.6∗
MolmoWeb-8B (Gupta et al., 2026)78.235.3
Fara1.5-4B (Awadallah et al., 2026)80.8
Fara1.5-9B (Awadallah et al., 2026)86.663.4
Fara1.5-27B (Awadallah et al., 2026)89.372.3
Ours
UI-Venus-2-9B90.874.076.977.362.0
UI-Venus-2-27B93.478.380.280.466.3
Figure 4: Examples of synthesized GUI Grounding training data. Our pipeline generates diverse and realistic interface screenshots spanning desktop (macOS, Windows), mobile (iOS), and web platforms, covering both professional software and consumer applications. All interfaces are rendered in a real headless Chromium browser via Playwright, ensuring high-fidelity visual output that closely mirrors authentic user environments.
Figure 4: Examples of synthesized GUI Grounding training data. Our pipeline generates diverse and realistic interface screenshots spanning desktop (macOS, Windows), mobile (iOS), and web platforms, covering both professional software and consumer applications. All interfaces are rendered in a real headless Chromium browser via Playwright, ensuring high-fidelity visual output that closely mirrors authentic user environments.
Table 4: Performance comparison on various Grounding Benchmarks. VenusBench-GD reports English-instruction micro-average point-in-box accuracy. “*” indicates baselines evaluated or reproduced by us.
ModelsGrounding Benchmarks
VenusBench-GDScreenSpot-ProOSworld-G-RUI-Vision
General VLMs
Qwen 3.7 Plus (Qwen Team, 2026c)75.2*68.978.268.0
Seed 2.1 Pro (ByteDance Seed, 2026)73.9*65.378.062.0
Kimi-K2.6 (Moonshot AI, 2026a)73.1*52.0*69.7*51.7*
Qwen3.6-27B (Qwen Team, 2026b)67.7*65.2*76.9*58.3*
GUI-specific Models
UI-Venus-Ground-72B (Gu et al., 2025)70.261.969.536.8
Holo2-30B-A3B (H-Company, 2025)59.5*66.176.140.9*
Step-GUI-4B (Yan et al., 2025)54.6*60.066.930.0*
MAI-UI-8B (Zhou et al., 2025b)65.2*65.868.640.7
MAI-UI-32B (Zhou et al., 2025b)-67.973.947.1
UI-Venus-1.5-30B-A3B (Team et al., 2026c)75.069.676.454.7
Qwen-UI-Agent-27B (Zhou et al., 2026)-76.678.570.0
Ours
UI-Venus-2-9B77.173.078.553.2
UI-Venus-2-27B80.174.179.166.9
Figure 5: Overview of VenusBench-CAPTCHA. Each panel shows one complete, uncropped representative screenshot. The OCR label gives the target transcription, numbered boxes indicate the required click order, and arrows visualize annotated drag trajectories. These annotations are added for presentation only and are not part of the model input.
Figure 5: Overview of VenusBench-CAPTCHA. Each panel shows one complete, uncropped representative screenshot. The OCR label gives the target transcription, numbered boxes indicate the required click order, and arrows visualize annotated drag trajectories. These annotations are added for presentation only and are not part of the model input.
Table 5: Performance Comparison across CAPTCHA Benchmarks. All results are Pass@1 percentages, and higher is better. We evaluate on VenusBench-CAPTCHA and four public benchmarks: MCA-Bench Wu et al. (2026b), Spatial-CAPTCHA-Bench Kharlamova et al. (2026), NextGen-CAPTCHAs Liu et al. (2026b), and Open CaptchaWorld Luo et al. (2025b). We use 1,000 sampled MCA-Bench examples, 15 NextGen-CAPTCHAs task types, and 16 Open CaptchaWorld task types; see the appendix for selection details.
ModelsVenusBench- CAPTCHAMCA-BenchSpatial- CAPTCHA-BenchNextGen- CAPTCHAsOpen CaptchaWorld
General VLMs
Qwen3.5-9B (Qwen Team, 2026a)28.330.44.92.836.4
Qwen3.6-27B (Qwen Team, 2026b)53.051.731.014.147.7
Seed-2.0-Pro (Seed, 2026)47.936.543.820.455.6
Kimi-K2.6 (Moonshot AI, 2026a)39.738.724.87.247.8
Claude-Opus-4.6 (Anthropic, 2026a)16.025.99.52.823.3
Ours
UI-Venus-2-9B78.175.742.847.650.7
UI-Venus-2-27B79.979.648.654.556.3
Table 6: Safety evaluation on OSHARM and OS-BLIND benchmarks. ASR denotes Attack Success Rate (%, lower is better).
ModelOSHarm (ASR ↓)OSBlind (ASR ↓)
General VLMs
Qwen3.5-9B (Qwen Team, 2026a)25.379.4
Qwen3.5-27B (Qwen Team, 2026a)18.089.3
Kimi K2.6 (Moonshot AI, 2026a)32.093.6
GUI-specific Models
EvoCUA-8B (Xue et al., 2026)39.385.3
EvoCUA-32B (Xue et al., 2026)33.390.7
UI-TARS-1.5 (Seed, 2025b)36.083.3
ScaleCUA (Lv et al., 2026)25.384.7
Ours
UI-Venus-2-9B11.348.8
UI-Venus-2-27B15.347.9
Table 7: All actions and their definitions used in UI-Venus-2. We unify the action space and map all the actions in the existing open-source dataset to this space.
ActionDefinition
Shared Actions (All Platforms)
Click(point=(x, y))Click at coordinates (x, y).
Drag(start=(x1, y1), end=(x2, y2))Drag from (x1, y1) to (x2, y2).
Swipe(start=(x1, y1), end=(x2, y2))Scroll by swiping from (x1, y1) to (x2, y2).
DoubleClick(point=(x, y))Perform a double-click/tap at coordinates (x, y).
LongPress(point=(x, y))Long press at coordinates (x, y) to trigger extra options.
Type(content=”)Type the specified content.
Wait()Wait for loading.
CallUser(content=”)Request user takeover or additional information.
Finished(content=”)Mark the task as completed, with optional information.
Mobile†† Coordinates are normalized to [0,999].
PressBack()Press the ‘back’ button.
PressHome()Press the ‘home’ button.
PressEnter()Press the ‘enter’ button.
PressRecent()Press the ‘recent’ button.
LaunchApp(app=”)Launch the specified app.
GetScreenshot()Take a screenshot and save it to the photo album.
Answer(content=”)Answer the user’s questions as requested.
Desktop†† Coordinates are normalized to [0,999].
RightClick(box=(x, y))Right-click at (x, y) to open context menus.
Hotkey(keys=[‘ctrl’, ‘c’])Press a keyboard shortcut, e.g., Ctrl+C for copy.
Web †† Coordinates are normalized to [0,999].
Scroll(point=(x, y), direction=‘up/down/left/right’)Scroll at (x, y) in the specified direction.
Launch(url=”)Launch the target URL.
GetUrl()Get the URL of the current browser tab.
TakeNote(content=”)Record important information from the screenshot.
Hover(point=(x, y))Move the mouse cursor to coordinates (x, y) without clicking.
Hotkey(keys=(‘ctrl’, ‘c’))Press combination keys (up to 3).
SelectOption(index=3)Choose an option from a native HTML select element.
PressBack()Return to the previous page.

실제로 확인된 결과

  • Figure 1과 Table 1~5는 UI-Venus-2-27B와 9B를 MobileWorld, WebVoyager, Odysseys, VenusBench-CAPTCHA, VenusBench-GD 등 여러 벤치마크에서 선별된 강력한 비교 모델들과 나란히 평가한 결과를 보여준다.
  • VenusBench-CAPTCHA는 219개 예시, 8가지 상호작용 유형(OCR 텍스트 입력, 순서 클릭, 회전, 드래그, 슬라이더 퍼즐 등)에 대해 Pass@1 지표로 평가됐다.
  • Table 6은 OSHARM, OS-BLIND 벤치마크에서 공격 성공률(ASR, 낮을수록 안전) 기준으로 안전성을 평가했다.
  • MOPD 단계에서 행동 유형별 조건부 지식증류 신호를 적용한 구조가 여러 GUI 도메인에서 더 견고한 통합 결과를 낸다고 보고됐다.

어디에 쓸 수 있나

  • 로그인·양식 작성 등 반복적인 모바일·웹 업무를 화면 조작 방식으로 자동화하는 실험
  • CAPTCHA 처리가 필요한 워크플로우 자동화 파이프라인 구축 시 참고 사례
  • 데스크톱 OS 작업(파일 관리, 오피스 소프트웨어 조작 등)까지 아우르는 범용 에이전트 설계 참고
  • 강화학습 보상 신호 검증 방법(궤적 단위·샘플 단위 다중 모델 투표) 설계 참고

한계와 남은 검증

  • 논문이 밝힌 비교는 각 소스가 보고한 행동 스캐폴드나 평가 프로토콜이 서로 다를 수 있어 완전히 동일 조건 비교는 아니다.
  • OSWorld-Verified 등 일부 벤치마크는 원 논문의 361개 과제 설정을 그대로 쓰지 않고 다른 과제 수(108개)로 평가돼 직접 비교에 제약이 있다.
  • 안전 메커니즘은 OSHARM, OS-BLIND 두 벤치마크에서만 확인됐고, 더 넓은 실제 위험 상황에서의 검증은 언급되지 않았다.
  • 다국어 모바일 앱 커버리지는 중국어 100여 개, 영어 70여 개 앱에 한정되어 있어 다른 언어권 앱에서의 성능은 확인되지 않았다.
  • 웹 벤치마크는 라이브 사이트 상태가 평가 시점에 따라 달라질 수 있다고 논문 스스로 명시했다.

왜 중요한가

GUI 에이전트가 특정 벤치마크에서만 잘 작동하고 실제 앱·웹·데스크톱에서는 쉽게 무너지는 문제는 자동화 도구를 실무에 쓰려는 개발자들에게 큰 걸림돌이었다. 이 연구는 환경 범위, 작업 생성, 검증 신뢰도를 동시에 늘리는 접근으로 그 격차를 줄이려 한 시도라는 점에서 참고할 가치가 있다.

이 논문의 용어

  • GUI 에이전트 · 화면을 보고 클릭·타이핑 같은 사람과 유사한 동작으로 프로그램을 조작하는 AI
  • 오프라인 강화학습 · 이미 수집된 상호작용 기록만으로 보상을 학습하는 강화학습 방식
  • 온폴리시 지식증류(MOPD) · 학생 모델이 스스로 만든 행동을 여러 교사 모델이 채점해 학생을 학습시키는 방식
  • 궤적 · 한 작업을 수행하는 동안의 화면-행동 연속 기록
  • VLM as a Judge · 비전-언어 모델이 다른 모델의 결과를 판정관처럼 평가하는 방법

저자 · Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jian

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Venus Team et al., arXiv:2609.00028, arxiv-nonexclusive