Preprint · 2026

HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Tongbo Chen1*, Junbo Niu2*, Zhengxi Lu1, Niu Lian3, Fei Tang1, Yuchen Yan1, Yike Hong1, Yong Du1, Yizhou Liu1, Bofan Chen1, Yongliang Shen1†

1Zhejiang University  ·  2Peking University  ·  3Tsinghua University *Equal contribution · †Corresponding author

Three panels contrasting GUI-only actions (17 steps, wrong result), GUI–tool interleaved actions (app-specific, costly to build), and GUI–CLI interleaved actions, which click once and then use two bash commands to compute a profit column correctly.
Figure 1. GUI-only agents spend long click-and-type sequences on work a command does in one step; GUI–tool agents are fast but need a per-application integration. The GUI–CLI hybrid keeps the generality of the screen and borrows the efficiency of code.

Abstract

Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.

Access to a shell is not the bottleneck

Handing a capable MLLM a shell does not help — it hurts. All four representative agents lose accuracy on OSWorld once the CLI is exposed, by 2.5 to 11.5 points, and their CLI usage scatters from 0.15% of steps to 67%: some grind through GUI sequences a single command would replace, others reach for the shell too eagerly. HybridCUA-9B uses the CLI about as much as the most CLI-heavy baseline, yet gains 14.8 points where that baseline loses 6.6. Two gaps explain the deficit — GUI, shell, and code corpora are siloed, so the data never shows hybrid behaviour; and step imitation plus outcome rewards are blind to interface choice, unable to tell a well-chosen interface from a successful but inefficient path.

Paired bar chart: Qwen3.5-27B drops 51.4 to 48.9, EvoCUA-32B 56.7 to 53.0, Kimi-K3 84.8 to 73.3, Claude-Opus5 83.4 to 76.8 once the CLI is exposed, while HybridCUA-9B rises from 38.8 to 53.6. A second panel shows CLI step shares of 15.0%, 0.15%, 59.0%, 67.0% and 64.0%.
Figure 2. (a) Adding a CLI interface to existing agents reduces OSWorld accuracy, whereas the trained HybridCUA-9B improves over its GUI-only counterpart. (b) CLI step share varies substantially across existing agents; HybridCUA learns to use the CLI selectively rather than merely maximising command usage.
View the underlying numbers as a table

Method

Hybrid supervision need not be collected from scratch: many GUI action sequences admit an equivalent shell command. HybridCUA exploits this to build HybridCUA-8K — 5,023 trajectories across 11 application domains in three interaction modes, plus 3,000 RLVR tasks that each ship an executable verifier — then makes interface choice explicit in the reward.

Pipeline diagram with three panels: (a) scalable data generation — GUI demonstration conversion, skill construction and rollout in a live environment producing CLI-only, GUI–CLI and GUI-to-CLI trajectories, alongside RLVR task synthesis with interface guides, verifiers and validation labelling; (b) supervised fine-tuning on the SFT corpus; (c) online agentic RL with CLI-aware rewards.
Figure 3. (a) Generation of GUI-only, CLI-only and interleaved trajectories, plus annotated RL tasks. (b) Supervised fine-tuning on 𝒟SFT. (c) Online agentic RL with CLI-aware reward signals.
Three panels: a donut of domain by modality composition, a bar chart of trajectory modality (GUI only 870 / 17.3%, CLI only 3,155 / 62.8%, hybrid 998 / 19.9%), and a histogram of logical steps per trajectory peaking at 5–6 steps with median 7, mean 10.4, p90 22.
Figure 4. Composition of the supervised corpus. (a) Application-domain totals, with the outer-ring shade indicating GUI-only, CLI-only or hybrid interaction within each domain. (b) Overall trajectory counts by modality. (c) Distribution of logical steps per trajectory, stacked by modality.
GUI only · 870

Open-source UI-MOPD trajectories, re-expressed action by action as equivalent pyautogui calls inside the unified wrapper.

CLI only · 3,155

Per-application CLI skills distilled from docs and reference code; a Qwen3.8-27B agent solves CUA-Gym tasks with the shell alone, keeping successful rollouts.

Interleaved · 998

Either the agent chooses between both interfaces live, or terminal-typing segments of GUI trajectories are folded back into single commands and replayed.

One action space, two interfaces

Every executable interaction goes through the same bash action: a direct shell command and a quoted Python heredoc of pyautogui calls are both just commands, so switching interfaces costs the model no extra grammar. The wrapper is a representation, not a reclassification — PyAutoGUI through bash is still a GUI action. Keeping the two in distinct tools instead costs 7.2 accuracy points under matched training.

At each step the agent sees the current screenshot and, after a direct command, its stdout/stderr. CLI-only trajectories therefore restrict the action interface, not the observation — they still contain screenshots.

GUI action
bash(command="""python3 <<'PY'
import pyautogui
pyautogui.click(500, 400)
pyautogui.hotkey('alt', 'f')
PY""")
CLI action
bash(command="""python3 <<'PY'
import docx
d = docx.Document(p)
run.font.size = Pt(12)
d.save(p)
PY""")

Plus the control actions wait, terminate and answer. Coordinates are predicted in a 1000×1000 space and mapped to the display at execution time.

Two CLI-aware rewards

Supervised warm-up on the mixed corpus teaches the unified format and interface switching. Online GRPO then fixes what a plain accuracy reward cannot: successful rollouts score the same despite unnecessary switches, and a failed command goes unpenalised if the task eventually completes.

Task level · when

RCLI = 𝟙[Success(τ)] · 𝟙[b(τ) = b★]

Every RLVR task carries a label b★ for whether CLI use has a clear execution advantage, obtained by sampling 16 rollouts per interface mode and ranking them. A successful rollout is rewarded only when its CLI usage matches that label — so the model learns selective, not maximal, shell use.

λCLI = 0.1, added before group normalisation.

Step level · how

rtexec = −1 if the command fails, else 0

Terminal feedback localises the blame. A shell-level failure penalises exactly the tokens of the action that caused it, so command reliability improves without muddying the trajectory-level signal. Non-CLI steps get zero.

λexec = 0.3, on the responsible action's tokens only.

Training and evaluation configuration

Results

HybridCUA-9B reaches 53.6% in 14.0 average steps — best on both metrics among comparably sized models. Merely exposing the CLI cuts steps by 9.5 but accuracy from 38.8% to 18.4%, so the shorter trajectories reflect early failure, not efficiency. Trained, the second interface pays off at both stages: 46.0% vs 44.2% after SFT, 53.6% vs 50.4% after RL, with the two branches sharing the same base model, comparable SFT corpora, and identical RLVR tasks and RL steps.

Out-of-distribution generalisation

On OSWorld-MCP, HybridCUA-9B scores 47.1%, 9.1 points over Qwen3.5-9B and on par with API-integrated ToolCUA-8B. On WindowsAgentArena it reaches 36.0%, +4.0 over its base model — issuing PowerShell commands despite training only on Linux shells, which suggests what transfers is when to delegate to a shell rather than memorised commands.

Ablations

SFT data composition

Mixing all three trajectory types beats any single type.

Action schema

A single bash action against two separate tools.

Four line panels over RL steps 0 to 120 comparing full HybridCUA-9B, a variant without the task-level CLI reward, and a variant without the step-level execution reward: accuracy rises to about 53.6, average steps fall to 14, CLI step ratio rises to 64 percent, and CLI execution error falls to 11.5 percent for the full model while the no-execution-reward variant climbs above 16 percent.
Figure 5. Removing RCLI costs little accuracy but stalls the efficiency gain (CLI usage 58.9% instead of 64.0%; trajectories shorten 18.2% rather than 29.3%). Removing rtexec leaves accuracy and steps nearly intact, yet execution errors climb back to 16.5% — above the SFT level — while the full model settles at 11.5%. The task-level reward decides when; the step-level reward governs how reliably.

What the model routes where

Interface choice tracks the interaction the task actually needs. The CLI dominates OS tasks (84% of steps), where commands reach files and system settings directly; the GUI dominates Chrome (74%), where the work lives on a rendered page. By operation, content editing (85.8% CLI) and result verification (98.4% CLI) suit direct manipulation and state inspection, while spatial adjustment (58.4% GUI) needs the visual feedback loop — and information gathering splits nearly evenly, which is the point: neither interface wins categorically, so routing has to be learned per situation.

Two panels: per-domain accuracy from OS 60.9 to 79.2 (+18.3) through Multi-apps 19.9 to 36.8 (+16.9), with Thunderbird the single regression at 66.7 to 57.1 (-9.6); and the matching GUI/CLI step share per domain, from 84% CLI on OS to 26% CLI on Chrome.
Figure 6. Domain-level accuracy and GUI/CLI step shares on OSWorld.
View the underlying numbers as a table
Four pie charts of GUI versus CLI step share: information gathering 36.1% CLI, content editing 85.8% CLI, spatial adjustment 41.6% CLI, result verification 98.4% CLI.
Figure 7. GUI/CLI step shares by operation category for HybridCUA-9B.

Cooperation, not just coexistence

In an essay-formatting task the interfaces cooperate at two levels. Across the trajectory, GUI steps focus the Writer window and open its menu (step 7) while CLI steps use python-docx to apply 12-point text and single, double and one-and-a-half line spacing before saving (step 10). Within a single step, step 13 dismisses the menu with pyautogui and then reopens the saved file through the shell to check font size and spacing — both in one bash call.

Case study: a user asks for an essay with single-spaced introduction, double-spaced body and one-and-a-half-spaced conclusion at 12 point. Step 7 uses pyautogui to focus the document and open the menu, step 10 uses python-docx to set sizes and spacing and save, step 13 combines a pyautogui escape keypress with a python-docx read-back that prints ALL CHECKS PASSED.
Figure 8. GUI–CLI cooperation in essay formatting. Code is excerpted and screenshots are cropped.

Trajectory types in the corpus

One representative example of each of the three construction routes behind HybridCUA-8K.

A GUI-only Impress trajectory that changes a slide background from the default colour to red, with each source action shown beside its equivalent pyautogui call.
Figure 9. GUI only. A representative trajectory paired with its direct PyAutoGUI representation: the slide background changes from the default colour to red, and each source operation — clicking, dragging, scrolling, shortcuts, text entry — becomes the equivalent PyAutoGUI call.
A CLI-only Impress trajectory using Python and UNO to identify slides 1 and 5 by their speaker-photo placeholders, recolour their backgrounds to pale yellow, save, and cross-check the live UNO state against the saved XML.
Figure 10. CLI only. Using Python/UNO, the agent identifies slides 1 and 5 from their speaker-photo placeholders, changes their backgrounds from white to pale yellow (#FFFFCC), saves the presentation, and cross-checks the live UNO state against the saved XML while confirming that the other four slides remain unchanged.
An interleaved trajectory: the CLI installs the Night Owl VS Code extension, GUI actions open the theme picker and preview and confirm the theme, then the CLI reads back the persisted workbench.colorTheme setting.
Figure 11. Interleaved GUI–CLI. The CLI installs the Night Owl extension, making it available to VS Code; GUI actions then open the theme picker, select and preview the exact theme, and confirm the choice; the trajectory returns to the CLI to read back the persisted workbench.colorTheme setting. Direct commands handle setup and state inspection, while the GUI is retained for the visually grounded selection step.

Case Studies

Three complete, successful OSWorld rollouts from the appendix, each pairing the pre-action screenshot with the action the model issued at that step. Use the step buttons to walk through a trajectory; every one of these scored 1.0 with the task evaluator.

BibTeX

@misc{chen2026hybridcualearningorchestrategui,
      title={HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents}, 
      author={Tongbo Chen and Junbo Niu and Zhengxi Lu and Niu Lian and Fei Tang and Yuchen Yan and Yike Hong and Yong Du and Yizhou Liu and Bofan Chen and Yongliang Shen},
      year={2026},
      eprint={2609.38008},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.38008}, 
}