1/ 今天我们推出 Cua-S1-4B-0.2,这是首个使用 RLOO 在实时计算机使用任务上训练的多模态决策模型,采用任务完成奖励。
文本和多模态适配器采用 Apache-2.0 许可提供:http://t.cn/A6rwf9oN
[视频]
2/ At each step, the CUA-S1 model receives the screen state, the task goal, and a fixed set of candidate actions.
It returns one action. The environment changes, and the next step starts from the new state.
3/ The training recipe has two stages.
Supervised training teaches the decision format. Agentic RL then runs the model in live cua-bench-basic environments, where only a completed task earns the environment's reward.
4/ On the same held-out tasks, Cua-S1-4B-0.2 completes 17/18 text episodes and 13/18 multimodal. Zero-shot djev completes 16/18 and 12/18.
5/ On a frozen GUI-360 split of 168 multimodal tasks, Cua-S1-4B-0.2 reaches 92.9% versus 60.1% for untrained djev.
Both use the same inputs and scoring, with no accessibility tree.
6/ The training code and benchmark results are now merged into Cua.
Explore the code: http://t.cn/AXOR6XS2…
Get the Apache-2.0 text and multimodal adapters: http://t.cn/AXOkMBuu…
http://t.cn/AXOkMBu3
