Papers
arxiv:2610.02320

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

Published on Oct 1
· Submitted by
Said Gurbuz
on Oct 6
Authors:
,
,
,

Abstract

Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: https://saidgurbuz.github.io/deskforge/

Community

Paper author Paper submitter

DeskForge is a controllable desktop environment that composes and explores real Linux applications to generate dense supervision for computer-use agents. Each scene varies the apps, their content, window layout, visual style and resolution. Every visible element is labeled with its type, text, occlusion-aware visible region, owning window and interaction properties, and every click is recorded with the screen before and after it.

DeskForge-1M: 1.2M annotated desktop screenshots, 159.7M element annotations and 917K recorded click transitions across 19 applications, 7 visual styles and 7 resolutions.

Results

  • Fine-tuning four VLMs (Qwen3.5-4B, Gemma4-E4B, InternVL3.5-8B, UI-R1-3B) on 200K DeskForge-1M grounding examples improves all of them on every held-out desktop condition and on all five external GUI grounding benchmarks. Qwen3.5-4B gains +11.5 points on ScreenSpot-Pro and +10.1 on OSWorld-G. On the held-out desktops, all four also beat every open model we tested (7B–32B).
  • With the same Qwen3.6-27B planner and only the action model swapped, Qwen3.5-4B solves 50 instead of 31 of 119 WebArena-Infinity tasks and 15 instead of 3 of 100 OpenApps tasks.
  • An RT-DETRv4-L detector trained on the dense labels outperforms the OmniParser v2 and ScreenParse detectors on out-of-domain GroundCUA screens.

🗂️ Dataset: docling-project/DeskForge-1M
🤖 Models: DeskForge collection

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.02320
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 2