A French AI developer has released a pair of compact models capable of navigating graphical user interfaces, a task that has proven difficult for most large language models despite their prowess at answering text prompts. H unveiled its Holo 4 family of models on Monday, designed to point, click, scroll, and type through desktop environments in Windows and Linux that were originally built for human users. The models aim to solve what the company calls "bots' blindness" when dealing with applications that prioritize visual design over functional simplicity.
The Holo 4 models are built on top of Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B foundations, then refined using supervised training and reinforcement learning. By optimizing for command line interfaces, application programming interfaces, and graphical interfaces simultaneously, H claims its models deliver far greater versatility than pure computer use models could provide. The company also updated its Holotron model, based on Nvidia's Nemotron 3, with similar capabilities. In demonstrations, Holo 4 27B used FreeCAD's macro function to programmatically design a 3D model of the Eiffel Tower rather than manually constructing it with basic shapes like cubes, and in another test created the company's logo using extruded shapes.
According to H, Holo 4 outperforms significantly larger frontier models from companies like OpenAI while using a fraction of the parameters, though the company cautions readers to view its benchmarks skeptically. However, despite higher performance scores, the models often cost more per task in many cases. The report attributes this to Holo 4 using substantially more "thinking" tokens to reach a final result compared to alternatives like GPT 6 Luna, which scores lower on the OSWorld 2.0 benchmark but costs substantially less. The open weights models' small size means researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware, with a 24 GB Nvidia RTX 3090 more than capable of running these models at 4-bit precision.
Throughout computing history, computer use has largely fallen into three categories: command line interfaces, application programming interfaces, and graphical user interfaces. AI agents can easily plug into the first two, but navigating desktop environments remains an ongoing challenge because these applications often prioritize form before function. The smaller parameter count of Holo 4 addresses this by making the technology accessible to users without massive computational resources, which explains why H released the model in multiple formats including BF16, FP8, and NVFP4 weights, plus a Llama.cpp-friendly GGUF version available on Hugging Face for use with LM Studio and Ollama. The company also plans to release DSpark draft weights to speed up inference using speculative decoding, a technique where a small model guesses the outputs of a larger model to accelerate token processing and generation, falling back to the base model when predictions fail to ensure no loss in output quality.
H has developed several agentic harnesses including its open source HAI-Agents harness, available for download on GitHub, though the models should theoretically work with third-party computer use harnesses as well. The report notes that H isn't alone in this space: AWS announced its own computer-use models at last year's Re:Invent conference, while OpenAI, Google, and Anthropic are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. The combination of open weights, modest hardware requirements, and compatibility with popular inference frameworks positions Holo 4 as an accessible option for organizations exploring AI agents that can interact with legacy desktop software. The tension between performance and cost efficiency will likely define which use cases justify deploying these models, particularly as enterprises weigh the trade-off between superior benchmark results and operational expense per task.

