H Company ships Holo4: open-weight agents that work any software interface — GUI, code, MCP and APIs
The most capable computer-use agents have a quiet limitation: they are either screen agents or tool agents, and real work refuses to pick a lane. H Company’s new model family, Holo4, is built for the tasks that start in a spreadsheet, hop into a browser, and finish through an API — and the company just made both models openly downloadable.

Paris-based H Company released Holo4 on September 28, 2026: two agentic models — a 27-billion-parameter dense model and a 35B-A3B mixture-of-experts model — trained to operate software through whatever interface is available. Screen clicks, handwritten code, MCP tool calls, direct API calls: the same model picks the approach that fits the task, with no need to swap models for desktops, the web, Android, a code sandbox, or business APIs.
Both are live on the H Models API, and the weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF formats. The dense model is research-only under CC BY-NC 4.0; the MoE goes out under Apache 2.0 — the commercially friendly one.
One model for every interface#
The pitch refuses a trade-off the field has quietly accepted: GUI-tuned agents go blind without a screen, while tool-calling agents stall in front of software that exposes no API. H Company's framing: a generalist agent needs every lever, and one model everywhere beats a fleet of specialists.
The 27B model is described in its model card as a vision-language model built on the Qwen3.8 dense architecture with a 262,144-token context window. Publishing weights in four precisions — including 4-bit GGUF — means anyone can run a credible computer-use agent locally or audit each of its steps, instead of renting time on a closed API.

The numbers, with a grain of salt#
The headline figure: Holo4 27B scores 85.2% on OSWorld, the standard desktop-automation benchmark, at a reported $0.08 per task. The MoE posts 80.8% at $0.05. The company's table puts the Qwen3.8 27B base at 84.3% and $0.22 per task, with frontier references of 86.0% for Fable 5 and 86.1% for Qwen3.8 Max: near-frontier desktop scores at single-digit cents per task.
The harder test is OSWorld 2.0, which measures long computer workflows. The 27B reports 61.7% there against 81.8% for Claude Opus 5.5 and 73.5% for GPT-6 Astra — a real gap, but at $1.22 per task versus $8.48 and $9.07. H Company's argument is the cost curve: it trails the best closed models while running on far fewer parameters for far less money.
A caveat the company discloses itself: on AutomationBench, 480 of the 600 public tasks fall within the split it used to collect training data. On the 120 held-out tasks it reports 49.3%. Comparison scores also come from mixed harnesses — not apples-to-apples. The reassuring move is a transparency one: the company open-sourced every trajectory behind its public benchmark scores, replayable step-by-step in a viewer and downloadable as a Hugging Face dataset.

How it was built#
Two phases: supervised fine-tuning on 127 billion tokens — roughly three quarters of them successful agentic trajectories across desktop, web, MCP/API and mobile environments — followed by asynchronous online reinforcement learning on long-horizon tasks, which produced two specialized LoRA experts merged back into the model.
More interesting is what fed the fine-tuning: the company's Agentic Task Factory, a pipeline that generates interactive environments and verifiable tasks from documentation alone — screenshots of real websites or open-source software. The company says the factory has produced roughly 10,000 tasks, and applying the same stack to NVIDIA's Nemotron 3 Nano Omni lifted Holotron4 Nano's OSWorld score from 21.0 to 76.3.
Why it matters#
Computer-use agents have lived behind closed APIs partly because driving a real desktop is high-stakes. An open-weight generalist changes who can build: startups escape per-task API economics, researchers can study failure modes closed providers keep opaque, and enterprises can run it inside their own perimeter where the screenshots it reads are sensitive data.
But at eight cents a task the bottleneck stops being the model bill and starts being reliability. An agent that finishes a workflow 85% of the time is not something you hand the company credit card to. The metric to watch is whether those percentages survive other people's software and the long tail of ugly real-world interfaces.
What to watch#
- Independent replication. The company published its benchmark trajectories; watch for third parties reproducing the OSWorld numbers on their own harnesses — the fastest way to separate a strong release from strong marketing.
- The license split. The more capable dense model is research-only; the Apache 2.0 MoE is the one commercial users can deploy. Whether the 27B gets a friendlier license decides how widely the best version spreads.
- Real-world reliability. Eight cents a task only matters if the agent finishes. Watch for production reports on mixed GUI-plus-API workflows
Sources
- H Company, "Holo4: powering generalist computer-use agents" (newsroom post and Hugging Face blog, September 28, 2026) — lineup, benchmarks, costs, training, licensing, Holotron4 Nano. All figures are company-reported from its own harness.
- Unite.AI, "H Company Releases Holo4, Open-Weight Models for Computer-Use Agents" — secondary coverage confirming lineup, licenses, benchmark details.
- NeoTeo, "Holo4 computer-use agents: H Company's two models" — third-party coverage confirming release details and the license split.