Paris-based H Company released Holo4 on September 28, 2026: two agentic models — a 27-billion-parameter dense model and a 35B-A3B mixture-of-experts model — trained to operate software through whatever interface is available. Screen clicks, handwritten code, MCP tool calls, direct API calls: the same model picks the approach that fits the task, with no need to swap models for desktops, the web, Android, a code sandbox, or business APIs.

Both are live on the H Models API, and the weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF formats. The dense model is research-only under CC BY-NC 4.0; the MoE goes out under Apache 2.0 — the commercially friendly one.

One model for every interface#

The pitch refuses a trade-off the field has quietly accepted: GUI-tuned agents go blind without a screen, while tool-calling agents stall in front of software that exposes no API. H Company's framing: a generalist agent needs every lever, and one model everywhere beats a fleet of specialists.

The 27B model is described in its model card as a vision-language model built on the Qwen3.8 dense architecture with a 262,144-token context window. Publishing weights in four precisions — including 4-bit GGUF — means anyone can run a credible computer-use agent locally or audit each of its steps, instead of renting time on a closed API.

Illustration of a desktop screen with a glowing cursor linked by threads of light to an API network, representing one model using GUIs and code together
GUI and code as one pipeline, not two separate agents. Credit: AI-generated illustration for AI Frontier Post.

The numbers, with a grain of salt#

The headline figure: Holo4 27B scores 85.2% on OSWorld, the standard desktop-automation benchmark, at a reported $0.08 per task. The MoE posts 80.8% at $0.05. The company's table puts the Qwen3.8 27B base at 84.3% and $0.22 per task, with frontier references of 86.0% for Fable 5 and 86.1% for Qwen3.8 Max: near-frontier desktop scores at single-digit cents per task.

The harder test is OSWorld 2.0, which measures long computer workflows. The 27B reports 61.7% there against 81.8% for Claude Opus 5.5 and 73.5% for GPT-6 Astra — a real gap, but at $1.22 per task versus $8.48 and $9.07. H Company's argument is the cost curve: it trails the best closed models while running on far fewer parameters for far less money.

A caveat the company discloses itself: on AutomationBench, 480 of the 600 public tasks fall within the split it used to collect training data. On the 120 held-out tasks it reports 49.3%. Comparison scores also come from mixed harnesses — not apples-to-apples. The reassuring move is a transparency one: the company open-sourced every trajectory behind its public benchmark scores, replayable step-by-step in a viewer and downloadable as a Hugging Face dataset.

Illustration of a small luminous AI chip brain above a laptop and smartphone surrounded by glowing checkmarks, representing benchmark wins
Near-frontier scores at cents per task — the company's headline claim. Credit: AI-generated illustration for AI Frontier Post.

How it was built#

Two phases: supervised fine-tuning on 127 billion tokens — roughly three quarters of them successful agentic trajectories across desktop, web, MCP/API and mobile environments — followed by asynchronous online reinforcement learning on long-horizon tasks, which produced two specialized LoRA experts merged back into the model.

More interesting is what fed the fine-tuning: the company's Agentic Task Factory, a pipeline that generates interactive environments and verifiable tasks from documentation alone — screenshots of real websites or open-source software. The company says the factory has produced roughly 10,000 tasks, and applying the same stack to NVIDIA's Nemotron 3 Nano Omni lifted Holotron4 Nano's OSWorld score from 21.0 to 76.3.

Why it matters#

Computer-use agents have lived behind closed APIs partly because driving a real desktop is high-stakes. An open-weight generalist changes who can build: startups escape per-task API economics, researchers can study failure modes closed providers keep opaque, and enterprises can run it inside their own perimeter where the screenshots it reads are sensitive data.

But at eight cents a task the bottleneck stops being the model bill and starts being reliability. An agent that finishes a workflow 85% of the time is not something you hand the company credit card to. The metric to watch is whether those percentages survive other people's software and the long tail of ugly real-world interfaces.

What to watch#

  • Independent replication. The company published its benchmark trajectories; watch for third parties reproducing the OSWorld numbers on their own harnesses — the fastest way to separate a strong release from strong marketing.
  • The license split. The more capable dense model is research-only; the Apache 2.0 MoE is the one commercial users can deploy. Whether the 27B gets a friendlier license decides how widely the best version spreads.
  • Real-world reliability. Eight cents a task only matters if the agent finishes. Watch for production reports on mixed GUI-plus-API workflows

Sources