Holo4 27B and 35B-A3B agentic models released
TL;DR
Hugging Face released the Holo4 series – a 27B dense model and a 35B‑A3B Mixture‑of‑Experts model – that can operate across graphical user interfaces, code execution, MCP, and APIs, delivering competitive scores on academic benchmarks while costing far less than comparable frontier models.
Unified agentic capabilities across every interface
Holo4 is designed to click, type, write code, run that code, and call MCP or API tools as needed. Unlike most agentic models that specialize in a single interaction mode, Holo4 uses the most appropriate interface for each sub‑task, enabling it to handle real‑world business workflows that blend GUI actions, tool calls, and code execution. The same model can be deployed on desktops, web browsers, Android devices, code sandboxes, and against business APIs without any model switching.
Competitive benchmark performance at low cost
Holo4 builds on a Qwen base and shows substantial improvements. On the OSWorld 2.0 benchmark, the 27B version scores 61.7 % (vs. 81.8 % for the closed‑source Opus 5.5) and the 35B‑A3B version reaches 30.9 %. Despite lower absolute scores, Holo4 achieves these results with orders of magnitude fewer parameters and at a much lower cost per task. Cost‑performance charts for OSWorld 2.0 and AutomationBench demonstrate that Holo4’s token‑based pricing on the H Models API undercuts Qwen 3.8 27B, Qwen 3.6 35B‑A3B, and leading closed models.
Real‑world professional software tasks
Holo4 was trained on ~10,000 tasks generated by the internal Agentic Task Factory, covering web apps, MCP servers, and desktop environments. In head‑to‑head comparisons with its Qwen 3.8 27B base model, Holo4 consistently required fewer calls and fewer tokens to complete complex tasks such as:
- 3D modeling of the Eiffel Tower in FreeCAD – 84 calls / 1.3 M tokens (Holo4) vs. 60 calls / 1.0 M tokens (Qwen).
- Creating the H‑company logo in FreeCAD – 94 calls / 1.5 M tokens (Holo4) vs. 118 calls / 1.9 M tokens (Qwen).
- Building a Pac‑Man‑style game in Godot – 68 calls / 2.4 M tokens (Holo4) vs. 197 calls / 11.4 M tokens (Qwen).
These examples illustrate Holo4’s ability to orchestrate GUI interactions, code generation, and execution in a single, coherent workflow.
Training pipeline and harness improvements
The Agentic Task Factory generated the training data, converting documentation and screenshots into interactive environments and verifiable tasks. Holo4 underwent supervised fine‑tuning on 127 B tokens, followed by reinforcement learning from two expert RL policies, then merged into a single model. Parallel to model training, the team rebuilt the harness—the execution loop that manages context over hundreds of steps. Key harness upgrades included:
- A reliable memory system capable of tracking hundreds of steps.
- Direct shell access on the desktop machine.
- Continuous feedback from OSWorld 2.0 runs, where agents tagged failures and engineers reviewed fixes.
Holotron4 Nano: extending the recipe to Nemotron 3 Nano Omni
Applying the same post‑training stack to the NVIDIA Nemotron 3 Nano Omni model produced Holotron4 Nano, a 30B‑A3B agentic model. Benchmarks across five tasks show absolute percentage‑point gains over the base Nemotron model, confirming that the recipe generalizes across model sizes and architectures.
Availability and next steps
Both Holo4 variants are live on the H Models API. Model weights are hosted on Hugging Face in BF16, FP8, NVFP4, and 4‑bit GGUF formats, alongside the smaller Holotron4 Nano. The team will soon release optimized DSpark drafter checkpoints to further speed up inference.
Resources
- Model cards – Holo4‑27B | Holo4‑35B‑A3B | Holotron4‑Nano
- Full collection – FP16, FP8, GGUF formats: https://huggingface.co/collections/Hcompany/holo4
- Trajectory viewer & dataset – https://trajectories.hcompany.ai/ | https://huggingface.co/datasets/Hcompany/trajectories
- API quick‑start – https://hub.hcompany.ai/models-api/quickstart
- Full blog post – https://hcompany.ai/newsroom/holo4