The most legible result is a game left running on its own
Hcompany’s Holo4 demo asks an agent to build a Pac-Man-style game in Godot, then leave it playing indefinitely. The requested game has a grid maze, pellets, a player, three chasing ghosts, a score and lives. The agent is told to move toward pellets and avoid ghosts without keyboard input. The published clip shows the finished game running inside the editor—not just a code answer. Hcompany says Holo4 27B completed the task in 68 calls and 2.4 million tokens.
That is a useful shift in what a coding demo can show: not only whether code was generated, but whether the agent can produce and start a small interactive project. It is a company demo, not an independent replication, and it does not establish that the model can build a production-ready commercial game from one prompt. Source video and task: https://huggingface.co/blog/Hcompany/holo4
The comparison makes the result more concrete
Hcompany ran the same Pac-Man prompt with Qwen3.8 27B as its base-model comparison. The company reports 197 calls and 11.4 million tokens for Qwen3.8, versus 68 calls and 2.4 million tokens for Holo4 27B. In this particular harness and task, the tuned agent used about one-third as many calls and tokens. That is a workflow result, not a general benchmark proving Holo4 is three times better across coding or computer use.
The source video makes the output inspectable: it shows the project being built and then the game playing. Hcompany also publishes its benchmark trajectories and screenshots, so readers can inspect what the agent did step by step: https://huggingface.co/datasets/Hcompany/trajectories
It also attempted precise 3D work in FreeCAD
The other demonstrations ask Holo4 to build an Eiffel Tower model and a simple extruded Hcompany logo in FreeCAD. The tower prompt is unusually specific: dimensions at multiple heights, four separate legs, three platforms, a mast, and a requirement that the model stay open between its legs. Hcompany reports 84 calls and 1.3 million tokens for the Holo4 tower run.
For the logo task, Hcompany reports 94 calls for Holo4 and 118 for Qwen3.8. The demos are useful because the desired geometry is explicit and the output is visible in a real desktop application. They also show the cost of open-ended work: even these constrained examples involve dozens of tool calls and long agent runs, not a single instant response.
A generalist agent needs a workflow, not just a model
Holo4 comes in 27B dense and 35B-A3B mixture-of-experts versions. Hcompany says the agent can use a graphical interface, code, MCP and APIs, and releases weights in several formats. On OSWorld 2.0, Hcompany reports 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B, compared with 81.8% for Opus 5.5. Those results come with vendor-specific harness and cost assumptions, which Hcompany documents; benchmark comparisons across different harnesses should be read cautiously.
The genuinely interesting part of these demos is the whole loop: translate a detailed request into actions, inspect the application, write or adjust code, and leave behind a result that can be opened and tested. The Pac-Man task makes that payoff immediately understandable, while the source videos and published trajectories let builders look past the headline.
Source published September 28, 2026. Coverage is based on the maker’s announcement and demonstration.
