Fine-tuning the controller and its guardrail

Humanoid robots often track a motion reference while a separate safety filter modifies actions that violate constraints. The CoFiT paper argues that training the tracker and filter independently can create a mismatch: the filter changes the actions the policy actually executes, which changes the states it experiences.

The authors propose constrained filter-aware tuning, or CoFiT, to adapt a pretrained tracker with the safety filter in the loop. The objective is to make the policy work with the corrections instead of relying on the filter to repair every action afterward.

The reported result on hardware

On Unitree G1 hardware, the researchers report that CoFiT reduced violation time by 83% for the TWIST2 task. Every CoFiT trial finished without operator intervention, while half of baseline trials required an operator stop. In simulation, the paper reports reductions of 91% on TWIST2 and 21% on SONIC versus filter-only training.

These are results from one research setup and specific constraint scenes, not evidence that the robot is safe for general deployment. The comparison is valuable because it exposes a concrete engineering issue: a safety layer and the policy beneath it need to be evaluated together.

Why the interface matters

A controller can look safe in isolation while behaving poorly after a runtime filter changes its output. Training with the filter active gives the policy experience closer to what it will encounter during execution.

For robot developers, the contribution is a practical way to reduce how often a separate safety mechanism must intervene. Larger tests across tasks, environments and robot platforms would show how broadly the result transfers.

Explore the original source ↗

Source published October 1, 2026. Coverage is based on the maker’s announcement and demonstration.