A small draft model helps a larger vision model

Liquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model. The smaller model proposes likely next tokens, and the larger target model checks them. When many proposals are accepted, the target can produce more output per verification step.

The drafter has about 279.5 million parameters, increasing the target model’s total parameter count by roughly 8.9%. Liquid AI says this adds a small memory cost while improving decoding throughput without changing the target model’s output quality.

The speed gains depend on the device

On an M5 Max MacBook Pro running MLX-VLM, Liquid AI reports decoding improvements from 2.30× to 3.13× across six vision tasks. End-to-end latency improvements in those tests ranged from 1.56× to 2.62×. On an M3 Ultra using llama.cpp, the reported decoding gains were lower.

For one H100 80GB GPU setup using SGLang, Liquid AI reports up to 2.66× faster decoding and up to 2.27× end-to-end improvement. These are maker-run benchmark results under specified settings, not a universal guarantee for every device or request.

Decoding speed is not the whole response time

Speculative decoding accelerates token generation after the input is processed. A vision-language model also needs to encode the image and process the prompt, so the overall experience can improve less than the raw decoding rate. Liquid AI publishes both measurements, which helps explain the difference.

The results covered six vision tasks, including general visual questions, chart understanding and image captioning. The company used FP16 or BF16 weights and did not test quantized models in this release, an important limit for edge deployments where compressed models are common.

The draft model is available in open tooling

Liquid AI says the vision DSpark drafter is available through Hugging Face and has support in llama.cpp, MLX-VLM and SGLang. That gives developers several routes to try the technique across Apple Silicon and GPU serving setups.

The useful question is whether the speed gain survives a developer’s actual prompts, image mix and runtime configuration. The experiment offers a compact way to test speculative decoding on a vision-language model, while the published ranges show that acceptance rate and hardware shape the payoff.

Explore the original source ↗

Source published 2026-09-24. Coverage is based on the maker’s announcement and demonstration.