Hardware

DeepGrove's Maple-Preview Runs a 20B Ternary MoE at 120 Tokens Per Second on an iPhone

The natively trained ternary-weight model solves IMO-level problems and claims a 5–16x speed advantage over efficient rivals like Gemma 4, Qwen3.5, and GPT-OSS.

J
By James Park Enterprise Tech Reporter
August 5, 2026 / 7 min read

On-device AI took a step forward on August 4 with the release of Maple-Preview, a 20-billion-parameter ternary-weight reasoning model from startup DeepGrove that the company says runs at more than 120 tokens per second on an iPhone.

The Specs

Maple-Preview is a 20B-A1B mixture-of-experts model with 1.49 billion active parameters and a 131,072-token context window. The checkpoint weighs 5.31 GB. DeepGrove says it solves IMO 2024 Problem 1 with a perfect 7/7 score at 281.5 tokens per second on a MacBook Pro with an M5 Pro chip, and that on an iPhone it completes a "make me a carrot cake" prompt in about 10 seconds versus more than six minutes for a 1-bit Bonsai 27B baseline.

Benchmarks

On DeepGrove's internal harness, Maple-Preview scores 75.1 on LCBv6, 87.5 on AIME 2026, 78.8 on HMMT 2026, and 73.5 on GPQA-Diamond, for a 78.7 average. The company acknowledges the preview is focused on raw reasoning and may underperform on agentic tasks, with broader training planned before the full Maple release.

The Architecture Bet

DeepGrove argues that ultra-low precision should be a first-class citizen, not an afterthought. At ternary bitwidths, matrix multiplication can be replaced with additions, lowering arithmetic workload. The model was trained natively at low precision rather than quantized down from a full-precision teacher, which DeepGrove says avoids the performance ceiling of post-training compression.

On-Device Learning

Beyond inference, DeepGrove demoed on-device adaptation: the model can "dream" overnight about a user preference — such as vegan dietary restrictions — by generating synthetic data and updating its weights. The company contrasts this with context-based memory, arguing that weight adaptation retains subtle preferences that text-based memory misses.

Why It Matters

If the numbers hold, Maple-Preview suggests that capable reasoning models no longer require cloud GPUs or subscription APIs. The next generation of AI assistants may run locally, learn locally, and remain private by default.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.