Check out Bonsai (27B) and Gemma 4 using the Locally AI by LM Studio App on your iPhone 😎
Both run in Locally AI on an iPhone 17 Pro, but they’re built for different jobs, so the “winner” depends on what you’re doing.
TL:DR – Bonsai is the stronger pick for reasoning and code; Gemma 4 is the safer one for images and function calling.
Bonsai (27B) is PrismML’s ternary-quantized build of Qwen3.6. The 27B weights occupy about 3.9 GB and generate roughly 11 tokens per second on an iPhone 17 Pro Max. Its strength is punching way above its size class. It clearly beats Gemma on math and coding benchmarks thanks to that 27B base.
The tradeoff: math and coding retain most of their baseline, while instruction following, vision, and multi-step tool use drop off more. On an iPhone it’s the heaviest option, so expect more RAM pressure and slower generation. (Locally AI also offers lighter 1-bit and Ternary Bonsai 8B variants if 27B feels sluggish.)
Gemma 4 in Locally AI comes as the E2B and E4B variants, Google’s small, on-device-native models. Gemma 4 is built specifically for on-device inference rather than adapted from a desktop-sized checkpoint, which gives faster cold-start loads and fewer out-of-memory crashes. Gemma holds up better on vision and tool calling, and its less aggressively compressed weights make it a steadier general-purpose choice on a phone.
So: pick Bonsai 27B if you want the smartest offline reasoning/coding help and don’t mind it being heavier and text-focused. Pick Gemma 4 if you want something lighter, faster to load, multimodal (image understanding), and more reliable for everyday chat. If you have the storage, keeping both and switching per task is the practical move. Locally AI makes swapping easy.
One note: well, I haven’t personally had any problems,
Both run in Locally AI on an iPhone 17 Pro, but they’re built for different jobs, so the “winner” depends on what you’re doing.
The quick version: Bonsai is the stronger pick for reasoning and code; Gemma 4 is the safer one for images and function calling.
Bonsai (27B) is PrismML’s ternary-quantized build of Qwen3.6. The 27B weights occupy about 3.9 GB and generate roughly 11 tokens per second on an iPhone 17 Pro Max. Its strength is punching way above its size class — it clearly beats Gemma on math and coding benchmarks thanks to that 27B base. The tradeoff: math and coding retain most of their baseline, while instruction following, vision, and multi-step tool use drop off more. On an iPhone it’s the heaviest option, so expect more RAM pressure and slower generation. (Locally AI also offers lighter 1-bit and Ternary Bonsai 8B variants if 27B feels sluggish.)
Gemma 4 in Locally AI comes as the E2B and E4B variants — Google’s small, on-device-native models. Gemma 4 is built specifically for on-device inference rather than adapted from a desktop-sized checkpoint, which gives faster cold-start loads and fewer out-of-memory crashes. Gemma holds up better on vision and tool calling, and its less aggressively compressed weights make it a steadier general-purpose choice on a phone.
So: pick Bonsai 27B if you want the smartest offline reasoning/coding help and don’t mind it being heavier and text-focused. Pick Gemma 4 if you want something lighter, faster to load, multimodal (image understanding), and more reliable for everyday chat. If you have the storage, keeping both and switching per task is the practical move — Locally AI makes swapping easy.
One note: While I have not had any problems, on an iPhone 17 Pro specifically, Bonsai 27B will be near the top of what the device’s RAM comfortably handles, so if you hit crashes or slowdowns, the Bonsai 8B or a Gemma 4 variant will feel much smoother.
