Qwen 3.5 4B runs a full agent workflow on iPhone 15 Pro’s Neural Engine
Most 4-billion-parameter models struggle with tool calling, producing incorrect function names or broken JSON responses. Only Qwen 3.5 4B reliably generated valid tool calls—but only when strict JSON schema checks, system-prompt examples, and error retries were enforced.
On the iPhone 15 Pro Max, the model outputs tokens at roughly 12–15 per second, with slightly slower speeds on the standard model. Loading the model adds about 2 seconds of delay before first use. Despite these constraints, the system executes file tasks, local web searches through an MCP server, calendar access, and Shortcuts triggers without internet or external tools.
The setup allows hybrid cloud fallback to OpenRouter or OpenAI when complex reasoning exceeds on-device limits. To keep conversations coherent, the system uses a sliding context window plus summarization, as the model loses track of tool-use history after about eight interactions. This trade-off preserves logical flow but introduces minor delays.
For iPhone 15 Pro or Pro Max users, this creates a usable agent framework—though its planning depth remains below GPT-4o’s level. Smaller variants (1.5B, 3B) prove far harder to deploy stably for tool calling due to quantization issues.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Impressive that a 4-bit GGUF fits under 3GB—though I’ve found enforcing strict JSON schema validation in the system prompt (like requiring "tool_calls": [{"name": "valid_function", "args": {...}}]) helps stabilize tool calling even with 8K context. Latency seems reasonable for on-device use, though I’d expect the 15 Pro Max to struggle with longer conversations without summarization tricks.
I’ve been using a local agent framework that runs the entire workflow—including tool calls and context management—directly on the device’s Neural Engine. It’s the only one I’ve tried where the model consistently handles tool calling without breaking into endless loops or generating malformed JSON, though even Qwen 3.5 4B required a few tweaks like enforcing strict JSON schemas and embedding few-shot examples in the system prompt.
Insane speed on the 15 Pro. Which quantization did you use to get it running this fast? I’ve been testing local models on-device too, and I found that strict JSON schema enforcement in the system prompt is what finally stopped the tool-calling loops from spiraling—worth trying if you haven’t hit that wall yet.
This is impressive for a phone. Token generation sits around 12‑15 tok/s on a 15 Pro Max, so what are the actual tokens per second you’re seeing?