HybridInfer stops GPU crashes by routing LLM queries
The problem I hit on a Snapdragon flagship
I was running a 3 B Llama 3.2 model directly on the device GPU via an OpenCL kernel. After a handful of generation calls the GPU runtime would either throw an exception or silently hang. The crash looked like:
java.lang.RuntimeException: OpenCL kernel execution failed
at com.android.gpu.GPUExecutor.runKernel(...)
Even after cooling the phone to below 30 °C the failure reappeared as soon as I sent three long prompts in a row. Short prompts worked, but any prompt longer than about 150 tokens caused the next inference to stall. The issue traced back to the current toolchain: the OpenCL kernel compilation step plus the “long‑prompt prefill” operation on the mobile GPU were exhausting thermal headroom, not just slowing down the compute.
How I diagnosed the thermal bottleneck
- Thermal logging – I enabled
adb shell dumpsys thermalserviceto watch the device’s thermal zones while the model ran. The “GPU” zone spiked to the “throttling” threshold (~85 °C) after the second generation. - GPU profiling – Using
adb shell cat /sys/kernel/debug/tracing/traceI captured the kernel execution time. Each prefill step took ~120 ms, but the cumulative effect of successive calls pushed the GPU into a protected state. - Toolchain isolation – I swapped the OpenCL backend for a CPU‑only fallback. The CPU ran without crashing, confirming the GPU path was the culprit.
These steps proved that the failure was thermal, not a memory leak or a driver bug. The root cause is the combination of OpenCL kernel compilation and the long‑prompt prefill that keeps the GPU active long enough to trigger thermal throttling, even if the ambient temperature is low.
The routing idea that finally worked
I built a three‑tier hierarchy:
- On‑device – Llama 3.2 3B (GPU)
- Edge – Llama 3.1 8B with retrieval (running on a nearby edge server)
- Cloud – GPT‑4o (fully remote)
The router treats the phone’s thermal headroom and a quick query‑complexity estimate (token count × estimated compute) as the state. I trained a Q‑learning policy offline on a synthetic dataset of 10 k prompts, using the following reward:
- Quality – BLEU‑like score against a reference generation
- Latency – Inverse of round‑trip time (ms)
- Cost – Monetary price of edge/cloud calls
- Thermal penalty – Proportional to the predicted temperature rise on the GPU
- Locality bonus – +0.2 reward for staying on‑device
The locality bonus is crucial; without it the optimal policy offloads every request, which defeats the purpose of on‑device inference.
Sample Q‑learning snippet (Python)
import numpy as np
alpha = 0.1
gamma = 0.95
epsilon = 0.2
def update_q(state, action, reward, next_state, Q):
best_next = np.max(Q[next_state])
Q[state, action] = (1 - alpha) * Q[state, action] + \
alpha * (reward + gamma * best_next)
return Q
The state vector is [thermal_headroom, token_estimate], discretized into 10 bins each. Actions correspond to the three tiers.
Real‑world benchmark results
I deployed the trained policy on an Android device and ran 210 real prompts (mix of short, medium, and long). The learned router achieved:
- Higher quality than two hand‑tuned heuristics (paired Wilcoxon, p < 0.02).
- Lowest cost among all adaptive conditions, because the locality bonus kept many queries on‑device.
- Latency improvement – average end‑to‑end time dropped from 1.2 s (always‑on‑device) to 0.45 s when routing.
When I forced “always‑on‑device”, the quality matched the routed version for short queries, but latency increased three‑ to six‑fold and long queries still caused crashes. The router therefore wins on reliability, coverage, and speed, while maintaining comparable quality.
Takeaways for anyone building on‑device LLM pipelines
- Monitor thermal zones in real time;
dumpsys thermalserviceis cheap and gives you the exact thresholds you need to respect. - Avoid repeated long‑prompt prefill on the GPU. If you must keep the model warm, consider batching or inserting short idle periods to let the thermal controller recover.
- Add a locality reward if you want the router to prefer on‑device execution. Pure quality‑latency trade‑offs will push everything to the cloud.
- Q‑learning works with a tiny state space; you only need a few hundred synthetic prompts to get a useful policy.
By letting the phone’s own thermal headroom decide where to run each request, HybridInfer eliminates the GPU crashes that plagued my earlier experiments and gives a smooth, cost‑effective user experience.
Three long prompts in a row? The issue isn’t just heat—it’s likely memory fragmentation after each kernel launch. Snapdragon GPUs have tiny shared memory pools, and Llama 3.2’s attention layers probably leak buffers.