Background
Not every AI application belongs in the cloud. When data cannot leave the premises, connectivity is unreliable or long-term cost matters, the model has to run locally. On a Radxa ROCK 5C (RK3588S, 8GB RAM) we deployed LLM, VLM and speech models on the built-in NPU and wrapped them in an OpenAI-compatible API — existing code switches from cloud to local by changing the base URL.
Endpoints
/v1/chat/completions: multi-turn chat, streaming and image understanding/v1/gui/ground: takes a screenshot and an instruction and returns click coordinates, for agents that operate a computer/v1/audio/transcriptions: speech-to-text (Qwen3-ASR / SenseVoice)
Benchmarks
Speech-to-text (FLEURS Chinese, 10 clips, 115.5 s)
| Option | CER | RTF | Punctuation |
|---|---|---|---|
| Qwen3-ASR-1.7B (NPU) | 2.9% | 0.48 | ✅ |
| Whisper large-v3-turbo (NPU) | 3.4% | 1.93 | Partial |
| Whisper large-v3-turbo q5_0 (CPU, 4 cores) | 3.7% | 4.17 | Partial |
| SenseVoice (NPU) | 4.7% | 0.076 | ❌ |
| SenseVoice int8 (CPU, 4 cores) | 4.9% | 0.046 | ❌ |
| Whisper small (CPU, 8 cores) | 7.6% | 0.62 | Partial |
Takeaway: for accuracy, Qwen3-ASR on the NPU runs faster than real time and adds punctuation; for speed, SenseVoice int8 transcribes a minute of audio in under 3 seconds.
GUI grounding (32 items sampled from ScreenSpot-v2)
| Setup | Accuracy | Per step |
|---|---|---|
| Qwen3.5-2B, full screenshot only | 19% | 7.4 s |
| 2B + one zoom | 41% | 17.2 s |
| 2B + two zooms | 53% | 21.0 s |
A small model looking at the whole screenshot does poorly; a coarse-then-zoom search nearly triples accuracy.
Deployment Notes
- Stability: running all 8 cores hard-resets this board, so heavy CPU work is pinned to the 4 big cores
- Runtime versions: RKLLM 1.2.3 and 1.3.0 are ABI-incompatible, so each model runs in its own subprocess behind one API
- Swappable backends: the OpenAI interface stays fixed, so models can change without touching the applications above
Good Fit For
- Healthcare, finance and factory settings where data cannot leave the network
- Devices and products that need offline speech recognition or vision
- Teams moving inference to the edge to cut long-term cloud API spend