AI Automation·2026

Edge AI Deployment: OpenAI-Compatible API on RK3588 NPU

Chat, vision, GUI grounding and speech-to-text served as an OpenAI-compatible API from the NPU of an 8GB RK3588 board, with seven speech-recognition options benchmarked

clientInternal R&D
durationSeptember 2026
categoryAI Automation
stack
RK3588RKLLMRKNNQwen3.5-VLQwen3-ASRSenseVoiceWhisperFastAPIOpenAI API

Background

Not every AI application belongs in the cloud. When data cannot leave the premises, connectivity is unreliable or long-term cost matters, the model has to run locally. On a Radxa ROCK 5C (RK3588S, 8GB RAM) we deployed LLM, VLM and speech models on the built-in NPU and wrapped them in an OpenAI-compatible API — existing code switches from cloud to local by changing the base URL.

Endpoints

  • /v1/chat/completions: multi-turn chat, streaming and image understanding
  • /v1/gui/ground: takes a screenshot and an instruction and returns click coordinates, for agents that operate a computer
  • /v1/audio/transcriptions: speech-to-text (Qwen3-ASR / SenseVoice)

Benchmarks

Speech-to-text (FLEURS Chinese, 10 clips, 115.5 s)

Option CER RTF Punctuation
Qwen3-ASR-1.7B (NPU) 2.9% 0.48 ✅
Whisper large-v3-turbo (NPU) 3.4% 1.93 Partial
Whisper large-v3-turbo q5_0 (CPU, 4 cores) 3.7% 4.17 Partial
SenseVoice (NPU) 4.7% 0.076 ❌
SenseVoice int8 (CPU, 4 cores) 4.9% 0.046 ❌
Whisper small (CPU, 8 cores) 7.6% 0.62 Partial

Takeaway: for accuracy, Qwen3-ASR on the NPU runs faster than real time and adds punctuation; for speed, SenseVoice int8 transcribes a minute of audio in under 3 seconds.

GUI grounding (32 items sampled from ScreenSpot-v2)

Setup Accuracy Per step
Qwen3.5-2B, full screenshot only 19% 7.4 s
2B + one zoom 41% 17.2 s
2B + two zooms 53% 21.0 s

A small model looking at the whole screenshot does poorly; a coarse-then-zoom search nearly triples accuracy.

Deployment Notes

  • Stability: running all 8 cores hard-resets this board, so heavy CPU work is pinned to the 4 big cores
  • Runtime versions: RKLLM 1.2.3 and 1.3.0 are ABI-incompatible, so each model runs in its own subprocess behind one API
  • Swappable backends: the OpenAI interface stays fixed, so models can change without touching the applications above

Good Fit For

  • Healthcare, finance and factory settings where data cannot leave the network
  • Devices and products that need offline speech recognition or vision
  • Teams moving inference to the edge to cut long-term cloud API spend

More work in AI Automation.