Skip to main content

TECH VEDA

Linux kernel & Device drivers starts on 24th Oct 2026 enrollingCorporate on-site training - Submit proposal Pick your modulesSharpen your kernel skills: deep dives, drivers, Yocto, CVEs, careers — updated daily. Read the blog →Embedded Linux fast track starts 30th sept 2026 enrollingEmbedded Linux Mastery track starts 30th sept 2026 enrollingLinux systems engineering starts 30th sept 2026 enrolling
Edge AI

6 TOPS on Paper. Which Inference Runtime Reaches It?

TensorRT, LiteRT, ONNX Runtime and OpenVINO on real embedded silicon: which ones reach the NPU on an i.MX, an AM62A, an RK3588 — and how to check yours.

6 TOPS on Paper. Which Inference Runtime Reaches It?

Put a quantised model on an RK3588 and none of TensorRT, LiteRT, ONNX Runtime or OpenVINO will touch the NPU in the die you paid for. On an i.MX95 one accelerates properly and another is marked experimental. On an AM62A two work. These four are not really competing for your project, because on most embedded silicon the inference runtime is settled by the part number — and what decides your product is which path your vendor’s BSP supports, at which version.

Every inference runtime comparison scores these four on model formats, quantisation, language bindings and operator coverage. Those tables are accurate and nearly useless for an embedded product, because they assume all four are available to you. On a real board they are not. What follows compares them on the only axis that changes an outcome: whether each inference runtime reaches the neural accelerator on the silicon you actually bought.

One naming note. TensorFlow Lite was renamed LiteRT in September 2024. The .tflite format and existing models are unchanged and the Interpreter API still works, so migrating Python is largely a package swap from tflite-runtime to ai-edge-litert. What changed is where acceleration lives: in LiteRT 2.x, GPU and NPU work runs through the newer CompiledModel API. The current release is LiteRT 2.2.0.

Who owns each inference runtime

RuntimeOwnerAccelerator reach
TensorRTNVIDIANVIDIA GPUs only, which on embedded means Jetson
OpenVINOIntelCPU on x86-64, ARM and RISC-V; GPU and NPU acceleration on Intel silicon only, in the official distribution
LiteRTGoogleCPU everywhere; first-party GPU and, on some phone-class silicon, first-party NPU; otherwise a vendor delegate
ONNX RuntimeMicrosoftCPU everywhere; accelerator only through a vendor execution provider

The OpenVINO row is the one most often stated wrongly, including in my own first draft. OpenVINO is not x86-only: its CPU plugin supports ARMv7, AArch64 and RISC-V, so it runs on an ARM board today. What is Intel-only is the acceleration. Take this inference runtime to a non-Intel SoC and you get CPU inference, with a feature set the documentation notes may differ from x86-64 — fair for a low-rate workload, a non-option if you bought the part for its NPU.

TensorRT has no equivalent nuance — NVIDIA’s runtime, NVIDIA’s silicon. JetPack 7.1 added TensorRT Edge-LLM, a C++ SDK that takes a model exported to ONNX, optimises it into a TensorRT engine and drives it from a C++ runtime with no Python interpreter in the inference path. Note the scope: that is Jetson Thor only. Orin modules stay on JetPack 6.

Three ways an inference runtime reaches an NPU

Across the platforms this article anchors on, every vendor solves the same problem one of three ways.

Pattern one: the vendor owns the inference runtime. NVIDIA and Intel build the accelerator and the software that drives it, and their own inference runtime is the reference path. You are not strictly locked in — both ship execution providers so ONNX Runtime can drive their hardware — but features land in the vendor runtime first and the conversion tooling is theirs either way.

Pattern two: the vendor plugs a backend into a standard inference runtime. NXP, TI and MediaTek. You keep the LiteRT or ONNX Runtime API, and the vendor supplies a delegate or execution provider that claims the parts of the model it can run and leaves the rest on the CPU.

Pattern three: the vendor replaces the inference runtime. Rockchip. You convert the model to a private format and link against a private runtime.

The pattern is not a matter of vendor taste. It follows who sells the margin. NVIDIA and Intel sell the accelerator itself, so the compiler is the moat, and handing that away through someone else’s API removes a reason to buy the part. NXP, TI and MediaTek sell an SoC where the NPU is one block among twenty; funding an inference runtime ecosystem against Google and Microsoft is not a business they are in, so they write a delegate and inherit one for free. Rockchip competes hard on unit price and does neither — a private format with a private runtime is the cheapest accelerator software a vendor can ship, and the saving comes out of your portability.

These are not exclusive boxes. MediaTek does both at once: an ONNX Runtime execution provider and a private compiled-DLA path through its own converter. And on parts built around licensed NPU IP, such as the i.MX93 with its Arm Ethos-U, the delegate and compiler come from the IP vendor rather than the SoC vendor. Read the three patterns as questions to ask, not as boxes every SoC sits in.

Pattern two in practice: NXP, TI and MediaTek

On the i.MX 8M Plus, the documented path to the NPU is LiteRT with NXP’s VX Delegate. The eIQ stack covers four inference engines — TensorFlow Lite, ONNX Runtime, PyTorch and OpenCV — and there is a VSI NPU execution provider for ONNX Runtime. Community reports call that provider’s quantised coverage patchy; NXP’s user guide lists it without qualification. Establish which on your own BSP, and treat LiteRT as the path with the production history.

The i.MX95 is where I had to correct myself. Its eIQ Neutron NPU has two documented paths. The mature one is TFLite with the Neutron Delegate. Since the LF6.18.20_2.0.0 BSP, NXP also documents a Neutron execution provider for ONNX Runtime — i.MX95 only, labelled experimental — covering quantised CNNs and 8-bit and 4-bit matmuls. So an ONNX-standardised team is not locked out; it is choosing an experimental path over a validated one. ONNX runs, ONNX does not yet reliably accelerate.

TI integrates TIDL with both ONNX Runtime and LiteRT, offloading supported subgraphs to the C7 accelerator across TDA4VM, AM62A, AM67A, AM68A and AM69A. This is the one case among the anchors where the team genuinely picks the inference runtime rather than inheriting it. Two caveats keep it honest. TIDL uses a TI fork of ONNX Runtime with its own compilation and execution providers, not upstream ONNX Runtime with a drop-in plugin. And the subgraph boundary, not the accelerator, is where throughput estimates go wrong — it moves whenever the model changes.

MediaTek’s Genio line exposes the NPU through a Neuron execution provider for ONNX Runtime, alongside CPU and XNNPACK. Before you budget on it: that provider is default only on Genio 720 and 520, needs an explicit flag on 700 and 510, and is absent on Genio 350. “Supported across the line” and “accelerated on your part” are different statements.

Pattern three in practice: the RK3588 has no seat at this table

The RK3588 carries a 6 TOPS NPU by Rockchip’s specification — three cores of roughly 2 TOPS each, so a single-core job gets about a third of the headline — and it is one of the most common parts in the Indian embedded market. None of the four runtimes in this article’s title will drive that NPU.

The supported flow is a conversion: PyTorch to ONNX, ONNX to RKNN with RKNN-Toolkit2, then execution through RKNN Runtime. ONNX appears in that chain, which is why people say the board “supports ONNX”, but it appears as an interchange format on a build machine, not as an inference runtime on the device. ONNX Runtime’s RKNPU execution provider is community-maintained, marked preview, and targets the older rknpu1 parts rather than the RK3588.

Rockchip recommends conversion on an x86_64 Linux PC and that is the tested path, though RKNN-Toolkit2 ships aarch64 wheels too, so an all-ARM build farm is possible if you are willing to be off the beaten track. The consequence that belongs in a schedule is the other one: because the accelerator is reached through a vendor format rather than a standard inference runtime API, moving this product to different silicon later means redoing the deployment path, not swapping a delegate.

What actually happens to your ONNX model

Most teams here export from PyTorch to ONNX and assume the hard part is done. Here is what that export meets on each pattern, and what tends to break first.

  • Pattern one, Jetson. ONNX is the expected input; TensorRT parses it and builds an engine. The usual failure is an unsupported operator or an unresolvable dynamic shape, and you find out at build time — the kindest of the three.
  • Pattern two, NXP and TI. Your graph is partitioned: the vendor compiler claims what it supports and the rest runs on CPU. The failure is silent — the model runs, the numbers look plausible, and half the graph never reaches the accelerator. This is the case worth engineering against.
  • Pattern three, Rockchip. ONNX is consumed on the host and discarded. The failure surfaces in the converter, on a machine that is not your target, and the error text refers to a graph you did not write.

Quantisation is a separate question from conversion, and a larger one. Every path here expects INT8 for real throughput, each vendor tool wants its own calibration data, and accuracy drift is measured by you or by nobody. That is the work that follows this decision rather than precedes it.

All three vendors ship a compiler that reports the partition on an x86 host, before you commit to a board. Ask for that output on your own model — a better procurement question than any TOPS figure.

The constraint nobody benchmarks

The interesting quantity is not any single version number, it is the gap between upstream and your vendor, and how fast it closes. As this is written, ONNX Runtime is at 1.30.0. MediaTek’s documented Genio build is 1.20.2. NXP’s current BSP carries 1.24.3. TI does not ship upstream ONNX Runtime at all — it ships a fork.

Those numbers will change. What will not is that a gap exists, because it is structural: an execution provider is built and validated against a pinned inference runtime release and a pinned driver, so the vendor cannot track upstream and you cannot get ahead of the vendor. You are not adopting ONNX Runtime. You are adopting your vendor’s build of ONNX Runtime, at their version, with their execution provider, on their release cadence.

Embedded teams know this shape from the kernel. The upstream project moves; the vendor tree is where you live. Treat the inference runtime the same way: find the version your BSP carries, look at what the previous three BSP releases carried rather than at the roadmap, and design as though you are on it for the life of the product. Measure the gap today and again in six months. The second number tells you what you bought.

What would break this argument

The strongest case against everything above is being built right now. Google declared LiteRT’s GPU and NPU acceleration production-ready in January 2026, and LiteRT 2.x ships first-party NPU backends for Qualcomm and MediaTek, with Intel NPU support in 2.2.0. The stated goal is that developers stop touching vendor SDKs — a direct attack on the fragmentation this article describes.

Two things temper it today. Those first-party accelerators are phone-SoC shaped: none reaches an i.MX, an AM62A or an RK3588. And on the RK3588, mainline is moving under the vendor stack through the kernel’s accel subsystem, which over time changes what a private runtime protects. If either lands on these boards, inference runtime choice becomes a real choice again and this piece needs rewriting. That is the honest boundary: silicon decides today, for merchant embedded SoCs, and that is what would change it.

Check your own board in about a minute

These commands tell you which inference runtime is installed and whether it can see an accelerator. Run them on the target. Start with the runtimes:

raghu@techveda.org:~$ python3 -c "import onnxruntime as o; print(o.__version__, o.get_available_providers())"
raghu@techveda.org:~$ python3 -c "import ai_edge_litert.interpreter as t; print('litert', t.__file__)"
raghu@techveda.org:~$ python3 -c "import tflite_runtime.interpreter as t; print('tflite-runtime', t.__file__)"

Run both LiteRT imports — vendor BSPs still ship the older tflite_runtime package, and which one is present tells you how current the image is. The provider list from the first command is the important output: if it shows only CPUExecutionProvider, nothing you do in application code will reach the NPU on that image.

Then find the vendor libraries and the device, remembering that Debian-based images use multiarch paths while Yocto images do not:

raghu@techveda.org:~$ find /usr/lib /usr/local/lib -maxdepth 2 \( -name 'lib*rknn*' -o -name '*vx_delegate*' -o -name 'lib*tidl*' -o -name '*neutron*' \) 2>/dev/null
raghu@techveda.org:~$ ls -l /dev/dri/ /dev/galcore /dev/accel/ 2>/dev/null
raghu@techveda.org:~$ ls /sys/class/remoteproc/ 2>/dev/null
raghu@techveda.org:~$ sudo dmesg | grep -iE "npu|galcore|rknpu|tidl|remoteproc"

Read those as evidence, not as pass or fail. The RK3588 NPU registers a DRM render node, so look in /dev/dri/ and for an Initialized rknpu line, not for /dev/rknpu. TI’s C7x is reached over remoteproc and appears in /sys/class/remoteproc/ with nothing in /dev. And dmesg needs privilege on most distributions and its ring buffer wraps, so a quiet result is not proof of absence.

A device node with no inference runtime that can open it, and a runtime with no device node, are two different problems with two different owners. Telling them apart takes a minute and saves a week.

If the answer is bad and the silicon is already chosen

Most readers of this article did not pick their SoC and cannot change it. If the provider list came back CPU-only, there are three moves and no others. Find the vendor’s delegate or execution provider package in their BSP feed and work out why your image did not pull it in. Rebuild the inference runtime against the vendor backend at the version the vendor validated, not the version you would prefer. Or accept CPU inference and re-scope the frame rate honestly. There is no fourth option in which upstream fixes this for you.

Choosing and bringing up this part of an embedded stack is something we work through in our embedded Linux and Yocto training.

If there is one thing to take from this, it is the command in the previous section. Run it on whatever board is on your desk. If it prints ['CPUExecutionProvider'], you have just learned that your NPU is decoration on this image, and you learned it in under a minute. Most teams find that out in integration week.

Was this worth your time?

Frequently asked questions

Is TensorFlow Lite the same as LiteRT?
Yes. TensorFlow Lite was renamed LiteRT in September 2024. The .tflite format and existing models are unchanged and the Interpreter API still works, so migrating Python code is largely a package swap from tflite-runtime to ai-edge-litert. The real change is that in LiteRT 2.x, GPU and NPU acceleration run through the newer CompiledModel API rather than the Interpreter.

Which inference runtime should I use on an RK3588?
None of the four, as far as the NPU is concerned. The supported path converts the model to Rockchip’s RKNN format with RKNN-Toolkit2 and runs it through RKNN Runtime. ONNX is an interchange format during conversion, not an inference runtime on the device.

Can I use ONNX Runtime on the i.MX95 NPU?
Partly. NXP documents a Neutron execution provider for ONNX Runtime on the i.MX95, labelled experimental and covering quantised CNNs and 8-bit and 4-bit quantised matmuls. The validated path for full NPU offload is still TFLite through the Neutron Delegate, so ONNX runs but does not yet reliably accelerate.

Does a supported-frameworks list mean the accelerator will be used?
No, and this is the most common misreading. An inference runtime can be supported in the sense that it executes on the CPU while the NPU sits idle. Look for the path the vendor documents and validates for accelerator offload, then confirm on hardware which execution providers or delegates are present.

Further reading

RB
Raghu Bharadwaj

Founder, TECH VEDA — 20+ years teaching the Linux kernel, device drivers and embedded systems.

Follow on LinkedIn

Get new posts by email

Kernel, embedded Linux and AI-era engineering — a few sharp reads a month. No spam.

We email occasionally and never share your address.