Edge AI6 TOPS on Paper. Which Inference Runtime Reaches It?
TensorRT, LiteRT, ONNX Runtime and OpenVINO on real embedded silicon: which ones reach the NPU on an i.MX, an AM62A, an RK3588…
On-device and edge AI for embedded Linux engineers: accelerators and NPUs, model runtimes and quantization, federated learning, and the systems engineering behind deploying AI at the edge.
Edge AITensorRT, LiteRT, ONNX Runtime and OpenVINO on real embedded silicon: which ones reach the NPU on an i.MX, an AM62A, an RK3588…
Edge AIConcurrent inference on an edge board collapses once processes exceed half the CPU cores. The kernel-side reason, and how to measure it…
Edge AIFile-backed weights let an edge device run a model from the page cache instead of a second private copy. What it saves,…
Edge AIChoosing an Edge AI SoC on TOPS alone goes wrong. What the NPU, memory, power and software stack really decide, with figures…
Edge AITo the kernel, an edge model is two mappings: a reclaimable file-backed one and a tensor arena that is anonymous and cannot…
Edge AIThe mainline NPU driver for RK3588 is real, in drivers/accel since Linux 6.18. What it does, what it does not, and where…
Edge AIOperator fallback pushes unsupported ops back to the CPU and fragments the graph. What it costs on real devices, and how to…
Edge AISpeculative decoding speeds up on-device LLM inference by amortizing the memory-bandwidth cost of decoding. Here is how it works and where it…
Edge AISelf-learning edge AI comes in three technical forms. How each one works, and where each stands today.
Edge AIEdge AI in automotive and industrial systems runs beside a deterministic, safety-critical control loop. How real-time Linux, TSN and mixed criticality help.
Edge AITinyML vs Edge AI on Linux is not a TOPS comparison. It is a choice of machine. A decision guide, the measured…
Edge AIWhy 4-bit weight quantization is a memory-system problem: DRAM bandwidth, unified memory on Jetson, kernel fusion, and the real accuracy and speed…
Edge AILLM inference from flash: how to run a model larger than your DRAM by streaming weights from storage, explained for embedded engineers.
Edge AIWhere edge AI heads by 2035: AI-native 6G networks, neuromorphic silicon, the EU AI Act, and a ten-year roadmap for embedded Linux…
Edge AIThe embedded Linux stack for edge AI — Yocto, runtimes, secure OTA — and federated learning with Flower to improve models across…
One email every two weeks — the best of our kernel & embedded writing, plus a hand-picked link or two. No spam.
// 1,200+ embedded engineers already read it
We email occasionally and never share your address.