SearcharxivSearch

arXiv subjects

Dimitrios Kafetzis

Publications and source records attributed to Dimitrios Kafetzis.

5 recordsLinked to original sources

SynapticOS: An Inference-First Runtime Architecture for Neural Processing Units on Resource-Constrained Microcontrollers

Microcontrollers with on-die neural processing units (NPUs) have become mainstream, but the system software hosting them has not: production combinations of Zephyr or FreeRTOS with TensorFlow Lite Micro treat AI inference as an application-layer library, leaving memory fragmentation, accelerator-state hygiene, and model-lifecycle guards as recurring application-developer concerns. We present the Phase 1 foundation of SynapticOS, an open-source runtime built on Zephyr that treats inference as a first-class workload. It contributes four cooperating subsystems: (1) a tensor-aware bump allocator with 16-byte DMA-aligned persistent and ephemeral lifetimes sharing a single arena, achieving constant-time allocation (~154 cycles per call, ~78,000 allocations per second at 150 MHz, invariant across tensor sizes) with zero fragmentation by construction; (2) a four-state hardware abstraction layer for the NPU and DSP, implemented by a deterministic software stub (for CI under QEMU) and a Neutron-flavoured backend (for the NXP MCXN947); (3) a three-state model lifecycle registry with duplicate-name detection, idempotent load/unload, and hot-swap guards; and (4) a four-mark cycle-accurate profiler. We evaluate on the NXP FRDM-MCXN947 (dual Cortex-M33 at 150 MHz) and the qemu_cortex_m3 emulator. Build footprints are 67 KB flash / 184 KB SRAM on FRDM (shell, 128 KB arena) and 24 KB flash / 28 KB SRAM on QEMU (no shell, 8 KB arena). End-to-end inference brackets through the deterministic stub kernel measure 1,038 us on FRDM and 781 us on QEMU for a 16x16x3 INT8 input; these are baseline overhead numbers, not Neutron silicon measurements, which arrive with the real SDK invoke path in Phase 2. A 61-test suite across 10 ZTEST suites passes 100% in 6.6 s on the CI emulator path. SynapticOS is released under Apache 2.0 at https://github.com/Dimitrios-Kafetzis/SynapticOS

eess.SY

Inference Pipelines as Operating-System Objects: Priority Scheduling and Constant-Footprint Streaming for Microcontroller Neural Inference

Microcontroller runtimes treat the inference pipeline -- pre-processing, accelerator invocation, post-processing -- as application code: every project re-implements stage sequencing, buffer sizing, and completion signalling around a library call. We argue these are operating-system concerns and present the Phase 2 inference engine of SynapticOS, an open-source Zephyr-based runtime that makes the pipeline a first-class OS object. A pipeline is drawn from a static pool, validated against a canonical stage order, and executed by a priority job scheduler (realtime > normal > best-effort, FIFO per class) with cancellation and a bounded job table; no heap on the inference path. Stage buffers are sized exactly from configuration and tensor geometry for the nine built-in processors (bounded 4x fallback for user stages); all intermediates live in an ephemeral arena reset per frame, so streaming footprint is constant. We evaluate on the NXP FRDM-MCXN947 (Cortex-M33, 150 MHz) and the qemu_cortex_m3 CI target, both running a deterministic stub NPU kernel: engine-overhead baselines, not silicon throughput. On the board the scheduler adds 92 us over the Phase 1 direct-HAL bracket (1,130 vs 1,038 us; dispatch 1 us); a 30-frame, six-stage face-detection pipeline averages 4.63 ms/frame (215.8 FPS, stub model included) vs 31.1 ms under QEMU soft-float, at a constant 2,784-byte arena peak returning to zero each frame. The PowerQuad DSP is routed and self-calibrated for FFT and Q15 matmul; end-to-end speedups are 5.51x (256-point FFT) and 1.66x (16x16 matmul), short of the plan's 10x target -- reported as missed, not re-scoped. Stage-boundary profiling now runs live on the board, closing a Phase 1 gap. The engine adds 3.8 KB flash on QEMU and 20.7 KB on FRDM. 99 tests across 13 ZTEST suites pass 100% under emulation. Released under Apache 2.0 at https://github.com/Dimitrios-Kafetzis/SynapticOS

cs.OS

Cross-Core Inference Offload as an Operating-System Service on Dual-Core Microcontrollers

Dual-core MCUs are asymmetric: on NXP's MCXN947, the second Cortex-M33 has no FPU, DSP extension, TrustZone, or MPU. We treat the asymmetry as a design input in the Phase 3 dual-core architecture of SynapticOS, an open-source Zephyr-based runtime: the AI runtime (models, NPU/DSP, scheduler) lives on the capable core, and the application core reaches inference only via a message-based OS service -- a remote system call. The transport is a pair of lock-free single-producer/single-consumer rings in shared SRAM: one writer per index, free-running 32-bit counters, ordering by data-memory barriers alone (the platform has no cross-core atomics). Because ring state is shared, a rebooting application core rejoins unaided. Requests carry priority classes, errors and timeouts propagate to the caller, and tensors stage zero-copy in a shared slot -- a 27 KB frame cannot exist twice in 64 KB of RAM. Measured on the FRDM-MCXN947 (both cores 150 MHz): the application core boots in 1,514 us and completes the handshake in 2,554 us, bit-identical over 11 boots; round trips are 15 us typical / 81 us worst-case (50 us budget); pushes cost 25 cycles; a 1,913-serve two-model soak had zero errors (stub-NPU latencies bracket transport, not silicon). An MPU region on the runtime core guards the application core's RAM (fault-injection verified); protection is one-directional -- the application core has no MPU, and ARMv8-M cannot block privileged reads. Two hardware-revealed defects are reported: releasing the second core into erased flash wedges the whole chip and its debug port (now prevented by a ROM-API blank check), and a Zephyr flash-driver Kconfig silently disarmed the devicetree MPU guard (now programmed at runtime). Firmware is 98.9 KB flash (runtime core) and 32.4 KB (application core, 42.6 of 64 KB RAM); 108 tests in 13 suites pass 100%. Apache 2.0: https://github.com/Dimitrios-Kafetzis/SynapticOS

cs.OS

Video Quality Monitoring for Remote Autonomous Vehicle Control

The delivery of high-quality, low-latency video streams is critical for remote autonomous vehicle control, where operators must intervene in real time. However, reliable video delivery over Fourth/Fifth-Generation (4G/5G) mobile networks is challenging due to signal variability, mobility-induced handovers, and transient congestion. In this paper, we present a comprehensive blueprint for an integrated video quality monitoring system, tailored to remote autonomous vehicle operation. Our proposed system includes subsystems for data collection onboard the vehicle, video capture and compression, data transmission to edge servers, real-time streaming data management, Artificial Intelligence (AI) model deployment and inference execution, and proactive decision-making based on predicted video quality. The AI models are trained on a hybrid dataset that combines field-trial measurements with synthetic stress segments and covers Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and encoder-only Transformer architectures. As a proof of concept, we benchmark 20 variants from these model classes together with feed-forward Deep Neural Network (DNN) and linear-regression baselines, reporting accuracy and inference latency. Finally, we study the trade-offs between onboard and edge-based inference. We further discuss the use of explainable AI techniques to enhance transparency and accountability during critical remote-control interventions. Our proactive approach to network adaptation and Quality of Experience (QoE) monitoring aims to enhance remote vehicle operation over next-generation wireless networks.

cs.NI

Large Language Model Partitioning for Low-Latency Inference at the Edge

Large Language Models (LLMs) based on autoregressive, decoder-only Transformers generate text one token at a time, where a token represents a discrete unit of text. As each newly produced token is appended to the partial output sequence, the length grows and so does the memory and compute load, due to the expanding key-value caches, which store intermediate representations of all previously generated tokens in the multi-head attention (MHA) layer. As this iterative process steadily increases memory and compute demands, layer-based partitioning in resource-constrained edge environments often results in memory overload or high inference latency. To address this and reduce inference latency, we propose a resource-aware Transformer architecture partitioning algorithm, where the partitioning decision is updated at regular intervals during token generation. The approach is myopic in that it is based on instantaneous information about device resource availability and network link bandwidths. When first executed, the algorithm places blocks on devices, and in later executions, it migrates these blocks among devices so that the sum of migration delay and inference delay remains low. Our approach partitions the decoder at the attention head level, co-locating each attention head with its key-value cache and allowing dynamic migrations whenever resources become tight. By allocating different attention heads to different devices, we exploit parallel execution of attention heads and thus achieve substantial reductions in inference delays. Our experiments show that in small-scale settings (3-5 devices), the proposed method achieves within 15 to 20 percent of an exact optimal solver's latency, while in larger-scale tests it achieves notable improvements in inference speed and memory usage compared to state-of-the-art layer-based partitioning approaches.

cs.DC