SearcharxivSearch

arXiv subjects

Luca Benini

Publications and source records attributed to Luca Benini.

At least 19 recordsLinked to original sources

MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion

The 3D center of mass (CoM) is a primary quantity in the biomechanical analysis of sport, rehabilitation, and clinical movement, yet existing 3D pose tracking, mesh recovery, and multi-view triangulation methods either optimize 3D keypoint accuracy without anatomical constraints or carry compute and capture infrastructure too heavy to deploy where CoM tracking is most useful. As a result, the metric CoM remains difficult for coaches and movement analysts to measure from a single camera where athletes train and compete. In this work, we introduce MuyBridge, an on-device system that estimates the athlete's segmental center of mass trajectory from a single phone camera video stream. MuyBridge couples a compact 2D pose network and a distilled single-step monocular depth network through an analytic metric fusion that uses anatomical and physical priors to anchor the metric CoM, requiring no 3D or task-specific supervision. Evaluated on the athletic movements of AthletePose3D (running, track and field, and figure skating), MuyBridge achieves 33-41 mm vertical CoM error and 2.3-6.6% absolute-relative range error (AbsRel) under a one-time calibration, and produces CoM estimates at the 63 FPS pose-estimation rate using asynchronous 2.86 Hz depth updates on iPhone 15. Code is available at: https://github.com/Abradshaw1/Muybridge

cs.CV

Scalable AXI4 Transaction Monitoring for Mixed-Criticality SoCs: From Phase-Level Precision to ID-Level Efficiency

Mixed-criticality Systems-on-Chip (SoCs) with on-chip interconnects based on the AXI4 open standard protocol lack a protocol-level timeout mechanism, exposing systems to deadlocks and missed real-time deadlines when subordinate devices or managers fail or stall due to hardware faults, radiation-induced upsets, or software errors. This work presents a configurable hardware intellectual property (IP), non-intrusive in fault-free operation, that detects AXI4 protocol violations and timing faults at runtime and restores interconnect liveness through a cut-and- drain isolation mechanism. To address the fundamental trade-off between monitoring granularity and area cost, we introduce three designs at decreasing monitoring granularity: Phase-Level Track-ing (PLT), which provides cycle-accurate fault localization across individual protocol phases; Channel-Level Tracking (CLT), which coalesces per-phase monitors into channel-level supervision; and ID-Level Tracking (ILT), which achieves sub-linear area scaling by monitoring only per-ID transaction boundaries. Synthesized in GlobalFoundries 12 nm technology, CLT reduces area by 36.7% relative to PLT while preserving worst-case detection bounds at a minimal detection latency overhead, whereas ILT achieves an 89.2% area reduction suitable for tightly constrained deployments at the cost of a 3.7x higher median detection latency with coarser fault localization. Fault injection campaigns on a RISC-V SoC across 1.2 million scenarios confirm that no fault manifesting as an AXI4 protocol or liveness violation escaped detection, with observed detection latencies consistently bounded by theoretical worst-case predictions.

cs.AR

An Energy-Proportional Multimodal and Context-Aware Vision IoT Node

While recent advancements in TinyML have significantly reduced the computational complexity of on-device vision pipelines, image acquisition remains a dominant contributor to system-level energy consumption and memory footprint. In vision-enabled IoT platforms, the image sensor consumes energy comparable to the inference engine, thereby offsetting algorithmic efficiency gains. Consequently, current designs face a fundamental trade-off: continuous and always-on sensing incurs prohibitive energy consumption, whereas aggressive duty cycling increases latency and risks missing transient events. This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture. Detection and recognition are decoupled by combining an event-based imager operating asynchronously in an energy-efficient always-on wake-on-motion mode together with an RGB imager. Deployed on a low-power microcontroller, a novel TinyissimoYOLOv12 is introduced for efficient and accurate object detection. By activating the high-power image acquisition and processing stages only upon sparse visual triggers, the proposed architecture improves efficiency and latency, eliminating redundant sensing while maintaining continuous monitoring coverage. Experimental results demonstrate an energy consumption of only 222$\mu$Wh. Upon a motion trigger, the system completes a full sense-to-report cycle-RGB acquisition, object detection across 80 classes, and LoRa telemetry-with a total energy consumption of 28.7mJ. The network achieves up to 32.3% mAP with a model size of 1 million parameters. At a 1% daily activity ratio, the platform achieves a three-month operational lifetime with a 1.85Wh battery, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.

eess.IV

Assessing Triple Modular Redundancy for Wide-Link, Low-Latency NoC Routers: Reliability and Physical Design Challenges

Protecting the Network-on-Chip (NoC) of physical-AI tile-based accelerators deployed in harsh environments against single-event effects (SEEs) is paramount for preventing NoC failures that can lead to deadlocks and silent data corruption (SDC). Prior work on reliable NoCs has largely focused on narrow links (e.g., 32-bit), deeply pipelined routers, and single-event upsets (SEUs). However, the state of the art has evolved toward low-latency NoC routers with ultra-wide links, implemented on advanced technology nodes and operating at frequencies above 1 GHz. We evaluate the cost and reliability trade-offs of implementing Triple Modular Redundancy (TMR) at three granularities (coarse, state-only, and full) for a 2-cycle-latency NoC router with 512-bit wide links. We carry out RTL-to-GDSII physical design in TSMC 7nm technology, as well as both RTL- and netlist-level SEU and SET fault injection campaigns. We evaluate the three TMR approaches in terms of reliability, cost, and physical design strategies, further extending the assessment from a standalone router to a full AI acceleration tile. Our results show that state-only and coarse-grained TMR do not provide sufficient protection against SEEs, whereas full TMR eliminates all observed failures across more than one million injected faults per experiment. Although the standalone full-TMR router incurs a 7.04x area overhead, this cost is drastically amortized once integrated into a complete AI accelerator tile with processors and local L1 memories: the same design adds only 16.8% area and 15.2% power consumption under a GEMM benchmark at the system level, with the critical path of the tile entirely unaffected. These results demonstrate that advanced technology nodes provide sufficient routing capacity to make full TMR a practical and deployable solution for protecting NoCs in Physical AI systems operating in harsh environments.

cs.AR

Multi-Sensor Edge Angle Detection for Performance Analysis in Ski Jumping

In ski jumping, performance during the gliding phase depends on achieving an aerodynamic posture that maximizes the lift-to-drag ratio. In the V-style technique, the ski edge angle is a key determinant. Reducing the edge angle flattens the skis, increases their effective surface area, and improves aerodynamic lift, ultimately contributing to longer flight distances. Ski edge angles are biomechanically constrained by the limited range of ankle inversion. Current sensing solutions widely quantify these angles using multi-system approaches that combine sensor signals through geometric relations. Such configurations require instrumentation on both the boot and the ski, altering mass distribution, affecting balance during flight, and increasing system complexity. To overcome these limitations, this work presents a wearable sensing system that measures both boot inclination and ski edge angle without modifying the ski surface. Two ultrasonic Time of Flight (ToF) sensors and an in-shoe Inertial Measurement Unit (IMU) are integrated into a single boot-mounted unit. Edge angles are estimated by combining ultrasonic distance measurements with IMU data through geometric reconstruction of the boot-ski configuration. Laboratory experiments demonstrate an angle resolution of 0.4500{\deg}, a Mean Absolute Error (MAE) of 0.2640{\deg}, and a coefficient of determination exceeding 99\% when compared with reference measurements, indicating strong linear agreement between the two modalities. The system achieves an end-to-end latency of 30.31 ms, enabling real-time feedback suitable for athlete training, while consuming 1.28 mW of power. With a total weight of only 18.6 g the proposed system enables unobtrusive measurement of ski edge angle and boot orientation.

eess.SP

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring. To address this, we introduce S-CEReBrO (Streaming CEReBrO), an evolution of the CEReBrO architecture designed for continuous monitoring. Our novel Windowed Alternating Attention mechanism factorizes attention computation into fixed-size spatiotemporal windows, guaranteeing constant KV cache memory as only the active window requires resident attention maps. Empirical scaling analysis confirms that windowed alternating attention can process signals 100X longer than full self-attention and 3X longer than low-rank linear attention. Compared to low-rank linear attention on long contexts, windowed alternating attention requires 55% of the memory while increasing inference throughput by 2.1X. Pre-trained on >25,000 hours of recordings from >12,000 subjects, S-CEReBrO achieves state-of-the-art performance on 7 of 11 downstream tasks, with up to 60% fewer parameters. This work represents a significant step toward the realization of efficient, generalizable, and continuous EEG monitoring. An accompanying code repository is available.

cs.LG

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. In the moderate-sparsity regime, Gustavson's dataflow provides a natural execution model for exploiting both activation and weight sparsity on vector processors through metadata-driven indexed accumulation. However, existing RVV architectures lack native support for this pattern, forcing kernels to rely on software index decoding and L1-backed indexed memory operations that keep sparse tensor contractions far below their roofline performance bound. We present Ventaglio, a runtime-configurable sparse execution unit coupled with RVV ISA extensions that drives sparse tensor contractions toward their roofline through indexed gather-accumulate-scatter support. Integrated into an open-source vector processing cluster and implemented in 12nm FinFET, Ventaglio accelerates sparse tensor contraction kernels by $6.9\text{--}7.4\times$ over optimized RVV baselines, with only $3.1\%$ area overhead for a cluster of tightly-L1 coupled vector processing elements. We build a performance-accurate instruction-level model of the Ventaglio extension, calibrate it against RTL implementation, and leverage it for scale-out performance analysis on a large $4\times4$ multi-cluster system. Using a DuoGPT-pruned LLaMA-3-8B model with practical $40\text{--}60\%$ dual sparsity, Ventaglio achieves $2.40\text{--}5.25\times$ and $2.06\text{--}3.16\times$ speedup over dense baselines during prefill and autoregressive decoding, respectively.

cs.AR

Scalable Attention for 5G NR Channel Estimation

Attention-based neural estimators achieve strong channel-estimation accuracy, but the computational cost of global attention over the time-frequency resource grid grows quadratically with the number of subcarriers, and these estimators are typically tied to a single resource allocation. This paper proposes Channel Estimation Attention (CHEA), a low-complexity channel estimator for 5G New Radio (5G NR) multi-user multiple-input multiple-output (MU-MIMO). CHEA replaces global attention with a multi-resolution windowed design: a high-resolution encoder preserves local pilot detail, a low-resolution encoder captures wider frequency-domain context, and a local cross-attention decoder transfers this coarse context back to the high-resolution pilot tokens. A per-Physical Resource Block (PRB) upsampling module then reconstructs the channel over the full slot. Because every attention operation is confined to a fixed-size window and reconstruction is performed per PRB, the cost of CHEA scales linearly with the number of subcarriers, and a single trained model supports different PRB allocations without retraining. On a standard-compliant Physical Uplink Shared Channel (PUSCH), CHEA achieves the lowest Mean Squared Error (MSE) among conventional and state-of-the-art neural estimators, while requiring 2.8\(\times\) to 22.0\(\times\) lower operations than existing attention-based estimators.

eess.SP

Improving Autonomous Nano-drones Performance via Automated End-to-End Optimization and Deployment of DNNs

The evolution of energy-efficient ultra-low-power (ULP) parallel processors and the diffusion of convolutional neural networks (CNNs) are fueling the advent of autonomous driving nano-sized unmanned aerial vehicles (UAVs). These sub-10 cm robotic platforms are envisioned as next-generation ubiquitous smart-sensors and unobtrusive robotic-helpers. However, the limited computational/memory resources available aboard nano-UAVs introduce the challenge of minimizing and optimizing vision-based CNNs -- which to date require error-prone, labor-intensive iterative development flows. This work explores methodologies and software tools to streamline and automate all the deployment of vision-based CNN navigation on a ULP multicore system-on-chip acting as a mission computer on a Crazyflie 2.1 nano-UAV. We focus on the deployment of PULP-Dronet, a state-of-the-art CNN for autonomous navigation of nano-UAVs, from the initial training to the final closed-loop evaluation. Compared to the original hand-crafted CNN, our results show a 2x reduction of memory footprint and a speedup of 1.6x in inference time while guaranteeing the same prediction accuracy and significantly improving the behavior in the field, achieving: i) obstacle avoidance with a peak braking-speed of 1.65 m/s and improving the speed/braking-space ratio of the baseline, ii) free flight in a familiar environment up to 1.96 m/s (0.5 m/s for the baseline), and iii) lane following on a path featuring a 90 deg turn -- all while using for computation less than 1.6% of the drone's power budget. To foster new applications and future research, we open-source all the software design in a ready-to-run project compatible with the Crazyflie 2.1

eess.IV

WULPUS PRO: Multi-mode Ultra-Low-Power Wearable Ultrasound and Array Imaging with CMUT Support

Wearable ultrasound enables continuous monitoring of physiological processes such as muscle dynamics, bladder volume, and cardiovascular activity. Existing fully wearable ultra-low-power platforms are limited to shallow, low-channel A-mode sensing, while larger multi-mode systems are too bulky and power-hungry for true wearability. We present WULPUS PRO, a runtime-programmable wearable ultrasound acquisition platform measuring $39\times21\times6 \mathrm{mm}$ and weighing $5 \mathrm{g}$. It integrates $30 \mathrm{V}$ excitation, 16 time-multiplexed channels, a low-noise receive front-end with up to $70 \mathrm{dB}$ gain, $9.9 \mathrm{MHz}$ bandwidth, time-gain compensation, and $32 \mathrm{dB}$ SNR. The platform supports deep-tissue echo acquisition up to $2.2 \mathrm{MHz}$ in RF-sampling mode and $8 \mathrm{MHz}$ in envelope-detection mode. We demonstrate B-mode imaging in a 16-channel ultra-low-power wearable with sub-millimeter axial and millimeter-scale lateral resolution in phantom experiments, while consuming $40 \mathrm{mW}$ at $50 \mathrm{Hz}$ PRF and under $60 \mathrm{mW}$ at $300 \mathrm{Hz}$ PRF. WULPUS PRO supports both piezoelectric and capacitive micromachined ultrasonic transducers, enabling integration with skin-conformal polymer-based CMUT arrays. As a host-agnostic acquisition front-end, it exposes standard data and power interfaces for BLE- and Wi-Fi-based wearable hosts. We demonstrate wireless transmission with external BLE and Wi-Fi modules and project 1-2 days of BLE operation at $50 \mathrm{Hz}$ PRF and over 3 h of Wi-Fi streaming at $300 \mathrm{Hz}$ PRF using a $300 \mathrm{mAh}$, $6.4 \mathrm{g}$ Li-Po cell. WULPUS PRO establishes a new class of fully programmable, B-mode-enabled, ultra-low-power wearable ultrasound platforms.

eess.SY

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however, effectively mapping attention mechanisms onto NPUs remains challenging due to architectural diversity and explicit data-movement programming models. In this work, we present STEEL, the first open-source implementation of FlashAttention targeting XDNA-like NPUs. STEEL introduces a dataflow formulation of prefill attention, enabling efficient exploitation of spatial parallelism and on-chip memory. Furthermore, STEEL addresses the load imbalance induced by the causal mask by leveraging a sparsity-aware pipeline placement onto the NPU array, reducing synchronization overhead and improving utilization. We evaluate STEEL on the AMD Ryzen AI 9 HX 370 SoC and compare its performance against optimized CPU and GPU implementations. Experimental results show that STEEL reduces energy consumption by an average of 9.17x and 1.75x relative to CPU and GPU baselines, respectively. On XDNA 1, STEEL achieves an average 9.6x latency reduction over the prior state of the art, and delivers a 22.8x speedup on average compared to a layer-by-layer attention implementation on XDNA 2.

cs.DC

Low-Power License Plate Detection and Recognition on a RISC-V Multi-Core MCU-Based Vision System

In this paper, we present the first (to the best of our knowledge) demonstration of a low-power MCU-based edge device for Automatic License Plate Recognition (ALPR). The design leverages on a 9-core RISC-V processor, GAP8, coupled with a QVGA ultra-low-power greyscale imager. The proposed visual processing pipeline uses a multi-model inference approach based on SSDlite-MobilenetV2 for license plate detection and LPRNet for optical character recognition, reaching a 38.9% mAP score for the first task and a recognition rate of >99.13% for the latter on public datasets. On real-world data, the pipeline recognizes registration numbers when the size of LP crops is as small as 30x5 pixels. Thanks to the applied compression and optimization strategies, the multi-model inference (687 MMAC) achieves a throughput of 1.09 FPS at a power cost of 117 mW when running on GAP8. Our solution is the first MCU-class device embedding such a level of network complexity, resulting to be 73x more energy-efficient w.r.t. precedent mobile-class ALPR system featuring a Raspberry Pi3. The proposed design does not resort to any hardwired acceleration engines, thus retaining full flexibility for future algorithmic improvements.

cs.CV

Physically-Aware Preemptive Virtual Channels for Deadlock-Free AXI Networks-on-Chip

As many-core Systems-on-Chip (SoCs) continue to scale, Networks-on-Chip (NoCs) must sustain increasingly high memory bandwidth while preserving deadlock freedom. In AXI4 systems, protocol-level dependencies between read and write traffic can create circular waits at the network endpoints, even when the routing algorithm itself is deadlock-free. Decoupling these traffic classes avoids such dependencies, but exposes a key implementation trade-off: multiplane NoCs duplicate wide physical links and increase routing pressure, whereas conventional Virtual Channel (VC) routers add substantial control complexity, area, and timing overhead. This work revisits this trade-off for modern wide-link NoCs. We evaluate four deadlock-free AXI4 traffic-class separation schemes: a multiplane baseline and three lightweight VC-based designs. Among these designs, we propose Preemptive VCs, a physically-aware architecture that can save up to 76% of link resources with comparable frequency and only 3% router area overhead relative to the multiplane design.

cs.AR

LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation

Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving as a catalyst for open-vocabulary tasks like ObjectNav. Yet, their computational footprint typically confines them to cloud execution, hindering low-latency inference with local deployment on resource-constrained robots. To address this challenge, we present a distillation strategy that transfers complex spatial-semantic reasoning from large frontier models into a lightweight, 4B-parameter local VLM for edge execution on embedded GPU devices (e.g., Jetson Orin). We first establish a State of the Art (SotA), Scene Graph (SG)-based pipeline using Claude Sonnet 4.6, achieving a 39.7% Success Rate (SR) on the HM3D OVON benchmark. We then demonstrate that fine-tuning Qwen3.5-4B on just 500 frontier reasoning traces effectively enables navigation capabilities, yielding a SR of 34.5%, narrowing the gap to the performance of large cloud models. Finally, we introduce E-RLVR with Token Generation (TG) regularization to compress output sequence lengths for physical deployment while grounding the agent in its task. This downstream optimization reduces TG overhead by 72.1% and latency by 71.8%. Combined with quantization, this joint strategy yields a cumulative 82.8% reduction in overall inference latency without significantly sacrificing performance, presenting a viable paradigm for local, low-latency VLM execution on mobile robots.

cs.RO

Quantizing Recursive Reasoning Models

Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are reused many times, quantizing them creates a unique dynamical problem: the quantization error is incurred at every step. While 8-bit quantization (integer or float) preserves accuracy, moving to a per-tensor 4-bit format causes a systematic bias to accumulate. The ensuing drift catastrophically degrades exact-solution accuracy on Sudoku from 84.1% to 0.0% (only ~25% of cells correct). In this work, we show that this collapse is caused by activation-scaling granularity rather than bit-width or number format. Crucially, moving to per-block scaling completely restores the transition. To implement this, we apply MXInt4, a blockwise integer activation format, to recursive reasoning models. It is competitive with blockwise float formats on our tasks, while keeping integer elements and power-of-two block scales. Finally, recursion depth and reuse modulate quantization sensitivity, with the deepest architecture we test (the EqR equilibrium model) the most sensitive. Yet blockwise scaling overcomes this vulnerability, staying robust across these architectures and transferring to the open-ended ARC-AGI benchmark.

cs.LG

Croc: Training the Next Generation Chip Designers on Domain-Specific End-to-End Open Source Silicon

The demand for domain-specific systems-on-chip (SoCs) in artificial intelligence, robotics, and automotive systems is increasing the need for engineers with hands-on expertise on very-large-scale integration (VLSI) design from architecture specification to fabricated silicon. Yet, most VLSI courses rely on restrictively licensed electronic design automation tools and process design kits (PDKs), as well as closed-source hardware designs. We present an end-to-end open-source domain-specific SoC design and fabrication flow built around Croc, a highly customizable RISC-V platform. Built from open-source SystemVerilog intellectual property blocks and integrated with an end-to-end open-source design flow in a 130nm open PDK, Croc enables tapeout projects supporting multiple domain customization options: instruction-set extensions, accelerator co-processors, and peripherals. In our first open-source course experience using Croc, 65 students completed 33 projects, 30 of which produced manufacturable layouts. 18 designs were selected as tapeout candidates, and five were fabricated. A first baseline chip has already been successfully characterized in silicon, demonstrating microcontroller-class functionality and implementation metrics comparable to those of products with similar functional complexity completed with closed-source toolchains and PDKs.

cs.AR

CVA6-RT: an Open-Source Time-Predictable RV64 Processor for Mixed-Criticality Systems

This work presents CVA6-RT, a real-time micro-architectural extension of the CVA6 core to bound worst-case latency and reduce task's timing execution variability. CVA6-RT implements the rv64gch ISA and features advanced support for real-time execution, including TLB partitioning and locking for predictable address translation, a dynamically reconfigurable scratchpad mode in the L1 caches for deterministic memory access, and low-latency interrupt handling via an enhanced interrupt controller combined with hardware-assisted context stacking. With real-time features enabled, CVA6-RT achieves an interrupt latency of 12 cycles, comparable to that of simpler Arm Cortex-M microcontrollers, and 10x lower than the baseline CVA6 core.

cs.AR

Towards Whole Hand and Wrist Kinematic Tracking with a Wearable A-Mode Ultrasound Probe

A-mode ultrasound (US) has emerged as a promising modality for hand and wrist motion tracking. Prior works have mainly addressed static gesture classification or regression of a few degrees of freedom (DoFs), typically relying on non-wearable systems and external computing devices, and highlight the need for strategies to ensure robustness to sensor repositioning. In this work, we propose a framework for robust whole-hand and wrist kinematic tracking via wearable A-mode US using the WULPUS platform, tackling the regression of 23 DoFs directly on the probe. First, we introduce a compact (11285 parameters) multi-output convolutional neural network combined with an incremental training strategy, which improves inter-session generalization and reduces mean absolute error by more than 17% compared to a non-incremental approach. Second, we demonstrate, for the first time, the feasibility of end-to-end hand and wrist kinematic tracking entirely on-device. We deploy the model on the WULPUS nRF52832 microcontroller, achieving 0.73 mJ per inference, 29.1 ms latency, and showing the feasibility of full operation (data acquisition, online inference, and BLE streaming of results) within 33 mW, enabling up to 36 hours of continuous use and an 88% reduction in wireless bandwidth compared to raw data transmission.

eess.SP