SearcharxivSearch

arXiv subjects

Riccardo Trivisonno

Publications and source records attributed to Riccardo Trivisonno.

7 recordsLinked to original sources

ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics

Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.

cs.RO

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.

cs.LG

A Goal-Oriented Networking Approach for Intelligent IoT Service Deployment

The first 6G standardization efforts are about to start, shaping the new generation of mobile networks. The IMT-2030 extends the IMT-2020 by expanding its usage scenarios to Immersive, Massive, and Hyper-Reliable and Low-Latency Communications. It also introduces novel scenarios by integrating Artificial Intelligence and Sensing with Communication and supporting Ubiquitous Connectivity. Compared to the previous generation, 6G is expected to improve not only throughput and latency, but also coverage and energy efficiency. A paradigm called Goal-Oriented (GO) communications has recently emerged as a promising solution to improve network efficiency. It relies on the fact that the goal of the communication network is to achieve a specific task with a defined accuracy, rather than creating perfect data delivery. Intelligent devices can pre-process data to send only what is relevant to achieve the task, thus saving precious network resources and energy. Recent works demonstrate that incorporating service- and application-level KPIs in the network allows to achieve higher communication efficiency for devices, but the consequence of using such techniques on the network itself has not yet been explored. This paper proposes a practical end-to-end framework to assess energy consumption, latency, and goal accuracy KPIs, which includes a Multi-Objective optimization model to evaluate the trade-offs between the multiple KPIs relevant to GO networking. We demonstrate, through simulation, that the network can benefit from the application of the GO paradigm, indicating its potential in future network architectures.

cs.NI

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions. However, training such agents with reinforcement learning in multi-session environments is challenging because memory turns the agent's past actions into part of its future environment. Once different rollouts write, update, or delete different memories, they no longer share the same intermediate memory state, making trajectory-level comparisons fundamentally unfair. This violates a key assumption behind group-relative methods such as GRPO, where rollouts are compared as if they were sampled from the same effective environment. Consequently, trajectory-level rewards provide noisy or biased credit signals for long-horizon memory operations. To address this challenge, we introduce Memory-R2, a training framework for long-horizon memory-augmented LLM agents. Its core algorithm, LoGo-GRPO, combines local and global group-relative optimization. The global objective preserves end-to-end learning from long-horizon trajectory-level rewards, while local rerollouts compare different memory-operation outcomes from the same intermediate memory state, yielding fairer group comparisons and more precise supervision for memory construction. Beyond credit assignment, Memory-R2 jointly optimizes memory formation and memory evolution with a shared-parameter co-learning design, where a fact extractor and a memory manager are instantiated from the same LLM backbone through role-specific prompts. To stabilize multi-step RL over long memory horizons, we adopt a progressive curriculum that increases the training horizon from 8 to 16 to 32 sessions. Together, these components provide an effective training paradigm for memory-augmented LLM agents in long-horizon multi-session settings.

cs.LG

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning

Memory has become an increasingly important component of agentic systems, as these systems are expected to reason over long-term experience. However, prior work has largely focused on unimodal memory, leaving multimodal memory relatively underexplored despite its central role in real-world applications. Compared with unimodal settings, multimodal memory introduces additional challenges, including heterogeneous input integration, person-centric information alignment, and evidence aggregation across different granularities. We present PyraVid, a hierarchical multimodal memory framework inspired by Event Segmentation Theory from cognitive science. PyraVid organizes long videos into a coarse-to-fine pyramid structure, enabling structured memory access and effective evidence aggregation. It further supports structure-guided memory expansion with pruning, allowing the retrieval of related events with strong causal connectivity but low semantic similarity while reducing noise. Experiments on multiple long-video understanding benchmarks show that PyraVid consistently improves performance across datasets, model scales, and question types, highlighting the effectiveness of hierarchical multimodal memory for long-horizon reasoning.

cs.MA

A Distributed Intelligence Architecture for B5G Network Automation

The management of networks is automated by closed loops. Concurrent closed loops aiming for individual optimization cause conflicts which, left unresolved, leads to significant degradation in performance indicators, resulting in sub-optimal network performance. Centralized optimization avoids conflicts, but impractical in large-scale networks for time-critical applications. Distributed, pervasive intelligence is therefore envisaged in the evolution to B5G networks. In this letter, we propose a Q-Learning-based distributed architecture (QLC), addressing the conflict issue by encouraging cooperation among intelligent agents. We design a realistic B5G network slice auto-scaling model and validate the performance of QLC via simulations, justifying further research in this direction.

cs.NI

End-to-End Architecture Modularisation and Slicing for Next Generation Networks

The journey towards the deployment of next generation networks has recently accelerated, driven by the joint effort of research and standards organisations. Despite this fact, the overall picture is still unclear as prioritization and understanding on several key concepts are not yet agreed by major vendors and network providers. Network Slicing is one of the central topics of the debate, and it is expected to become the key feature of next generation networks, providing the flexibility required to support the variety of 5G use cases and business. Network slices are seen as network operator business, offering the possibility to provide flexible services and even infrastructures to vertical industries and classical Telco customers alike. Another key ingredient is the Architecture Modularisation concept, discussed in this paper and regarded by the authors as the essential design principle to build a flexible network architecture natively supporting Network Slicing. According to this concept, conventional monolithic network functions, often corresponding to physical network elements in the existing systems, are to split into basic building blocks defined with the proper granularity, allowing the definition of different logical architectures (i.e. different Network Slices). In this paper, we further discuss a modularisation methodology as a criteria to define the right set of basic building blocks. Defined through this proposed methodology, the set of basic building blocks and the relating interfacing model are discussed. The paper concludes by proposing a modular 5G network architecture as candidate for next generation network standards.

cs.NI