SearcharxivSearch

arXiv subjects

Gang Feng

Publications and source records attributed to Gang Feng.

At least 19 recordsLinked to original sources

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehensive whole-body PET/CT analysis. In this work, we introduce MetaStructAtlas, a large-scale dataset for grounded whole-body PET/CT interpretation that synthesizes multimodal imaging with integrated anatomical, metabolic, and semantic annotations. MetaStructAtlas provides 490 co-registered 3D PET and CT volumes with 50,470 organ-level segmentation masks and grounded radiology reports. To facilitate interactive reasoning, we further developed MetaStructVQA, a standardized 3D grounded visual question-answering benchmark containing 100,565 QA pairs. This framework explicitly links diagnostic queries to visual evidence across modalities, encompassing anatomical, morphological, and metabolic characteristics. Finally, we evaluate state-of-the-art 3D medical VLMs on MetaStructVQA, establishing a robust foundation for multimodal representation learning and integrated whole-body reasoning in nuclear medicine.

cs.CV

Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret

This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.

cs.DC

Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach

Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead be shared while $A$ remains client-specific (Share-B/Local-A). With a least-squares surrogate, we reveal that Share-A/Local-B requires the client-specific LoRA update matrices to use a common rank-$r$ input-side space, whereas Share-B/Local-A requires a common rank-$r$ output-side space. The two strategies therefore incur different projection residuals, indicating that the preferred strategy is the one with the smaller aggregate residual across clients. With this insight, we propose Federated Adaptive Factor Sharing Low-Rank Adaptation (FedAS-LoRA), which selects the sharing side before training to enhance fine-tuning performance. To enable adaptive factor selection before training, we design a Rank-Aware Shared-Subspace Sufficiency (RSS) metric, which effectively assesses whether a shared rank-$r$ input subspace is sufficient for the local data distributions using representations extracted from a frozen LLM backbone. Experiments across different tasks, data distributions, LoRA ranks, and participation settings confirm the effectiveness of RSS and the superior performance of FedAS-LoRA.

cs.LG

Channel Knowledge Empowered Finite-Blocklength Rate-Splitting Transmission for High-Mobility Autonomous Driving

To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated channel state information (CSI) is prevalent.Moreover, the advanced sensing, localization, and on-board computation capabilities of autonomous driving vehicles facilitate the construction of a channel knowledge map (CKM), which is a key enabler for environment-aware communications in future 6G networks.In this work, we propose a CKM empowered finite-blocklength (FBL) RSMA for downlink autonomous driving system. The location-dependent large-scale channel information provided by CKM is exploited in RSMA to develop a refined rate-splitting design. The min-rate performance of FBL rate splitting is analyzed explicitly to ensure user fairness. We derive a new and tight closed-form bound for the private-stream ergodic rate. Combined with the closed-form common-stream expression, an efficient optimization design of rate-splitting ratios has been formulated. Numerical results show that the CKM empowered FBL RSMA outperforms space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), particularly in high-mobility scenarios. Its performance is improved by a data-based CKM, which provides more accurate large-scale channel information than model-based approaches and enables more precise common-stream allocation. The results also reveal that RSMA is sensitive to errors in large-scale channel knowledge, emphasizing the importance of accurate CKM information for optimal rate-splitting.

cs.IT

Distributed Resilient Fixed-Time Control for Cooperative Output Regulation of MASs over Directed Graphs under DoS Attacks

This paper addresses the problem of fixed-time cooperative output regulation for linear multi-agent systems over directed graphs under denial-of-service attacks. A novel distributed resilient fixed-time controller is developed that comprises a distributed resilient fixed-time observer taking general directed graphs into consideration, and a distributed resilient fixed-time control law for each agent. The proposed controller neither depends on Laplacian symmetry nor requires strong connectivity and detail-balanced condition, in contrast to existing distributed resilient fixed-time controllers. Under the proposed controller, the regulated outputs converge to zero in a fixed time with its upper bound independent of the initial states of the multi-agent system. Ultimately, the efficacy of the proposed controller is demonstrated via a simulation example.

eess.SY

Developing Foundation Models for Universal Segmentation from 3D Whole-Body Positron Emission Tomography

Positron emission tomography (PET) is a key nuclear medicine imaging modality that visualizes radiotracer distributions to quantify in vivo physiological and metabolic processes, playing an irreplaceable role in disease management. Despite its clinical importance, the development of deep learning models for quantitative PET image analysis remains severely limited, driven by both the inherent segmentation challenge from PET's paucity of anatomical contrast and the high costs of data acquisition and annotation. To bridge this gap, we develop generalist foundational models for universal segmentation from 3D whole-body PET imaging. We first build the largest and most comprehensive PET dataset to date, comprising 11041 3D whole-body PET scans with 59831 segmentation masks for model development. Based on this dataset, we present SegAnyPET, an innovative foundational model with general-purpose applicability to diverse segmentation tasks. Built on a 3D architecture with a prompt engineering strategy for mask generation, SegAnyPET enables universal and scalable organ and lesion segmentation, supports efficient human correction with minimal effort, and enables a clinical human-in-the-loop workflow. Extensive evaluations on multi-center, multi-tracer, multi-disease datasets demonstrate that SegAnyPET achieves strong zero-shot performance across a wide range of segmentation tasks, highlighting its potential to advance the clinical applications of molecular imaging.

cs.CV

Universality Reconsidered: Rethinking the Validation of Foundation Models for General-Purpose 3D Medical Segmentation

Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse targets and imaging modalities. Yet the prevailing conception of universality remains incomplete. Current models are predominantly developed and evaluated on datasets largely concentrated around a limited set of imaging modalities and anatomical regions. In this Perspective, we evaluate representative 3D segmentation foundation models using paired whole-body structural and functional imaging data. Our analysis reveals a substantial gap between benchmark-reported performance and real-world generalization, with marked degradation on previously unseen data and particularly severe failures on functional imaging modalities. These findings suggest that current foundation models remain far from achieving true universality. We argue that progress requires not only scaling models and datasets, but also a reconsideration of how universality is defined and validated, extending evaluation beyond regional structural benchmarks toward whole-body structural and functional imaging. Our observations highlight the need to distinguish benchmark success from genuine clinical generalization. Bridging this gap will be essential for translating foundation models from controlled evaluation settings to real-world medical practice.

cs.CV

PETWB-REP: A Multi-Cancer Whole-Body FDG PET/CT and Radiology Report Dataset for Medical Imaging Research

Publicly available, large-scale medical imaging datasets are crucial for developing and validating artificial intelligence models and conducting retrospective clinical research. However, datasets that combine functional and anatomical imaging with detailed clinical reports across multiple cancer types remain scarce. Here, we present PETWB-REP, a curated dataset comprising whole-body 18F-Fluorodeoxyglucose (FDG) Positron Emission Tomography/Computed Tomography (PET/CT) scans and corresponding radiology reports from 490 patients diagnosed with various malignancies. The dataset primarily includes common cancers such as lung cancer, liver cancer, breast cancer, prostate cancer, and ovarian cancer. This dataset includes paired PET and CT images, de-identified textual reports, and structured clinical metadata. It is designed to support research in medical imaging, radiomics, artificial intelligence, and multi-modal learning.

cs.CV

Resilient Global Practical Fixed-Time Cooperative Output Regulation of Uncertain Nonlinear Multi-Agent Systems Subject to Denial-of-Service Attacks

This paper investigates the problem of resilient global practical fixed-time cooperative output regulation of uncertain nonlinear multi-agent systems subject to denial-of-service attacks. A novel distributed resilient adaptive fixed-time control strategy is proposed, which consists of a novel distributed resilient fixed-time observer with a chain of nonlinear filters and a novel distributed resilient adaptive fixed-time controller. It is shown that the problem of resilient global practical fixed-time cooperative output regulation can be solved by the proposed control strategy. More specifically, the proposed {distributed} control strategy ensures the global boundedness of all the signals in the resulting closed-loop system and the global convergence of the regulated outputs to a {tunable} residual set in a fixed time. A simulation example is finally provided to illustrate the efficacy of the proposed control strategy.

eess.SY

PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography

Positron emission tomography (PET) is a cornerstone of modern oncologic and neurologic imaging, distinguished by its unique ability to illuminate dynamic metabolic processes that transcend the anatomical focus of traditional imaging technologies. Radiology reports are essential for clinical decision making, yet their manual creation is labor-intensive and time-consuming. Recent advancements of vision-language models (VLMs) have shown strong potential in medical applications, presenting a promising avenue for automating report generation. However, existing applications of VLMs in the medical domain have predominantly focused on structural imaging modalities, while the unique characteristics of molecular PET imaging have largely been overlooked. To bridge the gap, we introduce PET2Rep, a large-scale comprehensive benchmark for evaluation of general and medical VLMs for radiology report generation for PET images. PET2Rep stands out as the first dedicated dataset for PET report generation with metabolic information, uniquely capturing whole-body image-report pairs that cover dozens of organs to fill the critical gap in existing benchmarks and mirror real-world clinical comprehensiveness. In addition to widely recognized natural language generation metrics, we introduce a series of clinical efficacy metrics to evaluate the quality of radiotracer uptake pattern description in key organs in generated reports. We conduct a head-to-head comparison of 30 cutting-edge general-purpose and medical-specialized VLMs. The results show that the current state-of-the-art VLMs perform poorly on PET report generation task, falling considerably short of fulfilling practical needs. Moreover, we identify several key insufficiency that need to be addressed to advance the development in medical applications.

eess.IV

Global Regulation of Feedforward Nonlinear Systems: A Logic-Based Switching Gain Approach

In this article, we investigate the global regulation problem for a class of feedforward nonlinear systems. Notably, the systems under consideration allow unknown input-output-dependent nonlinear growth rates, which has not been considered in existing works. A novel logic-based switching (LBS) gain approach is proposed to counteract system uncertainties and nonlinearities. Furthermore, a tanh-type speed-regulation function is embedded into the switching mechanism for the first time to improve the convergence speed and transient performance. Then, a switching adaptive output feedback (SAOF) controller is proposed based on the developed switching mechanism, which is of a concise form and low-complexity characteristic. It is shown that the objective of global regulation is achieved with faster convergence speed and better transient performance under the proposed controller. Moreover, by strengthening the switching mechanism, the improved control approach can deal with feedforward nonlinear systems with external disturbances. Finally, representative examples are presented to demonstrate the effectiveness and advantages of our approach in comparison with the existing approaches.

eess.SY

TrustFed: A Reliable Federated Learning Framework with Malicious-Attack Resistance

As a key technology in 6G research, federated learning (FL) enables collaborative learning among multiple clients while ensuring individual data privacy. However, malicious attackers among the participating clients can intentionally tamper with the training data or the trained model, compromising the accuracy and trustworthiness of the system. To address this issue, in this paper, we propose a hierarchical audit-based FL (HiAudit-FL) framework, with the aim to enhance the reliability and security of the learning process. The hierarchical audit process includes two stages, namely model-audit and parameter-audit. In the model-audit stage, a low-overhead audit method is employed to identify suspicious clients. Subsequently, in the parameter-audit stage, a resource-consuming method is used to detect all malicious clients with higher accuracy among the suspicious ones. Specifically, we execute the model audit method among partial clients for multiple rounds, which is modeled as a partial observation Markov decision process (POMDP) with the aim to enhance the robustness and accountability of the decision-making in complex and uncertain environments. Meanwhile, we formulate the problem of identifying malicious attackers through a multi-round audit as an active sequential hypothesis testing problem and leverage a diffusion model-based AI-Enabled audit selection strategy (ASS) to decide which clients should be audited in each round. To accomplish efficient and effective audit selection, we design a DRL-ASS algorithm by incorporating the ASS in a deep reinforcement learning (DRL) framework. Our simulation results demonstrate that HiAudit-FL can effectively identify and handle potential malicious users accurately, with small system overhead.

cs.CR

Towards a model of "social touch'' for ubiquitous communication

One of the challenges of telepresence robotics is to provide ubiquitous social-interpersonalimmersion. In order to achieve this, there is a need to understand and model the factors that wouldallow the users to control the transmission of their vocal productions, and to give them perception,proprioception and inter-proprioception of this control. This model for transferring sociallyembodied vocal distance should take into account all parameters involved in the social-interpersonaleffect of vocal productions, especially intensity. It should also integrate the background informationwhich is relevant for the speakers to express their intentions. In this paper, we present a firstexperiment for measuring, analyzing and modeling how human beings perceive the distance to aninterlocutor, depending on socio-affective variations in the vocal productions of this interlocutor.These results will be the reference for the models which will be implanted in our telepresence robot:Robair Social Touch.

cs.RO

A study on the Lombard Effect in telepresence robotics

In this study, we present a new experiment in order to study the Lombard effect in telepresence robotics. In this experiment, one person talks with a robot controled remotely by someone in a different room. The remote pilot (R) is immersed in both environments, while the local interlocutor (L) interacts directly with the robot. In this context, the position of the noise source, in the remote or in the local room, may modify the subjects' voice adaptations. In order to study in details this phenomenon, we propose four particular conditions: no added noise, noise in room R heard only by R, virtual noise in room L heard only by R, and noise in room L heard by both R and L. We measured the variations of maximum intensity in order to quantify the Lombard effect. Our results show that there is indeed a modification of voice intensity in all noisy conditions. However, the amplitude of this modification varies depending on the condition.

cs.RO

Can we hear physical and social space together through prosody?

When human listeners try to guess the spatial position of a speech source, they are influenced by the speaker's production level, regardless of the intensity level reaching their ears. Because the perception of distance is a very difficult task, they rely on their own experience, which tells them that a whispering talker is close to them, and that a shouting talker is far away. This study aims to test if similar results could be obtained for prosodic variations produced by a human speaker in an everyday life environment. It consists in a localization task, during which blindfolded subjects had to estimate the incoming voice direction, speaker orientation and distance of a trained female speaker, who uttered single words, following instructions concerning intensity and social-affect to be performed. This protocol was implemented in two experiments. First, a complex pretext task was used in order to distract the subjects from the strange behavior of the speaker. On the contrary, during the second experiment, the subjects were fully aware of the prosodic variations, which allowed them to adapt their perception. Results show the importance of the pretext task, and suggest that the perception of the speaker's orientation can be influenced by voice intensity.

cs.RO

Automated Federated Learning in Mobile Edge Networks -- Fast Adaptation and Convergence

Federated Learning (FL) can be used in mobile edge networks to train machine learning models in a distributed manner. Recently, FL has been interpreted within a Model-Agnostic Meta-Learning (MAML) framework, which brings FL significant advantages in fast adaptation and convergence over heterogeneous datasets. However, existing research simply combines MAML and FL without explicitly addressing how much benefit MAML brings to FL and how to maximize such benefit over mobile edge networks. In this paper, we quantify the benefit from two aspects: optimizing FL hyperparameters (i.e., sampled data size and the number of communication rounds) and resource allocation (i.e., transmit power) in mobile edge networks. Specifically, we formulate the MAML-based FL design as an overall learning time minimization problem, under the constraints of model accuracy and energy consumption. Facilitated by the convergence analysis of MAML-based FL, we decompose the formulated problem and then solve it using analytical solutions and the coordinate descent method. With the obtained FL hyperparameters and resource allocation, we design a MAML-based FL algorithm, called Automated Federated Learning (AutoFL), that is able to conduct fast adaptation and convergence. Extensive experimental results verify that AutoFL outperforms other benchmark algorithms regarding the learning time and convergence performance.

cs.LG

Slicing4Meta: An Intelligent Integration Framework with Multi-dimensional Network Resources for Metaverse-as-a-Service in Web 3.0

As the next-generation Internet paradigm, Web 3.0 encapsulates the expectations of user-centric immersion and interaction experiences in a decentralized manner. Metaverse, a virtual world interacting with the physical world, is becoming one of the most potential technology to push forward with Web 3.0. In the Metaverse, users expect to tailor immersive and interactive experiences by customizing real-time Metaverse services (e.g., augmented/virtual reality and digital twin) in the three-dimensional virtual world. Nevertheless, there are still no unified solutions for the Metaverse services in terms of management and orchestration. It is calling for a continuous evolution of scalable solutions and mobile systems to support Metaverse services, and thus bring Metaverse into reality. In this paper, to provide scalable subscription solutions for Metaverse services, we propose a new concept, named Metaverse-as-a-Service (MaaS), in which various physical-virtual components and technologies in the Metaverse can be delivered as services. Furthermore, to unify the management and orchestration of MaaS, we propose a novel framework, called Slicing4Meta, to customize Metaverse services by intelligently integrating MaaS models and the associated multi-dimensional resources on the components and technologies. Additionally, we propose the classification for typical Metaverse services based on the quality of experience (QoE) requirements and illustrate how to fulfill the QoE requirements under the Slicing4Meta framework. We then illustrate a virtual traveling study case, in which we examine the relationship between the QoE and the multi-dimensional resources by quantitatively modeling the QoE of Metaverse users. Finally, we discuss some open challenges of Slicing4Meta and propose potential solutions to address the challenges.

cs.NI

Beam Management in Ultra-dense mmWave Network via Federated Reinforcement Learning: An Intelligent and Secure Approach

Deploying ultra-dense networks that operate on millimeter wave (mmWave) band is a promising way to address the tremendous growth on mobile data traffic. However, one key challenge of ultra-dense mmWave network (UDmmN) is beam management due to the high propagation delay, limited beam coverage as well as numerous beams and users. In this paper, a novel systematic beam control scheme is presented to tackle the beam management problem which is difficult due to the nonconvex objective function. We employ double deep Q-network (DDQN) under a federated learning (FL) framework to address the above optimization problem, and thereby fulfilling adaptive and intelligent beam management in UDmmN. In the proposed beam management scheme based on FL (BMFL), the non-rawdata aggregation can theoretically protect user privacy while reducing handoff cost. Moreover, we propose to adopt a data cleaning technique in the local model training for BMFL, with the aim to further strengthen the privacy protection of users while improving the learning convergence speed. Simulation results demonstrate the performance gain of our proposed scheme.

cs.IT