SearcharxivSearch

arXiv subjects

Huong Nguyen

Publications and source records attributed to Huong Nguyen.

13 recordsLinked to original sources

From Data Heterogeneity to Convergence: A Data-Centric Review of Federated Learning

Federated Learning (FL) has emerged as a promising solution for data hunger in centralized learning. This paradigm enables privacy with multiple clients to train a shared-task model collaboratively without exposing their local data. While being a key component in any learning system, data is also a primary source of vulnerabilities and challenges, and a major determinant of a stable and well-converged training. Existing FL reviews describe general foundations, security practices, opportunities, challenges, and applications, without delving into diverse aspects of data and considering problems from the data perspective. They rarely provide a data-lens synthesis that links concrete data properties, split protocols, and defenses to convergence speed and stability. This survey fills that gap with three advances. First, we analyze non-IID into measurable traits and rank their influence on convergence as strong, medium, or light, explaining the mechanisms behind each and reconciling evidence across images, texts, and graphs. Second, we connect experimental splitting practices to the real phenomena they emulate, expose the artifacts they introduce, and show how those artifacts affect target accuracy. Third, we analyze how data-related vulnerabilities and their proposed defenses affect convergence, reporting performance under clean and adversarial conditions to make the convergence-robustness trade-off explicit. To our knowledge, this is the first survey to provide a complete understanding of data-related challenges that govern FL. With clear takeaways distilled for each concern, our work serves as actionable guidance, helping practitioners design their system with predictable convergence and stability.

cs.CR

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present RedditPersona, a modular framework that standardizes these choices: it collects Reddit posts and comments, profiles active users, partitions them under five grouping strategies (subreddit-based, graph-structural, semantic, hybrid, and interaction-based), trains a parameter-efficient adapter per strategy via QLoRA, and evaluates them under a shared metric suite spanning fluency, fidelity, distributional alignment, and community identifiability. Applied to 112 subreddits in the urban well-being domain (301,429 user profiles, 16M+ comments), we find that adapters' behavioral identifiability tracks each strategy's agreement with the subreddit baseline, and that a consistent trade-off between identifiability and distributional similarity to real text holds across all five strategies. The code and configuration files are available at: https://github.com/Ahghaffari/redditpersona.

cs.AI

STM-Graph: A Python Framework for Spatio-Temporal Mapping and Graph Neural Network Predictions

Urban spatio-temporal data present unique challenges for predictive analytics due to their dynamic and complex nature. We introduce STM-Graph, an open-source Python framework that transforms raw spatio-temporal urban event data into graph representations suitable for Graph Neural Network (GNN) training and prediction. STM-Graph integrates diverse spatial mapping methods, urban features from OpenStreetMap, multiple GNN models, comprehensive visualization tools, and a graphical user interface (GUI) suitable for professional and non-professional users. This modular and extensible framework facilitates rapid experimentation and benchmarking. It allows integration of new mapping methods and custom models, making it a valuable resource for researchers and practitioners in urban computing. The source code of the framework and GUI are available at: https://github.com/Ahghaffari/stm_graph and https://github.com/tuminguyen/stm_graph_gui.

cs.LG

Development of Thin-Gap GEM-{\mu}RWELL Hybrid Detectors

Micro Pattern Gaseous Detectors (MPGDs) are used for tracking in High Energy Physics and Nuclear Physics because of their large area, excellent spatial resolution capabilities and low cost. However, for high energy charged particles impacting at a large angle with respect to the axis perpendicular to detector plane, the spatial resolution degrades significantly because of the long trail of ionization charges produced in clusters all along the track in the drift region of the detector. The long ionization charge trail results in registering hits from large number of strips in the readout plane which makes it challenging to precisely reconstruct the particle position using simple center of gravity algorithm. As a result, the larger the drift gap, the more severe the deterioration of spatial resolution for inclined tracks. For the same reason, the position resolution is also severely degraded in a large magnetic field, where the Lorentz E {\times} B effect causes the ionization charges to follow a curved and longer path in the detector gas volume. In this paper, we report on the development of thin-gap MPGDs as a way to maintain excellent spatial resolution capabilities of MPGD detectors over a wide angular range of incoming particles. In a thin-gap MPGD, the thickness of the gas volume in the drift region is reduced from typically {\sim} 3 mm to {\sim} 1 mm or less. We present preliminary test beam results demonstrating the improvement in spatial resolution from {\sim} 400 {\mu}m with a standard 3 mm gap {\mu}RWELL prototype to {\sim} 140 {\mu}m with a double amplification GEM-{\mu}RWELL thin-gap hybrid detector. We also discuss the impact of a thin-gap drift volume on other aspects of the performance of MPGD technologies such as the efficiency and detector stability.

physics.ins-det

Graph-based Gossiping for Communication Efficiency in Decentralized Federated Learning

Federated learning has emerged as a privacy-preserving technique for collaborative model training across heterogeneously distributed silos. Yet, its reliance on a single central server introduces potential bottlenecks and risks of single-point failure. Decentralizing the server, often referred to as decentralized learning, addresses this problem by distributing the server role across nodes within the network. One drawback regarding this pure decentralization is it introduces communication inefficiencies, which arise from increased message exchanges in large-scale setups. However, existing proposed solutions often fail to simulate the real-world distributed and decentralized environment in their experiments, leading to unreliable performance evaluations and limited applicability in practice. Recognizing the lack from prior works, this work investigates the correlation between model size and network latency, a critical factor in optimizing decentralized learning communication. We propose a graph-based gossiping mechanism, where specifically, minimum spanning tree and graph coloring are used to optimize network structure and scheduling for efficient communication across various network topologies and message capacities. Our approach configures and manages subnetworks on real physical routers and devices and closely models real-world distributed setups. Experimental results demonstrate that our method significantly improves communication, compatible with different topologies and data sizes, reducing bandwidth and transfer time by up to circa 8 and 4.4 times, respectively, compared to naive flooding broadcasting methods.

cs.DC

uRWELL detector developments at Jefferson Lab for high luminosity experiments

One of the future plans at Jefferson Lab is running electron scattering experiments with large acceptance detectors at luminosities $> 10^{37}cm^{-2}s^{-1}$. These experiments allow the measurements of the Double Deeply Virtual Compton Scattering (DDVCS) reaction, an important physics process in the formalism of Generalized Parton Distributions, which has never been measured because of its small cross-section. The luminosity upgrade of CLAS12 or the SOLID detector makes Jefferson Lab a unique place to measure DDVCS. One of the important components of these high luminosity detectors is a tracking system that can withstand high rates of $\approx 1MHz/cm^{2}$. The recently developed Micro-Resistive Well (uRWELL) detector technology is a promising option for such a tracking detector by combining good position resolutions, low material budget with simple mechanical construction, and low production costs. In this proceeding, we will discuss recent developments and studies with uRWELL detectors at Jefferson Lab for future upgrades of the CLAS12 detector to study the DDVCS reaction.

physics.ins-det

New Measurements of the Deuteron to Proton F2 Structure Function Ratio

Nucleon structure functions, as measured in lepton-nucleon scattering, have historically provided a critical observable in the study of partonic dynamics within the nucleon. However, at very large parton momenta it is both experimentally and theoretically challenging to extract parton distributions due to the probable onset of non-perturbative contributions and the unavailability of high precision data at critical kinematics. Extraction of the neutron structure and the d-quark distribution have been further challenging due to the necessity of applying nuclear corrections when utilizing scattering data from a deuteron target to extract free neutron structure. However, a program of experiments has been carried out recently at the energy-upgraded Jefferson Lab electron accelerator aimed at significantly reducing the nuclear correction uncertainties on the d-quark distribution function at large partonic momentum. This allows leveraging the vast body of deuterium data covering a large kinematic range to be utilized for d-quark parton distribution function extraction. We present new data from experiment E12-10-002 carried out in Jefferson Lab Hall C on the deuteron to proton cross-section ratio at large BJorken-x. These results significantly improve the precision of existing data, and provide a first look at the expected impact on quark distributions extracted from global parton distribution function fits.

hep-ex

Wait or Not to Wait: Evaluating Trade-Offs between Speed and Precision in Blockchain-based Federated Aggregation

This paper presents a fully coupled blockchain-assisted federated learning architecture that effectively eliminates single points of failure by decentralizing both the training and aggregation tasks across all participants. Our proposed system offers a high degree of flexibility, allowing participants to select shared models and customize the aggregation for local needs, thereby optimizing system performance, including accurate inference results. Notably, the integration of blockchain technology in our work is to promote a trustless environment, ensuring transparency and non-repudiation among participants when abnormalities are detected. To validate the effectiveness, we conducted real-world federated learning deployments on a private Ethereum platform, using two different models, ranging from simple to complex neural networks. The experimental results indicate comparable inference accuracy between centralized and decentralized federated learning settings. Furthermore, our findings indicate that asynchronous aggregation is a feasible option for simple learning models. However, complex learning models require greater training model involvement in the aggregation to achieve high model quality, instead of asynchronous aggregation. With the implementation of asynchronous aggregation and the flexibility to select models, participants anticipate decreased aggregation time in each communication round, while experiencing minimal accuracy trade-off.

cs.DC

Large language models in 6G security: challenges and opportunities

The rapid integration of Generative AI (GenAI) and Large Language Models (LLMs) in sectors such as education and healthcare have marked a significant advancement in technology. However, this growth has also led to a largely unexplored aspect: their security vulnerabilities. As the ecosystem that includes both offline and online models, various tools, browser plugins, and third-party applications continues to expand, it significantly widens the attack surface, thereby escalating the potential for security breaches. These expansions in the 6G and beyond landscape provide new avenues for adversaries to manipulate LLMs for malicious purposes. We focus on the security aspects of LLMs from the viewpoint of potential adversaries. We aim to dissect their objectives and methodologies, providing an in-depth analysis of known security weaknesses. This will include the development of a comprehensive threat taxonomy, categorizing various adversary behaviors. Also, our research will concentrate on how LLMs can be integrated into cybersecurity efforts by defense teams, also known as blue teams. We will explore the potential synergy between LLMs and blockchain technology, and how this combination could lead to the development of next-generation, fully autonomous security solutions. This approach aims to establish a unified cybersecurity strategy across the entire computing continuum, enhancing overall digital security infrastructure.

cs.CR

Situation Awareness for Autonomous Vehicles Using Blockchain-based Service Cooperation

Efficient Vehicle-to-Everything enabling cooperation and enhanced decision-making for autonomous vehicles is essential for optimized and safe traffic. Real-time decision-making based on vehicle sensor data, other traffic data, and environmental and contextual data becomes imperative. As a part of such Intelligent Traffic Systems, cooperation between different stakeholders needs to be facilitated rapidly, reliably, and securely. The Internet of Things provides the fabric to connect these stakeholders who share their data, refined information, and provided services with each other. However, these cloud-based systems struggle to meet the real-time requirements for smart traffic due to long distances across networks. Here, edge computing systems bring the data and services into the close proximity of fast-moving vehicles, reducing information delivery latencies and improving privacy as sensitive data is processed locally. To solve the issues of trust and latency in data sharing between these stakeholders, we propose a decentralized framework that enables smart contracts between traffic data producers and consumers based on blockchain. Autonomous vehicles connect to a local edge server, share their data, or use services based on agreements, for which the cooperating edge servers across the system provide a platform. We set up proof-of-concept experiments with Hyperledger Fabric and virtual cars to analyze the system throughput with secure unicast and multicast data transmissions. Our results show that multicast transmissions in such a scenario boost the throughput up to 2.5 times where the data packets of different sizes can be transmitted in less than one second.

cs.NI

Frequency shifts in the EPR spectrum of $^{39}$K due to spin-exchange collisions with polarized $^3$He and precise $^3$He polarimetry

The Zeeman splittings and EPR frequencies of alkali-metal atoms are shifted in the presence of a polarized noble gas. For a spherical geometry, the shift is enhanced over what is expected classically by a dimensionless atomic parameter $κ_0$ that is unique to each alkali-metal atom - noble-gas pair. We present a precise measurement of $κ_0$ for the $^{39}$K-$^3$He system with a relative accuracy of better than 1\%. A critical component of achieving sub-percent accuracy involved characterizing the shape of our samples using both MRI and CT medical-imaging techniques. The parameter $κ_0$ plays an important role in establishing the absolute polarization of $^3$He in a variety of contexts, including polarized targets for electron scattering experiments and MRI of the gas space of the lungs. Our measurement more than doubles the accuracy possible when using $κ_0$ for polarimetry purposes. Just as important, the work presented here represents the first {\it direct} measurement of $κ_0$ for the $^{39}$K-$^3$He system; previous values for $κ_0$ in the $^{39}$K-$^3$He system relied on a chain of measurements that were benchmarked by previous measurements of $κ_0$ in the Rb-$^3$He system.

physics.atom-ph

Reducing latency and bandwidth for video streaming using keypoint extraction and digital puppetry

COVID-19 has made video communication one of the most important modes of information exchange. While extensive research has been conducted on the optimization of the video streaming pipeline, in particular the development of novel video codecs, further improvement in the video quality and latency is required, especially under poor network conditions. This paper proposes an alternative to the conventional codec through the implementation of a keypoint-centric encoder relying on the transmission of keypoint information from within a video feed. The decoder uses the streamed keypoints to generate a reconstruction preserving the semantic features in the input feed. Focusing on video calling applications, we detect and transmit the body pose and face mesh information through the network, which are displayed at the receiver in the form of animated puppets. Using efficient pose and face mesh detection in conjunction with skeleton-based animation, we demonstrate a prototype requiring lower than 35 kbps bandwidth, an order of magnitude reduction over typical video calling systems. The added computational latency due to the mesh extraction and animation is below 120ms on a standard laptop, showcasing the potential of this framework for real-time applications. The code for this work is available at https://github.com/shubhamchandak94/digital-puppetry/.

eess.IV

Multilingual Schema Matching for Wikipedia Infoboxes

Recent research has taken advantage of Wikipedia's multilingualism as a resource for cross-language information retrieval and machine translation, as well as proposed techniques for enriching its cross-language structure. The availability of documents in multiple languages also opens up new opportunities for querying structured Wikipedia content, and in particular, to enable answers that straddle different languages. As a step towards supporting such queries, in this paper, we propose a method for identifying mappings between attributes from infoboxes that come from pages in different languages. Our approach finds mappings in a completely automated fashion. Because it does not require training data, it is scalable: not only can it be used to find mappings between many language pairs, but it is also effective for languages that are under-represented and lack sufficient training samples. Another important benefit of our approach is that it does not depend on syntactic similarity between attribute names, and thus, it can be applied to language pairs that have distinct morphologies. We have performed an extensive experimental evaluation using a corpus consisting of pages in Portuguese, Vietnamese, and English. The results show that not only does our approach obtain high precision and recall, but it also outperforms state-of-the-art techniques. We also present a case study which demonstrates that the multilingual mappings we derive lead to substantial improvements in answer quality and coverage for structured queries over Wikipedia content.

cs.DB