SearcharxivSearch

arXiv subjects

Xuhong Li

Publications and source records attributed to Xuhong Li.

At least 19 recordsLinked to original sources

Cellular Signal Constructed Convolutional Vision Transformer for High Accuracy Positioning

Modern cellular systems employ wide bandwidths and large antenna arrays to meet high data rate requirements. The high spatial and temporal resolution for communication also enables high-accuracy positioning as an ancillary benefit. Standard convolutional neural networks (CNNs) and vision Transformers have demonstrated excellent performance in positioning by leveraging delay-angle domain channel representations. However, they still face practical challenges in complicated cellular environments with low signal-to-noise ratios and severe inter-cell interference. This paper proposes a hybrid convolutional vision Transformer (ConViT) architecture that integrates the local receptive fields of CNNs to suppress local noise and employs Transformers to capture global attention among different multipath components. Various fusion strategies for combining signals from multiple distributed base stations are also evaluated. An extended Kalman filter with sensor fusion is applied to further mitigate long tail fluctuations of model estimates. Comprehensive validation is conducted with commercial long-term-evolution signals received by a large antenna array in urban environments with non line-of-sight signals and strong inter-cell interference. ConViT achieves a distance root mean square error (RMSE) of 3.46 meters and a yaw RMSE of 2.54 degrees, significantly outperforming benchmark models, while maintaining a lower parameter count and reduced computational complexity. Finally, a correspondence analysis between delay-angle power distributions and Transformer attention weights demonstrates the interpretability of the model.

eess.SP

Probabilistic Occupancy Grid for Radio-Based SLAM

Sensing is an integral part of 6G and beyond systems, providing exceptional environmental perception along with communication. Radio frequency (RF)-based sensing often relies on simplified geometric assumptions (e.g., point scatterers or planar surfaces) to model specular multipath and keep inference tractable. However, such representations are limited in their ability to capture extended objects with complex geometries and properties. This paper presents a probabilistic occupancy grid framework for radio-based simultaneous localization and mapping (SLAM), jointly reconstructing geometric structures and their RF-related properties. The proposed occupancy grid map representation is integrated into a multipath-based SLAM formulation to enable simultaneous mobile-agent localization and environment mapping using multipath measurements. To connect RF measurements with the grid map, a surface model is employed to describe candidate reflection paths, while occupancy grid cell states capture measurement uncertainties and fine-grained geometric details. RF-related object properties are represented through reflection coefficients. The proposed framework offers a principled, proof-of-concept approach to physically interpretable radio-based mapping, and simulation results demonstrate accurate reconstruction of geometry and material properties, as well as high-accuracy localization. In addition, the results highlight the potential to use prior occupancy maps obtained from other radio devices or complementary sensors for subsequent map extension and refinement.

eess.SP

Adaptive Multipath-Based SLAM for Distributed MIMO Systems

Localizing users and mapping the environment using radio signals is a key task in emerging applications such as low-latency communications and safety-critical navigation. Recently introduced multipath-based SLAM methods can jointly localize a mobile agent and map reflective surfaces in radio frequency (RF) environments. Most existing methods assume that map features and their corresponding RF propagation paths are statistically independent. This assumption neglects inherent dependencies that arise when a single reflective surface contributes to multiple propagation paths or when an agent communicates with multiple base stations. Existing approaches that aim to fuse information across propagation paths are further limited by their inability to perform ray tracing in RF environments with nonconvex geometries. In this paper, we propose a Bayesian multipath-based SLAM method for distributed MIMO systems that addresses these limitations. We exploit amplitude statistics to establish adaptive, time-varying detection probabilities. Based on the resulting 'soft' ray-tracing strategy, the proposed method can fuse information across propagation paths in RF environments with nonconvex geometries. A Bayesian estimation framework for the joint estimation of map features and agent state is developed by applying the message passing rules of the sum-product algorithm to a factor graph representation of the proposed statistical model. We further introduce a new initialization procedure for reflective surfaces that enables the introduction of new surface states even when measurements arise solely from double-bounce paths. The proposed method is validated using both synthetic and real RF measurements obtained in challenging scenarios with nonconvex geometries and OLoS conditions. The results demonstrate that it provides accurate localization and mapping performance and approaches the posterior CRLBs.

eess.SP

AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback

Large Language Models (LLMs) have achieved significant success in complex reasoning but remain bottlenecked by reliance on expert-annotated data and external verifiers. While existing self-evolution paradigms aim to bypass these constraints, they often fail to identify the optimal learning zone and risk reinforcing collective hallucinations and incorrect priors through flawed internal feedback. To address these challenges, we propose \underline{A}utonomous \underline{E}volutionary \underline{R}easoning \underline{O}ptimization (AERO), an unsupervised framework that achieves autonomous reasoning evolution by internalizing self-questioning, answering, and criticism within a synergistic dual-loop system. Inspired by the \textit{Zone of Proximal Development (ZPD)} theory, AERO utilizes entropy-based positioning to target the ``solvability gap'' and employs Independent Counterfactual Correction for robust verification. Furthermore, we introduce a Staggered Training Strategy to synchronize capability growth across functional roles and prevent curriculum collapse. Extensive evaluations across nine benchmarks spanning three domains demonstrate that AERO achieves average performance improvements of 4.57\% on Qwen3-4B-Base and 5.10\% on Qwen3-8B-Base, outperforming competitive baselines. Code is available at https://github.com/mira-ai-lab/AERO.

cs.CL

Posterior Cramér-Rao Bounds on Localization and Mapping Errors in Distributed MIMO SLAM

Radio-frequency simultaneous localization and mapping (RF-SLAM) methods jointly infer the position of mobile transmitters and receivers in wireless networks, together with a geometric map of the propagation environment. An inferred map of specular surfaces can be used to exploit non-line-of-sight components of the multipath channel to increase robustness, bypass obstructions, and improve overall communication and positioning performance. While performance bounds for user location are well established, the literature lacks performance bounds for map information. This paper derives the mapping error bound (MEB), i.e., the posterior Cramér-Rao lower bound on the position and orientation of specular surfaces, for RF-SLAM. In particular, we consider a very general scenario with single- and double-bounce reflections, as well as distributed anchors. We demonstrate numerically that a state-of-the-art RF-SLAM algorithm asymptotically converges to this MEB. The bounds assess not only the localization (position and orientation) but also the mapping performance of RF-SLAM algorithms in terms of global features.

eess.SP

Robust Localization in Modern Cellular Networks using Global Map Features

Radio frequency (RF) signal-based localization using modern cellular networks has emerged as a promising solution to accurately locate objects in challenging environments. One of the most promising solutions for situations involving obstructed-line-of-sight (OLoS) and multipath propagation is multipathbased simultaneous localization and mapping (MP-SLAM) that employs map features (MFs), such as virtual anchors. This paper presents an extended MP-SLAM method that is augmented with a global map feature (GMF) repository. This repository stores consistent MFs of high quality that are collected during prior traversals. We integrate these GMFs back into the MP-SLAM framework via a probability hypothesis density (PHD) filter, which propagates GMF intensity functions over time. Extensive simulations, together with a challenging real-world experiment using LTE RF signals in a dense urban scenario with severe multipath propagation and inter-cell interference, demonstrate that our framework achieves robust and accurate localization, thereby showcasing its effectiveness in realistic modern cellular networks such as 5G or future 6G networks. It outperforms conventional proprioceptive sensor-based localization and conventional MP-SLAM methods, and achieves reliable localization even under adverse signal conditions.

eess.SP

Low-latency D-MIMO Localization using Distributed Scalable Message-Passing Algorithm

Distributed MIMO and integrated sensing and communication are expected to be key technologies in future wireless systems, enabling reliable, low-latency communication and accurate localization. Dedicated localization solutions must support distributed architecture, provide scalability across different system configurations and meet strict latency requirements. We present a scalable message-passing localization method and architecture co-designed for a panel-based distributed MIMO system and network topology, in which interconnected units operate without centralized processing. This method jointly detects line-of-sight paths to distributed units from multipath measurements in dynamic scenarios, localizes the agent, and achieves very low latency. Additionally, we introduce a cycle-accurate system latency model based on implemented FPGA operations, and show important insights into processing latency and hardware utilization and system-level trade-offs. We compare our method to a multipath-based localization method and show that it can achieve similar localization performance, with wide enough distribution of array elements, while offering lower latency and computational complexity.

eess.SP

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning

Dataset diversity plays a pivotal role for the successful training of many machine learning models, particularly in the supervised fine-tuning (SFT) stage of large language model (LLM) development. Despite increasing recognition of its importance, systematic analyses of dataset diversity still remain underexplored. To address this gap, this work presents a systematic taxonomy of existing diversity-control strategies, which primarily focus on the instruction component, operating at either macroscopic (entire instruction semantics) or mesoscopic levels (instruction units), and furthermore introduces a novel analysis of microscopic diversity within the response component, specifically analyzing the statistical distribution of tokens in SFT training samples. In the experimental evaluation, we construct fixed-size datasets (e.g., 10,000 samples each) from a corpus of 117,000 open-source SFT samples, incorporating six distinct diversity-control strategies spanning macro-, meso-, and microscopic levels applied to both instructions and responses. We then fine-tune LLMs on these datasets to assess the six diversity-control strategies. Results reveal that while macroscopic and mesoscopic strategies lead to higher performance with increasing diversity, the microscopic strategy in responses exhibits both a stronger correlation between model performance and the degree of diversity and superior performance with maximum diversity across all strategies. These findings offer actionable insights for constructing high-performance SFT datasets.

cs.CL

SOLA-GCL: Subgraph-Oriented Learnable Augmentation Method for Graph Contrastive Learning

Graph contrastive learning has emerged as a powerful technique for learning graph representations that are robust and discriminative. However, traditional approaches often neglect the critical role of subgraph structures, particularly the intra-subgraph characteristics and inter-subgraph relationships, which are crucial for generating informative and diverse contrastive pairs. These subgraph features are crucial as they vary significantly across different graph types, such as social networks where they represent communities, and biochemical networks where they symbolize molecular interactions. To address this issue, our work proposes a novel subgraph-oriented learnable augmentation method for graph contrastive learning, termed SOLA-GCL, that centers around subgraphs, taking full advantage of the subgraph information for data augmentation. Specifically, SOLA-GCL initially partitions a graph into multiple densely connected subgraphs based on their intrinsic properties. To preserve and enhance the unique characteristics inherent to subgraphs, a graph view generator optimizes augmentation strategies for each subgraph, thereby generating tailored views for graph contrastive learning. This generator uses a combination of intra-subgraph and inter-subgraph augmentation strategies, including node dropping, feature masking, intra-edge perturbation, inter-edge perturbation, and subgraph swapping. Extensive experiments have been conducted on various graph learning applications, ranging from social networks to molecules, under semi-supervised learning, unsupervised learning, and transfer learning settings to demonstrate the superiority of our proposed approach over the state-of-the-art in GCL.

cs.LG

Pre-trained Molecular Language Models with Random Functional Group Masking

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to understand and predict molecular properties and activities, a critical step in fields like drug discovery and materials science. To further improve performance, researchers have introduced graph neural networks with graph-based molecular representations, such as GEM, incorporating the topology, geometry, 2D or even 3D structures of molecules into pre-training. While most of molecular graphs in existing studies were automatically converted from SMILES sequences, it is to assume that transformer-based language models might be able to implicitly learn structure-aware representations from SMILES sequences. In this paper, we propose \ours{} -- a SMILES-based \underline{\em M}olecular \underline{\em L}anguage \underline{\em M}odel, which randomly masking SMILES subsequences corresponding to specific molecular \underline{\em F}unctional \underline{\em G}roups to incorporate structure information of atoms during the pre-training phase. This technique aims to compel the model to better infer molecular structures and properties, thus enhancing its predictive capabilities. Extensive experimental evaluations across 11 benchmark classification and regression tasks in the chemical domain demonstrate the robustness and superiority of \ours{}. Our findings reveal that \ours{} outperforms existing pre-training models, either based on SMILES or graphs, in 9 out of the 11 downstream tasks, ranking as a close second in the remaining ones.

q-bio.BM

Converging Paradigms: The Synergy of Symbolic and Connectionist AI in LLM-Empowered Autonomous Agents

This article explores the convergence of connectionist and symbolic artificial intelligence (AI), from historical debates to contemporary advancements. Traditionally considered distinct paradigms, connectionist AI focuses on neural networks, while symbolic AI emphasizes symbolic representation and logic. Recent advancements in large language models (LLMs), exemplified by ChatGPT and GPT-4, highlight the potential of connectionist architectures in handling human language as a form of symbols. The study argues that LLM-empowered Autonomous Agents (LAAs) embody this paradigm convergence. By utilizing LLMs for text-based knowledge modeling and representation, LAAs integrate neuro-symbolic AI principles, showcasing enhanced reasoning and decision-making capabilities. Comparing LAAs with Knowledge Graphs within the neuro-symbolic AI theme highlights the unique strengths of LAAs in mimicking human-like reasoning processes, scaling effectively with large datasets, and leveraging in-context samples without explicit re-training. The research underscores promising avenues in neuro-vector-symbolic integration, instructional encoding, and implicit reasoning, aimed at further enhancing LAA capabilities. By exploring the progression of neuro-symbolic AI and proposing future research trajectories, this work advances the understanding and development of AI technologies.

cs.AI

Tokenization Falling Short: On Subword Robustness in Large Language Models

Language models typically tokenize raw text into sequences of subword identifiers from a predefined vocabulary, a process inherently sensitive to typographical errors, length variations, and largely oblivious to the internal structure of tokens--issues we term the curse of tokenization. In this study, we delve into these drawbacks and demonstrate that large language models (LLMs) remain susceptible to these problems. This study systematically investigates these challenges and their impact on LLMs through three critical research questions: (1) complex problem solving, (2) token structure probing, and (3) resilience to typographical variation. Our findings reveal that scaling model parameters can mitigate the issue of tokenization; however, LLMs still suffer from biases induced by typos and other text format variations. Our experiments show that subword regularization such as BPE-dropout can mitigate this issue. We release our evaluation code and data at https://github.com/FloatAI/TKEval.

cs.CL

Towards Automated Data Sciences with Natural Language and SageCopilot: Practices and Lessons Learned

While the field of NL2SQL has made significant advancements in translating natural language instructions into executable SQL scripts for data querying and processing, achieving full automation within the broader data science pipeline - encompassing data querying, analysis, visualization, and reporting - remains a complex challenge. This study introduces SageCopilot, an advanced, industry-grade system system that automates the data science pipeline by integrating Large Language Models (LLMs), Autonomous Agents (AutoAgents), and Language User Interfaces (LUIs). Specifically, SageCopilot incorporates a two-phase design: an online component refining users' inputs into executable scripts through In-Context Learning (ICL) and running the scripts for results reporting & visualization, and an offline preparing demonstrations requested by ICL in the online phase. A list of trending strategies such as Chain-of-Thought and prompt-tuning have been used to augment SageCopilot for enhanced performance. Through rigorous testing and comparative analysis against prompt-based solutions, SageCopilot has been empirically validated to achieve superior end-to-end performance in generating or executing scripts and offering results with visualization, backed by real-world datasets. Our in-depth ablation studies highlight the individual contributions of various components and strategies used by SageCopilot to the end-to-end correctness for data sciences.

cs.AI

When Search Engine Services meet Large Language Models: Visions and Challenges

Combining Large Language Models (LLMs) with search engine services marks a significant shift in the field of services computing, opening up new possibilities to enhance how we search for and retrieve information, understand content, and interact with internet services. This paper conducts an in-depth examination of how integrating LLMs with search engines can mutually benefit both technologies. We focus on two main areas: using search engines to improve LLMs (Search4LLM) and enhancing search engine functions using LLMs (LLM4Search). For Search4LLM, we investigate how search engines can provide diverse high-quality datasets for pre-training of LLMs, how they can use the most relevant documents to help LLMs learn to answer queries more accurately, how training LLMs with Learning-To-Rank (LTR) tasks can enhance their ability to respond with greater precision, and how incorporating recent search results can make LLM-generated content more accurate and current. In terms of LLM4Search, we examine how LLMs can be used to summarize content for better indexing by search engines, improve query outcomes through optimization, enhance the ranking of search results by analyzing document relevance, and help in annotating data for learning-to-rank tasks in various learning contexts. However, this promising integration comes with its challenges, which include addressing potential biases and ethical issues in training models, managing the computational and other costs of incorporating LLMs into search services, and continuously updating LLM training with the ever-changing web content. We discuss these challenges and chart out required research directions to address them. We also discuss broader implications for service computing, such as scalability, privacy concerns, and the need to adapt search engine architectures for these advanced models.

cs.IR

P2ANet: A Dataset and Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos

While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark \TheName{} for \emph{\underline{P}}ing \emph{\underline{P}}ong-\emph{\underline{A}}ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems -- \emph{action localization} and \emph{action recognition}. We evaluate a number of commonly-seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using \TheName{} for both problems, under various settings. These models can only achieve 48\% area under the AR-AN curve for localization and 82\% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that \TheName{} is still a challenging task and can be used as a special benchmark for dense action detection from videos.

cs.CV

HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization

Large language models (LLMs) have made significant progress in generating codes from textual prompts. However, existing benchmarks have mainly concentrated on translating English prompts to multilingual codes or have been constrained to very limited natural languages (NLs). These benchmarks have overlooked the vast landscape of massively multilingual NL to multilingual code, leaving a critical gap in the evaluation of multilingual LLMs. In response, we introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency. HumanEval-XL establishes connections between 23 NLs and 12 programming languages (PLs), and comprises of a collection of 22,080 prompts with an average of 8.33 test cases. By ensuring parallel data across multiple NLs and PLs, HumanEval-XL offers a comprehensive evaluation platform for multilingual LLMs, allowing the assessment of the understanding of different NLs. Our work serves as a pioneering step towards filling the void in evaluating NL generalization in the area of multilingual code generation. We make our evaluation code and data publicly available at \url{https://github.com/FloatAI/humaneval-xl}.

cs.CL

Interpretable Machine Learning for Weather and Climate Prediction: A Survey

Advanced machine learning models have recently achieved high predictive accuracy for weather and climate prediction. However, these complex models often lack inherent transparency and interpretability, acting as "black boxes" that impede user trust and hinder further model improvements. As such, interpretable machine learning techniques have become crucial in enhancing the credibility and utility of weather and climate modeling. In this survey, we review current interpretable machine learning approaches applied to meteorological predictions. We categorize methods into two major paradigms: 1) Post-hoc interpretability techniques that explain pre-trained models, such as perturbation-based, game theory based, and gradient-based attribution methods. 2) Designing inherently interpretable models from scratch using architectures like tree ensembles and explainable neural networks. We summarize how each technique provides insights into the predictions, uncovering novel meteorological relationships captured by machine learning. Lastly, we discuss research challenges around achieving deeper mechanistic interpretations aligned with physical principles, developing standardized evaluation benchmarks, integrating interpretability into iterative model development workflows, and providing explainability for large foundation models.

physics.ao-ph

A Wideband Distributed Massive MIMO Channel Sounder for Communication and Sensing

Channel sounding is a vital step in understanding wireless channels for the design and deployment of wireless communication systems. In this paper, we present the design and implementation of a coherent distributed massive MIMO channel sounder operating at 5-6 GHz with a bandwidth of 400 MHz based on the NI USRP X410. Through the integration of transceiver chains and RF switches, the design facilitates the use of a larger number of antennas without significant compromise in dynamic capability. Our current implementation is capable of measuring thousands of antenna combinations within tens of milliseconds. Every radio frequency switch is seamlessly integrated with a 16-element antenna array, making the antennas more practical to be transported and flexibly distributed. In addition, the channel sounder features real-time processing to reduce the data stream to the host computer and increase the signal-to-noise ratio. The design and implementation are verified through two measurements in an indoor laboratory environment. The first measurement entails a single-antenna robot as transmitter and 128 distributed receiving antennas. The second measurement demonstrates a passive sensing scenario with a walking person. We evaluate the results of both measurements using the super-resolution algorithm SAGE. The results demonstrate the great potential of the presented sounding system for providing high-quality radio channel measurements, contributing to high-resolution channel estimation, characterization, and active and passive sensing in realistic and dynamic scenarios.

eess.SP