Searcharxiv⌕ Search

arXiv subjects

He Chen

Publications and source records attributed to He Chen.

At least 37 records · Page 2Linked to original sources

EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting

Recent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direct transfer to remote sensing (RS) is hindered by heterogeneous sensing physics, diverse modalities, and unique spatial scales. Existing RS MLLMs are mainly limited to optical imagery and plain language interaction, preventing flexible and scalable real-world applications. In this article, EarthGPT-X is proposed, the first flexible spatial MLLM that unifies multi-source RS imagery comprehension and accomplishes both coarse-grained and fine-grained visual tasks under diverse visual prompts in a single framework. Distinct from prior models, EarthGPT-X introduces: 1) a dual-prompt mechanism combining text instructions with various visual prompts (i.e., point, box, and free-form) to mimic the versatility of referring in human life; 2) a comprehensive multi-source multi-level prompting dataset, the model advances beyond holistic image understanding to support hierarchical spatial reasoning, including scene-level understanding and fine-grained object attributes and relational analysis; 3) a cross-domain one-stage fusion training strategy, enabling efficient and consistent alignment across modalities and tasks. Extensive experiments demonstrate that EarthGPT-X substantially outperforms prior nature and RS MLLMs, establishing the first framework capable of multi-source, multi-task, and multi-level interpretation using visual prompting in RS scenarios.

cs.CV↗

Design of quasi phase matching crystal based on differential gray wolf algorithm

This paper focuses on the key problem in the development of nonlinear optical technology, the performance optimization of aperiodically polarized crystals. The performance of the crystal depends on the precise control of the micro distribution of crystal domains, but its optimization belongs to the high-dimensional discrete combination "NP hard" problem. The traditional algorithm has the bottleneck of slow convergence and easy to fall into local optimization, while the heuristic methods such as genetic algorithm are limited by the CPU serial calculation and inefficient. In order to solve the above challenges, this paper proposes the fusion scheme of hwsda hybrid optimization algorithm and GPU parallel acceleration technology: the differential evolution algorithm (DE) is used to realize the global search, and the gray wolf optimization algorithm (GWO) is used to strengthen the local search and convergence speed, and the two coordinate to balance the global and local optimization requirements; At the same time, it relies on GPU multi-core architecture to realize thread level parallel computing and improve optimization efficiency. This scheme effectively breaks through the optimization problem of high-dimensional discrete space, improves the accuracy of crystal domain control, improves the efficiency of quasi phase matching design by hundreds to thousands of times compared with traditional CPU serial computing, provides a new paradigm for the design of complex nonlinear optical devices, and helps promote the performance breakthrough and industrial application of related devices in the fields of quantum optics and laser processing.

cs.DC↗

Set Smoothness Unlocks Clarke Hyper-stationarity in Bilevel Optimization

Solving bilevel optimization (BLO) problems to global optimality is generally intractable. A common surrogate is to compute a hyper-stationary point -- a stationary point of the hyper-objective function obtained by minimizing or maximizing the upper-level objective over the lower-level solution set. Existing methods, however, either provide weak notions of stationarity or require restrictive assumptions to guarantee the smoothness of hyper-objective functions. In this paper, we eliminate these impractical assumptions and show that strong (Clarke) hyper-stationarity remains computable even when the hyper-objective is nonsmooth. Our key ingredient is a new structural property, called set smoothness, which captures the variational dependence of the lower-level solution set on the upper-level variable. We prove that this property holds for a broad class of BLO problems and ensures weak convexity (resp. concavity) of pessimistic (resp. optimistic) hyper-objective functions. Building on this foundation, we show that a zeroth-order algorithm that computes approximate Clarke hyper-stationary points with non-asymptotic convergence guarantees. To the best of our knowledge, this is the first computational guarantee for Clarke-type stationarity in nonsmooth BLO. Beyond this specific application, the set smoothness property emerges as a structural concept of independent interest, with potential to inform the analysis of broader classes of optimization and variational problems.

math.OC↗

CIRSense: Rethinking WiFi Sensing with Channel Impulse Response

WiFi sensing based on channel state information (CSI) collected from commodity WiFi devices has shown great potential across a wide range of applications, including vital sign monitoring and indoor localization. Existing WiFi sensing approaches typically estimate motion information directly from CSI. However, they often overlook the inherent advantages of channel impulse response (CIR), a delay-domain representation that enables more intuitive and principled motion sensing by naturally concentrating motion energy and separating multipath components. Motivated by this, we revisit WiFi sensing and introduce CIRSense, a new framework that enhances the performance and interpretability of WiFi sensing with CIR. CIRSense is built upon a new motion model that characterizes fractional delay effects, a fundamental challenge in CIR-based sensing. This theoretical model underpins technical advances for the three challenges in WiFi sensing: hardware distortion compensation, high-resolution distance estimation, and subcarrier aggregation for extended range sensing. CIRSense, operating with a 160 MHz channel bandwidth, demonstrates versatile sensing capabilities through its dual-mode design, achieving a mean error of approximately 0.25 bpm in respiration monitoring and 0.09 m in distance estimation. Comprehensive evaluations across residential spaces, far-range scenarios, and multi-target settings demonstrate CIRSense's superior performance over state-of-the-art CSI-based baselines. Notably, at a challenging sensing distance of 20 m, CIRSense achieves at least 3x higher average accuracy with more than 4.5x higher computational efficiency.

eess.SP↗

Accelerated Price Adjustment for Fisher Markets with Exact Recovery of Competitive Equilibrium

The canonical price-adjustment process, tâtonnement, typically fails to converge to the exact competitive equilibrium (CE) and requires a high iteration complexity of $\tilde{\mathcal{O}}(1/ε)$ to compute $ε$-CE prices in widely studied linear and quasi-linear Fisher markets. This paper proposes refined price-adjustment processes to overcome these limitations. By formulating the task of finding CE of a (quasi-)linear Fisher market as a strongly convex nonsmooth minimization problem, we develop a novel accelerated price-adjustment method (APM) that finds an $ε$-CE price in $\tilde{\mathcal{O}}(1/\sqrtε)$ lightweight iterations, which significantly improves upon the iteration complexities of tâtonnement methods. Furthermore, through our new formulation, we construct a recovery oracle that maps approximate CE prices to exact CE prices at a low computational cost. By coupling this recovery oracle with APM, we obtain an adaptive price-adjustment method whose iterates converge to CE prices in finite steps. To the best of our knowledge, this is the first convergence guarantee to exact CE for price-adjustment methods in linear and quasi-linear Fisher markets. Our developments pave the way for efficient lightweight computation of CE prices. We also present numerical results to demonstrate the fast convergence of the proposed methods and the efficient recovery of CE prices.

math.OC↗

Domino: Dominant Path-based Compensation for Hardware Impairments in Modern WiFi Sensing

WiFi sensing faces a critical reliability challenge due to hardware-induced RF distortions, especially with modern, market-dominant WiFi cards supporting 802.11ac/ax protocols. These cards employ sensitive automatic gain control and separate RF chains, introducing complex and dynamic distortions that render existing compensation methods ineffective. In this paper, we introduce Domino, a new framework that transforms channel state information (CSI) into channel impulse response (CIR) and leverages it for precise distortion compensation. Domino is built on the key insight that hardware-induced distortions impact all signal paths uniformly, allowing the dominant static path to serve as a reliable reference for effective compensation through delay-domain processing. Real-world respiration monitoring experiments show that Domino achieves at least 2x higher mean accuracy over existing methods, maintaining robust performance with a median error below 0.24 bpm, even using a single antenna in both direct line-of-sight and obstructed scenarios.

eess.SP↗

Can Large Language Models Identify Materials from Radar Signals?

Accurately identifying the material composition of objects is a critical capability for AI robots powered by large language models (LLMs) to perform context-aware manipulation. Radar technologies offer a promising sensing modality for material recognition task. When combined with deep learning, radar technologies have demonstrated strong potential in identifying the material of various objects. However, existing radar-based solutions are often constrained to closed-set object categories and typically require task-specific data collection to train deep learning models, largely limiting their practical applicability. This raises an important question: Can we leverage the powerful reasoning capabilities of pre-trained LLMs to directly infer material composition from raw radar signals? Answering this question is non-trivial due to the inherent redundancy of radar signals and the fact that pre-trained LLMs have no prior exposure to raw radar data during training. To address this, we introduce LLMaterial, the first study to investigate the feasibility of using LLM to identify materials directly from radar signals. First, we introduce a physics-informed signal processing pipeline that distills high-redundancy radar raw data into a set of compact intermediate parameters that encapsulate the material's intrinsic characteristics. Second, we adopt a retrieval-augmented generation (RAG) strategy to provide the LLM with domain-specific knowledge, enabling it to interpret and reason over the extracted intermediate parameters. Leveraging this integration, the LLM is empowered to perform step-by-step reasoning on the condensed radar features, achieving open-set material recognition directly from raw radar signals. Preliminary results show that LLMaterial can effectively distinguish among a variety of common materials, highlighting its strong potential for real-world material identification applications.

eess.SP↗

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation

Embodied agents exhibit immense potential across a multitude of domains, making the assurance of their behavioral safety a fundamental prerequisite for their widespread deployment. However, existing research predominantly concentrates on the security of general large language models, lacking specialized methodologies for establishing safety benchmarks and input moderation tailored to embodied agents. To bridge this gap, this paper introduces a novel input moderation framework, meticulously designed to safeguard embodied agents. This framework encompasses the entire pipeline, including taxonomy definition, dataset curation, moderator architecture, model training, and rigorous evaluation. Notably, we introduce EAsafetyBench, a meticulously crafted safety benchmark engineered to facilitate both the training and stringent assessment of moderators specifically designed for embodied agents. Furthermore, we propose Pinpoint, an innovative prompt-decoupled input moderation scheme that harnesses a masked attention mechanism to effectively isolate and mitigate the influence of functional prompts on moderation tasks. Extensive experiments conducted on diverse benchmark datasets and models validate the feasibility and efficacy of the proposed approach. The results demonstrate that our methodologies achieve an impressive average detection accuracy of 94.58%, surpassing the performance of existing state-of-the-art techniques, alongside an exceptional moderation processing time of merely 0.002 seconds per instance.

cs.AI↗

Optimizing Information Freshness of IEEE 802.11ax Uplink OFDMA-Based Random Access

The latest WiFi standard, IEEE 802.11ax (WiFi 6), introduces a novel uplink random access mechanism called uplink orthogonal frequency division multiple access-based random access (UORA). While existing work has evaluated the performance of UORA using conventional performance metrics, such as throughput and delay, its information freshness performance has not been thoroughly investigated in the literature. This is of practical significance as WiFi 6 and beyond are expected to support real-time applications. This paper presents the first attempt to fill this gap by investigating the information freshness, quantified by the Age of Information (AoI) metric, in UORA networks. We establish an analytical framework comprising two discrete-time Markov chains (DTMCs) to characterize the transmission states of stations (STAs) in UORA networks. Building on the formulated DTMCs, we derive an analytical expression for the long-term average AoI (AAoI), facilitating the optimization of UORA parameters for enhanced AoI performance through exhaustive search. To gain deeper design insights and improve the effectiveness of UORA parameter optimization, we derive a closed-form expression for the AAoI and its approximated lower bound for a simplified scenario characterized by a fixed backoff contention window and generate-at-will status updates. By analyzing the approximated lower bound of the AAoI, we propose efficient UORA parameter optimization algorithms that can be realized with only a few comparisons of different possible values of the parameters to be optimized. Simulation results validate our analysis and demonstrate that the AAoI achieved through our proposed parameter optimization algorithm closely approximates the optimal AoI performance obtained via exhaustive search, outperforming the round-robin and max-AoI policies in large and low-traffic networks.

cs.IT↗

FuseGrasp: Radar-Camera Fusion for Robotic Grasping of Transparent Objects

Transparent objects are prevalent in everyday environments, but their distinct physical properties pose significant challenges for camera-guided robotic arms. Current research is mainly dependent on camera-only approaches, which often falter in suboptimal conditions, such as low-light environments. In response to this challenge, we present FuseGrasp, the first radar-camera fusion system tailored to enhance the transparent objects manipulation. FuseGrasp exploits the weak penetrating property of millimeter-wave (mmWave) signals, which causes transparent materials to appear opaque, and combines it with the precise motion control of a robotic arm to acquire high-quality mmWave radar images of transparent objects. The system employs a carefully designed deep neural network to fuse radar and camera imagery, thereby improving depth completion and elevating the success rate of object grasping. Nevertheless, training FuseGrasp effectively is non-trivial, due to limited radar image datasets for transparent objects. We address this issue utilizing large RGB-D dataset, and propose an effective two-stage training approach: we first pre-train FuseGrasp on a large public RGB-D dataset of transparent objects, then fine-tune it on a self-built small RGB-D-Radar dataset. Furthermore, as a byproduct, FuseGrasp can determine the composition of transparent objects, such as glass or plastic, leveraging the material identification capability of mmWave radar. This identification result facilitates the robotic arm in modulating its grip force appropriately. Extensive testing reveals that FuseGrasp significantly improves the accuracy of depth reconstruction and material identification for transparent objects. Moreover, real-world robotic trials have confirmed that FuseGrasp markedly enhances the handling of transparent items. A video demonstration of FuseGrasp is available at https://youtu.be/MWDqv0sRSok.

cs.RO↗

99.9%-fidelity in measuring a superconducting qubit

Despite the significant progress in superconducting quantum computation over the past years, quantum state measurement still lags nearly an order of magnitude behind quantum gate operations in speed and fidelity. The main challenge is that the strong coupling and readout signal used to probe the quantum state may also introduce additional channels which may cause qubit state transitions. Here, we design a novel architecture to implement the long-sought longitudinal interaction scheme between qubits and resonators. This architecture not only provides genuine longitudinal interaction by eliminating residual transversal couplings, but also introduces proper nonlinearity to the resonator that can further minimize decay error and measurement-induced excitation error. Our experimental results demonstrate a measurement fidelity of 99.8% in 202 ns without the need for any first-stage amplification. After subtracting the residual preparation errors, the pure measurement fidelity is above 99.9%. Our scheme is compatible with the multiplexing readout scheme and can be used for quantum error correction.

quant-ph↗

Optimizing AoI at Query in Multiuser Wireless Uplink Networks: A Whittle Index Approach

In this paper, we explore how to schedule multiple users to optimize information freshness in a pull-based wireless network, where the status updates from users are requested by randomly arriving queries at the destination. We use the age of information at query (QAoI) to characterize the performance of information freshness. Such a decision-making problem is naturally modeled as a Markov decision process (MDP), which, however, is prohibitively high to be solved optimally by the standard method due to the curse of dimensionality. To address this issue, we employ Whittle index approach, which allows us to decouple the original MDP into multiple sub-MDPs by relaxing the scheduling constraints. However, the binary Markovian query arrival process results in a bi-dimensional state and complex state transitions within each sub-MDP, making it challenging to verify Whittle indexability using conventional methods. After a thorough analysis of the sub-MDP's structure, we show that it is unichain and its optimal policy follows a threshold-type structure. This facilitates the verification of Whittle indexability of the sub-MDP by employing an easy-to-verify condition. Subsequently, the steady-state probability distributions of the sub-MDP under different threshold-type policies are analyzed, constituting the analytical expressions of different Whittle indices in terms of the expected average QAoI and scheduling time of the sub-MDP. Building on these, we devise an efficient algorithm to calculate Whittle indices for the formulated sub-MDPs. The simulation results validate our analyses and show the proposed Whittle index policy outperforms baseline policies and achieves near-optimal performance.

cs.IT↗

DeepCRF: Deep Learning-Enhanced CSI-Based RF Fingerprinting for Channel-Resilient WiFi Device Identification

This paper presents DeepCRF, a new framework that harnesses deep learning to extract subtle micro-signals from channel state information (CSI) measurements, enabling robust and resilient radio-frequency fingerprinting (RFF) of commercial-off-the-shelf (COTS) WiFi devices across diverse channel conditions. Building on our previous research, which demonstrated that micro-signals in CSI, termed micro-CSI, most likely originate from RF circuitry imperfections and can serve as unique RF fingerprints, we develop a new approach to overcome the limitations of our prior signal space-based method. While the signal space-based method is effective in strong line-of-sight (LoS) conditions, we show that it struggles with the complexities of non-line-of-sight (NLoS) scenarios, compromising the robustness of CSI-based RFF. To address this challenge, DeepCRF incorporates a carefully trained convolutional neural network (CNN) with model-inspired data augmentation, supervised contrastive learning, and decision fusion techniques, enhancing its generalization capabilities across unseen channel conditions and resilience against noise. Our evaluations demonstrate that DeepCRF significantly improves device identification accuracy across diverse channels, outperforming both the signal space-based baseline and state-of-the-art neural network-based benchmarks. Notably, it achieves an average identification accuracy of 99.53% among 19 COTS WiFi network interface cards in real-world unseen scenarios using 4 CSI measurements per identification procedure.

eess.SP↗

Age of Information-Oriented Probabilistic Link Scheduling for Device-to-Device Networks

This paper focuses on optimizing the long-term average age of information (AoI) in device-to-device (D2D) networks through age-aware link scheduling. The problem is naturally formulated as a Markov decision process (MDP). However, finding the optimal policy for the formulated MDP in its original form is challenging due to the intertwined AoI dynamics of all D2D links. To address this, we propose an age-aware stationary randomized policy that determines the probability of scheduling each link in each time slot based on the AoI of all links and the statistical channel state information among all transceivers. By employing the Lyapunov optimization framework, our policy aims to minimize the Lyapunov drift in every time slot. Nonetheless, this per-slot minimization problem is nonconvex due to cross-link interference in D2D networks, posing significant challenges for real-time decision-making. After analyzing the permutation equivariance property of the optimal solutions to the per-slot problem, we apply a message passing neural network (MPNN), a type of graph neural network that also exhibits permutation equivariance, to optimize the per-slot problem in an unsupervised learning manner. Simulation results demonstrate the superior performance of the proposed age-aware stationary randomized policy over baselines and validate the scalability of our method.

eess.SP↗

Computing Competitive Equilibrium for Chores: Linear Convergence and Lightweight Iteration

Competitive equilibrium (CE) for chores has recently attracted significant attention, with many algorithms proposed to approximately compute it. However, existing algorithms either lack iterate convergence guarantees to an exact CE or require solving high-dimensional linear or quadratic programming subproblems. This paper overcomes these issues by proposing a novel unconstrained difference-of-convex formulation, whose stationary points correspond precisely to the CE for chores. We show that the new formulation possesses the local error bound property and the Kurdyka-Łojasiewicz property with an exponent of $1/2$. Consequently, we present the first algorithm whose iterates provably converge linearly to an exact CE for chores. Furthermore, by exploiting the max structure within our formulation and applying smoothing techniques, we develop a subproblem-free algorithm that finds an approximate CE in polynomial time. Numerical experiments demonstrate that the proposed algorithms outperform the state-of-the-art method.

math.OC↗

The Power of Large Language Models for Wireless Communication System Development: A Case Study on FPGA Platforms

Large language models (LLMs) have garnered significant attention across various research disciplines, including the wireless communication community. There have been several heated discussions on the intersection of LLMs and wireless technologies. While recent studies have demonstrated the ability of LLMs to generate hardware description language (HDL) code for simple computation tasks, developing wireless prototypes and products via HDL poses far greater challenges because of the more complex computation tasks involved. In this paper, we aim to address this challenge by investigating the role of LLMs in FPGA-based hardware development for advanced wireless signal processing. We begin by exploring LLM-assisted code refactoring, reuse, and validation, using an open-source software-defined radio (SDR) project as a case study. Through the case study, we find that an LLM assistant can potentially yield substantial productivity gains for researchers and developers. We then examine the feasibility of using LLMs to generate HDL code for advanced wireless signal processing, using the Fast Fourier Transform (FFT) algorithm as an example. This task presents two unique challenges: the scheduling of subtasks within the overall task and the multi-step thinking required to solve certain arithmetic problem within the task. To address these challenges, we employ in-context learning (ICL) and Chain-of-Thought (CoT) prompting techniques, culminating in the successful generation of a 64-point Verilog FFT module. Our results demonstrate the potential of LLMs for generalization and imitation, affirming their usefulness in writing HDL code for wireless communication systems. Overall, this work contributes to understanding the role of LLMs in wireless communication and motivates further exploration of their capabilities.

eess.SP↗

Towards Channel-Resilient CSI-Based RF Fingerprinting using Deep Learning

This work introduces DeepCRF, a deep learning framework designed for channel state information-based radio frequency fingerprinting (CSI-RFF). The considered CSI-RFF is built on micro-CSI, a recently discovered radio-frequency (RF) fingerprint that manifests as micro-signals appearing on the channel state information (CSI) curves of commercial WiFi devices. Micro-CSI facilitates CSI-RFF which is more streamlined and easily implementable compared to existing schemes that rely on raw I/Q samples. The primary challenge resides in the precise extraction of micro-CSI from the inherently fluctuating CSI measurements, a process critical for reliable RFF. The construction of a framework that is resilient to channel variability is essential for the practical deployment of CSI-RFF techniques. DeepCRF addresses this challenge with a thoughtfully trained convolutional neural network (CNN). This network's performance is significantly enhanced by employing effective and strategic data augmentation techniques, which bolster its ability to generalize to novel, unseen channel conditions. Furthermore, DeepCRF incorporates supervised contrastive learning to enhance its robustness against noises. Our evaluations demonstrate that DeepCRF significantly enhances the accuracy of device identification across previously unencountered channels. It outperforms both the conventional model-based methods and standard CNN that lack our specialized training and enhancement strategies.

eess.SP↗

Three-dimensional angular deviation and diffraction efficiency of a grating in Littrow-configuration ECDL

We consider in this paper the angular deviation and diffraction efficiency of the reflection gratings in Littrow-configuration for applications of external cavity diode laser using the rigorous coupled-wave analysis method. We consider the three-dimensional diffraction case in general, where the incidence plane is un-parallel with the grating vector, i.e. conical diffraction. The angular tolerance of arbitrary gratings under plane and conical diffraction are thus derived and presented. A typical blazed grating is chosen as an example to calculate its diffraction efficiency using the rigorous coupled-wave analysis method. Furthermore, we point out that the angular tolerance and reflection efficiency can be improved if the appropriate parameter settings are selected for Littrow-configuration devices, including incidence angle, diffraction order, grating period and blazed angle. Otherwise, a tiny slanting angle of the grating vector deviated from the incidence plane will deviate the feedback light away from entering the laser-diode-chip and halt laser oscillation in the external cavity. Finally, a general rule for the parameter settings in Littrow-configuration is provided as a benchmark.

physics.optics↗