SearcharxivSearch

arXiv subjects

Haidong Wang

Publications and source records attributed to Haidong Wang.

At least 19 recordsLinked to original sources

Event-RGB Adaptive Tracking for Nighttime Highway Perception

Intelligent Transportation Systems deployed on highways predominantly rely on conventional RGB cameras for traffic perception and vehicle tracking. However, highway environments present unique challenges: the absence of artificial lighting infrastructure, combined with high vehicle velocities, results in severely degraded perception performance under low-light conditions. Specifically, nighttime scenarios suffer from motion blur, insufficient exposure, and poor signal-to-noise ratios, which catastrophically impair the reliability of RGB-based sensing systems. To address these limitations, we propose a novel Joint Event-RGB Adaptive Tracking (JEAT) framework. Unlike existing multi-sensor trackers constrained by rigid, hard-coded prioritization, JEAT merges asynchronous event streams and RGB frames into a unified joint data association optimization. By employing an Adaptive Extended Kalman Filter to continuously estimate measurement noise via NIS statistics, the framework dynamically weights and fuses both modalities, optimally harnessing event streams during dark or high-speed motion while leveraging RGB frames under bright or static conditions. Furthermore, given the absence of publicly available datasets tailored for event-based highway perception with diverse environmental conditions, we present SEHN, a large-scale synthetic dataset generated using the CARLA simulator. Our dataset encompasses diverse environmental conditions (daytime, nighttime, nighttime with out artificial lighting) and varying traffic densities, providing synchronized RGB imagery and event streams to facilitate multi-modal fusion research. Our code and datasets will be available at https://github.com/haidongwang96/SEHN.

cs.CV

Learning to Optimize Joint Source and RIS-assisted Channel Encoding for Multi-User Semantic Communication Systems

In this paper, we explore a joint source and reconfigurable intelligent surface (RIS)-assisted channel encoding (JSRE) framework for multi-user semantic communications, where a deep neural network (DNN) extracts semantic features for all users and the RIS provides channel orthogonality, enabling a unified semantic encoding-decoding design. We aim to maximize the overall energy efficiency of semantic communications across all users by jointly optimizing the user scheduling, the RIS's phase shifts, and the semantic compression ratio. Although this joint optimization problem can be addressed using conventional deep reinforcement learning (DRL) methods, evaluating semantic similarity typically relies on extensive real environment interactions, which can incur heavy computational overhead during training. To address this challenge, we propose a truncated DRL (T-DRL) framework, where a DNN-based semantic similarity estimator is developed to rapidly estimate the similarity score. Moreover, the user scheduling strategy is tightly coupled with the semantic model configuration. To exploit this relationship, we further propose a semantic model caching mechanism that stores and reuses fine-tuned semantic models corresponding to different scheduling decisions. A Transformer-based actor network is employed within the DRL framework to dynamically generate action space conditioned on the current caching state. This avoids redundant retraining and further accelerates the convergence of the learning process. Numerical results demonstrate that the proposed JSRE framework significantly improves the system energy efficiency compared with the baseline methods. By training fewer semantic models, the proposed T-DRL framework significantly enhances the learning efficiency.

cs.NI

SA-GCS: Semantic-Aware Gaussian Curriculum Scheduling for UAV Vision-Language Navigation

Unmanned Aerial Vehicle (UAV) Vision-Language Navigation (VLN) aims to enable agents to accurately localize targets and plan flight paths in complex environments based on natural language instructions, with broad applications in intelligent inspection, disaster rescue, and urban monitoring. Recent progress in Vision-Language Models (VLMs) has provided strong semantic understanding for this task, while reinforcement learning (RL) has emerged as a promising post-training strategy to further improve generalization. However, existing RL methods often suffer from inefficient use of training data, slow convergence, and insufficient consideration of the difficulty variation among training samples, which limits further performance improvement. To address these challenges, we propose \textbf{Semantic-Aware Gaussian Curriculum Scheduling (SA-GCS)}, a novel training framework that systematically integrates Curriculum Learning (CL) into RL. SA-GCS employs a Semantic-Aware Difficulty Estimator (SA-DE) to quantify the complexity of training samples and a Gaussian Curriculum Scheduler (GCS) to dynamically adjust the sampling distribution, enabling a smooth progression from easy to challenging tasks. This design significantly improves training efficiency, accelerates convergence, and enhances overall model performance. Extensive experiments on the CityNav benchmark demonstrate that SA-GCS consistently outperforms strong baselines across all metrics, achieves faster and more stable convergence, and generalizes well across models of different scales, highlighting its robustness and scalability. The implementation of our approach is publicly available.

cs.CL

WebNovelBench: Placing LLM Novelists on the Web Novel Distribution

Robustly evaluating the long-form storytelling capabilities of Large Language Models (LLMs) remains a significant challenge, as existing benchmarks often lack the necessary scale, diversity, or objective measures. To address this, we introduce WebNovelBench, a novel benchmark specifically designed for evaluating long-form novel generation. WebNovelBench leverages a large-scale dataset of over 4,000 Chinese web novels, framing evaluation as a synopsis-to-story generation task. We propose a multi-faceted framework encompassing eight narrative quality dimensions, assessed automatically via an LLM-as-Judge approach. Scores are aggregated using Principal Component Analysis and mapped to a percentile rank against human-authored works. Our experiments demonstrate that WebNovelBench effectively differentiates between human-written masterpieces, popular web novels, and LLM-generated content. We provide a comprehensive analysis of 24 state-of-the-art LLMs, ranking their storytelling abilities and offering insights for future development. This benchmark provides a scalable, replicable, and data-driven methodology for assessing and advancing LLM-driven narrative generation.

cs.CL

FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models

Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing methods often struggle with insufficient multimodal fusion, weak generalization, and poor interpretability. To address these challenges, we propose FlightGPT, a novel UAV VLN framework built upon Vision-Language Models (VLMs) with powerful multimodal perception capabilities. We design a two-stage training pipeline: first, Supervised Fine-Tuning (SFT) using high-quality demonstrations to improve initialization and structured reasoning; then, Group Relative Policy Optimization (GRPO) algorithm, guided by a composite reward that considers goal accuracy, reasoning quality, and format compliance, to enhance generalization and adaptability. Furthermore, FlightGPT introduces a Chain-of-Thought (CoT)-based reasoning mechanism to improve decision interpretability. Extensive experiments on the city-scale dataset CityNav demonstrate that FlightGPT achieves state-of-the-art performance across all scenarios, with a 9.22\% higher success rate than the strongest baseline in unseen environments. Our implementation is publicly available.

cs.CL

Learning Joint Source-Channel Encoding in IRS-assisted Multi-User Semantic Communications

In this paper, we investigate a joint source-channel encoding (JSCE) scheme in an intelligent reflecting surface (IRS)-assisted multi-user semantic communication system. Semantic encoding not only compresses redundant information, but also enhances information orthogonality in a semantic feature space. Meanwhile, the IRS can adjust the spatial orthogonality, enabling concurrent multi-user semantic communication in densely deployed wireless networks to improve spectrum efficiency. We aim to maximize the users' semantic throughput by jointly optimizing the users' scheduling, the IRS's passive beamforming, and the semantic encoding strategies. To tackle this non-convex problem, we propose an explainable deep neural network-driven deep reinforcement learning (XD-DRL) framework. Specifically, we employ a deep neural network (DNN) to serve as a joint source-channel semantic encoder, enabling transmitters to extract semantic features from raw images. By leveraging structural similarity, we assign some DNN weight coefficients as the IRS's phase shifts, allowing simultaneous optimization of IRS's passive beamforming and DNN training. Given the IRS's passive beamforming and semantic encoding strategies, user scheduling is optimized using the DRL method. Numerical results validate that our JSCE scheme achieves superior semantic throughput compared to the conventional schemes and efficiently reduces the semantic encoder's mode size in multi-user scenarios.

eess.SP

Towards Intelligent Transportation with Pedestrians and Vehicles In-the-Loop: A Surveillance Video-Assisted Federated Digital Twin Framework

In intelligent transportation systems (ITSs), incorporating pedestrians and vehicles in-the-loop is crucial for developing realistic and safe traffic management solutions. However, there is falls short of simulating complex real-world ITS scenarios, primarily due to the lack of a digital twin implementation framework for characterizing interactions between pedestrians and vehicles at different locations in different traffic environments. In this article, we propose a surveillance video assisted federated digital twin (SV-FDT) framework to empower ITSs with pedestrians and vehicles in-the-loop. Specifically, SVFDT builds comprehensive pedestrian-vehicle interaction models by leveraging multi-source traffic surveillance videos. Its architecture consists of three layers: (i) the end layer, which collects traffic surveillance videos from multiple sources; (ii) the edge layer, responsible for semantic segmentation-based visual understanding, twin agent-based interaction modeling, and local digital twin system (LDTS) creation in local regions; and (iii) the cloud layer, which integrates LDTSs across different regions to construct a global DT model in realtime. We analyze key design requirements and challenges and present core guidelines for SVFDT's system implementation. A testbed evaluation demonstrates its effectiveness in optimizing traffic management. Comparisons with traditional terminal-server frameworks highlight SV-FDT's advantages in mirroring delays, recognition accuracy, and subjective evaluation. Finally, we identify some open challenges and discuss future research directions.

cs.ET

A Brain-Inspired Perception-Decision Driving Model Based on Neural Pathway Anatomical Alignment

In the realm of autonomous driving, conventional approaches for vehicle perception and decision-making primarily rely on sensor input and rule-based algorithms. However, these methodologies often suffer from lack of interpretability and robustness, particularly in intricate traffic scenarios. To tackle this challenge, we propose a novel brain-inspired driving (BID) framework. Diverging from traditional methods, our approach harnesses brain-inspired perception technology to achieve more efficient and robust environmental perception. Additionally, it employs brain-inspired decision-making techniques to facilitate intelligent decision-making. The experimental results show that the performance has been significantly improved across various autonomous driving tasks and achieved the end-to-end autopilot successfully. This contribution not only advances interpretability and robustness but also offers fancy insights and methodologies for further advancing autonomous driving technology.

cs.RO

BAN: Neuroanatomical Aligning in Auditory Recognition between Artificial Neural Network and Human Cortex

Drawing inspiration from neurosciences, artificial neural networks (ANNs) have evolved from shallow architectures to highly complex, deep structures, yielding exceptional performance in auditory recognition tasks. However, traditional ANNs often struggle to align with brain regions due to their excessive depth and lack of biologically realistic features, like recurrent connection. To address this, a brain-like auditory network (BAN) is introduced, which incorporates four neuroanatomically mapped areas and recurrent connection, guided by a novel metric called the brain-like auditory score (BAS). BAS serves as a benchmark for evaluating the similarity between BAN and human auditory recognition pathway. We further propose that specific areas in the cerebral cortex, mainly the middle and medial superior temporal (T2/T3) areas, correspond to the designed network structure, drawing parallels with the brain's auditory perception pathway. Our findings suggest that the neuroanatomical similarity in the cortex and auditory classification abilities of the ANN are well-aligned. In addition to delivering excellent performance on a music genre classification task, the BAN demonstrates a high BAS score. In conclusion, this study presents BAN as a recurrent, brain-inspired ANN, representing the first model that mirrors the cortical pathway of auditory recognition.

q-bio.NC

An Audio-Visual Fusion Emotion Generation Model Based on Neuroanatomical Alignment

In the field of affective computing, traditional methods for generating emotions predominantly rely on deep learning techniques and large-scale emotion datasets. However, deep learning techniques are often complex and difficult to interpret, and standardizing large-scale emotional datasets are difficult and costly to establish. To tackle these challenges, we introduce a novel framework named Audio-Visual Fusion for Brain-like Emotion Learning(AVF-BEL). In contrast to conventional brain-inspired emotion learning methods, this approach improves the audio-visual emotion fusion and generation model through the integration of modular components, thereby enabling more lightweight and interpretable emotion learning and generation processes. The framework simulates the integration of the visual, auditory, and emotional pathways of the brain, optimizes the fusion of emotional features across visual and auditory modalities, and improves upon the traditional Brain Emotional Learning (BEL) model. The experimental results indicate a significant improvement in the similarity of the audio-visual fusion emotion learning generation model compared to single-modality visual and auditory emotion learning and generation model. Ultimately, this aligns with the fundamental phenomenon of heightened emotion generation facilitated by the integrated impact of visual and auditory stimuli. This contribution not only enhances the interpretability and efficiency of affective intelligence but also provides new insights and pathways for advancing affective computing technology. Our source code can be accessed here: https://github.com/OpenHUTB/emotion}{https://github.com/OpenHUTB/emotion.

cs.HC

Thermoelectrical potential and derivation of Kelvin relation for thermoelectric materials

Current research on thermoelectricity is primarily focused on the exploration of materials with enhanced performance, resulting in a lack of fundamental understanding of the thermoelectric effect. Such circumstance is not conducive to the further improvement of the efficiency of thermoelectric conversion. Moreover, available physical images of the derivation of the Kelvin relations are ambiguous. Derivation processes are complex and need a deeper understanding of thermoelectric conversion phenomena. In this paper, a new physical quantity 'thermoelectrical potential' from the physical nature of the thermoelectric conversion is proposed. The quantity is expressed as the product of the Seebeck coefficient and the absolute temperature, i.e., ST. Based on the thermoelectrical potential, we clarify the conversion of the various forms of energy in the thermoelectric effect by presenting a clear physical picture. Results from the analysis of the physical mechanism of the Seebeck effect indicate that the thermoelectrical potential, rather than the temperature gradient field, exerts a force on the charge carriers in the thermoelectric material. Based on thermoelectric potential, the Peltier effects at different material interfaces can be macroscopically described. The Kelvin relation is rederived using the proposed quantity, which simplified the derivation process and elucidated the physical picture of the thermoelectrical conversion.

cond-mat.mtrl-sci

Role of Elastic Phonon Couplings in Dictating the Thermal Transport across Atomically Sharp SiC/Si Interfaces

Wide-bandgap (WBG) semiconductors have promising applications in power electronics due to their high voltages, radio frequencies, and tolerant temperatures. Among all the WBG semiconductors, SiC has attracted attention because of its high mobility, high thermal stability, and high thermal conductivity. However, the interfaces between SiC and the corresponding substrate largely affect the performance of SiC-based electronics. It is therefore necessary to understand and design the interfacial thermal transport across the SiC/substrate interfaces, which is critical for the thermal management design of these SiC-based power electronics. This work systematically investigates heat transfer across the 3C-SiC/Si, 4H-SiC/Si, and 6H-SiC/Si interfaces using non-equilibrium molecular dynamics simulations and diffuse mismatch model. We find that the room temperature ITC for 3C-SiC/Si, 4H-SiC/Si, and 6H-SiC/Si interfaces is 932 MW/m2K, 759 MW/m2K, and 697 MW/m2K, respectively. We also show the contribution of the ITC resulting from elastic scatterings at room temperature is 80% for 3C-SiC/Si interfaces, 85% for 4H-SiC/Si interfaces, and 82% for 6H-SiC/Si interfaces, respectively. We further find the ITC contributed by the elastic scattering decreases with the temperature but remains at a high ratio of 67%~78% even at an ultrahigh temperature of 1000 K. The reason for such a high elastic ITC is the large overlap between the vibrational density of states of Si and SiC at low frequencies (< ~ 18 THz), which is also demonstrated by the diffuse mismatch mode. It is interesting to find that the inelastic ITC resulting from the phonons with frequencies higher than the cutoff frequency of Si (i.e., ~18 THz) can be negligible. That may be because of the wide frequency gap between Si and SiC, which makes the inelastic scattering among these phonons challenging to meet the energy and momentum conservation rules.

cond-mat.mtrl-sci

PC-SNN: Predictive Coding-based Local Hebbian Plasticity Learning in Spiking Neural Networks

Spiking Neural Networks (SNNs), regarded as the third generation of neural networks, emulate the brain's information processing with unparalleled biological plausibility compared to traditional neural networks. However, their non-linear, event-driven dynamics pose significant challenges for training, and existing methods often deviate from neuroscientific principles of cortical learning. Drawing inspiration from predictive coding theory-a leading model of brain information processing-we propose PC-SNN, a novel learning framework that integrates predictive coding with SNNs to enable biologically plausible, local Hebbian plasticity without reliance on backpropagation. Unlike conventional SNN training approaches, PC-SNN leverages only local computations, aligning with the brain's distributed processing and overcoming the biological implausibility of global error propagation. Our classification model achieves competitive performance on the benchmark datasets, including Caltech Face/Motorbike, MNIST, and CIFAR10, surpassing state-of-the-art multi-layer SNNs. Furthermore, our predictive coding-based regression model outperforms backpropagation-based methods while adhering to local plasticity constraints, offering a scalable and biologically grounded alternative for SNN training. PC-SNN drives progress in neuromorphic computing through validating the adaptability of bio-inspired algorithms within spiking neural architectures, but also unveils novel understandings of neurocognitive learning processes, presenting a conceptual framework distinguished by its theoretical originality and functional efficacy.

cs.NE

Enhancing thermoelectric properties of isotope graphene nanoribbons via machine learning guided manipulation of disordered antidots and interfaces

Structural manipulation at the nanoscale breaks the intrinsic correlations among different energy carrier transport properties, achieving high thermoelectric performance. However, the coupled multifunctional (phonon and electron) transport in the design of nanomaterials makes the optimization of thermoelectric properties challenging. Machine learning brings convenience to the design of nanostructures with large degree of freedom. Herein, we conducted comprehensive thermoelectric optimization of isotopic armchair graphene nanoribbons (AGNRs) with antidots and interfaces by combining Green's function approach with machine learning algorithms. The optimal AGNR with ZT of 0.894 by manipulating antidots was obtained at the interfaces of the aperiodic isotope superlattices, which is 5.69 times larger than that of the pristine structure. The proposed optimal structure via machine learning provides physical insights that the carbon-13 atoms tend to form a continuous interface barrier perpendicular to the carrier transport direction to suppress the propagation of phonons through isotope AGNRs. The antidot effect is more effective than isotope substitution in improving the thermoelectric properties of AGNRs. The proposed approach coupling energy carrier transport property analysis with machine learning algorithms offers highly efficient guidance on enhancing the thermoelectric properties of low-dimensional nanomaterials, as well as to explore and gain non-intuitive physical insights.

cond-mat.mes-hall

STURE: Spatial-Temporal Mutual Representation Learning for Robust Data Association in Online Multi-Object Tracking

Online multi-object tracking (MOT) is a longstanding task for computer vision and intelligent vehicle platform. At present, the main paradigm is tracking-by-detection, and the main difficulty of this paradigm is how to associate current candidate detections with historical tracklets. However, in the MOT scenarios, each historical tracklet is composed of an object sequence, while each candidate detection is just a flat image, which lacks temporal features of the object sequence. The feature difference between current candidate detections and historical tracklets makes the object association much harder. Therefore, we propose a Spatial-Temporal Mutual Representation Learning (STURE) approach which learns spatial-temporal representations between current candidate detections and historical sequences in a mutual representation space. For historical trackelets, the detection learning network is forced to match the representations of sequence learning network in a mutual representation space. The proposed approach is capable of extracting more distinguishing detection and sequence representations by using various designed losses in object association. As a result, spatial-temporal feature is learned mutually to reinforce the current detection features, and the feature difference can be relieved. To prove the robustness of the STURE, it is applied to the public MOT challenge benchmarks and performs well compared with various state-of-the-art online MOT trackers based on identity-preserving metrics.

cs.CV

SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search

The in-memory algorithms for approximate nearest neighbor search (ANNS) have achieved great success for fast high-recall search, but are extremely expensive when handling very large scale database. Thus, there is an increasing request for the hybrid ANNS solutions with small memory and inexpensive solid-state drive (SSD). In this paper, we present a simple but efficient memory-disk hybrid indexing and search system, named SPANN, that follows the inverted index methodology. It stores the centroid points of the posting lists in the memory and the large posting lists in the disk. We guarantee both disk-access efficiency (low latency) and high recall by effectively reducing the disk-access number and retrieving high-quality posting lists. In the index-building stage, we adopt a hierarchical balanced clustering algorithm to balance the length of posting lists and augment the posting list by adding the points in the closure of the corresponding clusters. In the search stage, we use a query-aware scheme to dynamically prune the access of unnecessary posting lists. Experiment results demonstrate that SPANN is 2$\times$ faster than the state-of-the-art ANNS solution DiskANN to reach the same recall quality $90\%$ with same memory cost in three billion-scale datasets. It can reach $90\%$ recall@1 and recall@10 in just around one millisecond with only 32GB memory cost. Code is available at: {\footnotesize\color{blue}{\url{https://github.com/microsoft/SPTAG}}}.

cs.DB

Two-step Dual-wavelength Flash Raman Mapping Method for Measuring Thermophysical Properties of Supported 2D Nanomaterials

The thermophysical properties of supported and free-standing nanomaterials would be different due to the substrate effect. To determine the thermophysical properties of supported 2D nanomaterials, this paper developed a two-step dual-wavelength flash Raman (DFR) mapping method. The thermal conductivity of the supported 2D nanomaterial and the thermal contact conductance between the sample and the substrate can be determined by steady-state step. And then the thermal diffusivity of the sample can be characterized by the transient step. Two models are also proposed in this paper. When the substrate temperature rises obviously, a full model considered the substrate temperature distribution is used to decouple the thermal diffusivity, thermal conductivity, and the thermal contact conductance. When the maximum substrate temperature rise is less than 20% of the maximum sample temperature rise, the temperature distribution and variation of the substrate is assumed proportionate to the sample, and then a simplified model can be used to analyse the thermal diffusivity and the other parameters. For the thermal diffusivity, the system error caused by simplifying assumptions is less than 1%, when a dimensionless parameter related to the thermal contact conductance and the thermal conductivity less than 1 and the maximum temperature rise of the substrate less than 15% of that of the supported sample.

physics.ins-det

Substrate effect on thermal conductivity of monolayer WS2: Experimental measurement and theoretical analysis

Monolayer WS2 has been a competitive candidate in electrical and optoelectronic devices due to its superior optoelectronic properties. To tackle the challenge of thermal management caused by the decreased size and concentrated heat in modern ICs, it is of great significance to accurately characterize the thermal conductivity of the monolayer WS2, especially with substrate supported. In this work, the dual-wavelength flash Raman method is used to experimentally measure the thermal conductivity of the suspended and the Si/SiO2 substrate supported monolayer WS2 at a temperature range of 200 K - 400 K. The room-temperature thermal conductivity of suspended and supported WS2 are 28.45 W/mK and 15.39 W/mK, respectively, with a ~50% reduction due to substrate effect. To systematically study the underlying mechanism behind the striking reduction, we employed the Raman spatial mapping analysis combined with the molecular dynamics simulation. The analysis of Raman spectra showed the increase of doping level, reduction of phonon lifetime and suppression of out-of-plane vibration mode due to substrate effect. In addition, the phonon transmission coefficient was mutually verified with Raman spectra analysis and further revealed that the substrate effect significantly enhances the phonon scattering at the interface and mainly suppresses the acoustic phonon, thus leading to the reduction of thermal conductivity. The thermal conductivity of other suspended and supported monolayer TMDCs (e.g. MoS2, MoSe2 and WSe2) were also listed for comparison. Our researches can be extended to understand the substrate effect of other 2D TMDCs and provide guidance for future TMDCs-based electrical and optoelectronic devices.

cond-mat.mes-hall