SearcharxivSearch

arXiv subjects

Tian He

Publications and source records attributed to Tian He.

At least 19 recordsLinked to original sources

GenHAR: Generalizing Cross-domain Human Activity Recognition for Last-mile Delivery

Human Activity Recognition (HAR) has shown remarkable effectiveness in various applications, such as smart healthcare and intelligent manufacturing. However, a major challenge faced by HAR is the distribution shift across different sensor data domains, which often leads to decreased performance when deployed for real-world applications. To address this issue, this paper introduces GenHAR, a novel framework designed to mitigate the domain gap by learning domain-invariant sensor representations. GenHAR aims to enhance the generalization capabilities of HAR on target domains purely with data from the source domain. The key novelty of GenHAR lies in two aspects. Firstly, GenHAR tokenizes sensor data and learns correlations among frequency sensor channel dimensions to improve the robustness of HAR models. Secondly, GenHAR improves the efficiency via selective masking and an efficient attention mechanism. We conduct a systematic analysis of GenHAR by comparing it with state-of-the-art HAR methods on real-world human activity datasets. Results show that GenHAR outperforms state-of-the-art methods by 9.97% in accuracy, and reduces Floating Point Operations by 6.4 times. Moreover, we deploy GenHAR at a leading logistics company in 4 cities, and have detected 2.15 billion real-time activities. We release our code at: https://github.com/Sensor-FoundationModel/GenHAR.

cs.CV

Birds of a Feather Cluster Nearby: a Proximity-Aware Geo-Codebook for Local Service Recommendation

Generative recommendation systems are increasingly adopted in local service platforms, where semantic relevance alone is insufficient without strict geographic feasibility. A key technical challenge lies in semantic ID (SID) tokenization, which directly impacts the recommendation performance. However, existing semantic codebooks neglect geographic constraints, often resulting in recommendations that are semantically relevant yet geographically unreachable. To address this limitation, we propose Pro-GEO, a Proximity-aware GEO-codebook. Pro-GEO establishes a geo-centroid local coordinate system to capture intra-cluster spatial relationships and a geo-rotary position encoding mechanism that models geographic proximity as orthogonal rotational transformations in the high-dimensional embedding. This design enables semantic and spatial signals to be jointly modeled in a balanced manner, without reducing geographic information to a weak auxiliary feature. Extensive experiments conducted on a large-scale industrial dataset reveal that Pro-GEO significantly outperforms state-of-the-art methods. In particular, Pro-GEO reduces the average geographic clustering distance by 45.60% and achieves a 1.87% improvement in Hit@50, highlighting its effectiveness for real-world local service recommendation.

cs.IR

WM-DAgger: Enabling Efficient Data Aggregation for Imitation Learning with World Models

Imitation learning is a powerful paradigm for training robotic policies, yet its performance is limited by compounding errors: minor policy inaccuracies could drive robots into unseen out-of-distribution (OOD) states in the training set, where the policy could generate even bigger errors, leading to eventual failures. While the Data Aggregation (DAgger) framework tries to address this issue, its reliance on continuous human involvement severely limits scalability. In this paper, we propose WM-DAgger, an efficient data aggregation framework that leverages World Models to synthesize OOD recovery data without requiring human involvement. Specifically, we focus on manipulation tasks with an eye-in-hand robotic arm and only few-shot demonstrations. To avoid synthesizing misleading data and overcome the hallucination issues inherent to World Models, our framework introduces two key mechanisms: (1) a Corrective Action Synthesis Module that generates task-oriented recovery actions to prevent misleading supervision, and (2) a Consistency-Guided Filtering Module that discards physically implausible trajectories by anchoring terminal synthesized frames to corresponding real frames in expert demonstrations. We extensively validate WM-DAgger on multiple real-world robotic tasks. Results that our method significantly improves success rates, achieving a 93.3\% success rate in soft bag pushing with only five demonstrations. The source code is publicly available at https://github.com/czs12354-xxdbd/WM-Dagger.

cs.RO

Distinct Suppression Mechanisms of Superconductivity by Magnetic Domains and Spin Fluctuations in EuFe2(As1-xPx)2 superconductors

Using ac composite magnetoelectric technique, we map the phase diagrams of EuFe(As1-xPx)2 to resolve the interplay between superconductivity and ferromagnetism. For samples with Tc TFM, strong short-range spin correlations and phase boundaries within a multiphase coexistence regime near the triple point act as potent pair-breaking centers, leading to pronounced Hc2 suppression.

cond-mat.supr-con

Collaborate sim and real: Robot Bin Packing Learning in Real-world and Physical Engine

The 3D bin packing problem, with its diverse industrial applications, has garnered significant research attention in recent years. Existing approaches typically model it as a discrete and static process, while real-world applications involve continuous gravity-driven interactions. This idealized simplification leads to infeasible deployments (e.g., unstable packing) in practice. Simulations with physical engine offer an opportunity to emulate continuous gravity effects, enabling the training of reinforcement learning (RL) agents to address such limitations and improve packing stability. However, a simulation-to-reality gap persists due to dynamic variations in physical properties of real-world objects, such as various friction coefficients, elasticity, and non-uniform weight distributions. To bridge this gap, we propose a hybrid RL framework that collaborates with physical simulation with real-world data feedback. Firstly, domain randomization is applied during simulation to expose agents to a spectrum of physical parameters, enhancing their generalization capability. Secondly, the RL agent is fine-tuned with real-world deployment feedback, further reducing collapse rates. Extensive experiments demonstrate that our method achieves lower collapse rates in both simulated and real-world scenarios. Large-scale deployments in logistics systems validate the practical effectiveness, with a 35\% reduction in packing collapse compared to baseline methods.

cs.RO

Resource-Oriented Optimization of Electric Vehicle Systems: A Data-Driven Survey on Charging Infrastructure, Scheduling, and Fleet Management

Driven by growing concerns over air quality and energy security, electric vehicles (EVs) has experienced rapid development and are reshaping global transportation systems and lifestyle patterns. Compared to traditional gasoline-powered vehicles, EVs offer significant advantages in terms of lower energy consumption, reduced emissions, and decreased operating costs. However, there are still some core challenges to be addressed: (i) Charging station congestion and operational inefficiencies during peak hours, (ii) High charging cost under dynamic electricity pricing schemes, and (iii) Conflicts between charging needs and passenger service requirements.Hence, in this paper, we present a comprehensive review of data-driven models and approaches proposed in the literature to address the above challenges. These studies cover the entire lifecycle of EV systems, including charging station deployment, charging scheduling strategies, and large-scale fleet management. Moreover, we discuss the broader implications of EV integration across multiple domains, such as human mobility, smart grid infrastructure, and environmental sustainability, and identify key opportunities and directions for future research.

eess.SY

Phase-Incremented, Steady-State Solution NMR: Maximizing Spectral Sensitivity Without Compromising Resolution

NMR acquisitions based on Ernst-angle excitations are widely used in analytical spectroscopy, as for over half a century they have been considered the optimal way for maximizing spectral sensitivity without compromising bandwidth or peak resolution. However, if as often happens in liquid state NMR relaxation times T1, T2 are long and similar, steady-state free-precession (SSFP) experiments can actually provide higher signal-to-noise ratios per square root of acquisition time (SNRt) than Ernst-angle-based counterparts. Although a strong offset dependence and a requirement for pulsing at repetition times TR << T2 leading to poor spectral resolution have impeded widespread analytical applications of SSFP, phase-incremented (PI) SSFP schemes could overcome these drawbacks. The present study explores if, when and how, can this approach to high resolution NMR improve SNRt over the performance afforded by Ernst-angle-based FT acquisitions. It is found that PI-SSFP can indeed often provide a superior SNRt than FT-NMR, but that achieving this requires implementing the acquisitions using relatively large flip angles. As also explained, however, this can restrict PI-SSFP's spectral resolution, and lead to distorted line shapes. To deal with this problem we introduce here a new outlook on SSFP experiments that can overcome this dichotomy, and lead to high spectral resolution even when utilizing relatively the large flip angles that provide optimal sensitivity. This new outlook also leads to a processing pipeline for PI-SSFP acquisitions, which is here introduced and exemplified. The enhanced SNRt that the ensuing method can provide over FT-based NMR counterparts collected under Ernst-angle excitation conditions, is examined with a series of 13C and 15N natural abundance investigations on organic compounds.

physics.chem-ph

AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data

Textual description of a physical location, commonly known as an address, plays an important role in location-based services(LBS) such as on-demand delivery and navigation. However, the prevalence of abnormal addresses, those containing inaccuracies that fail to pinpoint a location, have led to significant costs. Address rewriting has emerged as a solution to rectify these abnormal addresses. Despite the critical need, existing address rewriting methods are limited, typically tailored to correct specific error types, or frequently require retraining to process new address data effectively. In this study, we introduce AddrLLM, an innovative framework for address rewriting that is built upon a retrieval augmented large language model. AddrLLM overcomes aforementioned limitations through a meticulously designed Supervised Fine-Tuning module, an Address-centric Retrieval Augmented Generation module and a Bias-free Objective Alignment module. To the best of our knowledge, this study pioneers the application of LLM-based address rewriting approach to solve the issue of abnormal addresses. Through comprehensive offline testing with real-world data on a national scale and subsequent online deployment, AddrLLM has demonstrated superior performance in integration with existing logistics system. It has significantly decreased the rate of parcel re-routing by approximately 43\%, underscoring its exceptional efficacy in real-world applications.

cs.CL

mmSpyVR: Exploiting mmWave Radar for Penetrating Obstacles to Uncover Privacy Vulnerability of Virtual Reality

Virtual reality (VR), while enhancing user experiences, introduces significant privacy risks. This paper reveals a novel vulnerability in VR systems that allows attackers to capture VR privacy through obstacles utilizing millimeter-wave (mmWave) signals without physical intrusion and virtual connection with the VR devices. We propose mmSpyVR, a novel attack on VR user's privacy via mmWave radar. The mmSpyVR framework encompasses two main parts: (i) A transfer learning-based feature extraction model to achieve VR feature extraction from mmWave signal. (ii) An attention-based VR privacy spying module to spy VR privacy information from the extracted feature. The mmSpyVR demonstrates the capability to extract critical VR privacy from the mmWave signals that have penetrated through obstacles. We evaluate mmSpyVR through IRB-approved user studies. Across 22 participants engaged in four experimental scenes utilizing VR devices from three different manufacturers, our system achieves an application recognition accuracy of 98.5\% and keystroke recognition accuracy of 92.6\%. This newly discovered vulnerability has implications across various domains, such as cybersecurity, privacy protection, and VR technology development. We also engage with VR manufacturer Meta to discuss and explore potential mitigation strategies. Data and code are publicly available for scrutiny and research at https://github.com/luoyumei1-a/mmSpyVR/

cs.CR

Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning

Skeleton-based action representation learning aims to interpret and understand human behaviors by encoding the skeleton sequences, which can be categorized into two primary training paradigms: supervised learning and self-supervised learning. However, the former one-hot classification requires labor-intensive predefined action categories annotations, while the latter involves skeleton transformations (e.g., cropping) in the pretext tasks that may impair the skeleton structure. To address these challenges, we introduce a novel skeleton-based training framework (C$^2$VL) based on Cross-modal Contrastive learning that uses the progressive distillation to learn task-agnostic human skeleton action representation from the Vision-Language knowledge prompts. Specifically, we establish the vision-language action concept space through vision-language knowledge prompts generated by pre-trained large multimodal models (LMMs), which enrich the fine-grained details that the skeleton action space lacks. Moreover, we propose the intra-modal self-similarity and inter-modal cross-consistency softened targets in the cross-modal representation learning process to progressively control and guide the degree of pulling vision-language knowledge prompts and corresponding skeletons closer. These soft instance discrimination and self-knowledge distillation strategies contribute to the learning of better skeleton-based action representations from the noisy skeleton-vision-language pairs. During the inference phase, our method requires only the skeleton data as the input for action recognition and no longer for vision-language prompts. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets demonstrate that our method outperforms the previous methods and achieves state-of-the-art results. Code is available at: https://github.com/cseeyangchen/C2VL.

cs.CV

Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action Recognition

Skeleton-based zero-shot action recognition aims to recognize unknown human actions based on the learned priors of the known skeleton-based actions and a semantic descriptor space shared by both known and unknown categories. However, previous works focus on establishing the bridges between the known skeleton representation space and semantic descriptions space at the coarse-grained level for recognizing unknown action categories, ignoring the fine-grained alignment of these two spaces, resulting in suboptimal performance in distinguishing high-similarity action categories. To address these challenges, we propose a novel method via Side information and dual-prompts learning for skeleton-based zero-shot action recognition (STAR) at the fine-grained level. Specifically, 1) we decompose the skeleton into several parts based on its topology structure and introduce the side information concerning multi-part descriptions of human body movements for alignment between the skeleton and the semantic space at the fine-grained level; 2) we design the visual-attribute and semantic-part prompts to improve the intra-class compactness within the skeleton space and inter-class separability within the semantic space, respectively, to distinguish the high-similarity actions. Extensive experiments show that our method achieves state-of-the-art performance in ZSL and GZSL settings on NTU RGB+D, NTU RGB+D 120, and PKU-MMD datasets.

cs.CV

Suppression of flux jumps in high-$J_c$ Nb$_3$Sn conductors by ferromagnetic layer

Flux jumps observed in high-$J_c$ Nb$_3$Sn conductors are urgent problems to construct high field superconducting magnets. The low-field instabilities usually reduce the current-carrying capability and thus cause the premature quench of Nb$_3$Sn coils at low magnetic field. In this paper, we explore suppressing the flux jumps by ferromagnetic (FM) layer. Firstly, we experimentally and theoretically investigate the flux jumps of Nb$_3$Sn/FM hybrid wires exposed to a magnetic field loop with constant sweeping rate. Comparing with bare Nb$_3$Sn and Nb$_3$Sn/Cu wires, we reveal two underlying mechanisms that the suppression of flux jumps is mainly attributed to the thermal effect of FM layer for the case of lower sweeping rate, whereas both thermal and electromagnetic effects play a crucial role for the case of higher sweeping rate. Furthermore, we explore the flux jumps of Nb$_3$Sn/FM hybrid wires exposed to AC magnetic fields with amplitude $B_{a0}$ and frequency $\rm\omega$. We build up the phase diagrams of flux jumps in the plane $\rm\omega$-$B_{a0}$ for bare Nb$_{3}$Sn wire, Nb$_{3}$Sn/Cu wire and Nb$_{3}$Sn/FM wire, respectively. We stress that the region of flux jumps of Nb$_{3}$Sn/FM wire is much smaller than the other two wires, which indicates that the Nb$_{3}$Sn/FM wire has significant advantage over merely increasing the heat capacity. The findings shed light on suppression of the flux jumps by utilizing FM materials, which is useful for developing new type of high-$J_c$ Nb$_{3}$Sn conductors.

cond-mat.mtrl-sci

Adaptive physics-informed neural networks for dynamic thermo-mechanical coupling problems in large-size-ratio functionally graded materials

In this paper, we present the adaptive physics-informed neural networks (PINNs) for resolving three dimensional (3D) dynamic thermo-mechanical coupling problems in large-size-ratio functionally graded materials (FGMs). The physical laws described by coupled governing equations and the constraints imposed by the initial and boundary conditions are leveraged to form the loss function of PINNs by means of the automatic differentiation algorithm, and an adaptive loss balancing scheme is introduced to improve the performance of PINNs. The adaptive PINNs are meshfree and trained on batches of randomly sampled collocation points, which is the key feature and superiority of the approach, since mesh-based methods will encounter difficulties in solving problems with large size ratios. The developed methodology is tested for several 3D thermo-mechanical coupling problems in large-size-ratio FGMs, and the numerical results demonstrate that the adaptive PINNs are effective and reliable for dealing with coupled problems in coating structures with large size ratios up to 109, as well as complex large-size-ratio geometries such as the electrostatic comb, the airplane and the submarine.

cs.CE

Partial Symbol Recovery for Interference Resilience in Low-Power Wide Area Networks

Recent years have witnessed the proliferation of Low-power Wide Area Networks (LPWANs) in the unlicensed band for various Internet-of-Things (IoT) applications. Due to the ultra-low transmission power and long transmission duration, LPWAN devices inevitably suffer from high power Cross Technology Interference (CTI), such as interference from Wi-Fi, coexisting in the same spectrum. To alleviate this issue, this paper introduces the Partial Symbol Recovery (PSR) scheme for improving the CTI resilience of LPWAN. We verify our idea on LoRa, a widely adopted LPWAN technique, as a proof of concept. At the PHY layer, although CTI has much higher power, its duration is relatively shorter compared with LoRa symbols, leaving part of a LoRa symbol uncorrupted. Moreover, due to its high redundancy, LoRa chips within a symbol are highly correlated. This opens the possibility of detecting a LoRa symbol with only part of the chips. By examining the unique frequency patterns in LoRa symbols with time-frequency analysis, our design effectively detects the clean LoRa chips that are free of CTI. This enables PSR to only rely on clean LoRa chips for successfully recovering from communication failures. We evaluate our PSR design with real-world testbeds, including SX1280 LoRa chips and USRP B210, under Wi-Fi interference in various scenarios. Extensive experiments demonstrate that our design offers reliable packet recovery performance, successfully boosting the LoRa packet reception ratio from 45.2% to 82.2% with a performance gain of 1.8 times.

cs.NI

Decision Models for Workforce and Technology Planning in Services

Today's service companies operate in a technology-oriented and knowledge-intensive environment while recruiting and training individuals from an increasingly diverse population. One of the resulting challenges is ensuring strategic alignment between their two key resources - technology and workforce - through the resource planning and allocation processes. The traditional hierarchical decision approach to resource planning and allocation considers only technology planning as a strategic-level decision, with workforce recruiting and training planning as a subsequent tactical-level decision. However, two other decision approaches - joint and integrated - elevate workforce planning to the same strategic level as technology planning. Thus we investigate the impact of strategically aligning technology and workforce decisions through the comparison of joint and integrated models to each other and to a baseline hierarchical model in terms of the total cost. Numerical experiments are conducted to characterize key features of solutions provided by these approaches under conditions typically found in this type of service company. Our results show that the integrated model is the lowest cost across all conditions. This is because the integrated approach maintains a small but skilled workforce that can operate new and more advanced technology with higher capacity. However, the cost performance of the joint model is very close to the integrated model under many conditions and is easier to implement computationally and managerially, making it a good choice in many environments. Managerial insights derived from this study can serve as a valuable guide for choosing the proper decision approach for technology-oriented and knowledge-intensive service companies.

econ.GN

FC$^2$N: Fully Channel-Concatenated Network for Single Image Super-Resolution

Most current image super-resolution (SR) methods based on convolutional neural networks (CNNs) use residual learning in network structural design, which favors to effective back propagation and hence improves SR performance by increasing model scale. However, residual networks suffer from representational redundancy by introducing identity paths that impede the full exploitation of model capacity. Besides, blindly enlarging network scale can cause more problems in model training, even with residual learning. In this paper, a novel fully channel-concatenated network (FC$^2$N) is presented to make further mining of representational capacity of deep models, in which all interlayer skips are implemented by a simple and straightforward operation, i.e., weighted channel concatenation (WCC), followed by a 1$\times$1 conv layer. Based on the WCC, the model can achieve the joint attention mechanism of linear and nonlinear features in the network, and presents better performance than other state-of-the-art SR models with fewer model parameters. To our best knowledge, FC$^2$N is the first CNN model that does not use residual learning and reaches network depth over 400 layers. Moreover, it shows excellent performance in both largescale and lightweight implementations, which illustrates the full exploitation of the representational capacity of the model.

eess.IV

Free Side-channel Cross-technology Communication in Wireless Networks

Enabling direct communication between wireless technologies immediately brings significant benefits including, but not limited to, cross-technology interference mitigation and context-aware smart operation. To explore the opportunities, we propose FreeBee -- a novel cross-technology communication technique for direct unicast as well as cross-technology/channel broadcast among three popular technologies of WiFi, ZigBee, and Bluetooth. The key concept of FreeBee is to modulate symbol messages by shifting the timings of periodic beacon frames already mandatory for diverse wireless standards. This keeps our design generically applicable across technologies and avoids additional bandwidth consumption (i.e., does not incur extra traffic), allowing continuous broadcast to safely reach mobile and/or duty-cycled devices. A new \emph{interval multiplexing} technique is proposed to enable concurrent bro\-adcasts from multiple senders or boost the transmission rate of a single sender. Theoretical and experimental exploration reveals that FreeBee offers a reliable symbol delivery under a second and supports mobility of 30mph and low duty-cycle operations of under 5%.

cs.NI

Achieving Spectrum Efficient Communication Under Cross-Technology Interference

In wireless communication, heterogeneous technologies such as WiFi, ZigBee and BlueTooth operate in the same ISM band.With the exponential growth in the number of wireless devices, the ISM band becomes more and more crowded. These heterogeneous devices have to compete with each other to access spectrum resources, generating cross-technology interference (CTI). Since CTI may destroy wireless communication, this field is facing an urgent and challenging need to investigate spectrum efficiency under CTI. In this paper, we introduce a novel framework to address this problem from two aspects. On the one hand, from the perspective of each communication technology itself, we propose novel channel/link models to capture the channel/link status under CTI. On the other hand, we investigate spectrum efficiency from the perspective by taking all heterogeneous technologies as a whole and building crosstechnology communication among them. The capability of direct communication among heterogeneous devices brings great opportunities to harmoniously sharing the spectrum with collaboration rather than competition.

cs.NI