SearcharxivSearch

arXiv subjects

Liyu Zhang

Publications and source records attributed to Liyu Zhang.

14 recordsLinked to original sources

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

Human Activity Recognition (HAR) using inertial measurement units (IMUs) enables a wide range of applications, yet the field still lacks a unified model that can generalize across diverse subjects, devices, and activities. Training such a model is difficult due to two key challenges: sensing heterogeneity -- differences in sampling rates, channel configurations, and sensor placements -- and poor generalization to unseen activities and label vocabularies. We introduce HALO (Heterogeneity-Aware Language-aligned Open-set model), a domain-specific IMU foundation model that addresses both challenges through a two-stage training framework. Stage 1 pretrains the IMU encoder with heterogeneity-aware self-supervised learning, including adaptive-pooling tokenization, channel-independent feature extraction, and contextualized sensor conditioning that injects natural-language sensor descriptions into each channel embedding. Stage 2 aligns this IMU encoder with text embeddings via synonym-aware soft contrastive learning, enabling open-set recognition via cosine-similarity retrieval without per-dataset classifiers. Trained on 10 public HAR datasets and evaluated on 7 held-out datasets, HALO outperforms five state-of-the-art baselines on all 8 aggregate metrics, and still leads on 3 of 4 settings under baseline-matched inputs. Despite using only ~35M trainable parameters -- 10x fewer than the latest foundation model MOMENT (341.2M) -- HALO improves zero-shot open-set accuracy, measured over all 87 training labels, by 13.7 percentage points. On two further datasets with severe distribution shift, every model including HALO collapses zero-shot. A video demonstration of HALO's performance in real world is available at https://youtu.be/rooVKragtFU

cs.LG

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems focus primarily on optimizing task accuracy and rely on cloud-hosted models for inference, which introduces privacy concerns and network-dependent latency. As a result, fully on-device deployment of mobile GUI agents remains underexplored. We propose MobileExplorer, a new framework that accelerates on-device inference for vision-based mobile GUI agents via online exploration. The key idea is to exploit the long per-step reasoning time of vision-language models (VLMs) by performing lightweight, parallel exploration of UI elements. During model inference, the agent proactively probes semantically relevant UI elements and records these exploration traces as structured memory. To ensure reliable execution in live mobile environments, we design a two-level rollback mechanism that robustly restores the initial UI state when a fast but naive backtracking strategy fails. The collected exploration traces are then summarized into concise contextual hints and injected into the prompt to enhance the subsequent reasoning step. We evaluate MobileExplorer on multiple off-the-shelf devices using the AndroidWorld benchmark, as well as newly designed, more complex tasks and dynamic on-device environments. MobileExplorer reduces the average number of reasoning steps and end-to-end latency by 23\%, while maintaining or improving task success rates by up to 5\%. A video demonstration of MobileExplorer performance in the real world is available at https://youtu.be/thK7MJmdlvM .

cs.AI

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models

Post training via GRPO has demonstrated remarkable effectiveness in improving the generation quality of flow-matching models. However, GRPO suffers from inherently low sample efficiency due to its on-policy training paradigm. To address this limitation, we present OP-GRPO, the first Off-Policy GRPO framework tailored for flow-matching models. First, we actively select high-quality trajectories and adaptively incorporate them into a replay buffer for reuse in subsequent training iterations. Second, to mitigate the distribution shift introduced by off-policy samples, we propose a sequence-level importance sampling correction that preserves the integrity of GRPO's clipping mechanism while ensuring stable policy updates. Third, we theoretically and empirically show that late denoising steps yield ill-conditioned off-policy ratios, and mitigate this by truncating trajectories at late steps. Across image and video generation benchmarks, OP-GRPO achieves comparable or superior performance to Flow-GRPO with only 34.2% of the training steps on average, yielding substantial gains in training efficiency while maintaining generation quality.

cs.CV

InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning

Supervised fine-tuning (SFT) is fundamental to adapting large language models, yet training on complete datasets incurs prohibitive costs with diminishing returns. Existing data selection methods suffer from severe domain specificity: techniques optimized for general instruction-following fail on reasoning tasks, and vice versa. We observe that measuring entropy differences between base models and minimally instruction-tuned calibrated models reveals a pattern -- samples with the lowest differential entropy consistently yield optimal performance across domains, yet this principle manifests domain-adaptively: reasoning tasks favor entropy increase (cognitive expansion), while general tasks favor entropy decrease (cognitive compression). We introduce InstructDiff, a unified framework that operationalizes differential entropy as a domain-adaptive selection criterion through warmup calibration, bi-directional NLL filtering, and entropy-based ranking. Extensive experiments show that InstructDiff achieves 17\% relative improvement over full data training on mathematical reasoning and 52\% for general instruction-following, outperforming prior baselines while using only 10\% of the data.

cs.CL

Chorus: Harmonizing Context and Sensing Signals for Data-Free Model Customization in IoT

A key bottleneck toward scalable IoT sensing is efficiently adapting trained AI models to new deployment conditions. Context shifts, such as changes in sensor placement or ambient environments, can substantially alter sensing patterns and degrade model performance. We present Chorus, a context-bridged, data-free post-deployment model customization approach that adapts sensing models to unseen contexts without requiring target-domain sensor data or post-deployment retraining. Chorus learns compact, transferable context representations and aligns them with the sensor latent space using unlabeled sensor-context pairs, bridging context generalization with sensing-data generalization. It then uses a lightweight gated prediction head to integrate context priors at inference and an adaptive caching mechanism to reuse context representations when no context shift is detected, reducing on-device overhead. Experiments on IMU sensing, speech enhancement, and WiFi sensing under diverse context shifts show that Chorus outperforms state-of-the-art baselines by up to 20.2% in unseen contexts, achieves inference latency comparable to sensor-only deployment, and remains stable under continuous context transitions and varied context descriptions. A video demonstration is available at https://youtu.be/yANTZsk0TVU.

cs.LG

Collaborate sim and real: Robot Bin Packing Learning in Real-world and Physical Engine

The 3D bin packing problem, with its diverse industrial applications, has garnered significant research attention in recent years. Existing approaches typically model it as a discrete and static process, while real-world applications involve continuous gravity-driven interactions. This idealized simplification leads to infeasible deployments (e.g., unstable packing) in practice. Simulations with physical engine offer an opportunity to emulate continuous gravity effects, enabling the training of reinforcement learning (RL) agents to address such limitations and improve packing stability. However, a simulation-to-reality gap persists due to dynamic variations in physical properties of real-world objects, such as various friction coefficients, elasticity, and non-uniform weight distributions. To bridge this gap, we propose a hybrid RL framework that collaborates with physical simulation with real-world data feedback. Firstly, domain randomization is applied during simulation to expose agents to a spectrum of physical parameters, enhancing their generalization capability. Secondly, the RL agent is fine-tuned with real-world deployment feedback, further reducing collapse rates. Extensive experiments demonstrate that our method achieves lower collapse rates in both simulated and real-world scenarios. Large-scale deployments in logistics systems validate the practical effectiveness, with a 35\% reduction in packing collapse compared to baseline methods.

cs.RO

Direct Mapping of Intrinsic Topology of Bound States in the Continuum via Nonlinear Emission

The direct mapping of the intrinsic topology in a leaky photonic band is crucial and challenging in topological photonics. For instance, observables in bound states in the continuum (BICs) feature complex topological textures such as a polarization vortex in momentum space, which nonetheless is difficult to be characterized in far-field scattering, especially considering the dominant direct channel. Here, we propose and experimentally demonstrate a hybrid nonlinear metasurface that enables a direct visualization of the intrinsic topology in BICs via second-harmonic generation (SHG). The enhanced local-source of SHG from the ultrathin indium tin oxide can effectively excite the emissions from the eigenmodes of a TiO2 photonics crystal slab, achieving three-order enhancement of SHG magnitudes. Importantly, these enhanced SH emissions carry topological polarization textures of BICs to the far field. With this, we can directly construct polarization vector maps of symmetry-protected BICs and chiral symmetry-broken quasi-BICs, clearly visualizing the winding structure around V points, the generation and evolution of chiral C points. This work provides a universal approach for characterizing topological photonic systems via coherent nonlinearity processes, opening new avenues for studying topological phenomena in non-Hermitian photonic systems.

physics.optics

Manipulating terahertz phonon-polariton in the ultrastrong coupling regime with bound states in the continuum

The strong coupling between photons and phonons in polar materials gives rise to phonon-polaritons that encapsulate a wealth of physical information, offering crucial tools for the ultrafast terahertz sources and the topological engineering of terahertz light. However, it is still quite challenging to form and manipulate the terahertz phonon-polaritons under the ultrastrong coupling regime till now. In this work, we demonstrate the ultrastrong coupling between the phonon (at 0.95 THz) in a MaPbI 3 film and the metallic bound states in the continuum (BICs) in Au metasurfaces. The Rabi splitting can be continuously tuned from 28% to 48.4% of the phonon frequency by adjusting the parameters (size, shape and period) of Au metasurfaces, reaching the ultrastrong coupling regime. By introducing wavelet transform, the mode evolution information of the terahertz phonon-polariton is successfully extracted. It indicates that the phonon radiation intensity of the MaPbI 3 film is enhanced as the coupling strength is increased. This work not only establishes a new platform for terahertz devices but also opens new avenues for exploring the intricate dynamics of terahertz phonon-polaritons.

physics.optics

Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models

Large Language Models (LLMs) have achieved remarkable success but remain data-inefficient, especially when learning from small, specialized corpora with limited and proprietary data. Existing synthetic data generation methods for continue pre-training focus on intra-document content and overlook cross-document knowledge associations, limiting content diversity and depth. We propose Synthetic-on-Graph (SoG), a synthetic data generation framework that incorporates cross-document knowledge associations for efficient corpus expansion. SoG constructs a context graph by extracting entities and concepts from the original corpus, representing cross-document associations, and employing a graph walk strategy for knowledge-associated sampling. This enhances synthetic data diversity and coherence, enabling models to learn complex knowledge structures and handle rare knowledge. To further improve the quality of synthetic data, we integrate two complementary strategies, Chain-of-Thought (CoT) and Contrastive Clarifying (CC), to enhance both reasoning capability and discriminative power. Extensive experiments demonstrate that SoG surpasses state-of-the-art (SOTA) methods on multi-hop and domain-specific question answering, while achieving competitive performance on long-context reading comprehension. These results highlight the superior generalization ability of SoG. Our work advances the paradigm of synthetic data generation and offers practical solutions for efficient knowledge acquisition in LLMs, particularly for downstream tasks and domains with limited training data.

cs.CL

On the Role of Entity and Event Level Conceptualization in Generalizable Reasoning: A Survey of Tasks, Methods, Applications, and Future Directions

Conceptualization, a fundamental element of human cognition, plays a pivotal role in human generalizable reasoning. Generally speaking, it refers to the process of sequentially abstracting specific instances into higher-level concepts and then forming abstract knowledge that can be applied in unfamiliar or novel situations. This enhances models' inferential capabilities and supports the effective transfer of knowledge across various domains. Despite its significance, the broad nature of this term has led to inconsistencies in understanding conceptualization across various works, as there exists different types of instances that can be abstracted in a wide variety of ways. There is also a lack of a systematic overview that comprehensively examines existing works on the definition, execution, and application of conceptualization to enhance reasoning tasks. In this paper, we address these gaps by first proposing a categorization of different types of conceptualizations into four levels based on the types of instances being conceptualized, in order to clarify the term and define the scope of our work. Then, we present the first comprehensive survey of over 150 papers, surveying various definitions, resources, methods, and downstream applications related to conceptualization into a unified taxonomy, with a focus on the entity and event levels. Furthermore, we shed light on potential future directions in this field and hope to garner more attention from the community.

cs.CL

Seismic resolution enhancement via deep Learning with Knowledge Distillation and Domain Adaptation

High-resolution processing of seismic signals is crucial for subsurface geological characterization and thin-layer reservoir identification. Traditional high-resolution algorithms can partially recover high-frequency information but often lack robustness, computational efficiency, and consideration of inter-trace structural relationships. Many deep learning methods use end-to-end architectures that do not incorporate prior knowledge or address data domain disparities, leading to limited generalization.To overcome these challenges, this paper presents the Domain-Adaptive Knowledge Distillation Network (DAKD-Net), which integrates a knowledge distillation strategy with a domain adaptation mechanism for high-resolution seismic data processing. Trained on datasets from forward modeling, DAKD-Net establishes physical relationships between low and high-resolution data, extracting high-frequency prior knowledge during a guided phase before detail restoration without prior conditions. Domain adaptation enhances the model's generalization to real seismic data, improving both generalization capability and structural expression accuracy.DAKD-Net employs a U-Net backbone to extract spatial structural information from multi-trace seismic profiles. The knowledge distillation mechanism enables prior knowledge transfer, allowing recovery of high-resolution data directly from low-resolution inputs. Domain-adaptive fine-tuning further enhances the network's performance in actual survey areas. Experimental results show that DAKD-Net outperforms traditional methods and classical deep networks in longitudinal resolution and complex structural detail restoration, demonstrating strong robustness and practicality.

physics.geo-ph

ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning

Knowledge Editing (KE) aims to adjust a Large Language Model's (LLM) internal representations and parameters to correct inaccuracies and improve output consistency without incurring the computational expense of re-training the entire model. However, editing commonsense knowledge still faces difficulties, including limited knowledge coverage in existing resources, the infeasibility of annotating labels for an overabundance of commonsense knowledge, and the strict knowledge formats of current editing methods. In this paper, we address these challenges by presenting ConceptEdit, a framework that integrates conceptualization and instantiation into the KE pipeline for LLMs to enhance their commonsense reasoning capabilities. ConceptEdit dynamically diagnoses implausible commonsense knowledge within an LLM using another verifier LLM and augments the source knowledge to be edited with conceptualization for stronger generalizability. Experimental results demonstrate that LLMs enhanced with ConceptEdit successfully generate commonsense knowledge with improved plausibility compared to other baselines and achieve stronger performance across multiple question answering benchmarks. Our data, code, and models are publicly available at https://github.com/HKUST-KnowComp/ConKE.

cs.CL

SAMG: Offline-to-Online Reinforcement Learning via State-Action-Conditional Offline Model Guidance

Offline-to-online (O2O) reinforcement learning (RL) pre-trains models on offline data and refines policies through online fine-tuning. However, existing O2O RL algorithms typically require maintaining the tedious offline datasets to mitigate the effects of out-of-distribution (OOD) data, which significantly limits their efficiency in exploiting online samples. To address this deficiency, we introduce a new paradigm for O2O RL called State-Action-Conditional Offline \Model Guidance (SAMG). It freezes the pre-trained offline critic to provide compact offline understanding for each state-action sample, thus eliminating the need for retraining on offline data. The frozen offline critic is incorporated with the online target critic weighted by a state-action-adaptive coefficient. This coefficient aims to capture the offline degree of samples at the state-action level, and is updated adaptively during training. In practice, SAMG could be easily integrated with Q-function-based algorithms. Theoretical analysis shows good optimality and lower estimation error. Empirically, SAMG outperforms state-of-the-art O2O RL algorithms on the D4RL benchmark.

cs.LG

Angle dependent field-driven reorientation transitions in uniaxial antiferromagnet MnBi$_2$Te$_4$ single crystal

MnBi$_2$Te$_4$, a two-dimensional magnetic topological insulator with a uniaxial antiferromagnetic structure, is an ideal platform to realize quantum anomalous Hall effect. However, the strength of magnetic interactions is not clear yet. We performed systematic studies on the magnetization and angle dependent magnetotransport of MnBi$_2$Te$_4$ single crystal. The results show that the direction of the magnetic field has significant effects on the critical field values and magnetic structure of this compound, which leads to different magnetotransport behaviors. The field-driven reorientation transitions can be utilized to estimate the AFM interlayer exchange interaction coupling and uniaxial magnetic anisotropy D. The obtained Hamiltonian can well explain the experimental data by Monte Carlo simulations. Our comprehensive studies on the field-driven magnetic transitions phenomenon in MnBi$_2$Te$_4$ provide a general approach for other topological systems with antiferromagnetism.

cond-mat.mtrl-sci