SearcharxivSearch

arXiv subjects

Bo Xiao

Publications and source records attributed to Bo Xiao.

At least 19 recordsLinked to original sources

Full-Body Golf Swing Kinematic Reconstruction From a Smartwatch IMU

Quantitative measurement of the golf swing is critical for evaluating technique and enabling individualized feedback. However, existing methods are impractical to use on the golf course: optical motion capture is laboratory-bound, camera-based methods require impractical camera placement, and multi-sensor inertial measurement unit (IMU) systems require multi-segment setup and calibration. We thus propose a single wrist-worn IMU approach for estimating full-body joint angles during golf swings. The proposed Wrist-IMU Temporal Kinematic Network (WIT-KinNet) leverages modality-specific IMU embeddings and temporal kinematic encoding to learn wrist-to-body motion dependencies and estimate full-body joint angles during golf swings. Thirty-six golfers spanning beginner and skilled players, performed full, half, and quarter swings using seven club types: driver, 3-wood, 5-hybrid, 5-iron, 7-iron, 9-iron, and sand wedge. The proposed WIT-KinNet was evaluated under subject-wise cross-validation using synchronized smartwatch IMU data and ground-truth kinematics derived from an optical motion capture system. The proposed approach achieved a mean absolute error of 8.11 $\pm$ 1.84$^\circ$ across full-body joint angles. High temporal correlation was observed for pelvic rotation and upper torso rotation (r = 0.98 and 0.97, respectively), with X-factor and S-factor also showing strong correlation (r = 0.96 and 0.96). Linear mixed-effects models of the error revealed that swing amplitude, skill level, and club type all significantly affected measurement differences (p $<$ 0.05). The results establish the first single wrist-worn IMU approach for estimating full-body golf swing kinematics, enabling practical swing analysis during real gameplay.

cs.CV

Understanding and Modeling Perceived Cognitive and Physical Strain Dynamics for Planning-Oriented Human-Robot Collaboration in Prefabricated Construction

Human-robot collaboration (HRC) in prefabricated construction requires planning approaches that consider not only productivity but also time-dependent worker states during repeated work and rest. Existing planning models often rely on simplified assumptions about fatigue, workload, or recovery, with limited domain-specific empirical evidence on how perceived strain evolves. This study develops an empirically grounded, planning-oriented approach to characterize perceived strain accumulation and recovery in prefabricated construction HRC. A controlled repeated work-rest experiment assessed perceived cognitive and physical strain using the Rating Scale for Mental Effort and Borg's Rating of Perceived Exertion. Linear and exponential functional forms were evaluated, followed by mixed-effects modeling to examine collaborative conditions, session effects, and inter-individual variability. Results indicate that cognitive strain accumulation is best represented by a linear mixed-effects model, whereas rest-phase recovery follows nonlinear decay. The resulting planning-oriented models may inform future human-state-aware task allocation and scheduling research.

cs.RO

The Role of Instructional Guidance in Generative AI-Assisted Learning: Empirical Evidence from Construction Engineering Education

Generative artificial intelligence (AI) is increasingly used to support self-directed learning, yet student interaction with such systems often remains unstructured, limiting engagement in deeper cognitive processes. This study examines how instructional guidance shapes student and AI interaction in construction education. A five-step prompting framework grounded in Generative Learning Theory (GLT) is introduced to guide learner interaction during review activities. A controlled experiment compares three learning conditions: slide-based learning, unprompted AI-supported learning, and prompted AI-supported learning. Learning performance is assessed using multiple-choice and open-ended tasks, and user experience is measured using the User Experience Questionnaire (UEQ). Performance differences are concentrated on tasks requiring explanation and reasoning. The prompted condition achieves higher open-ended scores, with an improvement of approximately 2 or 3 points on a scale of 18 (p < 0.01), while no significant differences are observed in multiple-choice performance. The unprompted condition remains comparable to slide-based learning. These findings indicate that the effectiveness of AI-supported learning depends on how interaction is structured. The proposed framework provides a basis for integrating learning science principles into generative AI systems for construction education.

cs.HC

A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition

Grounded Multimodal Named Entity Recognition (GMNER) identifies named entities, including their spans and types, in natural language text and grounds them to the corresponding regions in associated images. Most existing approaches split this task into two steps: they first detect objects using a pre-trained general-purpose detector and then match named entities to the detected objects. However, these methods face a major limitation. Because pre-trained general-purpose object detectors operate independently of textual entities, they tend to detect common objects and frequently overlook specific fine-grained regions required by named entities. This misalignment between object detectors and entities introduces imprecision and can impair overall system performance. In this paper, we propose a proposal-free Query-Guided Network (QGN) that unifies multimodal reasoning and decoding through text guidance and cross- modal interaction. QGN enables accurate grounding and robust performance in open-domain scenarios. Extensive experiments demonstrate that QGN achieves top performance among compared GMNER models on widely used benchmarks.

cs.CV

Kinetic obstruction to pairing in the doped Kitaev-Heisenberg ladder

We investigate the hole-doped Kitaev-Heisenberg ($t$-$J$-$K$) model on a two-leg ladder geometry using the density-matrix renormalization group (DMRG). We first consider the behavior of the antiferromagnetic Kitaev (AFK) spin-liquid phase as a function of hopping strength $t$ and doping level. This reveals intriguing pairing tendencies only for $\frac{t}{K} \lesssim 0.65$, consistent with prior results on three-leg ladders, and firmly supports the emerging picture that the physics of doped Kitaev spin liquids strongly depends on the kinetic energy of the doped holes. Analysis of one- and two-hole doping uncovers close links between the spatial profiles of the plaquette operator and the charge density. We construct a doping-dependent phase diagram for antiferromagnetic Heisenberg interactions and intermediate hopping $t=1$. Upon doping, the rung-singlet region develops dominant superconducting correlations. Charge-density-wave correlations dominate at weak doping near the transition to the stripy phase. Spin-density wave-like behavior is found in the AFK and ferromagnetic Kitaev limits, and in the stripy phase.

cond-mat.str-el

BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding

Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.

cs.CV

Investigating a Quantum-Inspired Method for Quantum Dynamics

Building on recent advances in quantum algorithms which measure and reuse qubits and in efficient classical simulation leveraging projective measurements, we extend these frameworks to real-time dynamics of quantum many-body systems undergoing discrete-time and continuous-time Hamiltonian evolution, and find improvements that significantly reduce sampling overhead. The approach exploits causal light-cone structure by interleaving time and space evolution and applying projective measurements as soon as local subsystems reach the target physical time, suppressing entanglement growth. Comparing to time-evolving block decimation, the method reaches longer times per sample for the same resources. We also gain the ability to study dynamics of entanglement that would be occurring on quantum hardware when following similar protocols, such as the holographic quantum dynamics simulation framework. We show how to efficiently obtain local observables as well as equal-time and time-dependent correlation functions. Our findings show how optimizations for quantum hardware can benefit classical tensor network simulations and how such classical methods can yield insights into the utility of quantum simulations.

quant-ph

Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems

Enhancing customer experience is essential for business success, particularly as service demands grow in scale and complexity. Generative artificial intelligence and Large Language Models (LLMs) have empowered intelligent interaction systems to deliver efficient, personalized, and 24/7 support. In practice, intelligent interaction systems encounter several challenges: (1) Constructing high-quality data for cold-start training is difficult, hindering self-evolution and raising labor costs. (2) Multi-turn dialogue performance remains suboptimal due to inadequate intent understanding, rule compliance, and solution extraction. (3) Frequent evolution of business rules affects system operability and transferability, constraining low-cost expansion and adaptability. (4) Reliance on a single LLM is insufficient in complex scenarios, where the absence of multi-agent frameworks and effective collaboration undermines process completeness and service quality. (5) The open-domain nature of multi-turn dialogues, lacking unified golden answers, hampers quantitative evaluation and continuous optimization. To address these challenges, we introduce WOWService, an intelligent interaction system tailored for industrial applications. With the integration of LLMs and multi-agent architectures, WOWService enables autonomous task management and collaborative problem-solving. Specifically, WOWService focuses on core modules including data construction, general capability enhancement, business scenario adaptation, multi-agent coordination, and automated evaluation. Currently, WOWService is deployed on the Meituan App, achieving significant gains in key metrics, e.g., User Satisfaction Metric 1 (USM 1) -27.53% and User Satisfaction Metric 2 (USM 2) +25.51%, demonstrating its effectiveness in capturing user needs and advancing personalized service.

cs.CL

From weakly interacting spinons to tightly bound triplons in the frustrated quantum spin-Peierls chain

Fractionalized quasiparticles and their confinement into emergent bound states lie at the heart of modern quantum magnetism. While the evolution into magnonic bound states has been well characterized, experimental insight into the analogous transition to triplons remains limited. Here, using high-resolution neutron spectroscopy and state-of-the-art spin dynamics simulations, we uncover the transformation from weakly interacting spinons to tightly bound triplons in the spin-Peierls compound CuGeO3. Quantitative comparisons between the measured spectra and tensor network simulations reveal substantial next-nearest-neighbor frustration and weak external dimerization, placing the system deep within the spontaneously dimerized regime and near the exactly solvable Majumdar-Ghosh point. We further show an energy- and temperature-dependent evolution between two contrasting quasiparticle regimes: deconfined spinons with markedly suppressed interactions by frustration, and coherent triplonic bound states with no observable spinon degrees of freedom. Remarkably, triplon character persists into the two-particle regime, forming a structured two-triplon continuum with a spectral feature associated with a van Hove singularity at its lower boundary. These findings challenge the conventional view that robust triplons require strong external dimerization and demonstrate how the interplay between frustration and dimerization can reshape fractionalization and confinement.

cond-mat.str-el

Robust Chiral Edge Dynamics of a Kitaev Honeycomb on a Trapped Ion Processor

Kitaev's honeycomb model is a paradigmatic exactly solvable system hosting a quantum spin liquid with non-Abelian anyons and topologically protected edge modes, offering a platform for fault-tolerant quantum computation. However, real candidate Kitaev materials invariably include complex secondary interactions that obscure the realization of spin-liquid behavior and demand novel quantum computational approaches for efficient simulation. Here we report quantum simulations of a 22-site Kitaev honeycomb lattice on a trapped-ion quantum processor, without and with non-integrable Heisenberg interactions that are present in real materials. We develop efficient quantum circuits for ground-state preparation, achieving high accuracy with energy errors equivalent to an effective temperature of 0.2 (in units of the Kitaev interactions), consistent with the experimentally relevant spin-liquid regime. Starting from these states, we apply controlled perturbations and measure time-dependent spin correlations along the system's edge. In the non-Abelian phase, we observe chiral edge dynamics consistent with a non-zero Chern number, a hallmark of topological order, which vanishes upon transition to the Abelian toric code phase. Extending to the non-integrable Kitaev-Heisenberg model, we find that weak Heisenberg interactions preserve chiral edge dynamics, while stronger couplings suppress them, signaling the breakdown of topological protection. Our work demonstrates a viable route for probing dynamical signatures of topological order in quantum spin liquids using programmable quantum hardware, opening new pathways for quantum simulation of strongly correlated materials.

quant-ph

DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration

Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation - especially under semi-structured sparsity constraints. Existing pruning methods primarily focus on estimating the importance of individual weights, which limits their ability to preserve critical capabilities of the model. In this work, we propose a new perspective: rather than merely selecting which weights to prune, we first redistribute parameter importance to make the model inherently more amenable to pruning. By minimizing the information entropy of normalized importance scores, our approach concentrates importance onto a smaller subset of weights, thereby enhancing pruning robustness. We instantiate this idea through DenoiseRotator, which applies learnable orthogonal transformations to the model's weight matrices. Our method can be seamlessly integrated with existing pruning techniques such as Magnitude, SparseGPT, and Wanda. Evaluated on LLaMA3, Qwen2.5, and Mistral models under 50% unstructured and 2:4 semi-structured sparsity, DenoiseRotator consistently improves perplexity and zero-shot accuracy. For instance, on LLaMA3-70B pruned with SparseGPT at 2:4 semi-structured sparsity, DenoiseRotator reduces the perplexity gap to the dense model by 58%, narrowing the degradation from 8.1 to 3.4 points. Codes are available at https://github.com/Axel-gu/DenoiseRotator.

cs.LG

Multi-Sensor Fusion-Based Mobile Manipulator Remote Control for Intelligent Smart Home Assistance

This paper proposes a wearable-controlled mobile manipulator system for intelligent smart home assistance, integrating MEMS capacitive microphones, IMU sensors, vibration motors, and pressure feedback to enhance human-robot interaction. The wearable device captures forearm muscle activity and converts it into real-time control signals for mobile manipulation. The wearable device achieves an offline classification accuracy of 88.33\%\ across six distinct movement-force classes for hand gestures by using a CNN-LSTM model, while real-world experiments involving five participants yield a practical accuracy of 83.33\%\ with an average system response time of 1.2 seconds. In Human-Robot synergy in navigation and grasping tasks, the robot achieved a 98\%\ task success rate with an average trajectory deviation of only 3.6 cm. Finally, the wearable-controlled mobile manipulator system achieved a 93.3\%\ gripping success rate, a transfer success of 95.6\%\, and a full-task success rate of 91.1\%\ during object grasping and transfer tests, in which a total of 9 object-texture combinations were evaluated. These three experiments' results validate the effectiveness of MEMS-based wearable sensing combined with multi-sensor fusion for reliable and intuitive control of assistive robots in smart home scenarios.

cs.RO

Robustness of Vacancy-Bound Non-Abelian Anyons in the Kitaev Model in a Magnetic Field

Non-Abelian anyons in quantum spin liquids (QSLs) provide a promising route to fault-tolerant topological quantum computation. In the exactly solvable Kitaev honeycomb model, such anyons of the QSL state can be bound to nonmagnetic spin vacancies and endowed with non-Abelian statistics by an infinitesimal magnetic field. Here, we investigate how this approach for stabilizing non-Abelian anyons extends to a finite magnetic field represented by a proper Zeeman term. Through large-scale density-matrix renormalization group (DMRG) simulations, we compute the vacancy-anyon binding energy as a function of magnetic field for both the ferromagnetic (FM) and antiferromagnetic (AFM) Kitaev models. We find that anyon binding remains robust within the entire QSL phase for the FM Kitaev model but breaks down already inside this phase for the AFM Kitaev model. To compute a binding energy several orders of magnitude below the magnetic energy scale, we introduce both a refined definition and an extrapolation scheme based on carefully tailored perturbations.

cond-mat.str-el

Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism

Larger model sizes and longer sequence lengths have empowered the Large Language Model (LLM) to achieve outstanding performance across various domains. However, this progress brings significant storage capacity challenges for LLM pretraining. High Bandwidth Memory (HBM) is expensive and requires more advanced packaging technologies for capacity expansion, creating an urgent need for memory-efficient scheduling strategies. Yet, prior pipeline parallelism schedules have primarily focused on reducing bubble overhead, often neglecting memory efficiency and lacking compatibility with other memory-efficient strategies. Consequently, these methods struggle to meet the storage demands of storage capacity for next-generation LLM. This work presents ChronosPipe, a Chronos-aware pipeline parallelism for memory-efficient LLM pretraining. The core insight of ChronosPipe is to treat HBM as a fast but small 'cache,' optimizing and exploiting temporal locality within LLM pretraining to enhance HBM utilization. ChronosPipe introduces a pipeline scheduling strategy, Chronos-Pipe, to reduce the extrinsic overhead that disrupts the temporal locality of activations. Additionally, it leverages Chronos-Recomp and Chronos-Offload to efficiently harness the intrinsic temporal locality of activations and weights in Deep Neural Networks. Experiment results show that ChronosPipe can expand the trainable model size by 2.4x while maintaining comparable throughput, achieving 1.5x better than the 1F1B strategy combined with recomputation.

cs.DC

T-Stars-Poster: A Framework for Product-Centric Advertising Image Design

Creating advertising images is often a labor-intensive and time-consuming process. Can we automatically generate such images using basic product information like a product foreground image, taglines, and a target size? Existing methods mainly focus on parts of the problem and lack a comprehensive solution. To bridge this gap, we propose a novel product-centric framework for advertising image design called T-Stars-Poster. It consists of four sequential stages to highlight product foregrounds and taglines while achieving overall image aesthetics: prompt generation, layout generation, background image generation, and graphics rendering. Different expert models are designed and trained for the first three stages: First, a visual language model (VLM) generates background prompts that match the products. Next, a VLM-based layout generation model arranges the placement of product foregrounds, graphic elements (taglines and decorative underlays), and various nongraphic elements (objects from the background prompt). Following this, an SDXL-based model can simultaneously accept prompts, layouts, and foreground controls to generate images. To support T-Stars-Poster, we create two corresponding datasets with over 50,000 labeled images. Extensive experiments and online A/B tests demonstrate that T-Stars-Poster can produce more visually appealing advertising images.

cs.CV

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllable driving video generation, especially for geometry control, rendering existing control methods ineffective. To address these issues, we propose MagicDrive-V2, a novel approach that integrates the MVDiT block and spatial-temporal conditional encoding to enable multi-view video generation and precise geometric control. Additionally, we introduce an efficient method for obtaining contextual descriptions for videos to support diverse textual control, along with a progressive training strategy using mixed video data to enhance training efficiency and generalizability. Consequently, MagicDrive-V2 enables multi-view driving video synthesis with $3.3\times$ resolution and $4\times$ frame count (compared to current SOTA), rich contextual control, and geometric controls. Extensive experiments demonstrate MagicDrive-V2's ability, unlocking broader applications in autonomous driving.

cs.CV

A Risk Sensitive Contract-unified Reinforcement Learning Approach for Option Hedging

We propose a new risk sensitive reinforcement learning approach for the dynamic hedging of options. The approach focuses on the minimization of the tail risk of the final P&L of the seller of an option. Different from most existing reinforcement learning approaches that require a parametric model of the underlying asset, our approach can learn the optimal hedging strategy directly from the historical market data without specifying a parametric model; in addition, the learned optimal hedging strategy is contract-unified, i.e., it applies to different options contracts with different initial underlying prices, strike prices, and maturities. Our approach extends existing reinforcement learning methods by learning the tail risk measures of the final hedging P&L and the optimal hedging strategy at the same time. We carry out comprehensive empirical study to show that, in the out-of-sample tests, the proposed reinforcement learning hedging strategy can obtain statistically significantly lower tail risk and higher mean of the final P&L than delta hedging methods.

q-fin.RM

CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs

Parameter quantization for Large Language Models (LLMs) has attracted increasing attentions recently in reducing memory costs and improving computational efficiency. Early approaches have been widely adopted. However, the existing methods suffer from poor performance in low-bit (such as 2 to 3 bits) scenarios. In this paper, we present a novel and effective Column-Level Adaptive weight Quantization (CLAQ) framework by introducing three different types of adaptive strategies for LLM quantization. Firstly, a K-Means clustering based algorithm is proposed that allows dynamic generation of quantization centroids for each column of a parameter matrix. Secondly, we design an outlier-guided adaptive precision search strategy which can dynamically assign varying bit-widths to different columns. Finally, a dynamic outlier reservation scheme is developed to retain some parameters in their original float point precision, in trade off of boosted model performance. Experiments on various mainstream open source LLMs including LLaMA-1, LLaMA-2 and Yi demonstrate that our methods achieve the state-of-the-art results across different bit settings, especially in extremely low-bit scenarios. Code is available at https://github.com/fayuge/CLAQ.

cs.LG