SearcharxivSearch

arXiv subjects

Hui Zhao

Publications and source records attributed to Hui Zhao.

At least 19 recordsLinked to original sources

Stacking Polarity-Controlled Interlayer Photocarrier Dynamics in MoSe2/MoS2 Heterostructures

Control of interlayer photocarrier dynamics is central to optoelectronic applications of van der Waals heterostructures, yet deterministic and spatially uniform tuning strategies remain limited. Here we show that stacking polarity provides a global control parameter for photocarrier dynamics in MoSe$_2$/MoS$_2$ heterostructures. By comparing hexagonal (2H) and rhombohedral (3R) MoS$_2$ bilayers and engineering opposite interface terminations in 3R stacking, we resolve stacking-dependent interlayer charge-transfer dynamics using ultrafast pump--probe spectroscopy. While charge transfer in the 2H heterostructure occurs faster than the experimental resolution, the 3R heterostructures show time-resolvable charge transfer that slows from $0.25$ to $0.37$~ps depending on stacking polarity. Furthermore, the interlayer exciton lifetime is tuned from $\sim$$40$ to $\sim$$170$~ps. These effects arise from stacking-induced layer polarization in 3R MoS$_2$, which modulates interfacial wavefunction overlap.

cond-mat.mtrl-sci

SVOM/VT: Flight Model Verification and Pre-launch Testing

This paper presents pre-launch testing and calibration results for the SVOM/VT (Space-based Variable Objects Monitor, Visible Telescope) Flight Model (FM), validating its performance under simulated space conditions through thermal vacuum cycling, energy concentration analysis, stray light suppression, and CCD/electronics calibrations (gain, noise, quantum efficiency). The results confirm full compliance with design requirements: stray light suppression achieves point-source transmittance $<10^{-7}$ at $30^\circ$ off-axis, thermal control maintains stable CCD temperatures ($-75^\circ$C for the red channel, $-65^\circ$C for the blue channel), and detection sensitivity meets the limiting magnitude of 22.50 (SNR $>$ 3 with 300 seconds exposure). Early in-orbit tests further validate performance, yielding limiting magnitudes of 22.70 (V-band, red) and 22.78 (blue), consistent with pre-launch specifications.

astro-ph.IM

Vector Coded Caching Multiplicatively Boosts MU-MIMO Systems Under Practical Considerations

This work presents a first comprehensive analysis of the impact of vector coded caching (VCC) in multi-user multiple-input multiple-output (MU-MIMO) systems with multiple receive antennas and variable pathloss -- two key factors that critically influence systems with inherent MU unicasting behavior. We investigate two widely adopted precoding strategies: (i) blockdiagonalization (BD) at the transmitter combined with maximal ratio combining (MRC) at the receivers, and (ii) zero-forcing (ZF) precoding. Our analysis explicitly accounts for practical considerations such as channel fading, channel state information (CSI) acquisition overhead, and fairness-oriented power allocation. Our contributions span both analytical and simulation-based fronts. On the analytical side, we derive analytical expressions for the achievable throughput under BD-MRC and ZF, highlighting the performance benefits of equipping multi-antenna users with cache-aided interference management. Specifically, we develop a low-complexity BD-MRC optimization method that leverages matrix structure to significantly reduce the dimensionality involved in precoding computation, followed by solving the associated maxmin fairness problem through an efficient one-dimensional search. In the massive MIMO regime, an asymptotic expression for the achievable throughput over Rayleigh fading channels is also derived. Simulations validate our theoretical results, confirming that VCC delivers substantial performance gains over optimized cacheless MU-MIMO systems. For example, with 32 transmit antennas and 2 receive antennas per user, VCC yields throughput improvements exceeding 300%. These gains are further amplified under imperfect CSI at the transmitter, where VCC's ability to offload interference mitigation to the receivers ensures robust performance even in the face of degraded CSI quality and elevated acquisition costs.

cs.IT

Caching Yields up to 5x Spectral Efficiency in Multi-Beam Satellite Communications

This paper examines the integration of vector coded caching (VCC) into multi-beam satellite communications (SATCOM) systems and demonstrates that even limited receiver-side caching can substantially enhance spectral efficiency. By leveraging cached content to suppress interference, VCC enables the concurrent transmission of multiple precoded signal vectors that would otherwise require separate transmission resources. This leads to a multiplicative improvement in resource utilization in SATCOM. To characterize this performance, we model the satellite-to-ground channel using Rician-shadowed fading and after incorporating practical considerations such as matched-filter precoding, channel state information (CSI) acquisition overhead as well as CSI imperfections at the transmitter, we here derive closed-form expressions for the average sum rate and spectral efficiency gain of VCC in SATCOM. Our analysis, tightly validated through numerical simulations, reveals that VCC can yield spectral efficiency gains of 300% to 550% over traditional multi-user MISO SATCOM with the same resources. These gains -- which have nothing to do with multicasting, prefetching gains nor file popularity -- highlight VCC as a pure physical-layer solution for future high-throughput SATCOM systems, significantly narrowing the performance gap between satellite and wired networks.

cs.IT

RAIR: A Rule-Aware Benchmark Uniting Challenging Long-Tail and Visual Salience Subset for E-commerce Relevance Assessment

Search relevance plays a central role in web e-commerce. While large language models (LLMs) have shown significant results on relevance task, existing benchmarks lack sufficient complexity for comprehensive model assessment, resulting in an absence of standardized relevance evaluation metrics across the industry. To address this limitation, we propose Rule-Aware benchmark with Image for Relevance assessment(RAIR), a Chinese dataset derived from real-world scenarios. RAIR established a standardized framework for relevance assessment and provides a set of universal rules, which forms the foundation for standardized evaluation. Additionally, RAIR analyzes essential capabilities required for current relevance models and introduces a comprehensive dataset consists of three subset: (1) a general subset with industry-balanced sampling to evaluate fundamental model competencies; (2) a long-tail hard subset focus on challenging cases to assess performance limits; (3) a visual salience subset for evaluating multimodal understanding capabilities. We conducted experiments on RAIR using 14 open and closed-source models. The results demonstrate that RAIR presents sufficient challenges even for GPT-5, which achieved the best performance. RAIR data are now available, serving as an industry benchmark for relevance assessment while providing new insights into general LLM and Visual Language Model(VLM) evaluation.

cs.IR

LORE: A Large Generative Model for Search Relevance

Achievement. We introduce LORE, a systematic framework for Large Generative Model-based relevance in e-commerce search. Deployed and iterated over three years, LORE achieves a cumulative +27\% improvement in online GoodRate metrics. This report shares the valuable experience gained throughout its development lifecycle, spanning data, features, training, evaluation, and deployment. Insight. While existing works apply Chain-of-Thought (CoT) to enhance relevance, they often hit a performance ceiling. We argue this stems from treating relevance as a monolithic task, lacking principled deconstruction. Our key insight is that relevance comprises distinct capabilities: knowledge and reasoning, multi-modal matching, and rule adherence. We contend that a qualitative-driven decomposition is essential for breaking through current performance bottlenecks. Contributions. LORE provides a complete blueprint for the LLM relevance lifecycle. Key contributions include: (1) A two-stage training paradigm combining progressive CoT synthesis via SFT with human preference alignment via RL. (2) A comprehensive benchmark, RAIR, designed to evaluate these core capabilities. (3) A query frequency-stratified deployment strategy that efficiently transfers offline LLM capabilities to the online system. LORE serves as both a practical solution and a methodological reference for other vertical domains.

cs.IR

Lessons from $\alpha$-RuCl3 for pursuing quantum spin liquid physics in atomically thin materials

Quantum spin liquids can arise from Kitaev magnetic interactions, and exhibit fractionalized excitations with the potential for a topological form of quantum computation. This review surveys recent experimental and theoretical progress on the pursuit of phenomena related to Kitaev magnetism in layered and exfoliatable materials, which offer numerous opportunities to apply powerful techniques from the field of atomically thin materials. We primarily focus on the antiferromagnetic Mott insulator $\alpha$-RuCl3, which exhibits Kitaev couplings and is readily exfoliated to single- or few-layer sheets, and thus serves as a test bed for developing probes of Kitaev phenomena in atomically thin materials and devices. We introduce the Kitaev model and how it is realized in $\alpha$-RuCl3 and other material candidates; and cover $\alpha$-RuCl3 synthesis and fabrication into van der Waals heterostructure devices. A key discovery is a work-function-mediated charge transfer that heavily dopes both the $\alpha$-RuCl3 and proximate materials, and can enhance Kitaev interactions by up to 50%. We further discuss a wide range of recent results in electronic transport and optical and tunneling spectroscopies of $\alpha$-RuCl3 devices. The experimental techniques and theoretical insights developed for $\alpha$-RuCl3 establish a framework for discovering and engineering superior two-dimensional Kitaev materials that may ultimately realize elusive quantum spin liquid phases.

cond-mat.str-el

Transient Absorption Spectroscopy of NbOI$_2$

NbOI$_2$ has recently emerged as a new van der Waals material combining semiconducting behavior with intrinsic in plane ferroelectricity and pronounced transport and optical anisotropy. However, its photocarrier dynamics remain largely unexplored. Here we report transient absorption spectroscopy of NbOI$_2$ using femtosecond pump probe reflectance measurements. A pronounced transient absorption feature is observed near the 2.34 eV excitonic resonance, arising from photocarrier induced excitonic energy shifts and saturation. The decay dynamics reveal an exciton lifetime of several tens of picoseconds and show density-dependent behavior consistent with exciton exciton annihilation and defect-assisted Auger recombination, yielding a rate of 0.4 cm$^2$ s$^{-1}$, which is comparable to that in monolayer transition metal dichalcogenides. Polarization resolved measurements further reveal a pronounced in-plane anisotropy in the transient response that follows the linear absorption anisotropy. These findings provide fundamental insight into photocarrier dynamics in NbOI$_2$ and establish key parameters for understanding and exploiting its optoelectronic behavior.

cond-mat.mtrl-sci

Charge-Unified Semiconductor Switching Theory

Semiconductors and their downstream applications sustain the electronic, information, energy and industrial systems underpinning modern society. Improving their sustainability is therefore an urgent global priority, particularly as global electricity generation is projected to increase more than 2.5 fold by 2050. Yet, since the invention of the transistor in 1947, a unified, global view of circuit elements as media for charge redistribution and transfer one that reveals switching inertia and the dynamical nature of switching while connecting microscopic and macroscopic domains across the semiconductor value chain through a common theoretical language has remained absent. Switching consequently lacks a unified mechanistic account of its physical origins and spatiotemporal evolution, with fundamental disconnects between charge- and energy-conservation frameworks, among carrier dynamic mechanisms and across equivalent-circuit formalisms. These limitations fragment research domains and impede sustainability gains, particularly those requiring cross-domain causal information. Here, we present Charge-Unified Semiconductor Switching Theory (CUSST), a general theory that unifies circuit elements through a charge-mediated view, reveals switching inertia and the dynamical nature of switching, bridges these long-standing disconnects and establishes a unified conceptual, mechanistic, formal and analytical framework. Through these unifications, CUSST provides an unusually simple representation of otherwise fragmented switching phenomena. It establishes a unified micro-macro spatiotemporal view of switching, generalizes circuit theory, extends the application of conservation laws and provides a foundation for developing new theoretical systems.

eess.SY

Online Estimation of Table-Top Grown Strawberry Mass in Field Conditions with Occlusions

Accurate mass estimation of table-top grown strawberries under field conditions remains challenging due to frequent occlusions and pose variations. This study proposes a vision-based pipeline integrating RGB-D sensing and deep learning to enable non-destructive, real-time and online mass estimation. The method employed YOLOv8-Seg for instance segmentation, Cycle-consistent generative adversarial network (CycleGAN) for occluded region completion, and tilt-angle correction to refine frontal projection area calculations. A polynomial regression model then mapped the geometric features to mass. Experiments demonstrated mean mass estimation errors of 8.11% for isolated strawberries and 10.47% for occluded cases. CycleGAN outperformed large mask inpainting (LaMa) model in occlusion recovery, achieving superior pixel area ratios (PAR) (mean: 0.978 vs. 1.112) and higher intersection over union (IoU) scores (92.3% vs. 47.7% in the [0.9-1] range). This approach addresses critical limitations of traditional methods, offering a robust solution for automated harvesting and yield monitoring with complex occlusion patterns.

cs.CV

A Scientist Question: Research on the Impact of Super Structured Quadrilateral Meshes on Convergence and Accuracy of Finite Element Analysis

In the current practices of both industry and academia, the convergence and accuracy of finite element calculations are closely related to the methods and quality of mesh generation. For years, the research on high-quality mesh generation in the domestic academic field has mainly referred to the local quality of quadrilaterals and hexahedrons approximating that of squares and cubes. The main contribution of this paper is to propose a brand-new research direction and content: it is necessary to explore and study the influence of the overall global arrangement structure and pattern of super structured quadrilateral meshes on the convergence and calculation accuracy of finite element calculations. Through the research in this new field, it can help solve the non-rigorous state of serious reliance on "experience" in the mesh generation stage during simulation in the current industry and academia, and make clear judgments on which global arrangements of mesh generation can ensure the convergence of finite element calculations. In order to generate and design super-structured quadrilateral meshes with controllable overall arrangement structures, a large number of modern two-dimensional and three-dimensional geometric topology theories are required, such as moduli space, Teichm\"uller space, harmonic foliations, dynamical systems, surface mappings, meromorphic quadratic differentials, surface mappings, etc.

cs.GR

SA-WiSense: A Blind-Spot-Free Respiration Sensing Framework for Single-Antenna Wi-Fi Devices

Wi-Fi sensing offers a promising technique for contactless human respiration monitoring. A key challenge, however, is the blind spot problem caused by random phase offsets that corrupt the complementarity of respiratory signals. To address the challenge, we propose a single-antenna-Wi-Fi-sensing (SA-WiSense) framework to improve accuracy of human respiration monitoring, robust against random phase offsets. The proposed SA-WiSense framework is cost-efficient, as only a single antenna is used rather than multiple antennas as in the previous works. Therefore, the proposed framework is applicable to Internet of Thing (IoT), where most of sensors are equipped with a single antenna. On one hand, we propose a cross-subcarrier channel state information (CSI) ratio (CSCR) based blind spot mitigation approach for IoT, where the ratios of two values of CSI between subcarriers are leveraged to mitigate random phase offsets. We prove that the random phase offsets can be cancelled by the proposed CSCR approach, thereby restoring the inherent complementarity of signals for blind-spot-free sensing. On the other hand, we propose a genetic algorithm (GA) based subcarrier selection (GASS) approach by formulating an optimization problem in terms of the sensing-signal-to-noise ratio (SSNR) of CSCR between subcarriers. GA is utilized to solve the formulated optimization problem. We use commodity ESP32 microcontrollers to build an experiment test. The proposed works are validated to achieve an detection rate of 91.2% for respiration monitoring at distances up to 8.0 meters, substantially more accurate than the state-of-the-art methods with a single antenna.

eess.SP

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a physically realistic VLN platform supporting humanoid, quadruped, and wheeled robots. For the first time, we systematically evaluate several ego-centric VLN methods in physical robotic settings across different technical pipelines, including classification models for single-step discrete action prediction, a diffusion model for dense waypoint prediction, and a train-free, map-based large language model (LLM) integrated with path planning. Our results reveal significant performance degradation due to limited robot observation space, environmental lighting variations, and physical challenges like collisions and falls. This also exposes locomotion constraints for legged robots in complex environments. VLN-PE is highly extensible, allowing seamless integration of new scenes beyond MP3D, thereby enabling more comprehensive VLN evaluation. Despite the weak generalization of current models in physical deployment, VLN-PE provides a new pathway for improving cross-embodiment's overall adaptability. We hope our findings and tools inspire the community to rethink VLN limitations and advance robust, practical VLN models. The code is available at https://crystalsixone.github.io/vln_pe.github.io/.

cs.RO

Multipartite entanglement based on realignment moments

Based on the realignment moments of density matrix, we study parameterized entanglement criteria for bipartite and multipartite states. By adjusting the different parameter values, our criterion can detect not only bound entangled states, but also non-positive partial transpose entangled states for bipartite quantum systems. Moreover, we propose the definition of multipartite realignment moments and generalize the result of bipartite systems to obtain a sufficient criterion to detect entanglement for multipartite quantum states in arbitrary dimensions. And we further improve the conclusion to obtain another new entanglement criterion. The new method can detect more entangled states than previous methods as backed by detailed examples.

quant-ph

Gradient Deconfliction via Orthogonal Projections onto Subspaces For Multi-task Learning

Although multi-task learning (MTL) has been a preferred approach and successfully applied in many real-world scenarios, MTL models are not guaranteed to outperform single-task models on all tasks mainly due to the negative effects of conflicting gradients among the tasks. In this paper, we fully examine the influence of conflicting gradients and further emphasize the importance and advantages of achieving non-conflicting gradients which allows simple but effective trade-off strategies among the tasks with stable performance. Based on our findings, we propose the Gradient Deconfliction via Orthogonal Projections onto Subspaces (GradOPS) spanned by other task-specific gradients. Our method not only solves all conflicts among the tasks, but can also effectively search for diverse solutions towards different trade-off preferences among the tasks. Theoretical analysis on convergence is provided, and performance of our algorithm is fully testified on multiple benchmarks in various domains. Results demonstrate that our method can effectively find multiple state-of-the-art solutions with different trade-off strategies among the tasks on multiple datasets.

cs.LG

MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling

Click-Through Rate (CTR) prediction is a crucial task in recommendation systems, online searches, and advertising platforms, where accurately capturing users' real interests in content is essential for performance. However, existing methods heavily rely on ID embeddings, which fail to reflect users' true preferences for content such as images and titles. This limitation becomes particularly evident in cold-start and long-tail scenarios, where traditional approaches struggle to deliver effective results. To address these challenges, we propose a novel Multi-modal Content Interest Modeling paradigm (MIM), which consists of three key stages: Pre-training, Content-Interest-Aware Supervised Fine-Tuning (C-SFT), and Content-Interest-Aware UBM (CiUBM). The pre-training stage adapts foundational models to domain-specific data, enabling the extraction of high-quality multi-modal embeddings. The C-SFT stage bridges the semantic gap between content and user interests by leveraging user behavior signals to guide the alignment of embeddings with user preferences. Finally, the CiUBM stage integrates multi-modal embeddings and ID-based collaborative filtering signals into a unified framework. Comprehensive offline experiments and online A/B tests conducted on the Taobao, one of the world's largest e-commerce platforms, demonstrated the effectiveness and efficiency of MIM method. The method has been successfully deployed online, achieving a significant increase of +14.14% in CTR and +4.12% in RPM, showcasing its industrial applicability and substantial impact on platform performance. To promote further research, we have publicly released the code and dataset at https://pan.quark.cn/s/8fc8ec3e74f3.

cs.IR

Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains. Math Word Problems (MWPs) serve as a crucial benchmark for evaluating LLMs' reasoning abilities. While most research primarily focuses on improving accuracy, it often neglects understanding and addressing the underlying patterns of errors. Current error classification methods rely on static and predefined categories, which limit their ability to capture the full spectrum of error patterns in mathematical reasoning. To enable systematic error analysis, we collect error samples from 15 different LLMs of varying sizes across four distinct MWP datasets using multiple sampling strategies. Based on this extensive collection, we introduce MWPES-300K, a comprehensive dataset containing 304,865 error samples that cover diverse error patterns and reasoning paths. To reduce human bias and enable fine-grained analysis of error patterns, we propose a novel framework for automated dynamic error classification in mathematical reasoning. Experimental results demonstrate that dataset characteristics significantly shape error patterns, which evolve from basic to complex manifestations as model capabilities increase. With deeper insights into error patterns, we propose Error-Aware Prompting (EAP) that incorporates common error patterns as explicit guidance, leading to significant improvements in mathematical reasoning performance.

cs.CL

Machine Learning Analysis of Anomalous Diffusion

The rapid advancements in machine learning have made its application to anomalous diffusion analysis both essential and inevitable. This review systematically introduces the integration of machine learning techniques for enhanced analysis of anomalous diffusion, focusing on two pivotal aspects: single trajectory characterization via machine learning and representation learning of anomalous diffusion. We extensively compare various machine learning methods, including both classical machine learning and deep learning, used for the inference of diffusion parameters and trajectory segmentation. Additionally, platforms such as the Anomalous Diffusion Challenge that serve as benchmarks for evaluating these methods are highlighted. On the other hand, we outline three primary strategies for representing anomalous diffusion: the combination of predefined features, the feature vector from the penultimate layer of neural network, and the latent representation from the autoencoder, analyzing their applicability across various scenarios. This investigation paves the way for future research, offering valuable perspectives that can further enrich the study of anomalous diffusion and advance the application of artificial intelligence in statistical physics and biophysics.

cs.LG