SearcharxivSearch

arXiv subjects

Pei Zhang

Publications and source records attributed to Pei Zhang.

At least 19 recordsLinked to original sources

Comparing Tobit and Two-Part Hurdle Models for Semi-Continuous Longitudinal Data with an Application to Clonal Hematopoiesis

Zero-inflated nonnegative continuous longitudinal data frequently arise in biomedical studies where outcomes consist of a mixture of excess zeros and positive continuous measurements. Two widely used approaches for analyzing such data are mixed-model versions of Tobit and the two-part hurdle models. The Tobit assumes a latent regression model that is censored below a specified threshold, while the hurdle separately models the continuous positive outcomes and the binary indicator of being positive. The choice between these models has rarely been systematically discussed, and inappropriate model choice may lead to biased estimation and misleading scientific interpretations. In this paper, we derive rigorous mathematical conditions under which the two models are equivalent and show that the Tobit can be viewed as a special case of the hurdle model when the link function for the binary process is probit. Based on simulation studies, we found that the hurdle is more flexible and robust than the Tobit model. On the other hand, if the assumptions of the Tobit are met, this model is easier to interpret since it does not require distinct inferences on both the continuous and binary processes. We therefore recommend that the Tobit model only be used when these assumptions are scientifically plausible and empirically supported; otherwise, the hurdle model is preferable. We applied both models to study the dynamics of somatic mosaicism using longitudinal clonal fraction measurements from the Prostate, Lung, Colorectal, and Ovarian study data while accounting for excess zero values. Estimates obtained from the Tobit and hurdle models were broadly consistent with those from a standard linear model that ignored zero inflation. These findings provide additional support for previously reported associations in studies of clonal hematopoiesis across different types of mosaic chromosomal alterations.

stat.AP

Accelerated iterative method for solving the steady-state Boltzmann equation

The efficient simulation of steady-state rarefied gas flows remains a significant computational challenge due to the high dimensionality of the collision integral and the severe numerical stiffness in the near-continuum regime. In this work, we propose a modified Newton method equipped with a macroscopic synthetic system (Newton-MS) for the steady-state Boltzmann equation with the quadratic collision operator. In Newton-MS, the modified Newton iteration is utilized as the outer nonlinear solver, while each Newton correction equation is solved by an inner source iteration, where the linearized collision operator is utilized to approximate the quadratic collision model, and it is reduced into a linear iteration. Moreover, a macroscopic synthetic system based on Chapman-Enskog closure is derived to accelerate the convergence of the linear inner iteration in the continuum limit. Besides, the fully discrete macroscopic synthetic system is deduced under the framework of the discontinuous Galerkin method to reduce computational cost compared to directly discretizing the continuous macroscopic synthetic system. Several numerical examples, including the 1D Fourier, Couette flow problem, and the 2D cavity flow and thermal-driven cavity flow, are studied to validate the high efficiency of Newton-MS.

math.NA

On-Device Robotic Planning: Eliminating Inference Redundancy for Efficient Decision-Making

Reasoning-based robotic policies using large language and vision-language models achieve strong semantic planning capabilities but mostly suffer from a high inference latency that limits practical real-time deployment. In this work, we observe that robotic reasoning workloads contain substantial temporal redundancy, where consecutive observations frequently produce identical actions and subgoals. Based on this insight, we present REIS, a human cognition inspired robotic decision-making framework that minimizes unnecessary reasoning while preserving semantic adaptability. REIS combines lightweight scene gating, KV-steered affordance routing, and deliberative reasoning to accelerate robotic control under embodied constraints. Experiments on ALFRED, and real-world robotic tasks demonstrate that REIS significantly suppresses reasoning overhead while maintaining competitive task performance.

cs.RO

See Silhouettes in Motion with Neuromorphic Vision

Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization uses a minimal language to convey essential vision cues for maximum downstream efficiency, especially for tasks that require simple geometric, topological reasoning rather than heavy appearance modeling. The catch is that frame-based imaging often struggles on mobile platforms like drones, self-driving cars, and underwater vehicles, in which rapid motion causes severe motion blur and harsh lighting washes out scene details. To overcome these physical limits, neuromorphic vision via event cameras, featuring microsecond time resolution and high dynamic range, steps in as a natural solution. Building upon this event-driven paradigm, we propose a simple yet effective dual-modal approach that harnesses the synergy between frames and events for training-free, real-time, high-frame-rate binarization on CPU-only devices. Extensive evaluations show that it earns competitive performance against leading techniques in reducing blur artifacts and delivers impressive improvements under challenging illumination at a lower computational cost. Besides, its asynchronous nature bypasses long-standing event-scarcity issues that break traditional time-binning reconstruction at fixed time slots, maintaining clear target shapes even at extreme kilohertz frame rates. Its binary results further serve as reliable representations to facilitate a range of downstream tasks. This work paves the way towards lightweight perception and interaction in embodied intelligence on resource-constrained edge platforms.

eess.IV

LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation

Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial-domain models struggle to disentangle authentic gaze cues from irrelevant visual attributes. In this paper, we propose LISA, a \textbf{L}anguage-guided \textbf{I}nterference-aware \textbf{S}patial-Frequency \textbf{A}ttention framework that combines frequency-domain priors with vision-language knowledge. Observing that the amplitude spectrum remains relatively stable even under spatial perturbations, we design a dual-domain fusion mechanism. It integrates stable low-frequency semantics into high-frequency details, employing spatial attention to precisely target ocular regions. To reduce semantic ambiguity, we also introduce a training-time disentanglement strategy. Using a frozen CLIP encoder and orthogonal regularization, we explicitly separate gaze features from appearance interference. Experiments on two benchmarks show that LISA achieves state-of-the-art performance, with significantly improved robustness against occlusions and lighting variations. The code repository is available at https://github.com/Mason-bupt/LISA.

cs.CV

Aquatic Neuromorphic Optical Flow

Underwater environments impose severe constraints on conventional imaging systems and demand solutions that balance high-quality sensing with strict resource efficiency. While emerging event cameras offer a promising alternative, their potential in aquatic scenarios remains largely unexplored. Through the lens of neuromorphic vision, this work pioneers the investigation of motion fields that serve as key media for agile underwater perception. Built upon spiking neural networks, we introduce a self-supervised framework to estimate per-pixel optical flow from asynchronous event streams, elegantly bypassing the long-standing bottleneck of underwater data scarcity. Extensive evaluations demonstrate that our method achieves competitive visual and quantitative results against leading techniques while operating with superior computational efficiency. By bridging neuromorphic sensing and aquatic intelligence, this work opens new frontiers for lightweight, real-time, and low-cost perception on resource-constrained underwater edge platforms.

cs.CV

Floquet quantum multiparameter estimation with periodic-driving-induced topological phase transition

Periodically driven systems provide a powerful platform for quantum multiparameter estimation. Constructing a static effective Hamiltonian in a proper rotating frame is commonly employed to assess the attainable precision. However, such an approach becomes nonfeasible for more general time-periodically driven systems. To tackle this dilemma, we develop a quantum multiparameter estimation strategy in the Floquet theory framework. The contributions of Floquet eigenmodes, quasienergies, and multi-photon processes to the quantum Fisher information matrix and measurement incompatibility are determined, respectively. Moreover, this approach is applied to a ring-shaped Rashba spin-orbit interferometer model exhibiting the topological phase transition (TPT). In the vicinity of the TPT boundary, we reveal a pronounced enhancement in the estimation precision of multiple parameters with the Heisenberg limit scaling and even higher. Meanwhile, the measurement incompatibility vanishes in an oscillatory manner, and the stroboscopic projective measurement enables the highest estimation precision achievable. This work provides a complete Floquet picture for time-dependent critical quantum multiparameter estimation.

quant-ph

Electric Vehicle User Charging Behavior Analysis Integrating Psychological and Environmental Factors: A Statistical-Driven LLM based Agent Approach

With the growing adoption of electric vehicles (EVs), understanding user charging behavior has become critical for grid stability and transportation planning. This study investigates the behavioral heterogeneity of EV taxi drivers by analyzing the interaction between psychological traits and situational triggers within dynamic travel contexts. Leveraging large language models (LLMs) as a core simulation tool, a novel framework with statistical enhancement is developed to replicate and analyze the charging behaviors of taxi drivers. LLMs simulate personalized decision-making processes by leveraging natural language reasoning and role-playing capabilities, accounting for factors such as time sensitivity, price awareness, and range anxiety. Simulation results indicate that the framework reliably reproduces real-world charging behaviors across multiple urban environments. his fidelity arises from integrating statistical priors into the reasoning process, allowing the model to anchor its decisions in empirical behavioral patterns. Further analysis highlights the joint influence of environmental and psychological variables on charging decisions and reveals the heterogeneity of different user groups. The findings provide new insights into EV user behavior, offering a foundation for optimizing charging infrastructure, informing energy policy, and advancing the integration of EV behavioral models into smart transportation and energy management systems.

cs.AI

Emergent aperiodicity in Bose-Bose mixtures induced by spin-dependent periodic potentials

We study the ground-state and low-lying metastable phases of repulsive binary Bose-Einstein condensates confined in twisted, spin-dependent periodic optical lattices. For balanced mixtures, weak intercomponent interactions yield a fourfold momentum-space symmetry dictated by the lattice geometry. Increasing the coupling strength leads to the emergence of additional momentum peaks that combine with the lattice-induced structure to produce an eightfold rotationally symmetric pattern, signaling quasicrystalline order. At intermediate interactions, global phase separation suppresses this quasicrystalline state; however, at stronger coupling, local phase separation gives rise to a long-lived metastable phase in which the eightfold symmetry is restored. In this regime, a secondary ring of dominant momentum peaks appears at smaller wave vectors, indicating longer-wavelength density modulations and a crossover from lattice-dominated to interaction-driven quasicrystalline order. In contrast, imbalanced mixtures form partially miscible density clusters with eightfold-symmetric aperiodic patterns only at intermediate coupling, while stronger interactions drive global phase separation and permanently destroy quasicrystalline order. Real-time simulations demonstrate that these aperiodic structures are dynamically stable and experimentally accessible. Our results show that quasicrystalline order can emerge in binary condensates without explicitly aperiodic lattices and reveal population balance as a key ingredient for stabilizing quantum quasicrystals.

cond-mat.quant-gas

Qwen3-ASR Technical Report

In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR models that support language identification and ASR for 52 languages and dialects. Both of them leverage large-scale speech training data and the strong audio understanding ability of their foundation model Qwen3-Omni. We conduct comprehensive internal evaluation besides the open-sourced benchmarks as ASR models might differ little on open-sourced benchmark scores but exhibit significant quality differences in real-world scenarios. The experiments reveal that the 1.7B version achieves SOTA performance among open-sourced ASR models and is competitive with the strongest proprietary APIs while the 0.6B version offers the best accuracy-efficiency trade-off. Qwen3-ASR-0.6B can achieve an average TTFT as low as 92ms and transcribe 2000 seconds speech in 1 second at a concurrency of 128. Qwen3-ForcedAligner-0.6B is an LLM based NAR timestamp predictor that is able to align text-speech pairs in 11 languages. Timestamp accuracy experiments show that the proposed model outperforms the three strongest force alignment models and takes more advantages in efficiency and versatility. To further accelerate the community research of ASR and audio understanding, we release these models under the Apache 2.0 license.

cs.CL

Qwen3-TTS Technical Report

In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.

cs.SD

Autoregressive long-horizon prediction of plasma edge dynamics

Accurate modeling of scrape-off layer (SOL) and divertor-edge dynamics is vital for designing plasma-facing components in fusion devices. High-fidelity edge fluid/neutral codes such as SOLPS-ITER capture SOL physics with high accuracy, but their computational cost limits broad parameter scans and long transient studies. We present transformer-based, autoregressive surrogates for efficient prediction of 2D, time-dependent plasma edge state fields. Trained on SOLPS-ITER spatiotemporal data, the surrogates forecast electron temperature, electron density, and radiated power over extended horizons. We evaluate model variants trained with increasing autoregressive horizons (1-100 steps) on short- and long-horizon prediction tasks. Longer-horizon training systematically improves rollout stability and mitigates error accumulation, enabling stable predictions over hundreds to thousands of steps and reproducing key dynamical features such as the motion of high-radiation regions. Measured end-to-end wall-clock times show the surrogate is orders of magnitude faster than SOLPS-ITER, enabling rapid parameter exploration. Prediction accuracy degrades when the surrogate enters physical regimes not represented in the training dataset, motivating future work on data enrichment and physics-informed constraints. Overall, this approach provides a fast, accurate surrogate for computationally intensive plasma edge simulations, supporting rapid scenario exploration, control-oriented studies, and progress toward real-time applications in fusion devices.

physics.plasm-ph

Improving Local Training in Federated Learning via Temperature Scaling

Federated learning is inherently hampered by data heterogeneity: non-i.i.d. training data over local clients. We propose a novel model training approach for federated learning, FLex&Chill, which exploits the Logit Chilling method. Through extensive evaluations, we demonstrate that, in the presence of non-i.i.d. data characteristics inherent in federated learning systems, this approach can expedite model convergence and improve inference accuracy. Quantitatively, from our experiments, we observe up to 6X improvement in the global federated learning model convergence time, and up to 3.37% improvement in inference accuracy.

cs.LG

A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that \textsc{Anaphor}, \textsc{Quantifiers}, and \textsc{Ellipsis} in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs.

cs.CL

TagLabel: RFID Based Orientation and Material Sensing for Automated Package Inspection

Modern logistics systems face increasing difficulty in identifying counterfeit products, fraudulent returns, and hazardous items concealed within packages, yet current package screening methods remain too slow, expensive, and impractical for widespread use. This paper presents TagLabel, an RFID based system that determines both the orientation and contents of packages using low cost passive UHF tags. By analyzing how materials change RSSI and phase, the system identifies the contents of a package without opening it. Using orientation inferred from phase differences, tag occlusion, and antenna gain patterns, the system selects the tag with the greatest occlusion for accurate material sensing. We evaluate two and three tag configurations, and show that both can deliver high orientation and material sensing performance through the use of machine learning classifiers, even in realistic RF environments. When combined into a unified pipeline, TagLabel achieves more than 80 percent accuracy across all package orientations. Because it requires only standard RFID hardware and offers fast scanning times, this approach provides a practical way to enhance package inspection and improve automation in logistics operations.

eess.SP

Nonlocal Nonlinear Control of Photonic Spin Hall Effect in Strongly Interacting Rydberg Media

We present a theoretical study demonstrating enhanced tunability of the photonic spin Hall effect (PSHE) using a strongly interacting Rydberg atomic medium under electromagnetically induced transparency (EIT) conditions. In contrast to conventional approaches that rely on static refractiveindex profiles or metamaterials, here the PSHE is controlled via a nonlocal third-order nonlinear susceptibility arising from long range Rydberg-Rydberg interactions. We show that this nonlocal nonlinearity enables dynamic modulation of spin-dependent light trajectories, amplifying the normally weak PSHE into a readily observable and adjustable effect. These results pave the way for new capabilities in photonic information processing and sensing. In particular, an adjustable PSHE may enable beam steering based on photon spin, improve the sensitivity of precision measurements, and support photonic devices whose functionality can be reconfigured in real time.

quant-ph

Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding

Test-time scaling enhances large language model performance by allocating additional compute resources during inference. Best-of-N (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost-performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling: (1) Generating N full samples consumes substantial GPU memory, reducing inference capacity under limited resources. (2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs. Although some studies have explored efficiency improvements, none have addressed both challenges at once. To address this gap, we propose Self-Truncation Best-of-N (ST-BoN), a decoding method that avoids fully generating all N samples and eliminates the need for reward models. It leverages early sampling consistency in the model's internal states to identify the most promising path and truncate suboptimal ones. In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%. In terms of cost-performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%-80%, and under the same cost, it can improve accuracy by 3-4 points.

cs.CL

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs. We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning: (1) Reasoning performance varies widely across languages for current LLMs; (2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance; (3) The thinking length differs significantly by language for current LLMs. Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs.

cs.CL