Searcharxiv⌕ Search

arXiv subjects

Zhiyong Zhao

Publications and source records attributed to Zhiyong Zhao.

15 recordsLinked to original sources

Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation

Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.

cs.CL↗

FAAR: Format-Aware Adaptive Rounding for NVFP4

Deploying large language models (LLMs) on edge devices requires extremely low-bit quantization. Ultra-low precision formats such as NVFP4 offer a promising solution for reducing memory footprint and accelerating computation. However, existing quantization methods typically rely on conventional rounding strategies and fail to account for the non-uniformity of the NVFP4 numerical grid, resulting in suboptimal rounding decisions and amplified quantization errors. To address this, we propose Format-Aware Adaptive Rounding (FAAR), a learnable rounding strategy tailored for the NVFP4 format. Unlike conventional quantization paradigms, FAAR explicitly incorporates the non-uniform NVFP4 grid into the optimization process. By adaptively adjusting rounding decisions guided by loss gradients, our method effectively approximates the theoretically optimal quantization. To complement FAAR, we introduce a 2-stages Format Alignment (2FA) fine-tuning scheme that aligns LLM parameters layer-by-layer to the NVFP4 numerical space, further narrowing the performance gap. Remarkably, this learnable optimization incurs a minimal training overhead of only 4 GPU hours on Llama3-1B. Extensive experiments demonstrate the effectiveness of our approach. Compared with Round-to-Nearest (RTN), our method reduces perplexity on WikiText-2 from 14.28 to 12.60 on Llama3-1B and from 23.06 to 21.27 on Qwen3-1.7B. Additionally, our method consistently outperforms state-of-the-art approaches across various zero-shot downstream tasks.

cs.LG↗

Efficient Token Pruning for LLaDA-V

Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional attention mechanism and diffusion-style iterative denoising paradigm introduce significant computational overhead, as visual tokens are repeatedly processed across all layers and denoising steps. In this work, we conduct an in-depth attention analysis and reveal that, unlike autoregressive decoders, LLaDA-V aggregates cross-modal information predominantly in middle-to-late layers, leading to delayed semantic alignment. Motivated by this observation, we propose a structured token pruning strategy inspired by FastV, selectively removing a proportion of visual tokens at designated layers to reduce FLOPs while preserving critical semantic information. To the best of our knowledge, this is the first work to investigate structured token pruning in diffusion-based large multimodal models. Unlike FastV, which focuses on shallow-layer pruning, our method targets the middle-to-late layers of the first denoising step to align with LLaDA-V's delayed attention aggregation to maintain output quality, and the first-step pruning strategy reduces the computation across all subsequent steps. Our framework provides an empirical basis for efficient LLaDA-V inference and highlights the potential of vision-aware pruning in diffusion-based multimodal models. Across multiple benchmarks, our best configuration reduces computational cost by up to 65% while preserving an average of 95% task performance.

cs.CV↗

Highly efficient wideband and polarization-insensitive SMF-ARF coupling strategy with low back-reflection

We propose a lensed-fiber based coupling strategy for low-loss interconnection between single-mode fibers and anti-resonant fibers. By optimizing structural and geometric parameters, the design simultaneously achieves high coupling efficiency and suppressed back-reflection. Experimental results demonstrate an insertion loss of 1.2 dB and back-reflection of -36.22 dB at 1550 nm, with excellent spectral stability (below 0.72 dB variation across 1500-1600 nm) and polarization insensitivity (below 0.4 dB polarization dependent loss). The compact structure not only facilitates the fabrication process, but also enables seamless ARF integration into existing optical networks, thereby addressing critical demands for high-capacity data transmission.

physics.app-ph↗

Ion-Exchange Doping of Semiconducting Single-Walled Carbon Nanotubes

Semiconducting single-walled carbon nanotubes (SWCNTs) are a promising thermoelectric material with high power factors after chemical p- or n-doping. Understanding the impact of dopant counterions on charge transport and thermoelectric properties of nanotube networks is essential to further optimize doping methods and to develop better dopants. Here, we utilize ion-exchange doping to systematically vary the size of counterions in thin films of small and large diameter, polymer-sorted semiconducting SWCNTs with AuCl3 as the initial p-dopant and investigate the impact of ion size on conductivity, Seebeck coefficients and power factors. Larger anions are found to correlate with higher electrical conductivities and improved doping stability, while no significant effect on the power factors is found. Importantly, the effect of counterion size on the thermoelectric properties of dense SWCNT networks is not obscured by morphological changes upon doping. The observed trends of carrier mobilities and Seebeck coefficients can be explained by a random resistor model for the nanotube network that accounts for overlapping Coulomb potentials leading to the formation of an impurity band whose depth depends on the carrier density and counterion size. These insights can be applied more broadly to understand the thermoelectric properties of doped percolating disordered systems, including semiconducting polymers.

cond-mat.mtrl-sci↗

Intelligent Reflecting Surfaces for Wireless Networks: Deployment Architectures, Key Solutions, and Field Trials

Intelligent reflecting surfaces (IRSs) have emerged as a transformative technology for wireless networks by improving coverage, capacity, and energy efficiency through intelligent manipulation of wireless propagation environments. This paper provides a comprehensive study on the deployment and coordination of IRSs for wireless networks. By addressing both single- and multi-reflection IRS architectures, we examine their deployment strategies across diverse scenarios, including point-to-point, point-to-multipoint, and point-to-area setups. For the single-reflection case, we highlight the trade-offs between passive and active IRS architectures in terms of beamforming gain, coverage extension, and spatial multiplexing. For the multi-reflection case, we discuss practical strategies to optimize IRS deployment and element allocation, balancing cooperative beamforming gains and path loss. The paper further discusses practical challenges in IRS implementation, including environmental conditions, system compatibility, and hardware limitations. Numerical results and field tests validate the effectiveness of IRS-aided wireless networks and demonstrate their capacity and coverage improvements. Lastly, promising research directions, including movable IRSs, near-field deployments, and network-level optimization, are outlined to guide future investigations.

eess.SP↗

7 Tesla multimodal MRI dataset of ex-vivo human brain

Ex-vivo MRI offers invaluable insights into the complexity of the human brain, enabling high-resolution anatomical delineation and integration with histopathology, and thus, contributes to both basic and clinical studies on normal and pathological brains. However, ex-vivo MRI is challenging in sample preparation, acquisition, and data analysis, and existing ex-vivo MRI datasets are often single image modality and lack of ethnic diversity. In our study, we aimed to address these limitations by constructing a comprehensive multimodal MRI database acquired from six ex-vivo Chinese human brains. This database included structural MRI, high-angular resolution diffusion MRI, quantitative susceptibility mapping, and quantitative T1 and T2 maps, which enabled multifaceted depiction of brain microstructure and connectivity. Furthermore, we generated population-averaged multimodal templates and the segmentation labels to facilitate analysis of ex-vivo brain MRI. This public database offers a collection of high-resolution and multi-parametric ex-vivo human brain MRI and filled the gap of lacking Asian brain samples in existing databases.

q-bio.QM↗

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging Vision-Language Models (VLMs) for enhanced scene understanding and planning capabilities. DriveVLM integrates a unique combination of reasoning modules for scene description, scene analysis, and hierarchical planning. Furthermore, recognizing the limitations of VLMs in spatial reasoning and heavy computational requirements, we propose DriveVLM-Dual, a hybrid system that synergizes the strengths of DriveVLM with the traditional autonomous driving pipeline. Experiments on both the nuScenes dataset and our SUP-AD dataset demonstrate the efficacy of DriveVLM and DriveVLM-Dual in handling complex and unpredictable driving conditions. Finally, we deploy the DriveVLM-Dual on a production vehicle, verifying it is effective in real-world autonomous driving environments.

cs.CV↗

From Plastic Waste to Treasure: Selective Upcycling through Catalytic Technologies

The huge amount of plastic wastes has become a pressing global environmental problem, leading to severe environmental pollution and resource depletion through conventional downcycling technologies like incineration and landfilling. In contrast, selective upcycling of various plastics offers a promising solution for converting waste plastics into valuable products. This review provides a comprehensive overview of the recent advancements in innovative catalytic technologies, including thermocatalysis, electrocatalysis, and photocatalysis. Special emphasis is placed on elucidating the reaction mechanisms, activating designated chemical bonds for high selectivity, and elaborating the above techniques in terms of reaction conditions and products. Finally, the application prospects and future development trends in plastic catalysis are discussed, providing valuable insights for realizing a sustainable circular plastic economy.

physics.chem-ph↗

Enabling variable high spatial resolution retrieval from a long pulse BOTDA sensor

In the field of Internet of Things, there is an urgent need for sensors with large-scale sensing capability for scenarios such as intelligent monitoring of production lines and urban infrastructure. Brillouin optical time domain analysis (BOTDA) sensors, which can monitor thousands of continuous points simultaneously, show great advantages in these applications. We propose a convolutional neural network (CNN) to process the data of conventional Brillouin optical time domain analysis (BOTDA) sensors, which achieves unprecedented performance improvement that allows to directly retrieve higher spatial resolution (SR) from the sensing system that use long pump pulses. By using the simulated Brillouin gain spectrums (BGSs) as the CNN input and the corresponding high SR BFS as the output target, the trained CNN is able to obtain a SR higher than the theoretical value determined by the pump pulse width. In the experiment, the CNN accurately retrieves 0.5-m hotspots from the measured BGS with pump pulses from 20 to 50 ns, and the acquired BFS is in great agreement with 45/40 ns differential pulse-width pair (DPP) measurement results. Compared with the DPP technique, the proposed CNN demonstrates a 2-fold improvement in BFS uncertainty with only half the measurement time. In addition, by changing the training datasets, the proposed CNN can obtain tunable high SR retrieval based on conventional BOTDA sensors that use long pulses without any requirement of hardware modifications. The proposed data post-processing approach paves the way to enable novel high spatial resolution BOTDA sensors, which brings substantial improvement over the state-of-the-art techniques in terms of system complexity, measurement time and reliability, etc.

eess.SP↗

Interference fading suppression in Phi-OTDR using space-division multiplexed probes

We propose and experimentally demonstrate a novel interference fading suppression method for phase-sensitive optical time domain reflectometry (Phi-OTDR) using space-division multiplexed (SDM) pulse probes in few-mode fiber. The SDM probes consist of multiple different modes, and three spatial modes (LP01, LP11a and LP11b) are used in this work for proof of concept. Firstly, the Rayleigh backscattering light of different modes is experimentally characterized, and it turns out that the waveforms of Phi-OTDR traces of distinct modes are all different from each other. Thanks to the spatial difference of fading positions of distinct modes, multiple probes from spatially multiplexed modes can be used to suppress the interference fading in Phi-OTDR. Then, the performances of the Phi-OTDR systems using single probe and multiple probes are evaluated and compared. Specifically, statistical analysis shows that both fading probabilities over fiber length and time are reduced significantly by using multiple SDM probes, which verifies the significant performance improvement on fading suppression. The proposed novel interference fading suppression method does not require complicated frequency or phase modulation, which has the advantages of simplicity, good effectiveness and high reliability.

physics.optics↗

Improving the spatial resolution of a BOTDA sensor using deconvolution algorithm

Spatial resolution improvement from an acquired measurement using long pulse is developed for Brillouin optical time domain analysis (BOTDA) systems based on the total variation deconvolution algorithm. The frequency dependency of Brillouin gain temporal envelope is investigated by simulation, and its impact on the recovered results of deconvolution algorithm is thoroughly analyzed. To implement a reliable deconvolution process, differential pulse-width pair (DPP) technique is utilized to effectively eliminate the systematic BFS distortion stemming from the frequency dependency of temporal envelope. The width of the pulse pairs should be larger than 40 ns as is analyzed theoretically and verified experimentally. It has been demonstrated that the proposed method can realize flexible adjustment of spatial resolution with enhanced signal-to-noise ratio (SNR) from an established measurement with long pump pulse. In the experiment, the spatial resolution is increased to 0.5 m and 1 m with high measurement accuracy by using the deconvolution algorithm from the measurement of 60/40 ns DPP signals. Compared with the raw DPP results with the same spatial resolution, 9.2 dB and 8.4 dB SNR improvements are obtained for 0.5 m and 1 m spatial resolution respectively, thanks to the denoising capability of the total variation deconvolution algorithm. The impact of sampling rate on the recovery results is also studied. The proposed sensing system allows for distortion-free Brillouin distributed sensing with higher spatial resolution and enhanced SNR from the conventional DPP setup with long pulse pairs.

eess.SP↗

Complex Engel Structures

We study the geometry of Engel structures, which are 2-plane fields on 4-manifolds satisfying a generic condition, that are compatible with other geometric structures. A complex Engel structure is an Engel 2-plane field on a complex surface for which the 2-planes are complex lines. We solve the equivalence problems for complex Engel structures and use the resulting structure equations to classify homogeneous complex Engel structures. This allows us to determine all compact, homogeneous examples. Compact manifolds that support homogeneous complex Engel structures are diffeomorphic to $S^1\times SU(2)$ or quotients of $\mathbb{C}^2$, $S^1\times SU(2)$, $S^1\times G$ or $H$ by co-compact lattices, where $G$ is the connected and simply-connected Lie group with Lie algebra $\mathfrak{sl}_2(\mathbb{R})$ and $H$ is a solvable Lie group.

math.DG↗

Lagrangian Engel Structures

We study the geometry of Engel structures, which are 2-plane fields on 4-manifolds satisfying a generic condition, that are compatible with other geometric structures. A \em{Lagrangian} Engel structure is an Engel 2-plane field on a symplectic 4-manifold for which the 2-planes are Lagrangian with respect to the symplectic structure. We solve the equivalence problems for Lagrangian Engel structures and use the resulting structure equations to classify homogeneous Lagrangian Engel structures. This allows us to determine all compact, homogeneous examples. Compact manifolds that support homogeneous Lagrangian Engel structures are diffeomorphic to quotients of one of a determined list of nilpotent or solvable 4-dimensional Lie groups by co-compact lattices.

math.DG↗

BOTDA using OFDM channel estimation

A novel Brillouin optical time-domain analysis (BOTDA) system is proposed using intensity-modulated optical orthogonal frequency division multiplexing probe signal and direct detection (IM-DD-OOFDM) without frequency sweep operation. The influence of peak to average power ratio (PAPR) of OFDM probe signal on the recovery of Brillouin gain spectrum (BGS) is analyzed in theory and experiment. The complex BGS is reconstructed by channel estimation algorithm and Brillouin frequency shift (BFS) is located by curve fitting of intensity spectrum. The IM-DD-OOFDM BOTDA is demonstrated experimentally with 25m spatial resolution over 2 km standard single mode fiber.

physics.ins-det↗