SearcharxivSearch

arXiv subjects

Hua Li

Publications and source records attributed to Hua Li.

At least 19 recordsLinked to original sources

Order elevation of directly self-starting sub-step implicit integrators for transient dynamics

Directly self-starting implicit methods are attractive for transient analysis because they avoid auxiliary starting procedures while retaining the original first- or second-order governing equations. However, most existing formulations usually fix the last sub-step at the end of each time interval, which restricts the attainable order. This study develops a generalized $s$-sub-step implicit framework by releasing this constraint and treating all sub-step locations as design variables. The resulting methods admit a unified Runge--Kutta representation for both first- and second-order transient systems and preserve identical effective matrices over all sub-steps. Accuracy conditions are derived by simultaneously matching the numerical amplification factor and load operator, thereby accounting for both homogeneous and forced responses. For $s=1,~\cdots,~6$, two complementary families are obtained: $s$th-order members with user-controllable high-frequency numerical dissipation and adjustable sub-step locations, and $(s+1)$th-order members obtained by selecting the sub-step locations, with fixed dissipation. The latter reach up to seventh-order accuracy without increasing the number of sub-steps, although some high-order members are $A(\alpha)$-stable with stability angles extremely close to $90^\circ$. Analytical amplitude and phase errors further reveal parity-dependent superconvergence in undamped systems, and appropriate parameter selections can substantially increase either phase or amplitude accuracy beyond the formal order. Numerical benchmarks confirm the predicted convergence orders and the controllable suppression of spurious high-frequency responses.

math.NA

Transient Chirp Dynamics in Terahertz Quantum Cascade Lasers

Laser frequency chirp is a ubiquitous dynamical process in semiconductor lasers, vital for frequency-modulated photonic systems. In the mid-infrared (MIR) and terahertz (THz) ranges, quantum cascade lasers (QCLs) are ideal sources with high power, narrow linewidth and compact size. While chirp dynamics in MIR QCLs have been studied, the transient chirp behavior of THz QCLs--particularly the thermal chirp on microsecond to millisecond timescales--remains largely unexplored. Here, we experimentally investigate transient thermal chirp dynamics in single-mode THz QCLs via an on-chip heterodyne scheme. Twin monolithically integrated single-mode QCLs are used: one pulsed QCL as the device under test, and one continuous-wave (CW) QCL serving as both local oscillator (LO) and ultrafast THz detector. The frequency chirp is mapped to the radio-frequency (RF) domain by heterodyne down-conversion. By varying current and temperature, we observe three distinct chirp features: unidirectional down-chirp, V-shaped chirp, and unidirectional up-chirp. A two-node thermal model reproduces the dynamics with good agreement with experiments. Chirp dynamics in the multi-mode regime are also identified, showing the potential for sensitive dynamic spectral characterization. These findings deepen the understanding of THz QCL thermal chirp mechanisms and support applications in THz frequency combs, frequency-modulated continuous-wave (FMCW) radar, and high-speed coherent communications.

physics.optics

Dual-polarization control of broadband nonreciprocal thermal radiation by combining local and nonlocal metasurfaces

Nonreciprocal thermal radiation offers a route to decouple spectral directional absorptivity and emissivity, thereby enabling new paradigms in thermal-photonic systems. However, in magneto-optical platforms, the intrinsic gyroelectric response generally confines observable nonreciprocity to transverse-magnetic (TM) polarization, while the transverse-electric (TE) response is absent. In this work, we experimentally demonstrate, for the first time, a local thermal metasurface strategy to activate TE-polarized nonreciprocity by creating artificial gyromagnetic response in a gyroelectric semiconductor platform. We further extend this mechanism to broadband dual-polarization operation employing a nonlocal thermal metasurface, which combines a resonator supercell with gradient-doped epsilon-near-zero magneto-optical multilayers. Pronounced absorptivity contrast is maintained over 22-27 {\mu}m for TE polarization and 19-27 {\mu}m for TM polarization. This platform provides a mechanism-based route to achieve broadband and dual-polarization nonreciprocal thermal absorption, opening new opportunities for advancing radiative energy-conversion devices.

physics.optics

Unveiling neuronal microstructure in the human brain in vivo with time-dependent radial diffusivity in MRI

Diffusion time-dependence, defined as variations in diffusivity and/or diffusional kurtosis with diffusion time, has emerged as a valuable non-invasive imaging marker for characterizing tissue microstructural features, such as cell size, density, packing disorder, and membrane permeability. In white matter, diffusion time-dependent changes between the short diffusion time and long diffusion time in radial diffusivity (RD), defined as the diffusivity perpendicular to fiber tracts, were demonstrated to correlate strongly with mean axon diameter in ex vivo spinal cord tissues, and to reveal demyelination in mouse corpus callosum. Despite their potential to non-invasively unveil neuronal microstructures to improve the assessment and targeted therapy of neurological diseases, these novel image contrasts obtained at short diffusion times using oscillating gradient spin echo (OGSE) have only recently become feasible for human in vivo studies with high-performance gradient MRI systems. In this preliminary study, we characterized time-dependent RD with OGSE encoding in the human brain in vivo. The change in radial diffusivity between short diffusion time and long diffusion time (delta_RD) consistently exhibited high values in the corticospinal tract, indicating high sensitivity of delta_RD to large axon diameter in human brains. Imaging at a high OGSE frequency of 100 Hz and a moderate b-value of 800 s/mm2 produced the highest delta_RD in the corticospinal tract. This study established a baseline for future investigations of neuronal microstructural alterations in neurological disorders and diseases.

physics.med-ph

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

cs.CV

Rethinking Conditional Generation for Underwater Salient Object Detection

Salient Object Detection in underwater images remains challenging due to low contrast, uneven illumination, and color distortion caused by scattering and absorption effects, which limit the effectiveness of conventional SOD methods in underwater environments. To address these challenges, we propose a Degradation-aware Conditional Generation Network (DCGNet), specifically designed to construct reliable conditional features for underwater saliency generation. First, we design a Dynamic Multi-Granularity module (DMG) grounded in the human visual system to robustly detect salient objects of varying scales with blurred boundaries. Then, we develop an Underwater Physics-Prior module (UPP), which utilizes pseudo-depth guidance to estimate underwater light attenuation and backscatter, thereby restoring degradation-aware RGB features and mitigating color distortion and boundary ambiguity. Based on the physics-guided representation, we introduce an Underwater Spatial Gaussian module (USG), which constructs a spatial Gaussian saliency prior from the strongest guided response to enhance object-centered salient regions and suppress cluttered underwater backgrounds. In addition, a lightweight timestep-adaptive Diffusion Transformer (DiT) bottleneck is inserted into the denoising decoder to refine fused features at different diffusion timesteps. Comprehensive experiments on USOD10K, USOD, CSOD10K, MAS3K, and RMAS demonstrate that DCGNet significantly outperforms existing state-of-the-art methods, verifying its potential for complex underwater visual applications.

cs.CV

Correlating Quasi-Optical Coupling Efficiency with Measured Receiver Noise Temperature in Metalens Coupled THz HEB Mixer

Quasi-optical coupling serves as the critical interface in terahertz (THz) heterodyne receiver systems, enabling efficient transfer of incident radiation to photomixers through a focusing element and a planar microwave antenna. With recent advances in nanofabrication, planar dielectric metalenses have emerged as promising alternatives to conventional refractive optics due to their compactness and scalability. However, unlike conventional elliptical silicon lenses that are often treated as nearly ideal optical components, the focusing characteristics of metalenses, including both phase and amplitude, strongly depend on the local deflection angle across the aperture, creating an urgent need to quantitatively understand the coupling between a dielectric metalens and a planar antenna. In this work, we present a quasi-optical coupling analysis between a planar Si metalens and a logarithmic spiral antenna integrated with a THz superconducting NbN hot-electron bolometer (HEB) mixer operating at 1.63 THz using spherical-coordinate vectorial integration. By combining the angular radiation profile of the spiral antenna with the propagated complex electric-field profile from metalens numerical simulations, the calculated coupling efficiency accounts for angular power distribution, phase-front matching, and polarization-dependent vectorial overlap. The calculated coupling efficiency is then directly correlated with experimentally measured double-sideband receiver noise temperatures through comparison with a conventional elliptical Si lens measured under the same receiver configuration. The analysis establishes a quantitative relationship between metalens focusing efficiency, vectorial antenna coupling, and receiver noise temperature, providing guidance for optimizing metalens design and improving the overall performance of metalens-integrated THz heterodyne receivers.

physics.optics

Gradient Perturbation: Learning to Perturb Gradients for Adaptive Training

Deep neural network training involves both forward propagation (from features through logits to loss) and backward propagation (from loss through gradients to parameter updates). While perturbations along the forward chain, including feature perturbation, logit perturbation, and label perturbation, have been extensively studied, the backward chain's gradient perturbation has received little systematic investigation. In this paper, we establish a unified framework for gradient perturbation, revealing that existing methods such as Sharpness-Aware Minimization (SAM), gradient clipping, and gradient noise injection can all be interpreted as imposing specific forms of gradient perturbation. Analogous to the recently proposed Logit Perturbation Learning (LPL), we conjecture that amplifying the gradient norm for a class acts as positive augmentation (enhancing learning), while dampening it acts as negative augmentation (suppressing overfitting). Based on these observations, we propose Learning to Perturb Gradients (LPG), which adaptively perturbs logit-level gradients at the class level to achieve category-aware training. We also establish theoretical connections between gradient perturbation bounds and generalization guarantees via PAC-Bayesian analysis. Experiments on balanced classification, long-tail classification, and noisy label learning demonstrate that LPG consistently outperforms existing methods and can be combined with them as a plug-in module.

cs.LG

Learning to Perturb Hidden Representations for Generalizable Deep Learning

Deep neural networks process data through a cascade of representations: input features, hidden activations, logits, and loss. While perturbations at the input, logit, and label levels have been systematically studied, the intermediate hidden activations, which constitute the bulk of the network's computation, have received no unified perturbation analysis. In this paper, we establish a unified framework for hidden activation perturbation, revealing that Dropout, Manifold Mixup, adversarial feature perturbation, and related methods all impose specific forms of activation perturbation but with class-agnostic or random strategies. We conjecture that expansive perturbation (increasing activation norm) acts as positive augmentation, while contractive perturbation (decreasing activation norm) acts as negative augmentation, and that the perturbation layer determines whether the effect resembles input-level augmentation (shallow layers) or logit-level manipulation (deep layers). We propose Learning to Perturb Activations (LPA), which adaptively perturbs activations at a selected hidden layer with class-level perturbations learned via PGD. We further provide theoretical analysis connecting activation perturbation to flat minima and perturbation amplification through layers. Experiments on balanced classification, long-tail classification, and domain generalization demonstrate that LPA consistently outperforms existing methods and provides complementary benefits to logit perturbation methods such as LPL.

cs.LG

Clustering based on Stochastic Dominance with application for risk averters and risk seekers

Stochastic Dominance (SD) theory provides a rigorous framework for selecting superior assets tailored to the asset allocation needs of investors with varying risk preferences (i.e., risk-averse, risk-seeking, and risk-neutral). However, traditional stock clustering methods typically rely on geometric metrics such as Euclidean distance, which often fail to effectively capture the intrinsic risk dominance relationships among assets. To address this limitation, this paper proposes an innovative clustering analysis framework based on SD test statistics. Methodologically, this study deeply integrates SD theory with machine learning algorithms. Transcending the limitations of traditional reliance on geometric distance, we innovatively utilize test statistics from first-, second-, and third-order SD to construct a "Stochastic Dominance Coefficient Matrix." Building upon this matrix, we modify the classic K-means and Hierarchical Clustering algorithms. Specifically, we derive 12 distinct algorithm variants tailored to different orders of SD relationships. Simultaneously, we construct the SD-SC coefficient and the SD-DBI index as specialized validity indices to evaluate the clustering performance. Empirically, we analyze constituent stock data from a representative developed market (the US NASDAQ Index) and an emerging market (China's CSI 100 Index). The results verify the effectiveness and robustness of the proposed method. Furthermore, we apply the clustering results to the modification of the Single Index Model and the construction of Global Minimum Variance Portfolios (GMVP). The findings demonstrate that the proposed method effectively facilitates customized asset allocation for investors, holding significant theoretical value and practical implications.

stat.ML

Real-space imaging reveals symmetry-selected nonlinear energy routing in a mechanical resonator

Nonlinear energy routing among modes underlies phenomena ranging from internal resonance and wave mixing to frequency-comb generation in micro- and nanoelectromechanical resonators, yet modal interactions are typically inferred from spectra rather than imaged in real space. This leaves unresolved how energy is spatially routed and what determines which pathways are selected. Here, we use phase-locked multi-harmonic stroboscopic interferometry to reconstruct harmonic-resolved differential displacement maps in a nearly mirror-symmetric microelectromechanical resonator. These maps reveal that harmonics generated by a driven mode can be carried by distinct spatial eigenmodes, directly resolving pathways of nonlinear energy transfer. We further show that such mode-selective routing occurs even away from integer frequency matching: generated harmonics are dominated by eigenmodes sharing the driven mode's mirror parity, whereas spectrally closer opposite-parity modes remain strongly suppressed. A nonlinear modal framework links this hierarchy to symmetry-dependent modal-overlap integrals. These results identify spatial symmetry as a selection rule for nonlinear energy routing.

physics.optics

DongYuan: An LLM-Based Framework for Integrative Chinese and Western Medicine Spleen-Stomach Disorders Diagnosis

The clinical burden of spleen-stomach disorders is substantial. While large language models (LLMs) offer new potential for medical applications, they face three major challenges in the context of integrative Chinese and Western medicine (ICWM): a lack of high-quality data, the absence of models capable of effectively integrating the reasoning logic of traditional Chinese medicine (TCM) syndrome differentiation with that of Western medical (WM) disease diagnosis, and the shortage of a standardized evaluation benchmark. To address these interrelated challenges, we propose DongYuan, an ICWM spleen-stomach diagnostic framework. Specifically, three ICWM datasets (SSDF-Syndrome, SSDF-Dialogue, and SSDF-PD) were curated to fill the gap in high-quality data for spleen-stomach disorders. We then developed SSDF-Core, a core diagnostic LLM that acquires robust ICWM reasoning capabilities through a two-stage training regimen of supervised fine-tuning. tuning (SFT) and direct preference optimization (DPO), and complemented it with SSDF-Navigator, a pluggable consultation navigation model designed to optimize clinical inquiry strategies. Additionally, we established SSDF-Bench, a comprehensive evaluation benchmark focused on ICWM diagnosis of spleen-stomach disorders. Experimental results demonstrate that SSDF-Core significantly outperforms 12 mainstream baselines on SSDF-Bench. DongYuan lays a solid methodological foundation and provides practical technical references for the future development of intelligent ICWM diagnostic systems.

cs.CL

PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing

This paper primarily investigates the task of expression-only portrait video performance editing based on a driving video, which plays a crucial role in animation and film industries. Most existing research mainly focuses on portrait animation, which aims to animate a static portrait image according to the facial motion from the driving video. As a consequence, it remains challenging for them to disentangle the facial expression from head pose rotation and thus lack the ability to edit facial expression independently. In this paper, we propose PerformRecast, a versatile expression-only video editing method which is dedicated to recast the performance in existing film and animation. The key insight of our method comes from the characteristics of 3D Morphable Face Model (3DMM), which models the face identity, facial expression and head pose of 3D face mesh with separate parameters. Therefore, we improve the keypoints transformation formula in previous methods to make it more consistent with 3DMM model, which achieves a better disentanglement and provides users with much more fine-grained control. Furthermore, to avoid the misalignment around the boundary of face in generated results, we decouple the facial and non-facial regions of input portrait images and pre-train a teacher model to provide separate supervision for them. Extensive experiments show that our method produces high-quality results which are more faithful to the driving video, outperforming existing methods in both controllability and efficiency. Our code, data and trained models are available at https://youku-aigc.github.io/PerformRecast.

cs.CV

Broadband terahertz comb with sub-Hz comb linewidth

Terahertz (THz) frequency combs are increasingly essential for spectroscopy, metrology, and quantum science. However, generating a dense array of evenly spaced ultra-narrow THz comb lines is challenging. Here, we demonstrate broadband THz comb generation using a photoconductive antenna that transfers a noise-suppressed near-infrared electro-optical (EO) comb into the THz domain. Our noise-suppression strategy, leveraging soliton self-frequency shift and spectral filtering, effectively suppresses EO comb phase noise without requiring active stabilization. The resulting THz comb exhibits broad spectral coverage (0.05-4 THz), narrow comb linewidths (0.3 Hz at the Fourier-transform limit), and excellent frequency stability (8.6*10^-14 at 1-second integration). We further demonstrate asynchronous THz time-domain spectroscopy, resolving ~36,000 comb lines with 50 MHz spacing. Crucially, the inherent frequency agility of the EO comb enables rapid and wide-range tuning of the THz comb line spacing. These attributes position our THz comb as a versatile tool for high-resolution molecular spectroscopy and precision THz metrology.

physics.optics

Reconstructive comb spectroscopy: A single-pixel detection paradigm beyond dual-comb limitations

Frequency comb spectroscopy has revolutionized broadband molecular fingerprinting with mode-defined resolution. While dual-comb spectroscopy stands as a dominant paradigm for high-resolution measurements, it relies on mutually coherent dual combs, and its applicability to non-cooperative sensing is limited by the requirement for phase-sensitive detection and controlled optical returns. Here, we introduce reconstructive comb spectroscopy, a fundamentally different paradigm that eliminates these constraints. By integrating a mode-programmable optical comb with a computational sensing scheme based on single-pixel detection, our method achieves picometer-level spectral resolution over a 10-nm (1.27-THz) instantaneous bandwidth, with single-photon sensitivity down to 10^-4 photons per pulse, and compressed spectral acquisition at 2.5% sampling while maintaining reconstruction errors below 10%. We demonstrate robust performance through scattering media and from non-cooperative targets. These capabilities establish reconstructive comb spectroscopy as a new platform for gas sensing, with broad applicability in remote atmospheric monitoring, industrial leak detection, and standoff chemical-threat identification.

physics.optics

Fine-Tuning Diffusion-Based Recommender Systems via Reinforcement Learning with Reward Function Optimization

Diffusion models recently emerged as a powerful paradigm for recommender systems, offering state-of-the-art performance by modeling the generative process of user-item interactions. However, training such models from scratch is both computationally expensive and yields diminishing returns once convergence is reached. To remedy these challenges, we propose ReFiT, a new framework that integrates Reinforcement learning (RL)-based Fine-Tuning into diffusion-based recommender systems. In contrast to prior RL approaches for diffusion models depending on external reward models, ReFiT adopts a task-aligned design: it formulates the denoising trajectory as a Markov decision process (MDP) and incorporates a collaborative signal-aware reward function that directly reflects recommendation quality. By tightly coupling the MDP structure with this reward signal, ReFiT empowers the RL agent to exploit high-order connectivity for fine-grained optimization, while avoiding the noisy or uninformative feedback common in naive reward designs. Leveraging policy gradient optimization, ReFiT maximizes exact log-likelihood of observed interactions, thereby enabling effective post hoc fine-tuning of diffusion recommenders. Comprehensive experiments on wide-ranging real-world datasets demonstrate that the proposed ReFiT framework (a) exhibits substantial performance gains over strong competitors (up to 36.3% on sequential recommendation), (b) demonstrates strong efficiency with linear complexity in the number of users or items, and (c) generalizes well across multiple diffusion-based recommendation scenarios. The source code and datasets are publicly available at https://anonymous.4open.science/r/ReFiT-4C60.

cs.IR

Expose Camouflage in the Water: Underwater Camouflaged Instance Segmentation and Dataset

With the development of underwater exploration and marine protection, underwater vision tasks are widespread. Due to the degraded underwater environment, characterized by color distortion, low contrast, and blurring, camouflaged instance segmentation (CIS) faces greater challenges in accurately segmenting objects that blend closely with their surroundings. Traditional camouflaged instance segmentation methods, trained on terrestrial-dominated datasets with limited underwater samples, may exhibit inadequate performance in underwater scenes. To address these issues, we introduce the first underwater camouflaged instance segmentation (UCIS) dataset, abbreviated as UCIS4K, which comprises 3,953 images of camouflaged marine organisms with instance-level annotations. In addition, we propose an Underwater Camouflaged Instance Segmentation network based on Segment Anything Model (UCIS-SAM). Our UCIS-SAM includes three key modules. First, the Channel Balance Optimization Module (CBOM) enhances channel characteristics to improve underwater feature learning, effectively addressing the model's limited understanding of underwater environments. Second, the Frequency Domain True Integration Module (FDTIM) is proposed to emphasize intrinsic object features and reduce interference from camouflage patterns, enhancing the segmentation performance of camouflaged objects blending with their surroundings. Finally, the Multi-scale Feature Frequency Aggregation Module (MFFAM) is designed to strengthen the boundaries of low-contrast camouflaged instances across multiple frequency bands, improving the model's ability to achieve more precise segmentation of camouflaged objects. Extensive experiments on the proposed UCIS4K and public benchmarks show that our UCIS-SAM outperforms state-of-the-art approaches.

cs.CV

WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation

Underwater Salient Object Detection (USOD) faces significant challenges, including underwater image quality degradation and domain gaps. Existing methods tend to ignore the physical principles of underwater imaging or simply treat degradation phenomena in underwater images as interference factors that must be eliminated, failing to fully exploit the valuable information they contain. We propose WaterFlow, a rectified flow-based framework for underwater salient object detection that innovatively incorporates underwater physical imaging information as explicit priors directly into the network training process and introduces temporal dimension modeling, significantly enhancing the model's capability for salient object identification. On the USOD10K dataset, WaterFlow achieves a 0.072 gain in S_m, demonstrating the effectiveness and superiority of our method. https://github.com/Theo-polis/WaterFlow.

cs.CV