SearcharxivSearch

arXiv subjects

Zhihong Huang

Publications and source records attributed to Zhihong Huang.

13 recordsLinked to original sources

Experimental Validation of Skull Acoustic Modelling Strategies for Transcranial Focused Ultrasound Simulation: A Cross-Comparison Study

Accurate acoustic modelling of the skull is essential for simulation-guided transcranial focused ultrasound (tFUS), but commonly used skull parameterisation strategies differ in complexity and reported accuracy. This study experimentally compared five k-Wave skull models: two voxel-wise linear mapping models, two three-layer models, and one single-layer fixed-parameter model. Nineteen regions of interest from five historical and two Thiel-embalmed human skulls were tested at 220 kHz, 680 kHz, and 1000 kHz. Bowl-surface source fields were reconstructed using acoustic holography, and simulated intracranial pressure fields were benchmarked against needle-hydrophone measurements. Across frequencies, mean peak-pressure errors ranged from 20% to 31%, whereas intensity errors reached 41% to 77%. Errors in -6 dB focal volume ranged from 11% to 67%, and focal-position discrepancies were typically several millimetres. Simulations generally predicted smaller insertion losses than measured, indicating a tendency to underestimate skull-related attenuation and overestimate transmitted intracranial exposure. The linear mapping model with fixed attenuation gave the lowest frequency-averaged pressure error, but no model showed a consistent advantage across all metrics. These results show that current skull models can reproduce gross intracranial beam patterns while retaining substantial quantitative uncertainty in exposure, focal coverage, and target localisation.

physics.med-ph

MR-ImagenTime: Multi-Resolution Time Series Generation through Dual Image Representations

Time series forecasting is vital across many domains, yet existing models struggle with fixed-length inputs and inadequate multi-scale modeling. We propose MR-CDM, a framework combining hierarchical multi-resolution trend decomposition, an adaptive embedding mechanism for variable-length inputs, and a multi-scale conditional diffusion process. Evaluations on four real-world datasets demonstrate that MR-CDM significantly outperforms state-of-the-art baselines (e.g., CSDI, Informer), reducing MAE and RMSE by approximately 6-10 to a certain degree.

cs.LG

Beyond Quantity: Trajectory Diversity Scaling for Code Agents

As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic data and the diminishing returns of quantity scaling. Moreover, quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data. We propose TDScaling, a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. Under a fixed training budget, increasing trajectory diversity yields larger gains than adding more trajectories, improving the performance-cost trade-off for agent training. TDScaling integrates four innovations: (1) a Business Cluster mechanism that captures real-service logical dependencies; (2) a blueprint-driven multi-agent paradigm that enforces trajectory coherence; (3) an adaptive evolution mechanism that steers synthesis toward long-tail scenarios using Domain Entropy, Reasoning Mode Entropy, and Cumulative Action Complexity to prevent mode collapse; and (4) a sandboxed code tool that mitigates catastrophic forgetting of intrinsic coding capabilities. Experiments on general tool-use benchmarks (BFCL, tau^2-Bench) and code agent tasks (RebenchT, CodeCI, BIRD) demonstrate a win-win outcome: TDScaling improves both tool-use generalization and inherent coding proficiency. We plan to release the full codebase and the synthesized dataset (including 30,000+ tool clusters) upon publication.

cs.AI

Robust and Efficient Zeroth-Order LLM Fine-Tuning via Adaptive Bayesian Subspace Optimizer

Fine-tuning large language models (LLMs) with zeroth-order (ZO) optimization reduces memory by approximating gradients through function evaluations. However, existing methods essentially perform updates in a one-dimensional space, and suffer from collapse or substantial performance degradation under low-precision training. We introduce BSZO, an adaptive \textbf{B}ayesian \textbf{S}ubspace \textbf{Z}eroth-Order \textbf{O}ptimizer, which applies Kalman filtering to combine finite-difference information across multiple perturbation directions within a subspace. By treating each finite-difference measurement as a noisy observation, BSZO builds a posterior distribution over the subspace-projected gradient and updates it through Bayesian inference, with a residual-based adaptive mechanism to adapt to noise variations. Theoretical analysis shows that BSZO improves the convergence rate by a factor of $k/γ$ compared to standard ZO methods. Experiments on RoBERTa, Mistral, and OPT models show that BSZO outperforms the baselines across various tasks, achieving up to 6.67\% absolute average improvement on OPT-13B while remaining robust under fp16/bf16 precision and keeping memory usage close to inference-only baselines (1.00$\times$--1.08$\times$ of MeZO).

cs.LG

Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost

Recent advancements in large reasoning models (LRMs) have introduced an intermediate "thinking" process prior to generating final answers, improving their reasoning capabilities on complex downstream tasks. However, the potential of LRMs as evaluators for machine translation (MT) quality remains underexplored. We provides the first systematic analysis of LRM-as-a-judge in MT evaluation. We identify key challenges, revealing LRMs require tailored evaluation materials, tend to "overthink" simpler instances and have issues with scoring mechanisms leading to overestimation. To address these, we propose to calibrate LRM thinking by training them on synthetic, human-like thinking trajectories. Our experiments on WMT24 Metrics benchmarks demonstrate that this approach largely reduces thinking budgets by ~35x while concurrently improving evaluation performance across different LRM scales from 7B to 32B (e.g., R1-Distill-Qwen-7B achieves a +8.7 correlation point improvement). These findings highlight the potential of efficiently calibrated LRMs to advance fine-grained automatic MT evaluation.

cs.CL

Effects of skull properties on long-pulsed transcranial focused ultrasound transmission

Transcranial low-intensity focused ultrasound can deliver energy to the brain in a minimally invasive manner for neuromodulation applications. However, continuous sonication through the skull introduces significant wave interactions, complicating precise energy delivery to the target. We present a comprehensive examination of intracranial acoustic fields generated by focused ultrasound transducers and assess the characteristics of cranial bone that affect acoustic transmission. Acoustic field maps were generated at 88 regions of interest across 10 historical and 2 Thiel-embalmed human skull specimens with sonication at frequencies of 220 kHz, 650 kHz, and 1000 kHz. The average peak pressure insertion losses for historical were 3.6$\pm$3.4 dB, 9.3$\pm$3.3 dB, and 14.8$\pm$5.8 dB, respectively, and for Thiel skulls, the respective losses were 2.9$\pm$1.8 dB, 9.4$\pm$2.6 dB, and 17.0$\pm$5.5 dB. The effect of skull thickness, skull density ratio, and skull curvature on intracranial peak pressure, power and focal area was investigated and linear fits produced. Several unfavorable focusing performances were observed in regions with excessive thickness variation. The effects of angulation and spacing between the transducer and the skull were also investigated. Preliminary findings indicate that wave superposition resulting from skull and transducer spacing could lead to a 30-40% uncertainty in peak recorded intracranial pressure.

physics.med-ph

Silicon Optical Memory: Non-Volatile Optoelectronic Devices via Si-SiO$_2$ Hysteresis Effect

Implementing on-chip non-volatile optical memories has long been an actively pursued goal, promising significant enhancements in the capability and energy efficiency of photonic integrated circuits. Here, a novel optical memory has been demonstrated exclusively using the semiconductor primary material, silicon. By manipulating the optoelectronic effect of this device, we introduce a hysteresis effect at the silicon-silicon oxide interface, which in turn demonstrates multi-level, non-volatile optical data storage with robust retention and endurance. This new silicon optical memory provides a distinctively simple and accessible route to realize optical data storage in standard silicon foundry processes.

physics.app-ph

How Does Pretraining Improve Discourse-Aware Translation?

Pretrained language models (PLMs) have produced substantial improvements in discourse-aware neural machine translation (NMT), for example, improved coherence in spoken language translation. However, the underlying reasons for their strong performance have not been well explained. To bridge this gap, we introduce a probing task to interpret the ability of PLMs to capture discourse relation knowledge. We validate three state-of-the-art PLMs across encoder-, decoder-, and encoder-decoder-based models. The analysis shows that (1) the ability of PLMs on discourse modelling varies from architecture and layer; (2) discourse elements in a text lead to different learning difficulties for PLMs. Besides, we investigate the effects of different PLMs on spoken language translation. Through experiments on IWSLT2017 Chinese-English dataset, we empirically reveal that NMT models initialized from different layers of PLMs exhibit the same trends with the probing task. Our findings are instructive to understand how and when discourse knowledge in PLMs should work for downstream tasks.

cs.CL

Deep-Learning-based Vasculature Extraction for Single-Scan Optical Coherence Tomography Angiography

Optical coherence tomography angiography (OCTA) is a non-invasive imaging modality that extends the functionality of OCT by extracting moving red blood cell signals from surrounding static biological tissues. OCTA has emerged as a valuable tool for analyzing skin microvasculature, enabling more accurate diagnosis and treatment monitoring. Most existing OCTA extraction algorithms, such as speckle variance (SV)- and eigen-decomposition (ED)-OCTA, implement a larger number of repeated (NR) OCT scans at the same position to produce high-quality angiography images. However, a higher NR requires a longer data acquisition time, leading to more unpredictable motion artifacts. In this study, we propose a vasculature extraction pipeline that uses only one-repeated OCT scan to generate OCTA images. The pipeline is based on the proposed Vasculature Extraction Transformer (VET), which leverages convolutional projection to better learn the spatial relationships between image patches. In comparison to OCTA images obtained via the SV-OCTA (PSNR: 17.809) and ED-OCTA (PSNR: 18.049) using four-repeated OCT scans, OCTA images extracted by VET exhibit moderate quality (PSNR: 17.515) and higher image contrast while reducing the required data acquisition time from ~8 s to ~2 s. Based on visual observations, the proposed VET outperforms SV and ED algorithms when using neck and face OCTA data in areas that are challenging to scan. This study represents that the VET has the capacity to extract vascularture images from a fast one-repeated OCT scan, facilitating accurate diagnosis for patients.

eess.IV

A Multiple Market Trading Mechanism for Electricity, Renewable Energy Certificate and Carbon Emission Right of Virtual Power Plants

A multiple market trading mechanism for the VPP to participate in electricity, renewable energy certificate (REC) and carbon emission right (CER) markets is proposed. With the introduction of the inventory mechanism of REC and CER, the profit of the VPP increases and better trading decisions with multiple markets are made under the requirements of renewable portfolio standard (RPS) and carbon emission (CE) quota requirements. According to the Karush-Kuhn-Tucker (KKT) conditions of the proposed model, properties of the multiple market trading mechanism are discussed. Results from case studies verify the effectiveness of the proposed model.

eess.SY

Spatially controlled electrostatic doping in graphene p-i-n junction for hybrid silicon photodiode

Sufficiently large depletion region for photocarrier generation and separation is a key factor for two-dimensional material optoelectronic devices, but few device configurations has been explored for a deterministic control of a space charge region area in graphene with convincing scalability. Here we investigate a graphene-silicon p-i-n photodiode defined in a foundry processed planar photonic crystal waveguide structure, achieving visible - near-infrared, zero-bias and ultrafast photodetection. Graphene is electrically contacting to the wide intrinsic region of silicon and extended to the p an n doped region, functioning as the primary photocarrier conducting channel for electronic gain. Graphene significantly improves the device speed through ultrafast out-of-plane interfacial carrier transfer and the following in-plane built-in electric field assisted carrier collection. More than 50 dB converted signal-to-noise ratio at 40 GHz has been demonstrated under zero bias voltage, with quantum efficiency could be further amplified by hot carrier gain on graphene-i Si interface and avalanche process on graphene-doped Si interface. With the device architecture fully defined by nanomanufactured substrate, this study is the first demonstration of post-fabrication-free two-dimensional material active silicon photonic devices.

physics.app-ph

Absolute magnetometry based on quantum beats in diamond nitrogen-vacancy centers

We demonstrate an absolute magnetometer immune to temperature fluctuation and strain inhomogeneity, based on quantum beats in the ground state of nitrogen-vacancy centers in diamond. We apply this technique to measure low-frequency magnetic field noise using a single nitrogen-vacancy center located within 500 nm of the surface of an isotopically-pure (99.99% C12) diamond. The photon-shot-noise limited sensitivity achieves 38 nT/Hz^1/2 for 4.45 s acquisition time, a factor of 2^1/2 better than the implementation which uses only two spin levels. For long acquisition times (>10 s), we realize up to a factor of 15 improvement in magnetic sensitivity, which demonstrates the robustness of our technique against thermal drifts. Applying our technique to nitrogen-vacancy center ensembles, we eliminate dephasing from longitudinal strain inhomogeneity, resulting in a factor of 2.3 improvement in sensitivity.

quant-ph

Coupling of Nitrogen-Vacancy Centers to Photonic Crystal Cavities in Monocrystalline Diamond

The zero-phonon transition rate of a nitrogen-vacancy center is enhanced by a factor of ~70 by coupling to a photonic crystal resonator fabricated in monocrystalline diamond using standard semiconductor fabrication techniques. Photon correlation measurements on the spectrally filtered zero-phonon line show antibunching, a signature that the collected photoluminescence is emitted primarily by a single nitrogen-vacancy center. The linewidth of the coupled nitrogen-vacancy center and the spectral diffusion are characterized using high-resolution photoluminescence and photoluminescence excitation spectroscopy.

physics.optics