Searcharxiv⌕ Search

arXiv subjects

Yuanjin Zheng

Publications and source records attributed to Yuanjin Zheng.

17 recordsLinked to original sources

SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents

Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.

cs.CR↗

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenTSE, a two-stage decoder-only generative LM for TSE: Stage-1 predicts coarse semantic tokens, and Stage-2 generates fine acoustic tokens. Separating semantics and acoustics stabilizes decoding and yields more accurate target speech. Both stages use continuous SSL or codec embeddings, offering richer context than discretized-prompt methods. To reduce exposure bias, we employ a Frozen-LM Conditioning training strategy that conditions the LMs on predicted tokens from earlier checkpoints to reduce the gap between teacher-forcing training and autoregressive inference. We further apply DPO to better align outputs with perceptual preferences. Experiments on Libri2Mix show that GenTSE surpasses previous LM-based systems in speech quality, intelligibility, and speaker consistency.

eess.AS↗

CO-QLink: Cryogenic Optical Link for Scalable Quantum Computing Systems and High-Performance Cryogenic Computing Systems

Cryogenic systems necessitate extensive data transmission between room-temperature and cryogenic environments, as well as within the cryogenic temperature domain. High-speed, low-power data transmission is pivotal to enabling the deployment of larger-scale cryogenic systems, including the scalable quantum computing systems and the high-performance cryogenic computing systems fully immersed in liquid nitrogen. In contrast to wireline and microwave links, optical communication links are emerging as a solution characterized by high data rates, high energy efficiency, low signal attenuation, absence of thermal conduction, and superior scalability. In this work, a 4K heat-insulated high-speed (56Gbps) low-power (1.6pJ/b) transceiver (TRX) that achieves a complete link between 4K systems and room temperature (RT) equipment is presented. Copackaged with a PIN photodiode (PD), the RX uses an inverter-based analog front-end and an analog half-rate clock data recovery loop. Connecting to a Mach-Zehnder modulator (MZM), the TX contains a voltage-mode driver with current-mode injection for low-power output-swing-boosting and 3-tap feed-forward equalization (FFE). This link has been demonstrated in the control and readout of a complete superconducting quantum computing system.

quant-ph↗

MTA: A Merge-then-Adapt Framework for Personalized Large Language Model

Personalized Large Language Models (PLLMs) aim to align model outputs with individual user preferences, a crucial capability for user-centric applications. However, the prevalent approach of fine-tuning a separate module for each user faces two major limitations: (1) storage costs scale linearly with the number of users, rendering the method unscalable; and (2) fine-tuning a static model from scratch often yields suboptimal performance for users with sparse data. To address these challenges, we propose MTA, a Merge-then-Adapt framework for PLLMs. MTA comprises three key stages. First, we construct a shared Meta-LoRA Bank by selecting anchor users and pre-training meta-personalization traits within meta-LoRA modules. Second, to ensure scalability and enable dynamic personalization combination beyond static models, we introduce an Adaptive LoRA Fusion stage. This stage retrieves and dynamically merges the most relevant anchor meta-LoRAs to synthesize a user-specific one, thereby eliminating the need for user-specific storage and supporting more flexible personalization. Third, we propose a LoRA Stacking for Few-Shot Personalization stage, which applies an additional ultra-low-rank, lightweight LoRA module on top of the merged LoRA. Fine-tuning this module enables effective personalization under few-shot settings. Extensive experiments on the LaMP benchmark demonstrate that our approach outperforms existing SOTA methods across multiple tasks.

cs.CL↗

Unsupervised Attention-Based Multi-Source Domain Adaptation Framework for Drift Compensation in Electronic Nose Systems

Continuous, long-term monitoring of hazardous, noxious, explosive, and flammable gases in industrial environments using electronic nose (E-nose) systems faces the significant challenge of reduced gas identification accuracy due to time-varying drift in gas sensors. To address this issue, we propose a novel unsupervised attention-based multi-source domain shared-private feature fusion adaptation (AMDS-PFFA) framework for gas identification with drift compensation in E-nose systems. The AMDS-PFFA model effectively leverages labeled data from multiple source domains collected during the initial stage to accurately identify gases in unlabeled gas sensor array drift signals from the target domain. To validate the model's effectiveness, extensive experimental evaluations were conducted using both the University of California, Irvine (UCI) standard drift gas dataset, collected over 36 months, and drift signal data from our self-developed E-nose system, spanning 30 months. Compared to recent drift compensation methods, the AMDS-PFFA model achieves the highest average gas recognition accuracy with strong convergence, attaining 83.20% on the UCI dataset and 93.96% on data from our self-developed E-nose system across all target domain batches. These results demonstrate the superior performance of the AMDS-PFFA model in gas identification with drift compensation, significantly outperforming existing methods.

eess.SP↗

Machine Learning with Real-time and Small Footprint Anomaly Detection System for In-Vehicle Gateway

Anomaly Detection System (ADS) is an essential part of a modern gateway Electronic Control Unit (ECU) to detect abnormal behaviors and attacks in vehicles. Among the existing attacks, ``one-time`` attack is the most challenging to be detected, together with the strict gateway ECU constraints of both microsecond or even nanosecond level real-time budget and limited footprint of code. To address the challenges, we propose to use the self-information theory to generate values for training and testing models, aiming to achieve real-time detection performance for the ``one-time`` attack that has not been well studied in the past. Second, the generation of self-information is based on logarithm calculation, which leads to the smallest footprint to reduce the cost in Gateway. Finally, our proposed method uses an unsupervised model without the need of training data for anomalies or attacks. We have compared different machine learning methods ranging from typical machine learning models to deep learning models, e.g., Hidden Markov Model (HMM), Support Vector Data Description (SVDD), and Long Short Term Memory (LSTM). Experimental results show that our proposed method achieves 8.7 times lower False Positive Rate (FPR), 1.77 times faster testing time, and 4.88 times smaller footprint.

cs.CR↗

HSD-PAM: High Speed Super Resolution Deep Penetration Photoacoustic Microscopy Imaging Boosted by Dual Branch Fusion Network

Photoacoustic microscopy (PAM) is a novel implementation of photoacoustic imaging (PAI) for visualizing the 3D bio-structure, which is realized by raster scanning of the tissue. However, as three involved critical imaging parameters, imaging speed, lateral resolution, and penetration depth have mutual effect to one the other. The improvement of one parameter results in the degradation of other two parameters, which constrains the overall performance of the PAM system. Here, we propose to break these limitations by hardware and software co-design. Starting with low lateral resolution, low sampling rate AR-PAM imaging which possesses the deep penetration capability, we aim to enhance the lateral resolution and up sampling the images, so that high speed, super resolution, and deep penetration for the PAM system (HSD-PAM) can be achieved. Data-driven based algorithm is a promising approach to solve this issue, thereby a dedicated novel dual branch fusion network is proposed, which includes a high resolution branch and a high speed branch. Since the availability of switchable AR-OR-PAM imaging system, the corresponding low resolution, undersample AR-PAM and high resolution, full sampled OR-PAM image pairs are utilized for training the network. Extensive simulation and in vivo experiments have been conducted to validate the trained model, enhancement results have proved the proposed algorithm achieved the best perceptual and quantitative image quality. As a result, the imaging speed is increased 16 times and the imaging lateral resolution is improved 5 times, while the deep penetration merit of AR-PAM modality is still reserved.

eess.IV↗

Speckle-based optical cryptosystem and its application for human face recognition via deep learning

Face recognition has recently become ubiquitous in many scenes for authentication or security purposes. Meanwhile, there are increasing concerns about the privacy of face images, which are sensitive biometric data that should be carefully protected. Software-based cryptosystems are widely adopted nowadays to encrypt face images, but the security level is limited by insufficient digital secret key length or computing power. Hardware-based optical cryptosystems can generate enormously longer secret keys and enable encryption at light speed, but most reported optical methods, such as double random phase encryption, are less compatible with other systems due to system complexity. In this study, a plain yet high-efficient speckle-based optical cryptosystem is proposed and implemented. A scattering ground glass is exploited to generate physical secret keys of gigabit length and encrypt face images via seemingly random optical speckles at light speed. Face images can then be decrypted from the random speckles by a well-trained decryption neural network, such that face recognition can be realized with up to 98% accuracy. The proposed cryptosystem has wide applicability, and it may open a new avenue for high-security complex information encryption and decryption by utilizing optical speckles.

cs.CR↗

Adaptive optical focusing through perturbed scattering media with dynamic mutation algorithm

Optical focusing through/inside scattering media, like multimode fiber and biological tissues, has significant impact in biomedicine yet considered challenging due to strong scattering nature of light. Previously, promising progress has been made, benefiting from the iterative optical wavefront shaping, with which deep-tissue high-resolution optical focusing becomes possible. Most of iterative algorithms can overcome noise perturbations but fail to effectively adapt beyond the noise, e.g. sudden strong perturbations. Re-optimizations are usually needed for significant decorrelated medium since these algorithms heavily rely on the optimization in the previous iterations. Such ineffectiveness is probably due to the absence of a metric that can gauge the deviation of the instant wavefront from the optimum compensation based on the concurrently measured optical focusing. In this study, a square rule of binary-amplitude modulation, directly relating the measured focusing performance with the error in the optimized wavefront, is theoretically proved and experimentally validated. With this simple rule, it is feasible to quantify how many pixels on the spatial light modulator incorrectly modulate the wavefront for the instant status of the medium or the whole system. As an example of application, we propose a novel algorithm, dynamic mutation algorithm, with high adaptability against perturbations by probing how far the optimization has gone toward the theoretically optimum. The diminished focus of scattered light can be effectively recovered when perturbations to the medium cause significant drop of the focusing performance, which no existing algorithms can achieve due to their inherent strong dependence on previous optimizations. With further improvement, this study may boost or inspire many applications, like high-resolution imaging and stimulation, in instable scattering environments.

physics.optics↗

Large-scale Huygens metasurfaces for holographic 3D near-eye displays

Novel display technologies aim at providing the users with increasingly immersive experiences. In this regard, it is a long-sought dream to generate three-dimensional (3D) scenes with high resolution and continuous depth, which can be overlaid with the real world. Current attempts to do so, however, fail in providing either truly 3D information, or a large viewing area and angle, strongly limiting the user immersion. Here, we report a proof-of-concept solution for this problem, and realize a compact holographic 3D near-eye display with a large exit pupil of 10mm x 8.66mm. The 3D image is generated from a highly transparent Huygens metasurface hologram with large (>10^8) pixel count and subwavelength pixels, fabricated via deep-ultraviolet immersion photolithography on 300 mm glass wafers. We experimentally demonstrate high quality virtual 3D scenes with ~50k active data points and continuous depth ranging from 0.5m to 2m, overlaid with the real world and easily viewed by naked eye. To do so, we introduce a new design method for holographic near-eye displays that, inherently, is able to provide both parallax and accommodation cues, fundamentally solving the vergence-accommodation conflict that exists in current commercial 3D displays.

physics.optics↗

Learning-based super interpolation and extrapolation for speckled image reconstruction

Speckles arise when coherent light interacts with biological tissues. Information retrieval from speckles is desired yet challenging, requiring understanding or mapping of the multiple scattering process, or reliable capability to reverse or compensate for the scattering-induced phase distortions. In whatever situation, insufficient sampling of speckles undermines the encoded information, impeding successful object reconstruction from speckle patterns. In this work, we propose a deep learning method to combat the physical limit: the sub-Nyquist sampled speckles (~14 below the Nyquist criterion) are interpolated up to a well-resolved level (1024 times more pixels to resolve the same FOV) with smoothed morphology fine-textured. More importantly, the lost information can be retraced, which is impossible with classic interpolation or any existing methods. The learning network inspires a new perspective on the nature of speckles and a promising platform for efficient processing or deciphering of massive scattered optical signals, enabling widefield high-resolution imaging in complex scenarios.

eess.IV↗

Towards smart optical focusing: Deep learning-empowered wavefront shaping in nonstationary scattering media

Optical focusing at depths in tissue is the Holy Grail of biomedical optics that may bring revolutionary advancement to the field. Wavefront shaping is a widely accepted approach to solve this problem, but most implementations thus far have only operated with stationary media which, however, are scarcely existent in practice. In this article, we propose to apply a deep convolutional neural network named as ReFocusing-Optical-Transformation-Net (RFOTNet), which is a Multi-input Single-output network, to tackle the grand challenge of light focusing in nonstationary scattering media. As known, deep convolutional neural networks are intrinsically powerful to solve inverse scattering problems without complicated computation. Considering the optical speckles of the medium before and after moderate perturbations are correlated, an optical focus can be rapidly recovered based on fine-tuning of pre-trained neural networks, significantly reducing the time and computational cost in refocusing. The feasibility is validated experimentally in this work. The proposed deep learning-empowered wavefront shaping framework has great potentials in facilitating optimal optical focusing and imaging in deep and dynamic tissue.

physics.app-ph↗

A Noninvasive Magnetic Stimulator Utilizing Secondary Ferrite Cores and Resonant Structures for Field Enhancement

In this paper, secondary ferrite cores and resonant structures have been used for field enhancement. The tissue was placed between the double square source coil and the secondary ferrite core. Resonant coils were added which aided in modulating the electric field in the tissue. The field distribution in the tissue was measured using electromagnetic simulations and ex-vivo measurements with tissue. Calculations involve the use of finite element analysis (Ansoft HFSS) to represent the electrical properties of the physical structure. The setup was compared to a conventional design in which the secondary ferrite cores were absent. It was found that the induced electric field could be increased by 122%, when ferrite cores were placed below the tissue at 450 kHz source frequency. The induced electric field was found to be localized in the tissue, verified using ex-vivo experiments. This preliminary study maybe further extended to establish the verified proposed concept with different complicated body parts modelled using the software and in-vivo experiments as required to obtain the desired induced field.

physics.med-ph↗

A hierarchical neural stimulation model for pain relief by variation of coil design parameters

Neural stimulation represents a powerful technique for neural disorder treatment. This paper deals with optimization of coil design parameters to be used during stimulation for modulation of neuronal firing to achieve pain relief. Pain mechanism is briefly introduced and a hierarchical stimulation model from coil stimulation to neuronal firing is proposed. Electromagnetic field distribution for circular, figure of 8 and Magnetic resonance coupling figure of 8 coils are analyzed with respect to the variation of stimulation parameters such as distance between coils, stimulation frequency, number of turns and radius of coils. MRC figure of 8 coils were responsible for inducing the maximum Electric field for same amount of driving current in coils. Variation of membrane potential, ion channel conductance and neuronal firing frequency in a pyramidal neuronal model due to magnetic and acoustic stimulation are studied. The frequency of neuronal firing for cortical neurons is higher during pain state, compared to no pain state. Lowest neuronal firing frequency 18 Hz was found for MRC figure of 8 coils, compared to 30 Hz for circular coils. Therefore, MRC figure of 8 coils are most effective for modulation of neuronal firing, thereby achieving pain relief in comparison to other coils considered in this study

q-bio.NC↗

One laser pulse generates two photoacoustic signals

Photoacoustic sensing and imaging techniques have been studied widely to explore optical absorption contrast based on nanosecond laser illumination. In this paper, we report a long laser pulse induced dual photoacoustic (LDPA) nonlinear effect, which originates from unsatisfied stress and thermal confinements. Being different from conventional short laser pulse illumination, the proposed method utilizes a long square-profile laser pulse to induce dual photoacoustic signals. Without satisfying the stress confinement, the dual photoacoustic signals are generated following the positive and negative edges of the long laser pulse. More interestingly, the first expansion-induced photoacoustic signal exhibits positive waveform due to the initial sharp rising of temperature. On the contrary, the second contraction-induced photoacoustic signal exhibits exactly negative waveform due to the falling of temperature, as well as pulse-width-dependent signal amplitude which is caused by the concurrent heat accumulation and thermal diffusion during the long laser illumination. An analytical model is derived to describe the generation of the dual photoacoustic pulses, incorporating Gruneisen saturation and thermal diffusion effect, which is experimentally proved. Lastly, an alternate of LDPA technique using quasi-CW laser excitation is also introduced and demonstrated for both super-contrast in vitro and in vivo imaging. Compared with existing nonlinear PA techniques, the proposed LDPA nonlinear effect could enable a much broader range of potential applications.

physics.optics↗

Photoacoustics meets ultrasound: micro-Doppler photoacoustic effect and detection by ultrasound

In recent years, photoacoustics has attracted intensive research for both anatomical and functional biomedical imaging. However, the physical interaction between photoacoustic generated endogenous waves and an exogenous ultrasound wave is a largely unexplored area. Here, we report the initial results about the interaction of photoacoustic and external ultrasound waves leading to a micro-Doppler photoacoustic (mDPA) effect, which is experimentally observed and consistently modelled. It is based on a simultaneous excitation on the target with a pulsed laser and continuous wave (CW) ultrasound. The thermoelastically induced expansion will modulate the CW ultrasound and leads to transient Doppler frequency shift. The reported mDPA effect can be described as frequency modulation of the intense CW ultrasound carrier through photoacoustic vibrations. This technique may open the possibility to sensitively detect the photoacoustic vibration in deep optically and acoustically scattering medium, avoiding acoustic distortion that exists in state-of-the-art pulsed photoacoustic imaging systems.

physics.optics↗

Photoacoustic elastic oscillation and characterization

Photoacoustic imaging and sensing have been studied extensively to probe the optical absorption of biological tissue in multiple scales ranging from large organs to small molecules. However, its elastic oscillation characterization is rarely studied and has been an untapped area to be explored. In literature, photoacoustic signal induced by pulsed laser is commonly modelled as a bipolar "N-shape" pulse from an optical absorber. In this paper, the photoacoustic damped oscillation is predicted and modelled by an equivalent mass-spring system by treating the optical absorber as an elastic oscillator. The photoacoustic simulation incorporating the proposed oscillation model shows better agreement with the measured signal from an elastic phantom, than conventional photoacoustic simulation model. More interestingly, the photoacoustic damping oscillation effect could potentially be a useful characterization approach to evaluate biological tissue's mechanical properties in terms of relaxation time, peak number and ratio beyond optical absorption only, which is experimentally demonstrated in this paper.

physics.optics↗