SearcharxivSearch

arXiv subjects

Weibin Zhang

Publications and source records attributed to Weibin Zhang.

At least 19 recordsLinked to original sources

D-ADD: An Effective Plug-In for Defending Against Model Stealing

Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. Timely prevention of such model-stealing attacks is challenging, as it requires achieving robust protection, maintaining utility, and ensuring low deployment overhead at the same time. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local query dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. Together with shift adjustment, the ADD$^+$ detector is enhanced with the capability to manage domain shifts. When integrated with random-based prediction poisoning, it serves as a practical plug-and-play defense module, termed D-ADD, for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings. Codes are available from https://github.com/AI-EXP-group/D-ADD.

cs.CR

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks emerge as the most severe threat, where adversaries exploit prediction outputs to train surrogate models that illegally replicate the original model's functionality. In this work, we propose a rehearsal-based watermark embedding framework to enhance the robustness of model watermarks against model extraction attacks. By simulating the extraction process, our method leverages the loss of a \textit{simulated stolen model} on a trigger set as a training signal to fine-tune the watermark knowledge within the target model. This fine-tuning step encourages the watermark to be embedded in a way that boosts transferability, thereby increasing its chances of persisting and remaining detectable in stolen models. Comprehensive experiments conducted under diverse settings demonstrate that the proposed method significantly improves the robustness of model watermarks against both model extraction and subsequent watermark removal attacks.

cs.CR

Calibration of an Irradiated Prototype for the EIC Zero-Degree Calorimeter

We study the response of a prototype Zero-Degree Calorimeter (ZDC) detector to irradiation equivalent to 10$^{11}$ 1-MeV $n_{\text{eq}}/\text{cm}^{2}$, which matches the expected exposure after one year of operation at full nominal luminosity at the future Electron-Ion Collider (EIC). The prototype, which consists of 563 channels and represents about 10 percent of the final ZDC design in terms of both channel count and detector volume, was irradiated at the NASA Space Radiation Laboratory (NSRL) at Brookhaven National Laboratory (BNL) with proton beams. We demonstrate that, despite significant radiation damage to the SiPMs and non-uniform degradation across the detector volume, the detector can be successfully calibrated on a channel-by-channel basis using cosmic-ray data. The damage profile, similar to what is expected in the experiment, varies by an order of magnitude or more across the detector. Even for the most heavily damaged channels, the signal-to-noise ratio for a MIP signal remains above 5. This study provides a realistic test of the system's performance under irradiation. It complements previous SiPM-specific irradiation studies and will inform the future operation of the ZDC and other detectors that use SiPM-on-tile technology.

physics.ins-det

Beam Test of a SiPM-on-Tile ZDC Prototype with 5.3 GeV Positrons at Jefferson Laboratory

We report on a beam test of a Silicon Photo-Multiplier (SiPM)-on-tile Zero Degree Calorimeter (ZDC) prototype developed for the future Electron-Ion Collider (EIC). The detector implements the staggered scintillator-tile geometry envisioned for the final detector and includes 370 instrumented channels, corresponding to O(10%) of the full ZDC. The detector was tested using a 5.3 GeV positron beam at Jefferson Laboratory. We measure the energy response, shower shape, and spatial reconstruction performance of the detector and compare these results with simulation. These studies provide key input for the optimization of the final ePIC ZDC design.

physics.ins-det

SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition

In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from variable-length speech segments. However, this method treats different frame-level features equally and discards covariance information. In this paper, we propose the Semi-orthogonal parameter pooling of Covariance matrix (SoCov) method. The SoCov pooling computes the covariance matrix from the self-attentive frame-level features and compresses it into a vector using the semi-orthogonal parametric vectorization, which is then concatenated with the weighted standard deviation vector to form inputs to the segment-level layers. Deep embedding based on SoCov is called ``sc-vector''. The proposed sc-vector is compared to several different baselines on the SRE21 development and evaluation sets. The sc-vector system significantly outperforms the conventional x-vector system, with a relative reduction in EER of 15.5% on SRE21Eval. When using self-attentive deep feature, SoCov helps to reduce EER on SRE21Eval by about 30.9% relatively to the conventional ``mean + standard deviation'' statistics.

eess.AS

First-Ever Deployment of a SiPM-on-Tile Calorimeter in a Collider: A Parasitic Test with 200 GeV $pp$ Collisions at RHIC

We describe the testing of a prototype SiPM-on-tile iron-scintillator calorimeter at the Relativistic Heavy Ion Collider (RHIC) during its 200 GeV $pp$ run in 2024. The prototype, measuring $20 \times 20 \, \text{cm}^{2}$ and 24 radiation lengths in depth, was positioned in the STAR experimental hall, approximately 8 m from the interaction point and 65 cm from the beam line, covering a pseudorapidity range of about $3.1<η<3.4$. By using the dark current of a reference SiPM as a radiation monitor, we estimate that the prototype was exposed to a fluence of about $10^{10}$ 1-MeV $n_{\mathrm{eq}}$/cm$^2$. Channel-by-channel calibration was performed in a data-driven way with the signature from minimum-ionizing particles during beam-on conditions. A Geant4 detector simulation, with inputs from the Pythia8 event generator, describes measurements of energy spectra and hit multiplicities reasonably well. These results mark the first deployment, commissioning, calibration, and long-term operation of a SiPM-on-tile calorimeter in a collider environment. This experimental campaign will guide detector designs and operational strategies for the ePIC detector at the future EIC, as well as other applications.

physics.ins-det

A CT Image Denoising Method Based on Projection Domain Feature

In order to improve image quality of projection in industrial applications, generally, a standard method is to increase the current or exposure time, which might cause overexposure of detector units in areas of thin objects or backgrounds. Increasing the projection sampling is a better method to address the issue, but it also leads to significant noise in the reconstructed image. This paper proposed a projection domain denoising algorithm based on the features of the projection domain for this case. This algorithm utilized the similarity of projections of neighboring veiws to reduce image noise quickly and effectively. The availability of the algorithm proposed in this work has been conducted by numerical simulation and practical data experiments.

eess.IV

Content Caching-Assisted Vehicular Edge Computing Using Multi-Agent Graph Attention Reinforcement Learning

In order to avoid repeated task offloading and realize the reuse of popular task computing results, we construct a novel content caching-assisted vehicular edge computing (VEC) framework. In the face of irregular network topology and unknown environmental dynamics, we further propose a multi-agent graph attention reinforcement learning (MGARL) based edge caching scheme, which utilizes the graph attention convolution kernel to integrate the neighboring nodes' features of each agent and further enhance the cooperation among agents. Our simulation results show that our proposed scheme is capable of improving the utilization of caching resources while reducing the long-term task computing latency compared to the baselines.

cs.MA

Multi-Scale Temporal Transformer For Speech Emotion Recognition

Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures focus more on global information and require large computation. On the other hand, abundant speech emotional representations exist locally on different parts of the input speech. To tackle these problems, we propose a Multi-Scale TRansfomer (MSTR) for speech emotion recognition. It comprises of three main components: (1) a multi-scale temporal feature operator, (2) a fractal self-attention module, and (3) a scale mixer module. These three components can effectively enhance the transformer's ability to learn multi-scale local emotion representations. Experimental results demonstrate that the proposed MSTR model significantly outperforms a vanilla Transformer and other state-of-the-art methods across three speech emotion datasets: IEMOCAP, MELD and, CREMAD. In addition, it can greatly reduce the computational cost.

eess.AS

Co-learning-aided Multi-modal-deep-learning Framework of Passive DOA Estimators for a Heterogeneous Hybrid Massive MIMO Receiver

Due to its excellent performance in rate and resolution, fully-digital (FD) massive multiple-input multiple-output (MIMO) antenna arrays has been widely applied in data transmission and direction of arrival (DOA) measurements, etc. But it confronts with two main challenges: high computational complexity and circuit cost. The two problems may be addressed well by hybrid analog-digital (HAD) structure. But there exists the problem of phase ambiguity for HAD, which leads to its low-efficiency or high-latency. Does exist there such a MIMO structure of owning low-cost, low-complexity and high time efficiency at the same time. To satisfy the three properties, a novel heterogeneous hybrid MIMO receiver structure of integrating FD and heterogeneous HAD ($\rm{H}^2$AD-FD) is proposed and corresponding multi-modal (MD)-learning framework is developed. The framework includes three major stages: 1) generate the candidate sets via root multiple signal classification (Root-MUSIC) or deep learning (DL); 2) infer the class of true solutions from candidate sets using machine learning (ML) methods; 3) fuse the two-part true solutions to achieve a better DOA estimation. The above process form two methods named MD-Root-MUSIC and MDDL. To improve DOA estimation accuracy and reduce the clustering complexity, a co-learning-aided MD framework is proposed to form two enhanced methods named CoMDDL and CoMD-RootMUSIC. Moreover, the Cramer-Rao lower bound (CRLB) for the proposed $\rm{H}^2$AD-FD structure is also derived. Experimental results demonstrate that our proposed four methods could approach the CRLB for signal-to-noise ratio (SNR) > 0 dB and the proposed CoMDDL and MDDL perform better than CoMD-RootMUSIC and MD-RootMUSIC, particularly in the extremely low SNR region.

eess.SP

Federated Unlearning for Human Activity Recognition

The rapid evolution of Internet of Things (IoT) technology has spurred the widespread adoption of Human Activity Recognition (HAR) in various daily life domains. Federated Learning (FL) is frequently utilized to build a global HAR model by aggregating user contributions without transmitting raw individual data. Despite substantial progress in user privacy protection with FL, challenges persist. Regulations like the General Data Protection Regulation (GDPR) empower users to request data removal, raising a new query in FL: How can a HAR client request data removal without compromising other clients' privacy? In response, we propose a lightweight machine unlearning method for refining the FL HAR model by selectively removing a portion of a client's training data. Our method employs a third-party dataset unrelated to model training. Using KL divergence as a loss function for fine-tuning, we aim to align the predicted probability distribution on forgotten data with the third-party dataset. Additionally, we introduce a membership inference evaluation method to assess unlearning effectiveness. Experimental results across diverse datasets show our method achieves unlearning accuracy comparable to \textit{retraining} methods, resulting in speedups ranging from hundreds to thousands.

cs.LG

An inspection technology of inner surface of the fine hole based on machine vision

Fine holes are an important structural component of industrial components, and their inner surface quality is closely related to their function.In order to detect the quality of the inner surface of the fine hole,a special optical measurement system was investigated in this paper. A sight pipe is employed to guide the external illumination light into the fine hole and output the relevant images simultaneously. A flexible light array is introduced to suit the narrow space, and the effective field of view is analyzed. Besides, the arc surface projection error and manufacturing assembly error of the device are analyzed, then compensated or ignored if small enough. In the test of prefabricated circular defects with the diameter ϕ0.1mm, ϕ0.2mm, 0.4mm distance distribution and the fissure defects with the width 0.3mm, the maximum measurement error standard deviation are all about 10μm. The minimum diameter of the measured fine hole is 4mm and the depth can reach 47mm.

cs.CV

Beam Test of the First Prototype of SiPM-on-Tile Calorimeter Insert for the Electron-Ion Collider Using 4 GeV Positrons at Jefferson Laboratory

We recently proposed a high-granularity calorimeter insert for the Electron-Ion Collider (EIC) that uses plastic scintillator tiles read out by SiPMs. Among its innovative features are an ASIC-away-of-SiPM strategy for reducing cooling requirements and minimizing space use, along with employing 3D-printed frames to reduce optical crosstalk and dead areas. To evaluate these features, we built a 40-channel prototype and tested it using a 4 GeV positron beam at Jefferson Laboratory. The measured energy spectra and 3D shower shapes are well described by simulations, confirming the effectiveness of the design, construction techniques, and calibration strategy. This constitutes the first use of SiPM-on-tile technology in EIC detector designs.

physics.ins-det

A Few-Degree Calorimeter for the future Electron-Ion Collider

Measuring the region $0.1 < Q^{2} < 1.0$ GeV$^{2}$ is essential to support searches for gluon saturation at the future Electron-Ion Collider. Recent studies have revealed that covering this region at the highest beam energies is not feasible with current detector designs, resulting in the so-called $Q^{2}$ gap. In this work, we present a design for the Few-Degree Calorimeter (FDC), which addresses this issue. The FDC uses SiPM-on-tile technology with tungsten absorber and covers the range of $-4.6 < η< -3.6$. It offers fine transverse and longitudinal granularity, along with excellent time resolution, enabling standalone electron tagging. Our design represents the first concrete solution to bridge the $Q^{2}$ gap at the EIC.

physics.ins-det

DWFormer: Dynamic Window transFormer for Speech Emotion Recognition

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although transformer-based models have made progress in this field, the existing models could not precisely locate important regions at different temporal scales. To address the issue, we propose Dynamic Window transFormer (DWFormer), a new architecture that leverages temporal importance by dynamically splitting samples into windows. Self-attention mechanism is applied within windows for capturing temporal important information locally in a fine-grained way. Cross-window information interaction is also taken into account for global communication. DWFormer is evaluated on both the IEMOCAP and the MELD datasets. Experimental results show that the proposed model achieves better performance than the previous state-of-the-art methods.

cs.SD

Internal Language Model Estimation Through Explicit Context Vector Learning for Attention-based Encoder-decoder ASR

An end-to-end (E2E) ASR model implicitly learns a prior Internal Language Model (ILM) from the training transcripts. To fuse an external LM using Bayes posterior theory, the log likelihood produced by the ILM has to be accurately estimated and subtracted. In this paper we propose two novel approaches to estimate the ILM based on Listen-Attend-Spell (LAS) framework. The first method is to replace the context vector of the LAS decoder at every time step with a vector that is learned with training transcripts. Furthermore, we propose another method that uses a lightweight feed-forward network to directly map query vector to context vector in a dynamic sense. Since the context vectors are learned by minimizing the perplexities on training transcripts, and their estimation is independent of encoder output, hence the ILMs are accurately learned for both methods. Experiments show that the ILMs achieve the lowest perplexity, indicating the efficacy of the proposed methods. In addition, they also significantly outperform the shallow fusion method, as well as two previously proposed ILM Estimation (ILME) approaches on several datasets.

eess.AS

Fast Iterative Reconstruction for Multi-spectral CT by a Schmidt Orthogonal Modification Algorithm (SOMA)

Multi-spectral CT (MSCT) is increasingly used in industrial non-destructive testing and medical diagnosis because of its outstanding performance like material distinguishability. The process of obtaining MSCT data can be modeled as nonlinear equations and the basis material decomposition comes down to the inverse problem of the nonlinear equations. For different spectra data, geometric inconsistent parameters cause geometrical inconsistent rays, which will lead to mismatched nonlinear equations. How to solve the mismatched nonlinear equations accurately and quickly is a hot issue. This paper proposes a general iterative method to invert the mismatched nonlinear equations and develops Schmidt orthogonalization to accelerate convergence. The validity of the proposed method is verified by MSCT basis material decomposition experiments. The results show that the proposed method can decompose the basis material images accurately and improve the convergence speed greatly.

math.NA

Predicted high-temperature superconductivity in rare earth hydride ErH2 at moderate pressure

Hydrides offer an opportunity to study high-temperature (Tc) superconductivity at experimentally achievable pressures. However, they remained extremely high. Using density functional theory calculations, herein we demonstrated that a newly rare earth hydride, namely bulk ErH2, could be superconducting with a Tc around 80 K at 14.5 GPa. To date, the drived pressure is the lowest reported value for compressed hydrides. Besides superconductivity, Fermi Surface nesting and Kondo effect were manifested at this pressure. Intriguingly, due to Kondo destruction, superconductivity was prone to exist at 15 GPa. Under the rest of applied pressures, we also revealed a gap of band structure at 20 GPa on the background of normal metallic states. At 20 GPa, this compressed system could act as a host of superconductor being judged from a sharp jump of spontaneous magnetic susceptibility with an evanescent spin density of state at Fermi level along with the competition between spin density wave and superconductivity. Finally, electron pairing glue for ErH2 at these three typical pressures was attributed to the antiferromagnetic spin fluctuation.

cond-mat.supr-con