Searcharxiv⌕ Search

arXiv subjects

Lin Zhang

Publications and source records attributed to Lin Zhang.

At least 181 records · Page 10Linked to original sources

$\text{H}^2\text{TNE}$: Temporal Heterogeneous Information Network Embedding in Hyperbolic Spaces

Temporal heterogeneous information network (temporal HIN) embedding, aiming to represent various types of nodes of different timestamps into low dimensional spaces while preserving structural and semantic information, is of vital importance in diverse real-life tasks. Researchers have made great efforts on temporal HIN embedding in Euclidean spaces and got some considerable achievements. However, there is always a fundamental conflict that many real-world networks show hierarchical property and power-law distribution, and are not isometric of Euclidean spaces. Recently, representation learning in hyperbolic spaces has been proved to be valid for data with hierarchical and power-law structure. Inspired by this character, we propose a hyperbolic heterogeneous temporal network embedding ($\text{H}^2\text{TNE}$) model for temporal HINs. Specifically, we leverage a temporally and heterogeneously double-constrained random walk strategy to capture the structural and semantic information, and then calculate the embedding by exploiting hyperbolic distance in proximity measurement. Experimental results show that our method has superior performance on temporal link prediction and node classification compared with SOTA models.

cs.SI↗

Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio

This paper defines Spoof Diarization as a novel task in the Partial Spoof (PS) scenario. It aims to determine what spoofed when, which includes not only locating spoof regions but also clustering them according to different spoofing methods. As a pioneering study in spoof diarization, we focus on defining the task, establishing evaluation metrics, and proposing a benchmark model, namely the Countermeasure-Condition Clustering (3C) model. Utilizing this model, we first explore how to effectively train countermeasures to support spoof diarization using three labeling schemes. We then utilize spoof localization predictions to enhance the diarization performance. This first study reveals the high complexity of the task, even in restricted scenarios where only a single speaker per audio file and an oracle number of spoofing methods are considered. Our code is available at https://github.com/nii-yamagishilab/PartialSpoof.

eess.AS↗

Adapter-X: A Novel General Parameter-Efficient Fine-Tuning Framework for Vision

Parameter-efficient fine-tuning (PEFT) has become increasingly important as foundation models continue to grow in both popularity and size. Adapter has been particularly well-received due to their potential for parameter reduction and adaptability across diverse tasks. However, striking a balance between high efficiency and robust generalization across tasks remains a challenge for adapter-based methods. We analyze existing methods and find that: 1) parameter sharing is the key to reducing redundancy; 2) more tunable parameters, dynamic allocation, and block-specific design are keys to improving performance. Unfortunately, no previous work considers all these factors. Inspired by this insight, we introduce a novel framework named Adapter-X. First, a Sharing Mixture of Adapters (SMoA) module is proposed to fulfill token-level dynamic allocation, increased tunable parameters, and inter-block sharing at the same time. Second, some block-specific designs like Prompt Generator (PG) are introduced to further enhance the ability of adaptation. Extensive experiments across 2D image and 3D point cloud modalities demonstrate that Adapter-X represents a significant milestone as it is the first to outperform full fine-tuning in both 2D image and 3D point cloud modalities with significantly fewer parameters, i.e., only 0.20% and 1.88% of original trainable parameters for 2D and 3D classification tasks. Our code will be publicly available.

cs.CV↗

How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?

Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding of the decision-making process of CMs is limited. We utilize Grad-CAM and introduce a quantitative analysis metric to interpret CMs' decisions. We find that CMs prioritize the artifacts of transition regions created when concatenating bona fide and spoofed audio. This focus differs from that of CMs trained on fully spoofed audio, which concentrate on the pattern differences between bona fide and spoofed parts. Our further investigation explains the varying nature of CMs' focus while making correct or incorrect predictions. These insights provide a basis for the design of CM models and the creation of datasets. Moreover, this work lays a foundation of interpretability in the field of partial spoofed audio detection that has not been well explored previously.

eess.AS↗

XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference

Length generalization failure problem, namely the large language model (LLM) fails to generalize to texts longer than its maximum training length, greatly restricts the application of LLM in the scenarios with streaming long inputs. To address this problem, the existing methods either require substantial costs or introduce precision loss. In this paper, we empirically find that the accuracy of the LLM's prediction is highly correlated to its certainty. Based on this, we propose an efficient training free framework, named XL3M (it means extra-long large language model), which enables the LLMs trained on short sequences to reason extremely long sequence without any further training or fine-tuning. Under the XL3M framework, the input context will be firstly decomposed into multiple short sub-contexts, where each sub-context contains an independent segment and a common ``question'' which is a few tokens from the end of the original context. Then XL3M gives a method to measure the relevance between each segment and the ``question'', and constructs a concise key context by splicing all the relevant segments in chronological order. The key context is further used instead of the original context to complete the inference task. Evaluations on comprehensive benchmarks show the superiority of XL3M. Using our framework, a Llama2-7B model is able to reason 20M long sequences on an 8-card Huawei Ascend 910B NPU machine with 64GB memory per card.

cs.CL↗

Diffusion Model Driven Test-Time Image Adaptation for Robust Skin Lesion Classification

Deep learning-based diagnostic systems have demonstrated potential in skin disease diagnosis. However, their performance can easily degrade on test domains due to distribution shifts caused by input-level corruptions, such as imaging equipment variability, brightness changes, and image blur. This will reduce the reliability of model deployment in real-world scenarios. Most existing solutions focus on adapting the source model through retraining on different target domains. Although effective, this retraining process is sensitive to the amount of data and the hyperparameter configuration for optimization. In this paper, we propose a test-time image adaptation method to enhance the accuracy of the model on test data by simultaneously updating and predicting test images. We modify the target test images by projecting them back to the source domain using a diffusion model. Specifically, we design a structure guidance module that adds refinement operations through low-pass filtering during reverse sampling, regularizing the diffusion to preserve structural information. Additionally, we introduce a self-ensembling scheme automatically adjusts the reliance on adapted and unadapted inputs, enhancing adaptation robustness by rejecting inappropriate generative modeling results. To facilitate this study, we constructed the ISIC2019-C and Dermnet-C corruption robustness evaluation benchmarks. Extensive experiments on the proposed benchmarks demonstrate that our method makes the classifier more robust across various corruptions, architectures, and data regimes. Our datasets and code will be available at \url{https://github.com/minghu0830/Skin-TTA_Diffusion}.

eess.IV↗

Polarization dependent non-Hermitian atomic grating controlled by dipole blockade effect

We propose a theoretical scheme for a non-Hermitian atomic grating within an ultra-cold rubidium-87 ($^{87}Rb$) atomic ensemble. The grating's diffraction properties depend on the polarization states of incident photons and are controlled non-locally through Rydberg interactions. Multiple types of polarization-dependent diffraction modes are generated, benefiting from no crosstalk atomic transition channels based on transition selection rules. Those polarization-dependent diffraction modes can be switched using dynamic optical pulse trains, exploiting the Rydberg blockade effect, and are tunable by non-Hermitian optical modulation. Our work will advance the application of asymmetric optical scattering by utilizing the polarization degree of freedom within continuous media and benefit the application of versatile non-Hermitian/asymmetric optical devices.

physics.atom-ph↗

A characterization of entangled two-qubit states via partial-transpose-moments

Although quantum entanglement is an important resource, its characterization is quite challenging. The partial transposition is a common method to detect bipartite entanglement. In this paper, the authors study the partial-transpose(PT)-moments of two-qubit states,and completely describe the whole region, composed of the second and third PT-moments, for all two-qubit states. Furthermore, they determine the accurate region corresponding to all entangled two-qubit states. The states corresponding to those boundary points of the whole region, and to the border lines between separable and entangled states are analyzed. As an application, they characterize the entangled region of PT-moments for the two families of Werner states and Bell-diagonal states. The relations between entanglement and the pairs of PT-moments are revealed from these typical examples. They also numerically plot the whole region of possible PT-moments for all two-qubit X-states, and find that this region is almost the same as the whole region of PT-moments for all two-qubit states. Moreover, they extend their results to detect the entanglement of multiqubit states. By utilizing the PT-moment-based method to characterize the entanglement of the multiqubit states mixed by the GHZ and W states, they propose an operational way of verifying the genuine entanglement in such states.

quant-ph↗

Uncertainty relation and the constrained quadratic programming

The uncertainty relation is a fundamental concept in quantum theory, plays a pivotal role in various quantum information processing tasks. In this study, we explore the additive uncertainty relation pertaining to two or more observables, in terms of their variance,by utilizing the generalized Gell-Mann representation in qudit systems. We find that the tight state-independent lower bound of the variance sum can be characterized as a quadratic programming problem with nonlinear constraints in optimization theory. As illustrative examples, we derive analytical solutions for these quadratic programming problems in lower-dimensional systems, which align with the state-independent lower bounds. Additionally, we introduce a numerical algorithm tailored for solving these quadratic programming instances, highlighting its efficiency and accuracy. The advantage of our approach lies in its potential ability to simultaneously achieve the optimal value of the quadratic programming problem with nonlinear constraints but also precisely identify the extremal state where this optimal value is attained. This enables us to establish a tight state-independent lower bound for the sum of variances, and further identify the extremal state at which this lower bound is realized.

quant-ph↗

Decentralised, Collaborative, and Privacy-preserving Machine Learning for Multi-Hospital Data

Machine Learning (ML) has demonstrated its great potential on medical data analysis. Large datasets collected from diverse sources and settings are essential for ML models in healthcare to achieve better accuracy and generalizability. Sharing data across different healthcare institutions is challenging because of complex and varying privacy and regulatory requirements. Hence, it is hard but crucial to allow multiple parties to collaboratively train an ML model leveraging the private datasets available at each party without the need for direct sharing of those datasets or compromising the privacy of the datasets through collaboration. In this paper, we address this challenge by proposing Decentralized, Collaborative, and Privacy-preserving ML for Multi-Hospital Data (DeCaPH). It offers the following key benefits: (1) it allows different parties to collaboratively train an ML model without transferring their private datasets; (2) it safeguards patient privacy by limiting the potential privacy leakage arising from any contents shared across the parties during the training process; and (3) it facilitates the ML model training without relying on a centralized server. We demonstrate the generalizability and power of DeCaPH on three distinct tasks using real-world distributed medical datasets: patient mortality prediction using electronic health records, cell-type classification using single-cell human genomes, and pathology identification using chest radiology images. We demonstrate that the ML models trained with DeCaPH framework have an improved utility-privacy trade-off, showing it enables the models to have good performance while preserving the privacy of the training data points. In addition, the ML models trained with DeCaPH framework in general outperform those trained solely with the private datasets from individual parties, showing that DeCaPH enhances the model generalizability.

cs.LG↗

Resonances involving integer magnons and spin-1/2 excitations in a magnetism modulated two-dimensional electron gas

We conduct an experimental study of high-mobility two-dimensional electron gas (2DEG) in GaAs/AlGaAs quantum wells modulated by strong magnetism at an in-plane magnetic field ($B$). The modulated $B$-fields are performed via the single stripe and gratings which are made of the heavy rare earth metal Terbium (Tb) thin films on the sample surface. The robust ferromagnetic resonances (FMRs) persist to the temperature of 50-70 K in both the stripe and grating samples, for the ferromagnetism (FM) phase of Tb exists at above 100 K. The high-order (with integer numbers of $j \equiv \hbarω/gμ_{B}B = 1, 2,$...) magnetic resonances can also be observed in the stripe structure via the microwave (MW) photovoltaic detection and magnetoresistance under microwave irradiation. In addition, the resonance features around $j = 1/2$ are robust in the single-stripe modulated sample, which suggests the spinons with spin-1/2 collective excitations in a 1D Heisenberg model.

cond-mat.mes-hall↗

Controllable quantum scars induced by spin-orbit couplings in quantum dots

Spin-orbit couplings (SOCs), originating from the relativistic corrections in the Dirac equation, offer nonlinearity in the classical limit and are capable of driving chaotic dynamics. In a nanoscale quantum dot confined by a two-dimensional parabolic potential with SOCs, various quantum scar states emerge quasi-periodically in the eigenstates of the system, when the ratio of confinement energies in the two directions is nearly commensurable. The scars, displaying both quantum interference and classical trajectory features on the electron density, due to relativistic effects, serve as a bridge between the classical and quantum behaviors of the system. When the strengths of Rashba and Dresselhaus SOCs are identical, the chaos in the classical limit is eliminated as the classical Hamilton's equations become linear, leading to the disappearance of all quantum scar states. Importantly, the quantum scars induced by SOCs are robust against small perturbations of system parameters. With precise control achievable through external gating, the quantum scar induced by Rashba SOC is fully controllable and detectable.

cond-mat.mes-hall↗

Recovery from Adversarial Attacks in Cyber-physical Systems: Shallow, Deep and Exploratory Works

Cyber-physical systems (CPS) have experienced rapid growth in recent decades. However, like any other computer-based systems, malicious attacks evolve mutually, driving CPS to undesirable physical states and potentially causing catastrophes. Although the current state-of-the-art is well aware of this issue, the majority of researchers have not focused on CPS recovery, the procedure we defined as restoring a CPS's physical state back to a target condition under adversarial attacks. To call for attention on CPS recovery and identify existing efforts, we have surveyed a total of 30 relevant papers. We identify a major partition of the proposed recovery strategies: shallow recovery vs. deep recovery, where the former does not use a dedicated recovery controller while the latter does. Additionally, we surveyed exploratory research on topics that facilitate recovery. From these publications, we discuss the current state-of-the-art of CPS recovery, with respect to applications, attack type, attack surfaces and system dynamics. Then, we identify untouched sub-domains in this field and suggest possible future directions for researchers.

eess.SY↗

Dynamic solar Primakoff process

The Primakoff mechanism is one of the primary channels for the production of solar axion. In canonical estimation of the Primakoff photon-axion conversion rate, the recoil effect is neglected and a static structure factor is adopted. By use of the linear response theory, we provide a dynamic description of the solar Primakoff process. It is found that the collective electrons overtake ions as the dominant factor, in contrast to the static screening picture where ions contribute more to the photon-axion conversion. Nonetheless, the resulting axion flux is only 1-2% lower than the standard estimate based on the static structure factor.

hep-ph↗

AgentGroupChat: An Interactive Group Chat Simulacra For Better Eliciting Emergent Behavior

Language significantly influences the formation and evolution of Human emergent behavior, which is crucial in understanding collective intelligence within human societies. Considering that the study of how language affects human behavior needs to put it into the dynamic scenarios in which it is used, we introduce AgentGroupChat in this paper, a simulation that delves into the complex role of language in shaping collective behavior through interactive debate scenarios. Central to this simulation are characters engaging in dynamic conversation interactions. To enable simulation, we introduce the Verbal Strategist Agent, utilizing large language models to enhance interaction strategies by incorporating elements of persona and action. We set four narrative scenarios based on AgentGroupChat to demonstrate the simulation's capacity to mimic complex language use in group dynamics. Evaluations focus on aligning agent behaviors with human expectations and the emergence of collective behaviors within the simulation. Results reveal that emergent behaviors materialize from a confluence of factors: a conducive environment for extensive information exchange, characters with diverse traits, high linguistic comprehension, and strategic adaptability. During discussions on ``the impact of AI on humanity'' in AgentGroupChat simulation, philosophers commonly agreed that ``AI could enhance societal welfare with judicious limitations'' and even come to a conclusion that ``the essence of true intelligence encompasses understanding the necessity to constrain self abilities''. Additionally, in the competitive domain of casting for primary roles in films in AgentGroupChat, certain actors were ready to reduce their remuneration or accept lesser roles, motivated by their deep-seated desire to contribute to the project.

cs.AI↗

Dynamical characterization of $Z_{2}$ Floquet topological phases via quantum quenches

The complete characterization of a generic $d$-dimensional Floquet topological phase is usually hard for the requirement of information about the micromotion throughout the entire driving period. In a recent work [L. Zhang et al., Phys. Rev. Lett. 125, 183001 (2020)], an experimentally feasible dynamical detection scheme was proposed to characterize the integer Floquet topological phases using quantum quenches. However, this theory is still far away from completion, especially for free-fermion Floquet topological phases, where the states can also be characterized by $Z_{2}$ invariants. Here we develop the first full and unified dynamical characterization theory for the $Z_{2}$ Floquet topological phases of different dimensionality and tenfold-way symmetry classes by quenching the system from a trivial and static initial state to the Floquet topological regime through suddenly changing the parameters and turning on the periodic driving. By measuring the minimal information of Floquet bands via the stroboscopic time-averaged spin polarizations, we show that the topological spin texture patterns emerging on certain discrete momenta of Brillouin zone called the $0$ or $π$ gap highest-order band-inversion surfaces provide a measurable dynamical $Z_{2}$ Floquet invariant, which uniquely determines the Floquet boundary modes in the corresponding quasienergy gap and characterizes the $Z_{2}$ Floquet topology. The applications of our theory are illustrated via one- and two-dimensional models that are accessible in current quantum simulation experiments. Our work provides a highly feasible way to detect the $Z_{2}$ Floquet topology and completes the dynamical characterization for the full tenfold classes of Floquet topological phases, which shall advance the research in theory and experiments.

cond-mat.quant-gas↗

Dynamic motion trajectory control with nanoradian accuracy for multi-element X-ray optical systems via laser interferometry

The past decades have witnessed the development of new X-ray beam sources with brightness growing at a rate surpassing Moore's law. Current and upcoming diffraction limited and fully coherent X-ray beam sources, including multi-bend achromat based synchrotron sources and high repetition rate X-ray free electron lasers, puts increasingly stringent requirements on stability and accuracy of X-ray optics systems. Parasitic motion errors at sub-micro radian scale in beam transport and beam conditioning optics can lead to significant loss of coherence and brightness delivered from source to experiment. To address this challenge, we incorporated optical metrology based on interferometry and differential wavefront sensing as part of the X-ray optics motion control system. A prototype X-ray optics system was constructed following the optical layout of a tunable X-ray cavity. On-line interferometric metrology enabled dynamical feedback to a motion control system to track and compensate for motion errors. The system achieved sub-microradian scale performance, as multiple optical elements are synchronously and continuously adjusted. This first proof of principle measurement demonstrated both the potential and necessity of incorporating optical metrology as part of the motion control architecture for large scale X-ray optical systems such as monochromators, delay lines, and in particular, X-ray cavity systems to enable the next generation cavity-based X-ray free electron lasers.

physics.optics↗

Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models

Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of large language models~(LLMs), evaluating how these models identify key information and reason to solve questions becomes increasingly relevant. We introduces the DetectBench, a reading comprehension dataset designed to assess a model's ability to jointly ability in key information detection and multi-hop reasoning when facing complex and implicit information. The DetectBench comprises 3,928 questions, each paired with a paragraph averaging 190 tokens in length. To enhance model's detective skills, we propose the Detective Thinking Framework. These methods encourage models to identify all possible clues within the context before reasoning. Our experiments reveal that existing models perform poorly in both information detection and multi-hop reasoning. However, the Detective Thinking Framework approach alleviates this issue.

cs.CL↗