SearcharxivSearch

arXiv subjects

Rui Lin

Publications and source records attributed to Rui Lin.

At least 37 records · Page 2Linked to original sources

Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation

Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source multi-scale representations module to obtain complementary segment- and track-level features, (2) a hierarchical augmentation strategy to mitigate overfitting, and (3) a hybrid training objective that integrates regression and ranking losses for accurate scoring and reliable top-tier song identification. Experiments demonstrate that HEAR consistently outperforms the baseline across all metrics on both tracks of the ICASSP 2026 SongEval benchmark. The code and trained model weights are available at https://github.com/Eps-Acoustic-Revolution-Lab/EAR_HEAR.

cs.SD

Back to Ear: Perceptually Driven High Fidelity Music Reconstruction

Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual aspects during training, leading to weaknesses in phase accuracy and stereophonic spatial representation. To address these challenges, we propose εar-VAE, an open-source music signal reconstruction model that rethinks and optimizes the VAE training paradigm. Our contributions are threefold: (i) A K-weighting perceptual filter applied prior to loss calculation to align the objective with auditory perception. (ii) Two novel phase losses: a Correlation Loss for stereo coherence, and a Phase Loss using its derivatives--Instantaneous Frequency and Group Delay--for precision. (iii) A new spectral supervision paradigm where magnitude is supervised by all four Mid/Side/Left/Right components, while phase is supervised only by the LR components. Experiments show εar-VAE at 44.1kHz substantially outperforms leading open-source models across diverse metrics, showing particular strength in reconstructing high-frequency harmonics and the spatial characteristics.

cs.SD

Adjusting the Output of Decision Transformer with Action Gradient

Decision Transformer (DT), which integrates reinforcement learning (RL) with the transformer model, introduces a novel approach to offline RL. Unlike classical algorithms that take maximizing cumulative discounted rewards as objective, DT instead maximizes the likelihood of actions. This paradigm shift, however, presents two key challenges: stitching trajectories and extrapolation of action. Existing methods, such as substituting specific tokens with predictive values and integrating the Policy Gradient (PG) method, address these challenges individually but fail to improve performance stably when combined due to inherent instability. To address this, we propose Action Gradient (AG), an innovative methodology that directly adjusts actions to fulfill a function analogous to that of PG, while also facilitating efficient integration with token prediction techniques. AG utilizes the gradient of the Q-value with respect to the action to optimize the action. The empirical results demonstrate that our method can significantly enhance the performance of DT-based algorithms, with some results achieving state-of-the-art levels.

cs.LG

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Large language models (LLMs) hold promise in clinical decision support but face major challenges in safety evaluation and effectiveness validation. We developed the Clinical Safety-Effectiveness Dual-Track Benchmark (CSEDB), a multidimensional framework built on clinical expert consensus, encompassing 30 criteria covering critical areas like critical illness recognition, guideline adherence, and medication safety, with weighted consequence measures. Thirty-two specialist physicians developed and reviewed 2,069 open-ended Q&A items aligned with these criteria, spanning 26 clinical departments to simulate real-world scenarios. Benchmark testing of six LLMs revealed moderate overall performance (average total score 57.2%, safety 54.7%, effectiveness 62.3%), with a significant 13.3% performance drop in high-risk scenarios (p < 0.0001). Domain-specific medical LLMs showed consistent performance advantages over general-purpose models, with relatively higher top scores in safety (0.912) and effectiveness (0.861). The findings of this study not only provide a standardized metric for evaluating the clinical application of medical LLMs, facilitating comparative analyses, risk exposure identification, and improvement directions across different scenarios, but also hold the potential to promote safer and more effective deployment of large language models in healthcare environments.

cs.CL

Pauli crystal superradiance

Pauli crystals are unique geometric structures of non-interacting fermions, resembling crystals, that emerge solely from Fermi statistics and confinement. Unlike genuine quantum crystals that arise from interparticle interactions, Pauli crystals do not break translation symmetry but nonetheless exhibit nontrivial many-body correlations. In this Letter, we explore Pauli crystal formation in a cavity-fermion setup. We analytically show that when coupled to a cavity, degeneracy in Pauli crystals can trigger zero-threshold transitions to superradiance. This superradiance is accompanied by the emergence of a genuine quantum crystalline state, wherein the atomic density is periodically modulated. We substantiate our findings using state-of-the-art numerical simulations. The combined interplay between statistics, confinement geometry and interactions mediated by light thus facilitates a novel pathway to quantum crystallization.

cond-mat.quant-gas

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually comprehensible components. Lorsa is designed to address the challenge of attention superposition to understand attention-mediated interaction between features in different token positions. We show that Lorsa heads find cleaner and finer-grained versions of previously discovered MHSA behaviors like induction heads, successor heads and attention sink behavior (i.e., heavily attending to the first token). Lorsa and Sparse Autoencoder (SAE) are both sparse dictionary learning methods applied to different Transformer components, and lead to consistent findings in many ways. For instance, we discover a comprehensive family of arithmetic-specific Lorsa heads, each corresponding to an atomic operation in Llama-3.1-8B. Automated interpretability analysis indicates that Lorsa achieves parity with SAE in interpretability while Lorsa exhibits superior circuit discovery properties, especially for features computed collectively by multiple MHSA heads. We also conduct extensive experiments on architectural design ablation, Lorsa scaling law and error analysis.

cs.LG

Quantum adiabatic optimization with Rydberg arrays: localization phenomena and encoding strategies

Quantum adiabatic optimization seeks to solve combinatorial problems using quantum dynamics, requiring the Hamiltonian of the system to align with the problem of interest. However, these Hamiltonians are often incompatible with the native constraints of quantum hardware, necessitating encoding strategies to map the original problem into a hardware-conformant form. While the classical overhead associated with such mappings is easily quantifiable and typically polynomial in problem size, it is much harder to quantify their overhead on the quantum algorithm, e.g., the transformation of the adiabatic timescale. In this work, we address this challenge on the concrete example of the encoding scheme proposed in [Nguyen et al., PRX Quantum 4, 010316 (2023)], which is designed to map optimization problems on arbitrarily connected graphs into Rydberg atom arrays. We consider the fundamental building blocks underlying this encoding scheme and determine the scaling of the minimum gap with system size along adiabatic protocols. Even when the original problem is trivially solvable, we find that the encoded problem can exhibit an exponentially closing minimum gap. We show that this originates from a quantum coherent effect, which gives rise to an unfavorable localization of the ground-state wavefunction. On the QuEra Aquila neutral atom machine, we observe such localization and its effect on the success probability of finding the correct solution to the encoded optimization problem. Finally, we propose quantum-aware modifications of the encoding scheme that avoid this quantum bottleneck and lead to an exponential improvement in the adiabatic performance. This highlights the crucial importance of accounting for quantum effects when designing strategies to encode classical problems onto quantum platforms.

quant-ph

AI-Enabled Rapid Assembly of Thousands of Defect-Free Neutral Atom Arrays with Constant-time-overhead

Assembling increasingly larger-scale defect-free optical tweezer-trapped atom arrays is essential for quantum computation and quantum simulations based on atoms. Here, we propose an AI-enabled, rapid, constant-time-overhead rearrangement protocol, and we experimentally assemble defect-free 2D and 3D atom arrays with up to 2024 atoms with a constant time cost of 60 ms. The AI model calculates the holograms for real-time atom rearrangement. With precise controls over both position and phase, a high-speed spatial light modulator moves all the atoms simultaneously. This protocol can be readily used to generate defect-free arrays of tens of thousands of atoms with current technologies, and become a useful toolbox for quantum error correction.

quant-ph

Tunable Einstein-Bohr recoiling-slit gedankenexperiment at the quantum limit

In 1927, during the fifth Solvay Conference, Einstein and Bohr described a double-slit interferometer with a "movable slit" that can detect the momentum recoil of one photon. Here, we report a faithful realization of the Einstein-Bohr interferometer using a single atom in an optical tweezer, cooled to the motional ground state in three dimensions. The single atom has an intrinsic momentum uncertainty comparable to a single photon, which serves as a movable slit obeying the minimum Heisenberg uncertainty principle. The atom's momentum wavefunction is dynamically tunable by the tweezer laser power, which enables observation of an interferometric visibility reduction at a shallower trap, demonstrating the quantum nature of this interferometer. We further identify classical noise due to atom heating and precession, illustrating a quantum-to-classical transition.

quant-ph

A Unifying Tensor View for Lightweight CNNs

Despite the decomposition of convolutional kernels for lightweight CNNs being well studied, existing works that rely on tensor network diagrams or hyperdimensional abstraction lack geometry intuition. This work devises a new perspective by linking a 3D-reshaped kernel tensor to its various slice-wise and rank-1 decompositions, permitting a straightforward connection between various tensor approximations and efficient CNN modules. Specifically, it is discovered that a pointwise-depthwise-pointwise (PDP) configuration constitutes a viable construct for lightweight CNNs. Moreover, a novel link to the latest ShiftNet is established, inspiring a first-ever shift layer pruning that achieves nearly 50% compression with < 1% drop in accuracy for ShiftResNet.

cs.CV

Decoding the drive-bath interplay: A guideline to enhance superconductivity

Driven-dissipative physics lie at the core of quantum optics. However, the full interplay between a driven quantum many-body system and its environment remains relatively unexplored in the solid state realm. In this work, we inspect this interplay beyond the commonly employed stroboscopic Hamiltonian picture based on the specific example of a driven superconductor. Using the Shirley-Floquet and Keldysh formalisms as well as a generalization of the notion of superconducting fitness to the driven case, we show how a drive which anti-commutes with the superconducting gap operator generically induces an unusual particle-hole structure in the spectral functions from the perspective of the thermal bath. Concomitant with a driving frequency which is near resonant with the intrinsic cutoff frequency of the underlying interaction, this spectral structure can be harnessed to enhance the superconducting transition temperature. Our work paves the way for further studies for driven-dissipative engineering of exotic phases of matter in solid-state systems.

cond-mat.supr-con

Lite it fly: An All-Deformable-Butterfly Network

Most deep neural networks (DNNs) consist fundamentally of convolutional and/or fully connected layers, wherein the linear transform can be cast as the product between a filter matrix and a data matrix obtained by arranging feature tensors into columns. The lately proposed deformable butterfly (DeBut) decomposes the filter matrix into generalized, butterflylike factors, thus achieving network compression orthogonal to the traditional ways of pruning or low-rank decomposition. This work reveals an intimate link between DeBut and a systematic hierarchy of depthwise and pointwise convolutions, which explains the empirically good performance of DeBut layers. By developing an automated DeBut chain generator, we show for the first time the viability of homogenizing a DNN into all DeBut layers, thus achieving an extreme sparsity and compression. Various examples and hardware benchmarks verify the advantages of All-DeBut networks. In particular, we show it is possible to compress a PointNet to < 5% parameters with < 5% accuracy drop, a record not achievable by other compression schemes.

cs.LG

Crystal Facet Effect in Plasmonic Catalysis

In the realm of plasmonic catalytic systems, much attention has been devoted to the plasmon-derived mechanisms, yet the influence of nanoparticles' crystal facets in this type of processes has been sparsely investigated. In this work, we study the plasmon-assisted electrocatalytic CO2 reduction reaction using three different shapes of plasmonic Au nanoparticles - nanocube (NC), rhombic dodecahedron (RD) and octahedron (OC) - with three different exposed facets: {100}, {110} and {111}, respectively. These particles were synthesized with similar sizes and LSPR wavelengths to reveal the role of the facet more than other contributions to the plasmon-assisted reaction. Upon plasmon excitation, Au OCs exhibited nearly a doubling in the Faradaic efficiency of CO (FE(CO)) and a remarkable threefold enhancement in the partial current density of CO (j(CO)) compared to the non-illuminated response, NCs also demonstrated an improved performance under illumination. In contrast, Au RDs showed nearly the same performance in dark or light conditions. Temperature-dependent experiments ruled out heat as the main factor in the enhanced response of Au OCs and NCs. Large-scale atomistic simulations of the nanoparticles' electronic structure and electromagnetic modeling revealed higher hot carrier abundance and electric field enhancement on Au OCs and NCs compared to RDs. Abundant hot carriers on edges facilitate molecular activation, leading to enhanced selectivity and activity. Thus, OCs with the highest edge/facet ratio exhibited the strongest enhancement in FE(CO) and j(CO) upon illumination. This observation is further supported by plasmon-assisted H2 evolution reaction experiments. Our findings highlight the dominance of low coordinated sites over facets in plasmonic catalytic processes, providing valuable insights for designing more efficient catalysts for solar fuels production.

physics.optics

Cluster-based Method for Eavesdropping Identification and Localization in Optical Links

We propose a cluster-based method to detect and locate eavesdropping events in optical line systems characterized by small power losses. Our findings indicate that detecting such subtle losses from eavesdropping can be accomplished solely through optical performance monitoring (OPM) data collected at the receiver. On the other hand, the localization of such events can be effectively achieved by leveraging in-line OPM data.

stat.ML

A Spectral Perspective towards Understanding and Improving Adversarial Robustness

Deep neural networks (DNNs) are incredibly vulnerable to crafted, imperceptible adversarial perturbations. While adversarial training (AT) has proven to be an effective defense approach, the AT mechanism for robustness improvement is not fully understood. This work investigates AT from a spectral perspective, adding new insights to the design of effective defenses. In particular, we show that AT induces the deep model to focus more on the low-frequency region, which retains the shape-biased representations, to gain robustness. Further, we find that the spectrum of a white-box attack is primarily distributed in regions the model focuses on, and the perturbation attacks the spectral bands where the model is vulnerable. Based on this observation, to train a model tolerant to frequency-varying perturbation, we propose a spectral alignment regularization (SAR) such that the spectral output inferred by an attacked adversarial input stays as close as possible to its natural input counterpart. Experiments demonstrate that SAR and its weight averaging (WA) extension could significantly improve the robust accuracy by 1.14% ~ 3.87% relative to the standard AT, across multiple datasets (CIFAR-10, CIFAR-100 and Tiny ImageNet), and various attacks (PGD, C&W and Autoattack), without any extra data.

cs.CV

Frequency Regularization for Improving Adversarial Robustness

Deep neural networks are incredibly vulnerable to crafted, human-imperceptible adversarial perturbations. Although adversarial training (AT) has proven to be an effective defense approach, we find that the AT-trained models heavily rely on the input low-frequency content for judgment, accounting for the low standard accuracy. To close the large gap between the standard and robust accuracies during AT, we investigate the frequency difference between clean and adversarial inputs, and propose a frequency regularization (FR) to align the output difference in the spectral domain. Besides, we find Stochastic Weight Averaging (SWA), by smoothing the kernels over epochs, further improves the robustness. Among various defense schemes, our method achieves the strongest robustness against attacks by PGD-20, C\&W and Autoattack, on a WideResNet trained on CIFAR-10 without any extra data.

cs.CV

PECAN: A Product-Quantized Content Addressable Memory Network

A novel deep neural network (DNN) architecture is proposed wherein the filtering and linear transform are realized solely with product quantization (PQ). This results in a natural implementation via content addressable memory (CAM), which transcends regular DNN layer operations and requires only simple table lookup. Two schemes are developed for the end-to-end PQ prototype training, namely, through angle- and distance-based similarities, which differ in their multiplicative and additive natures with different complexity-accuracy tradeoffs. Even more, the distance-based scheme constitutes a truly multiplier-free DNN solution. Experiments confirm the feasibility of such Product-Quantized Content Addressable Memory Network (PECAN), which has strong implication on hardware-efficient deployments especially for in-memory computing.

cs.LG

Observing dynamical currents in a non-Hermitian momentum lattice

We report on the experimental realization and detection of dynamical currents in a spin-textured lattice in momentum space. Collective tunneling is implemented via cavity-assisted Raman scattering of photons by a spinor Bose-Einstein condensate into an optical cavity. The photon field inducing the tunneling processes is subject to cavity dissipation, resulting in effective directional dynamics in a non-Hermitian setting. We observe that the individual tunneling events are superradiant in nature and locally resolve them in the lattice by performing real-time, frequency-resolved measurements of the leaking cavity field. The results can be extended to a regime exhibiting a cascade of currents and simultaneous coherences between multiple lattice sites, where numerical simulations provide further understanding of the dynamics. Our observations showcase dynamical tunneling in momentum-space lattices and provide prospects to realize dynamical gauge fields in driven-dissipative settings.

cond-mat.quant-gas