SearcharxivSearch

arXiv subjects

Ning Wu

Publications and source records attributed to Ning Wu.

At least 19 recordsLinked to original sources

Magic-free coexisting photonic and phononic moir\'e flat bands

Moir\'e flat bands enhance localization and interactions through suppressed group velocity, but existing approaches largely target a single physical field because distinct excitations generally require different, finely tuned magic configurations. Here we introduce a flat-band mechanism based on strong diffractive hybridization among moir\'e-folded bands. Period-mismatched modulations open distinct coupling channels whose hybridization renormalizes the band dispersion. An effective Hamiltonian shows that increasing the diffractive coupling progressively suppresses the group velocity, driving the system toward a flat-band regime without field-specific magic configurations. This coupling-induced mechanism enables band flattening across distinct physical excitations. We demonstrate this mechanism in a single-layer moir\'e optomechanical crystal, where photonic and phononic flat bands are simultaneously realized, and their localized modes and optomechanical interaction are experimentally observed. Beyond photonic and phononic systems, this mechanism may extend to other wave and quasiparticle platforms, providing a general route to co-localizing and coupling distinct physical fields in moir\'e systems.

physics.optics

Analytical diagonalization of the open-boundary bosonic Kitaev chain: An asymmetric plane-wave ansatz approach

The bosonic Kitaev chain under open boundary conditions has attracted recent attention due to its realization in driven-dissipative systems and its intriguing non-Hermitian boundary physics. The model is known to be solvable via local squeezing transformations in the position-momentum representation. In this paper, we present an alternative, purely algebraic solution that relies entirely on the standard bosonic Bogoliubov transformation. For an $N$-site chain, we propose an asymmetric plane-wave ansatz with unequal left- and right-moving momenta to analytically solve the associated $2N\times 2N$ non-Hermitian ``associated matrix". The left eigenvalue problem yields $N$ distinct eigenvalues, each of which is twofold degenerate. By carefully resolving these degeneracies using the bosonic commutation relations, we construct the $N$ physical Bogoliubov quasiparticle operators. The construction reveals non-uniquenesses that in special cases exactly mirror the freedom in the local squeezing transformations of the original approach. The diagonal form of the Hamiltonian is obtained explicitly and is shown to be equivalent to the original Hamiltonian. The proposed asymmetric plane-wave ansatz and degeneracy-resolution technique are not limited to the present model and can be generalized to other bosonic pairing systems, including those with inhomogeneous pairing or hopping terms.

quant-ph

Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation

Distilling demonstration effects into hidden-space interventions offers a lightweight alternative to full finetuning. However, existing multimodal variants are mostly evaluated on short-form tasks, where outputs end after a few tokens. Extending these methods to long-form generation exposes a fundamental yet underexamined limitation: token-level distillation implicitly treats all output tokens as equally informative, but long-form outputs are dominated by high-frequency template and grammatical tokens, while the tokens that actually determine output quality are sparsely distributed. In medical report generation (MRG), two such decisive tokens stand out: pathology-related tokens that determine diagnostic content, and the end-of-sequence (EOS) event that determines termination. Both receive insufficient supervision under uniform cross-entropy, and autoregressive decoding further compounds the problem by drifting away from teacher-forced trajectories. We propose DIVE, a frozen-backbone distillation framework that addresses long-form report generation through two complementary mechanisms matched to these failures. Decisive-token supervision restores supervision balance by upweighting the cross-entropy contribution of pathology-related tokens and the EOS event, ensuring that content fidelity and termination are learned during training rather than imposed at decoding time. State-conditioned dynamic steering replaces fixed open-loop residuals with hidden-state-dependent adapters, allowing the injected signal to adapt as decoding drifts. Experiments on MIMIC-CXR and CheXpert Plus with two medical VLM backbones show that DIVE consistently ranks among the strongest methods across lexical and clinical-proxy metrics. Our method achieves the best BLEU-4, ROUGE-L, and RadGraph F1 in all dataset--backbone settings, while remaining competitive on coarse label-level CheXbert F1.

cs.CL

Exact momentum-space analysis of small spin-1/2 $J_1$-$J_2$ rings

This paper considers an $N$-site spin-1/2 $J_1$-$J_2$ ring with $N=6$ and $8$. With the help of a set of exact few-magnon Bloch states, we obtain the block-diagonalized Hamiltonian consisting of block matrices of at most four dimensions. Partial of the eigenstates are analytically solved. For the six-site anisotropic ring, we reveal a subset of eigenstates that are simultaneous eigenstates of the Hamiltonian and the total angular momentum operator, even though the latter is not conserved. For both the six- and eight-site isotropic rings, we achieve momentum-space manifestations of several important states, including the famous Majumdar-Ghosh (MG) ground states and the Hamada-Kane-Nakagawa-Natsume (HKNN) ground state. The equivalence of these states with their real-space counterparts is explicitly shown for $N=6$. The structure of the HKNN ground state for small rings suggests that for any even number $N$ this state might behave like a ``bound state" with $N/2$ successive down spins binding together.

cond-mat.str-el

Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning

Large language models (LLMs) have made impressive strides in mathematical reasoning, often fine-tuned using rejection sampling that retains only correct reasoning trajectories. While effective, this paradigm treats supervision as a binary filter that systematically excludes teacher-generated errors, leaving a gap in how reasoning failures are modeled during training. In this paper, we propose TrajFusion, a fine-tuning strategy that reframes rejection sampling as a structured supervision construction process. Specifically, TrajFusion forms fused trajectories that explicitly model trial-and-error reasoning by interleaving selected incorrect trajectories with reflection prompts and correct trajectories. The length of each fused sample is adaptively controlled based on the frequency and diversity of teacher errors, providing richer supervision for challenging problems while safely reducing to vanilla rejection sampling fine-tuning (RFT) when error signals are uninformative. TrajFusion requires no changes to the architecture or training objective. Extensive experiments across multiple math benchmarks demonstrate that TrajFusion consistently outperforms RFT, particularly on challenging and long-form reasoning problems.

cs.CL

Arbitrary-order exceptional points in a nanomechanical cavity

Higher-order exceptional points (EPs) govern non-Hermitian system dynamics through their enriched and sharpened spectral topology, yet the intrinsic topological fragility hinders robust experimental realization. Here, we present a scalable architecture that implements arbitrary-order EPs via a recurrent network comprising a single nanomechanical resonator and unlimited virtual resonators. We experimentally realize mechanical EPs up to the seventh order and confirm this architecture's scalability. Moreover, we reveal that the fundamental noise component and the measured signal share the same system coupling channel and thus undergo identical root-response amplification near EPs of arbitrary order, consistent with our signal-to-noise ratio measurements. Our work establishes a general platform for exploring higher-order EP-based phenomena while clarifying the fundamental boundary of non-Hermitian sensitivity enhancement across diverse physical systems.

physics.optics

Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme

Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have demonstrated effectiveness, their reliance on high-performance graphical computing limits deployment on resource-constrained devices such as personal computers or home audiovisual systems. To address this limitation, this study proposes a novel audio-only ensemble learning framework capable of classifying movie scenes into three emotional categories: Good, Neutral, and Bad. The model integrates ten support vector machines and six neural networks within a stacking ensemble architecture to enhance classification performance. A tailored data preprocessing pipeline, including feature extraction, outlier handling, and feature engineering, is designed to optimize emotional information from audio inputs. Experiments on a simulated dataset achieve 67% accuracy, while a real-world dataset collected from 15 diverse films yields an impressive 86% accuracy. These results underscore the potential of audio-based, lightweight emotion recognition methods for broader consumer-level applications, offering both computational efficiency and robust classification capabilities.

cs.SD

Exact solution of the two-magnon problem in the $k=-\pi/2$ sector of a finite-size anisotropic spin-1/2 frustrated ferromagnetic chain

The two-magnon problem in the $k=-\pi/2$ sector of a \emph{finite-size} spin-1/2 chain with ferromagnetic nearest-neighbor (NN) interaction $(J_1>0)$ and antiferromagnetic next-nearest-neighbor (NNN) interaction $(J_2<0)$ and anisotropy parameters $\Delta_1$ and $\Delta_2$ is solved exactly by combining a set of exact two-magnon Bloch states and a plane-wave ansatz. Two types of two-magnon bound states (BSs), i.e., the NN and NNN exchange BSs, are revealed. We establish a phase diagram in the $J_1/(|J_2|\Delta_2)$-$\Delta_1$ plane where regions supporting different types of BSs are analytically identified. It is found that no BSs exist (the two types of BSs coexist) when both $\Delta_1$ and $\Delta_2$ are small (large) enough. Our results for the isotropic case $\Delta_1=\Delta_2=1$ are consistent with an early work [Ono I, Mikado S and Oguchi T 1971 \emph{J. Phys. Soc. Japan} \textbf{30} 358].

cond-mat.str-el

AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing

Unmanned aerial vehicles (UAVs) are increasingly employed in diverse applications such as land surveying, material transport, and environmental monitoring. Following missions like data collection or inspection, UAVs must land safely at docking stations for storage or recharging, which is an essential requirement for ensuring operational continuity. However, accurate landing remains challenging due to factors like GPS signal interference. To address this issue, we propose a deviation warning system for UAV landings, powered by a novel vision-based model called AeroLite-MDNet. This model integrates a multiscale fusion module for robust cross-scale object detection and incorporates a segmentation branch for efficient orientation estimation. We introduce a new evaluation metric, Average Warning Delay (AWD), to quantify the system's sensitivity to landing deviations. Furthermore, we contribute a new dataset, UAVLandData, which captures real-world landing deviation scenarios to support training and evaluation. Experimental results show that our system achieves an AWD of 0.7 seconds with a deviation detection accuracy of 98.6\%, demonstrating its effectiveness in enhancing UAV landing reliability. Code will be available at https://github.com/ITTTTTI/Maskyolo.git

cs.RO

BPO: Revisiting Preference Modeling in Direct Preference Optimization

Direct Preference Optimization (DPO) have emerged as a popular method for aligning Large Language Models (LLMs) with human preferences. While DPO effectively preserves the relative ordering between chosen and rejected responses through pairwise ranking losses, it often neglects absolute reward magnitudes. This oversight can decrease the likelihood of chosen responses and increase the risk of generating out-of-distribution responses, leading to poor performance. We term this issue Degraded Chosen Responses (DCR).To address this issue, we propose Balanced Preference Optimization (BPO), a novel framework that dynamically balances the optimization of chosen and rejected responses through two key components: balanced reward margin and gap adaptor. Unlike previous methods, BPO can fundamentally resolve DPO's DCR issue, without introducing additional constraints to the loss function. Experimental results on multiple mathematical reasoning tasks show that BPO significantly outperforms DPO, improving accuracy by +10.1% with Llama-3.1-8B-Instruct (18.8% to 28.9%) and +11.7% with Qwen2.5-Math-7B (35.0% to 46.7%). It also surpasses DPO variants by +3.6% over IPO (43.1%), +5.0% over SLiC (41.7%), and +3.1% over Cal-DPO (43.6%) on the same model. Remarkably, our algorithm requires only a single line of code modification, making it simple to implement and fully compatible with existing DPO-based frameworks.

cs.CL

FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Recent advancements in Large Language Models (LLMs) have demonstrated that Process Reward Models (PRMs) play a crucial role in enhancing model performance. However, training PRMs typically requires step-level labels, either manually annotated or automatically generated, which can be costly and difficult to obtain at scale. To address this challenge, we introduce FreePRM, a weakly supervised framework for training PRMs without access to ground-truth step-level labels. FreePRM first generates pseudo step-level labels based on the correctness of final outcome, and then employs Buffer Probability to eliminate impact of noise inherent in pseudo labeling. Experimental results show that FreePRM achieves an average F1 score of 53.0% on ProcessBench, outperforming fully supervised PRM trained on Math-Shepherd by +24.1%. Compared to other open-source PRMs, FreePRM outperforms upon RLHFlow-PRM-Mistral-8B (28.4%) by +24.6%, EurusPRM (31.3%) by +21.7%, and Skywork-PRM-7B (42.1%) by +10.9%. This work introduces a new paradigm in PRM training, significantly reducing reliance on costly step-level annotations while maintaining strong performance.

cs.CL

Collective charging of an organic quantum battery

We study the collective charging of a quantum battery (QB) consisting of a one-dimensional molecular aggregate and a coupled single-mode cavity, to which we refer as an ``organic quantum battery" since the battery part is an organic material. The organic QB can be viewed as an extension of the so-called Dicke QB [D. Ferraro \emph{et al}., Phys. Rev. Lett. \textbf{120}, 117702 (2018)] by including finite exciton hopping and exciton-exciton interaction within the battery. We consider two types of normalizations of the exciton-cavity coupling when the size of the aggregate $N$ is increased: (I) The cavity length also increases to keep the density of monomers constant, (II) The cavity length does not change. Our main findings are that: (i) For fixed $N$ and exciton-cavity coupling, there exist optimal exciton-exciton interactions at which the maximum stored energy density and the maximum charging power density reach their respective maxima that both increase with increasing exciton-cavity coupling. The existence of such maxima for weak exciton-cavity coupling is argued to be due to the non-monotonic behavior of the one-exciton to two-exciton transition probability in the framework of second-order time-dependent perturbation theory. (ii) Under normalization I, no quantum advantage is observed in the scaling of the two quantities with varying $N$. Under normalization II, it is found that both the maximum stored energy density and the maximum charging power density exhibit quantum advantages compared with the Dicke QB.

quant-ph

Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data

Robustness has become a critical attribute for the deployment of RAG systems in real-world applications. Existing research focuses on robustness to explicit noise (e.g., document semantics) but overlooks implicit noise (spurious features). Moreover, previous studies on spurious features in LLMs are limited to specific types (e.g., formats) and narrow scenarios (e.g., ICL). In this work, we identify and study spurious features in the RAG paradigm, a robustness issue caused by the sensitivity of LLMs to semantic-agnostic features. We then propose a novel framework, SURE, to empirically quantify the robustness of RALMs against spurious features. Beyond providing a comprehensive taxonomy and metrics for evaluation, the framework's data synthesis pipeline facilitates training-based strategies to improve robustness. Further analysis suggests that spurious features are a widespread and challenging problem in the field of RAG. Our code is available at https://github.com/maybenotime/RAG-SpuriousFeatures .

cs.CL

MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads

Large Language Models (LLMs) frequently show distracted attention due to irrelevant information in the input, which severely impairs their long-context capabilities. Inspired by recent studies on the effectiveness of retrieval heads in long-context factutality, we aim at addressing this distraction issue through improving such retrieval heads directly. We propose Multi-Document Attention Focusing (MuDAF), a novel method that explicitly optimizes the attention distribution at the head level through contrastive learning. According to the experimental results, MuDAF can significantly improve the long-context question answering performance of LLMs, especially in multi-document question answering. Extensive evaluations on retrieval scores and attention visualizations show that MuDAF possesses great potential in making attention heads more focused on relevant information and reducing attention distractions.

cs.CL

Zero-Shot NAS via the Suppression of Local Entropy Decrease

Architecture performance evaluation is the most time-consuming part of neural architecture search (NAS). Zero-Shot NAS accelerates the evaluation by utilizing zero-cost proxies instead of training. Though effective, existing zero-cost proxies require invoking backpropagations or running networks on input data, making it difficult to further accelerate the computation of proxies. To alleviate this issue, architecture topologies are used to evaluate the performance of networks in this study. We prove that particular architectural topologies decrease the local entropy of feature maps, which degrades specific features to a bias, thereby reducing network performance. Based on this proof, architectural topologies are utilized to quantify the suppression of local entropy decrease (SED) as a data-free and running-free proxy. Experimental results show that SED outperforms most state-of-the-art proxies in terms of architecture selection on five benchmarks, with computation time reduced by three orders of magnitude. We further compare the SED-based NAS with state-of-the-art proxies. SED-based NAS selects the architecture with higher accuracy and fewer parameters in only one second. The theoretical analyses of local entropy and experimental results demonstrate that the suppression of local entropy decrease facilitates selecting optimal architectures in Zero-Shot NAS.

cs.LG

ControlMath: Controllable Data Generation Promotes Math Generalist Models

Utilizing large language models (LLMs) for data augmentation has yielded encouraging results in mathematical reasoning. However, these approaches face constraints in problem diversity, potentially restricting them to in-domain/distribution data generation. To this end, we propose ControlMath, an iterative method involving an equation-generator module and two LLM-based agents. The module creates diverse equations, which the Problem-Crafter agent then transforms into math word problems. The Reverse-Agent filters and selects high-quality data, adhering to the "less is more" principle, achieving better results with fewer data points. This approach enables the generation of diverse math problems, not limited to specific domains or distributions. As a result, we collect ControlMathQA, which involves 190k math word problems. Extensive results prove that combining our dataset with in-domain datasets like GSM8K can help improve the model's mathematical ability to generalize, leading to improved performances both within and beyond specific domains.

cs.LG

Evolution of two-magnon bound states in a higher-spin ferromagnetic chain with single-ion anisotropy: A complete solution

Few-magnon bound states in quantum spin chains have been long studied and attracted much recent attentions. For a higher-spin ferromagnetic XXZ chain with single-ion anisotropy, several features regarding the evolution of the low-lying two-magnon bound states with varying wave number were observed in the literature. However, most of these observations are only qualitatively understood due to the lack of analytical tools. By combining a set of exact two-magnon Bloch states and a plane-wave ansatz, we achieve a complete solution of the two-magnon problem in such a system. We identify parameter regions that support different types of two-magnon bound states, with the boundaries defined by algebraic equations. We discover for the first time a narrow region in which two single-ion bound states coexist. We show that the phase diagrams for distinct wave numbers are similar to each other, which enables us to map the evolution of the bound states to the rectilinear movement of a representative point for given parameters in a rescaled phase diagram. This dynamic picture provides quantitative interpretations of the observed features.

cond-mat.str-el

Relaxing towards generalized one-body Boltzmann states

Isolated quantum systems follow the reversible unitary evolution; if we focus on the dynamics of local states and observables, they exhibit the irreversible relaxation behaviors. Here we study the local relaxation process in an isolated chain consisting of \emph{N} three level systems. Though the entropy of the full many body state keeps a constant, it turns out the total correlation of this system approximately exhibits a monotonically increasing behavior. More importantly, a variation analysis shows that, the total correlation entropy would achieve its theoretical maximum when each site stays in a generalized one-body Boltzmann state, which is not solely determined by the energy but also depends on the spin value of each onsite level. It turns out such a theoretical correlation maximum is highly coincident with the result obtained from the exact time dependent evolution. In this sense, the total correlation entropy well serves as an indicator for the dynamical irreversibility of the nonequilibrium relaxation in this isolated system.

quant-ph