SearcharxivSearch

arXiv subjects

Yadong Wang

Publications and source records attributed to Yadong Wang.

At least 19 recordsLinked to original sources

Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning

Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent transitions. However, existing latent reasoning methods usually rely on dense and entangled transitions, making their reasoning trajectories difficult to inspect or intervene on. We introduce LSTR (Latent Sparse Transcoder Reasoning), a framework that turns sparse transcoders from post-hoc diagnostic tools into in-loop, intervenable transition components for latent reasoning. At each latent step, a Latent Transition Transcoder (LTT) combines a linear skip path with a Top-k sparse innovation path, exposing a small set of active sparse features. Under matched compression settings, LSTR offers a mechanistically inspectable alternative to dense latent reasoning. On GSM8K-Aug, ablating only a few top-active sparse features reduces accuracy by up to 16.5%, whereas analogous interventions have much smaller effects in dense latent baselines. These results indicate that the active sparse features are causally involved in the latent transition process, rather than merely post-hoc descriptors. Additional experiments on mathematical benchmarks and StrategyQA suggest that sparse latent transitions can preserve the compression benefits of latent reasoning while making the resulting trajectories more inspectable and intervenable.

cs.AI

Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation

Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic forgetting, yet open-ended factual failures do not necessarily imply that the underlying facts have been erased. In this work, we identify a more specific phenomenon, factual access failure: after domain SFT, models can still recognize or rank the correct answer under constrained evaluation, while failing to produce it in closed-book generation. Through benchmark-level comparisons, same-fact multiple-choice and generation probes, and failure-mode analysis, we show that SFT-induced factual degradation reflects both genuine wrong-answer generations and expression-level failures such as verbosity, formatting mismatch, and exact-match artifacts. To address this problem, we introduce Recall-Anchored Distillation (RAD), a base-anchored self-distillation objective that preserves out-of-distribution generation behavior by aligning the adapted model with the original base model's soft continuation distribution on unlabeled OOD text. RAD requires no gold OOD answers, external judges, or labeled factual data. Across three backbones fine-tuned on MedMCQA, RAD recovers a consistent portion of the lost OOD recall while preserving target-domain adaptation. Compared with replay on the same OOD text, RAD shows that the key preservation signal is the base model's soft distribution rather than additional text exposure alone.

cs.AI

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.

cs.AI

Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion

Large Language Models (LLMs) show remarkable potential for few-shot information extraction (IE), yet their performance is highly sensitive to the choice of in-context examples. Conventional selection strategies often fail to provide informative guidance, as they overlook a key source of model fallibility: confusion stemming not just from semantic content, but also from the generation of well-structured formats required by IE tasks. To address this, we introduce Active Prompting for Information Extraction (APIE), a novel active prompting framework guided by a principle we term introspective confusion. Our method empowers an LLM to assess its own confusion through a dual-component uncertainty metric that uniquely quantifies both Format Uncertainty (difficulty in generating correct syntax) and Content Uncertainty (inconsistency in extracted semantics). By ranking unlabeled data with this comprehensive score, our framework actively selects the most challenging and informative samples to serve as few-shot exemplars. Extensive experiments on four benchmarks show that our approach consistently outperforms strong baselines, yielding significant improvements in both extraction accuracy and robustness. Our work highlights the critical importance of a fine-grained, dual-level view of model uncertainty when it comes to building effective and reliable structured generation systems.

cs.CL

SynLeaF: A Dual-Stage Multimodal Fusion Framework for Synthetic Lethality Prediction Across Pan- and Single-Cancer Contexts

Accurate prediction of synthetic lethality (SL) is important for guiding the development of cancer drugs and therapies. SL prediction faces significant challenges in the effective fusion of heterogeneous multi-source data. Existing multimodal methods often suffer from "modality laziness" due to disparate convergence speeds, which hinders the exploitation of complementary information. This is also one reason why most existing SL prediction models cannot perform well on both pan-cancer and single-cancer SL pair prediction. In this study, we propose SynLeaF, a dual-stage multimodal fusion framework for SL prediction across pan- and single-cancer contexts. The framework employs a VAE-based cross-encoder with a product of experts mechanism to fuse four omics data types (gene expression, mutation, methylation, and CNV), while simultaneously utilizing a relational graph convolutional network to capture structured gene representations from biomedical knowledge graphs. To mitigate modality laziness, SynLeaF introduces a dual-stage training mechanism employing featurelevel knowledge distillation with adaptive uni-modal teacher and ensemble strategies. In extensive experiments across eight specific cancer types and a pancancer dataset, SynLeaF achieves superior performance in 17 out of 19 scenarios. Ablation studies and gradient analyses further validate the critical contributions of the proposed fusion and distillation mechanisms to model robustness and generalization. To facilitate community use, a web server is available at https://synleaf.bioinformatics-lilab.cn.

q-bio.GN

SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment

Despite the intrinsic risk-awareness of Large Language Models (LLMs), current defenses often result in shallow safety alignment, rendering models vulnerable to disguised attacks (e.g., prefilling) while degrading utility. To bridge this gap, we propose SafeThinker, an adaptive framework that dynamically allocates defensive resources via a lightweight gateway classifier. Based on the gateway's risk assessment, inputs are routed through three distinct mechanisms: (i) a Standardized Refusal Mechanism for explicit threats to maximize efficiency; (ii) a Safety-Aware Twin Expert (SATE) module to intercept deceptive attacks masquerading as benign queries; and (iii) a Distribution-Guided Think (DDGT) component that adaptively intervenes during uncertain generation. Experiments show that SafeThinker significantly lowers attack success rates across diverse jailbreak strategies without compromising utility, demonstrating that coordinating intrinsic judgment throughout the generation process effectively balances robustness and practicality.

cs.CR

Magnetic switching of exciton lifetime in CrSBr

Exciton dynamics in layered magnetic semiconductors provide a sensitive probe of the interplay between spin order and light-matter interaction. Here, we study thin CrSBr layers using time-resolved photoluminescence spectroscopy in an external magnetic field, revealing a step-like reduction in the exciton lifetime from 11 to 7 ps, during the magnetization flip from the antiferromagnetic to the ferromagnetic phase. The reduction of the exciton lifetime in the ferromagnetic phase persists below the Néel temperature, as evidenced by its strong magnetic-field dependence that disappears in the paramagnetic phase. Ab initio calculations reveal a one-dimensional nature of free excitons accompanied by a pronounced change in the oscillator strength across the magnetic phase transition predicting a shorter radiative lifetime of free excitons in the antiferromagnetic phase of CrSBr contradicting the experimental observations. This discrepancy is explained by strong localization of excitons at low tempature. We show both experimentally and theoretically that the observed magnetic switching of the exciton lifetime is attributed to a larger exciton localization volume leading to a larger oscillator strength in the ferromagnetic phase. The results show that disorder-induced localization effects play a key role in exciton dynamics in CrSBr.

cond-mat.mes-hall

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical QA tasks, little attention has been paid to KE in multimodal medical scenarios. Unlike text-only settings, medical KE demands integrating updated knowledge with visual reasoning to support safe and interpretable clinical decisions. To address this gap, we propose MultiMedEdit, the first benchmark tailored to evaluating KE in clinical multimodal tasks. Our framework spans both understanding and reasoning task types, defines a three-dimensional metric suite (reliability, generality, and locality), and supports cross-paradigm comparisons across general and domain-specific models. We conduct extensive experiments under single-editing and lifelong-editing settings. Results suggest that current methods struggle with generalization and long-tail reasoning, particularly in complex clinical workflows. We further present an efficiency analysis (e.g., edit latency, memory footprint), revealing practical trade-offs in real-world deployment across KE paradigms. Overall, MultiMedEdit not only reveals the limitations of current approaches but also provides a solid foundation for developing clinically robust knowledge editing techniques in the future.

cs.AI

Photonics in Flatland: Challenges and Opportunities for Nanophotonics with 2D Semiconductors

Two-dimensional (2D) semiconductors are emerging as a versatile platform for nanophotonics, offering unprecedented tunability in optical properties through exciton resonance engineering, van der Waals heterostructuring, and external field control. These materials enable active optical modulation, single-photon emission, quantum photonics, and valleytronic functionalities, paving the way for next-generation optoelectronic and quantum photonic devices. However, key challenges remain in achieving large-area integration, maintaining excitonic coherence, and optimizing amplitude-phase modulation for efficient light manipulation. Advances in fabrication, strain engineering, and computational modelling will be crucial to overcoming these limitations. This perspective highlights recent progress in 2D semiconductor-based nanophotonics, emphasizing opportunities for scalable integration into photonics.

physics.optics

Exciton-polaritons in a monolayer semiconductor coupled to van der Waals dielectric nanoantennas on a metallic mirror

Polaritons in nanophotonic structures have attracted long-standing interest owing to their fundamental importance and potential for applications in nonlinear and quantum optics. Nanoantennas (NAs) made from high refractive index dielectrics offer a suitable platform for polariton physics thanks to the strongly confined optical Mie resonances and low optical losses in contrast to metallic NAs. However, Mie modes are mainly confined within the NA, making inefficient their coupling with excitons in materials deposited externally. Here, we overcome this limitation by using a high-refractive index van der Waals material WS$_2$, which allows straightforward fabrication of NAs on gold. The combination of a 27 nm tall WS$_2$ NA and a gold substrate enables strong modification of the Mie mode distribution and field enhancement inside and in the vicinity of the NA. This allows observation of room-temperature Mie-polaritons (with a Rabi splitting above 80 meV) arising from the strong coupling between Mie modes and the exciton in a monolayer WSe$_2$ placed on WS$_2$/gold NAs. We demonstrate strong nonlinearity of Mie-polaritons, one order of magnitude higher than for excitons in monolayer WSe$_2$ on gold. Our results highlight applicability of van der Waals materials for the realisation of hybrid dielectric-metallic nanophotonics for the study of the strong light-matter interaction.

physics.optics

Infrared Image Deturbulence Restoration Using Degradation Parameter-Assisted Wide & Deep Learning

Infrared images captured under turbulent conditions are degraded by complex geometric distortions and blur. We address infrared deturbulence as an image restoration task, proposing DparNet, a parameter-assisted multi-frame network with a wide & deep architecture. DparNet learns a degradation prior (key parameter matrix) directly from degraded images without external knowledge. Its wide & deep architecture uses these learned parameters to directly modulate restoration, achieving spatially and intensity adaptive results. Evaluated on dedicated infrared deturbulence (49,744 images) and visible image denoising (109,536 images) datasets, DparNet significantly outperforms State-of-the-Art (SOTA) methods in restoration performance and efficiency. Notably, leveraging these parameters improves PSNR by 0.6-1.1 dB with less than 2% increase in model parameters and computational complexity. Our work demonstrates that degraded images hide key degradation information that can be learned and utilized to boost adaptive image restoration.

cs.CV

MSNGO: multi-species protein function annotation based on 3D protein structure and network propagation

Motivation: In recent years, protein function prediction has broken through the bottleneck of sequence features, significantly improving prediction accuracy using high-precision protein structures predicted by AlphaFold2. While single-species protein function prediction methods have achieved remarkable success, multi-species protein function prediction methods are still in the stage of using PPI networks and sequence features. Providing effective cross-species label propagation for species with sparse protein annotations remains a challenging issue. To address this problem, we propose the MSNGO model, which integrates structural features and network propagation methods. Our validation shows that using structural features can significantly improve the accuracy of multi-species protein function prediction. Results: We employ graph representation learning techniques to extract amino acid representations from protein structure contact maps and train a structural model using a graph convolution pooling module to derive protein-level structural features. After incorporating the sequence features from ESM-2, we apply a network propagation algorithm to aggregate information and update node representations within a heterogeneous network. The results demonstrate that MSNGO outperforms previous multi-species protein function prediction methods that rely on sequence features and PPI networks. Availability: https://github.com/blingbell/MSNGO.

cs.LG

Fabrication of ultra-smooth, high-aspect ratio, sub-10 nanometer nanostructures

Deterministic and versatile approaches to sample preparation on nanoscopic scales are important in many fields including photonics, electronics, biology and material science. However, challenges exist in meeting many nanostructuring demands--particularly in emerging optical materials and component architectures. Here, we report a nanofabrication workflow that overcomes long-standing challenges in deterministic and top-down sample preparation procedures. The salient feature is a carbon mask with a low sputter yield that can be readily shaped using high resolution electron beam processing techniques. When combined with focused ion beam processing, the masking technique yields structures with ultra-smooth, near-vertical side walls. We target different material platforms to showcase the broad utility of the technique. As a first test case, we prepared nanometric gaps in evaporated Au. Gap widths of 7 plus/minus 2 nm, aspect ratios of 17, and line edge roughness values of 3sigma = 2.04 nm are achieved. Furthermore, the gap widths represent an order of magnitude improvement on system resolution limits. As a second test case, we designed and fabricated dielectric resonators in the ternary compounds MnPSe3 and NiPS3; a class of van der Waals material resistant to chemical etch approaches. Nanoantenna arrays with incrementally increasing diameter were fabricated in crystalline, exfoliated flakes. The optical response was measured by dark field spectroscopy and is in agreement with simulations. The workflow reported here leverages established techniques in material processing without the need for custom or specialized hardware. It is broadly applicable to functional materials and devices, and extends high speed focused ion beam milling to true sub-10 nm length scales.

physics.optics

Van der Waals Nanoantennas on Gold as Hosts for Hybrid Mie-Plasmonic Resonances

Dielectric nanoresonators have been shown to circumvent the heavy optical losses associated with plasmonic devices, however they suffer from less confined resonances. By constructing a hybrid system of both dielectric and metallic materials, one can retain the low losses of dielectric resonances, whilst gaining additional control over the tuning of the modes with the metal, and achieving stronger mode confinement. In particular, multi-layered van der Waals materials are emerging as promising candidates for integration with metals owing to their weak attractive forces, which enable deposition onto such substrates without the requirement of lattice matching. Here we use layered, high refractive index WS$_2$ exfoliated on gold, to fabricate and optically characterize a hybrid nanoantenna-on-gold system. We experimentally observe a hybridization of Mie resonances, Fabry-Pérot modes, and surface plasmon-polaritons launched from the nanoantennas into the substrate. We achieve experimental quality factors of Mie-plasmonic modes of up to 20 times that of Mie resonances in nanoantennas on silica, and observe signatures of a supercavity mode with a Q factor of 263 $\pm$ 28, resulting from strong mode coupling between a higher-order anapole and Fabry-Pérot-plasmonic mode. We further simulate WS$_2$ nanoantennas on gold with an hBN spacer, resulting in calculated electric field enhancements exceeding 2600, and a Purcell factor of 713. Our results demonstrate dramatic changes in the optical response of dielectric nanophotonic structures placed on gold, opening new possibilities for nanophotonics and sensing with simple-to-fabricate devices.

cond-mat.mes-hall

Probing Electronic States in Monolayer Semiconductors through Static and Transient Third-Harmonic Spectroscopy

Electronic states and their dynamics are of critical importance for electronic and optoelectronic applications. Here, we probe various relevant electronic states in monolayer MoS2, such as multiple excitonic Rydberg states and free-particle energy bands, with a high relative contrast of up to >200 via broadband (from ~1.79 to 3.10 eV) static third-harmonic spectroscopy, which is further supported by theoretical calculations. Moreover, we introduce transient third-harmonic spectroscopy to demonstrate that third-harmonic generation can be all-optically modulated with a modulation depth exceeding ~94% at ~2.18 eV, providing direct evidence of dominant carrier relaxation processes, associated with carrier-exciton and carrier-phonon interactions. Our results indicate that static and transient third-harmonic spectroscopies are not only promising techniques for the characterization of monolayer semiconductors and their heterostructures, but also a potential platform for disruptive photonic and optoelectronic applications, including all-optical modulation and imaging.

physics.optics

Community Question Answering Entity Linking via Leveraging Auxiliary Data

Community Question Answering (CQA) platforms contain plenty of CQA texts (i.e., questions and answers corresponding to the question) where named entities appear ubiquitously. In this paper, we define a new task of CQA entity linking (CQAEL) as linking the textual entity mentions detected from CQA texts with their corresponding entities in a knowledge base. This task can facilitate many downstream applications including expert finding and knowledge base enrichment. Traditional entity linking methods mainly focus on linking entities in news documents, and are suboptimal over this new task of CQAEL since they cannot effectively leverage various informative auxiliary data involved in the CQA platform to aid entity linking, such as parallel answers and two types of meta-data (i.e., topic tags and users). To remedy this crucial issue, we propose a novel transformer-based framework to effectively harness the knowledge delivered by different kinds of auxiliary data to promote the linking performance. We validate the superiority of our framework through extensive experiments over a newly released CQAEL data set against state-of-the-art entity linking methods.

cs.CL

Second-harmonic generation in germanium-on-insulator from visible to telecom wavelengths

The second-order $χ^{2}$ process underpins many important nonlinear optical applications in the field of classical and quantum optics. Generally, the $χ^{2}$ process manifests itself only in a non-centrosymmetric dielectric medium via an anharmonic electron oscillation when driven by an intense optical field. Due to inversion symmetry, group-IV semiconductors like silicon (Si) and germanium (Ge) are traditionally not considered as ideal candidates for second-order nonlinear optics applications. Here, we report the experimental observation of the second-harmonic generation (SHG) in a Ge-on-insulator (GOI) sample under femtosecond optical pumping. Specially, we report the first-time measurement of the SHG signal from a GOI sample in the telecom S-band by pumping at $\sim$$3000$ nm.

physics.optics

Topologically Protected Ferroelectric Domain Wall Memory with Large Readout Current

The discovery and precise manipulation of atomic-size conductive ferroelectric domain defects, such as geometrically confined walls, offer new opportunities for a wide range of prospective electronic devices, and the so-called walltronics is emerging consequently. Here we demonstrate the highly stable and fatigue-resistant nonvolatile ferroelectric memory device based on deterministic creation and erasure of conductive domain wall geometrically confined inside a topological domain structure. By introducing a pair of delicately designed co-axial electrodes onto the epitaxial BiFeO3 film, one can easily create quadrant center topological polar domain structure. More importantly, a reversible switching of such center topological domain structure between the convergent state with highly conductive confined wall and the divergent state with insulating confined wall can be realized, hence resulting in an apparent resistance change with a large On/Off ratio > 104 and a technically preferred readout current (up to 40 nA). Owing to the topological robustness of the center domain structure, the device exhibits the excellent restoration repeatability over 106 cycles and a long retention over 12 days (> 106 s). This work demonstrates a good example for implementing the exotic polar topologies in high-performance nanoscale devices, and would spur more interest in exploring the rich emerging applications of these exotic topological states.

physics.app-ph