Searcharxiv⌕ Search

arXiv subjects

Yan Meng

Publications and source records attributed to Yan Meng.

At least 37 records · Page 2Linked to original sources

Pressure-induced reentrant superconductivity in a misfit layered compound $\mathrm{(SnS)_{1.15}(TaS_2)}$

Misfit layered compounds are natural van der Waals heterostructures in which electronically active transition-metal dichalcogenide layers are decoupled by incommensurate blocking layers, enabling bulk realization of quasi-two-dimensional quantum states. Here we investigate the superconducting, transport,and structural properties of the misfit compound $\mathrm{(SnS)_{1.15}(TaS_2)}$ under pressures up to 150 GPa. The low-pressure superconducting phase is gradually suppressed and disappears near 14.7 GPa,accompanied by increasing residual resistance. Remarkably, a distinct superconducting phase reemerges above 80 GPa and persists to the highest pressures achieved. This reentrant superconductivity follows a pressure-induced sign reversal of the Hall coefficient near 60 GPa and a nonmonotonic evolution of the normal-state resistance, indicating an electronic reconstruction. No structural phase transition is detected over the entire pressure range. Our results demonstrate a pressure-driven electronic reconstruction leading to reentrant superconductivity in a misfit layered compound, establishing pressure as an effective route to engineer superconductivity and electronic states in natural van der Waals heterostructures.

cond-mat.supr-con↗

An AI Framework for Microanastomosis Motion Assessment

Proficiency in microanastomosis is a fundamental competency across multiple microsurgical disciplines. These procedures demand exceptional precision and refined technical skills, making effective, standardized assessment methods essential. Traditionally, the evaluation of microsurgical techniques has relied heavily on the subjective judgment of expert raters. They are inherently constrained by limitations such as inter-rater variability, lack of standardized evaluation criteria, susceptibility to cognitive bias, and the time-intensive nature of manual review. These shortcomings underscore the urgent need for an objective, reliable, and automated system capable of assessing microsurgical performance with consistency and scalability. To bridge this gap, we propose a novel AI framework for the automated assessment of microanastomosis instrument handling skills. The system integrates four core components: (1) an instrument detection module based on the You Only Look Once (YOLO) architecture; (2) an instrument tracking module developed from Deep Simple Online and Realtime Tracking (DeepSORT); (3) an instrument tip localization module employing shape descriptors; and (4) a supervised classification module trained on expert-labeled data to evaluate instrument handling proficiency. Experimental results demonstrate the effectiveness of the framework, achieving an instrument detection precision of 97%, with a mean Average Precision (mAP) of 96%, measured by Intersection over Union (IoU) thresholds ranging from 50% to 95% (mAP50-95).

cs.CV↗

Excitation Energy Transfer in Nanohybrid System of Organic Molecule and Inorganic Transition Metal Dichalcogenides Nanoflake

Excitation energy transfer (EET) in an organic/inorganic nanohybrid system, composed of a single \textit{para}-sexiphenyl (6P) molecule physisorbed on a finite-sized MoS$_2$ nanoflake, is investigated theoretically. % The electronic structure of the MoS$_2$ nanoflake is described by using an 11-band tight-binding model, in which edge states are passivated with H atoms to restore a well-defined bandgap. % Within a configuration-interaction scheme, excitonic states are constructed and, for computational efficiency, approximated by uncorrelated electron-hole pairs in the relevant high-energy window. % The EET rates are evaluated via Fermi's golden rule, incorporating Coulomb coupling, thermal broadening, and spectral overlap between the molecular excitation and the MoS$_2$ nanoflake's electron-hole pairs. % Our results reveal that energy transfer from the molecule to the nanoflake is the dominant process, and its efficiency depends strongly on the size of the MoS$_2$ nanoflake, as well as the molecule's vertical distance and lateral position relative to the nanoflake.

cond-mat.mes-hall↗

Kinematic-Based Assessment of Surgical Actions in Microanastomosis

Proficiency in microanastomosis is a critical surgical skill in neurosurgery, where the ability to precisely manipulate fine instruments is crucial to successful outcomes. These procedures require sustained attention, coordinated hand movements, and highly refined motor skills, underscoring the need for objective and systematic methods to evaluate and enhance microsurgical training. Conventional assessment approaches typically rely on expert raters supervising the procedures or reviewing surgical videos, which is an inherently subjective process prone to inter-rater variability, inconsistency, and significant time investment. These limitations highlight the necessity for automated and scalable solutions. To address this challenge, we introduce a novel AI-driven framework for automated action segmentation and performance assessment in microanastomosis procedures, designed to operate efficiently on edge computing platforms. The proposed system comprises three main components: (1) an object tip tracking and localization module based on YOLO and DeepSORT; (2) an action segmentation module leveraging self-similarity matrix for action boundary detection and unsupervised clustering; and (3) a supervised classification module designed to evaluate surgical gesture proficiency. Experimental validation on a dataset of 58 expert-rated microanastomosis videos demonstrates the effectiveness of our approach, achieving a frame-level action segmentation accuracy of 92.4% and an overall skill classification accuracy of 85.5% in replicating expert evaluations. These findings demonstrate the potential of the proposed method to provide objective, real-time feedback in microsurgical education, thereby enabling more standardized, data-driven training protocols and advancing competency assessment in high-stakes surgical environments.

cs.CV↗

AI-Driven Evaluation of Surgical Skill via Action Recognition

The development of effective training and evaluation strategies is critical. Conventional methods for assessing surgical proficiency typically rely on expert supervision, either through onsite observation or retrospective analysis of recorded procedures. However, these approaches are inherently subjective, susceptible to inter-rater variability, and require substantial time and effort from expert surgeons. These demands are often impractical in low- and middle-income countries, thereby limiting the scalability and consistency of such methods across training programs. To address these limitations, we propose a novel AI-driven framework for the automated assessment of microanastomosis performance. The system integrates a video transformer architecture based on TimeSformer, improved with hierarchical temporal attention and weighted spatial attention mechanisms, to achieve accurate action recognition within surgical videos. Fine-grained motion features are then extracted using a YOLO-based object detection and tracking method, allowing for detailed analysis of instrument kinematics. Performance is evaluated along five aspects of microanastomosis skill, including overall action execution, motion quality during procedure-critical actions, and general instrument handling. Experimental validation using a dataset of 58 expert-annotated videos demonstrates the effectiveness of the system, achieving 87.7% frame-level accuracy in action segmentation that increased to 93.62% with post-processing, and an average classification accuracy of 76% in replicating expert assessments across all skill aspects. These findings highlight the system's potential to provide objective, consistent, and interpretable feedback, thereby enabling more standardized, data-driven training and evaluation in surgical education.

cs.CV↗

Boost of critical current density near quantum critical points in FeSe-Based superconductors with two superconducting domes

Recent studies have identified two superconducting domes in FeSe-based superconductors. It was discovered that each dome is accompanied by a distinct nematic quantum critical point (QCP): one associated with a pure nematic QCP, and the other with a nematic QCP entangled with antiferromagnetism (AFM). In this study, we delve into the evolution of the critical current density ($J_{\rm{c}}$) with doping in FeSe${_{1-x}}$(Te/S)${_{x}}$ single crystals, focusing on the behavior within the two superconducting domes. Surprisingly, three maxima of $J_{\rm{c}}$ were found in the two superconducting domes, with two sharp peaks in $J_{\rm{c}}$ observed precisely at the endpoints of the nematic phases, at $x$(Te) $\sim$ 0.5 for Te-doped and $x$(S) $\sim$ 0.17 for S-doped FeSe. The mechanisms of vortex pinning and the influence of quantum critical fluctuations have been extensively explored, emphasizing the contribution of quantum critical fluctuations in modulating $J_{\rm{c}}$. Additionally, an increase in $J_{\rm{c}}$ was also noted near FeSe$_{0.1}$Te$_{0.9}$, where its origin has been explored and discussed. This finding provides crucial clues about the existence of an ordered phase endpoint beneath the superconducting dome, offering an initial basis for further investigation into the potential presence of a QCP beneath it.

cond-mat.supr-con↗

Deep Learning-Accelerated Shapley Value for Fair Allocation in Power Systems: The Case of Carbon Emission Responsibility

Allocating costs, benefits, and emissions fairly among power system participant entities represents a persistent challenge. The Shapley value provides an axiomatically fair solution, yet computational barriers have limited its adoption beyond small-scale applications. This paper presents SurroShap, a scalable Shapley value approximation framework combining efficient coalition sampling with deep learning surrogate models that accelerate characteristic function evaluations. Exemplified through carbon emission responsibility allocation in power networks, SurroShap enables Shapley-based fair allocation for power systems with thousands of entities for the first time. We derive theoretical error bounds proving that time-averaged SurroShap allocations converge to be $\varepsilon$-close to exact Shapley values. Experiments on nine systems ranging from 26 to 1,951 entities demonstrate completion within the real-time operational window even at maximum scale, achieving 10^4-10^5 speedups over other sampling-based methods while maintaining tight error bounds. The resulting Shapley-based carbon allocations possess six desirable properties aligning individual interests with decarbonization goals. Year-long simulations on the Texas 2000-bus system validate real-world applicability, with regional analysis revealing how renewable-rich areas offset emission responsibility through exports while load centers bear responsibility for driving system-wide generation.

eess.SY↗

Your Microphone Array Retains Your Identity: A Robust Voice Liveness Detection System for Smart Speakers

Though playing an essential role in smart home systems, smart speakers are vulnerable to voice spoofing attacks. Passive liveness detection, which utilizes only the collected audio rather than the deployed sensors to distinguish between live-human and replayed voices, has drawn increasing attention. However, it faces the challenge of performance degradation under the different environmental factors as well as the strict requirement of the fixed user gestures. In this study, we propose a novel liveness feature, array fingerprint, which utilizes the microphone array inherently adopted by the smart speaker to determine the identity of collected audios. Our theoretical analysis demonstrates that by leveraging the circular layout of microphones, compared with existing schemes, array fingerprint achieves a more robust performance under the environmental change and user's movement. Then, to leverage such a fingerprint, we propose ARRAYID, a lightweight passive detection scheme, and elaborate a series of features working together with array fingerprint. Our evaluation on the dataset containing 32,780 audio samples and 14 spoofing devices shows that ARRAYID achieves an accuracy of 99.84%, which is superior to existing passive liveness detection schemes.

cs.CR↗

Gate Voltage Tunable Second Harmonic Generation in Mono- and Bi-layer Black Phosphene

Black phosphorene (BP) has emerged as a promising platform for tunable nonlinear photonics due to its layer-dependent bandgap, high carrier mobility, and remarkable in-plane anisotropy. This study investigates the second-harmonic generation (SHG) of monolayer and bilayer BP under an external static electric field, with describing the electronic states by a tight-binding model and the dynamics by semiconductor Bloch equations. Our results reveal that BP exhibits large second-order nonlinear optical response along the armchair direction, with significant resonant enhancement when the incident photon energy approaches half of its bandgap. Under an applied electric field of $10^7$ V/m, the effective second-order nonlinear susceptibility of BP can be as large as $10^3$ pm/V, surpassing that of the conventional nonlinear crystal AgGaSe$_2$ by more than an order of magnitude. With respect to the static electric field induced by gate voltage, we discuss the relation between the electric-field-induced second harmonic (EFISH) generation and conventional SHG -- under lower gate voltage, the EFISH approach agrees well with the SHG solutions, whereas the former is no longer applicable under higher gate voltage. Specifically, as the increasing gate voltage, monolayer BP exhibits the bandgap expansion and the corresponding blue-shift in the SHG resonant peak. In contrast, bilayer BP undergoes a semiconductor-to-semimetal transition, forming Dirac cone and generating divergent SHG spectra as photon energy goes to zero. Additionally, the chemical potential allows for precise control over interband and intraband nonlinear responses. This work provides important theoretical foundations for the development of BP-based tunable nonlinear photonic devices and expands the application potential of anisotropic two-dimensional materials in nonlinear optics.

cond-mat.mes-hall↗

Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs

Diversity is an essential metric for evaluating the creativity of outputs generated by language models. Temperature-based sampling is a common strategy to increase diversity. However, for tasks that require high precision, e.g., mathematical reasoning, uncontrolled high temperature sampling, e.g., min-$p$ or top-$p$, degrades reasoning quality. We demonstrate that the loss of accuracy is caused by sampling incorrect continuations in sensitive decoding positions. To address this, in this paper, we propose \textbf{selective sampling}, a method that dynamically switches between greedy and high-temperature sampling based on a sampling risk metric. This risk metric estimates the likelihood of output errors when applying high-temperature sampling on the current token position. To predict sampling risk, we train a lightweight classifier on a small subset of verifiable problems. The trained classifier can be integrated with the base language model with minimal latency overhead. Experiments on mathematical reasoning tasks demonstrate that selective sampling enhances the quality-diversity trade-off, even in high-temperature settings.

cs.LG↗

Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Large language models (LLMs) excel at mathematical reasoning and logical problem-solving. The current popular training paradigms primarily use supervised fine-tuning (SFT) and reinforcement learning (RL) to enhance the models' reasoning abilities. However, when using SFT or RL alone, there are respective challenges: SFT may suffer from overfitting, while RL is prone to mode collapse. The state-of-the-art methods have proposed hybrid training schemes. However, static switching faces challenges such as poor generalization across different tasks and high dependence on data quality. In response to these challenges, inspired by the curriculum learning-quiz mechanism in human reasoning cultivation, We propose SASR, a step-wise adaptive hybrid training framework that theoretically unifies SFT and RL and dynamically balances the two throughout optimization. SASR uses SFT for initial warm-up to establish basic reasoning skills, and then uses an adaptive dynamic adjustment algorithm based on gradient norm and divergence relative to the original distribution to seamlessly integrate SFT with the online RL method GRPO. By monitoring the training status of LLMs and adjusting the training process in sequence, SASR ensures a smooth transition between training schemes, maintaining core reasoning abilities while exploring different paths. Experimental results demonstrate that SASR outperforms SFT, RL, and static hybrid training methods.

cs.LG↗

Depth Gives a False Sense of Privacy: LLM Internal States Inversion

Large Language Models (LLMs) are increasingly integrated into daily routines, yet they raise significant privacy and safety concerns. Recent research proposes collaborative inference, which outsources the early-layer inference to ensure data locality, and introduces model safety auditing based on inner neuron patterns. Both techniques expose the LLM's Internal States (ISs), which are traditionally considered irreversible to inputs due to optimization challenges and the highly abstract representations in deep layers. In this work, we challenge this assumption by proposing four inversion attacks that significantly improve the semantic similarity and token matching rate of inverted inputs. Specifically, we first develop two white-box optimization-based attacks tailored for low-depth and high-depth ISs. These attacks avoid local minima convergence, a limitation observed in prior work, through a two-phase inversion process. Then, we extend our optimization attack under more practical black-box weight access by leveraging the transferability between the source and the derived LLMs. Additionally, we introduce a generation-based attack that treats inversion as a translation task, employing an inversion model to reconstruct inputs. Extensive evaluation of short and long prompts from medical consulting and coding assistance datasets and 6 LLMs validates the effectiveness of our inversion attacks. Notably, a 4,112-token long medical consulting prompt can be nearly perfectly inverted with 86.88 F1 token matching from the middle layer of Llama-3 model. Finally, we evaluate four practical defenses that we found cannot perfectly prevent ISs inversion and draw conclusions for future mitigation design.

cs.CR↗

A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM Method

In the field of multi-sensor fusion for simultaneous localization and mapping (SLAM), monocular cameras and IMUs are widely used to build simple and effective visual-inertial systems. However, limited research has explored the integration of motor-encoder devices to enhance SLAM performance. By incorporating such devices, it is possible to significantly improve active capability and field of view (FOV) with minimal additional cost and structural complexity. This paper proposes a novel visual-inertial-encoder tightly coupled odometry (VIEO) based on a ViDAR (Video Detection and Ranging) device. A ViDAR calibration method is introduced to ensure accurate initialization for VIEO. In addition, a platform motion decoupled active SLAM method based on deep reinforcement learning (DRL) is proposed. Experimental data demonstrate that the proposed ViDAR and the VIEO algorithm significantly increase cross-frame co-visibility relationships compared to its corresponding visual-inertial odometry (VIO) algorithm, improving state estimation accuracy. Additionally, the DRL-based active SLAM algorithm, with the ability to decouple from platform motion, can increase the diversity weight of the feature points and further enhance the VIEO algorithm's performance. The proposed methodology sheds fresh insights into both the updated platform design and decoupled approach of active SLAM systems in complex environments.

cs.RO↗

SPPSFormer: High-quality Superpoint-based Transformer for Roof Plane Instance Segmentation from Point Clouds

Transformers have been seldom employed in point cloud roof plane instance segmentation, which is the focus of this study, and existing superpoint Transformers suffer from limited performance due to the use of low-quality superpoints. To address this challenge, we establish two criteria that high-quality superpoints for Transformers should satisfy and introduce a corresponding two-stage superpoint generation process. The superpoints generated by our method not only have accurate boundaries, but also exhibit consistent geometric sizes and shapes, both of which greatly benefit the feature learning of superpoint Transformers. To compensate for the limitations of deep learning features when the training set size is limited, we incorporate multidimensional handcrafted features into the model. Additionally, we design a decoder that combines a Kolmogorov-Arnold Network with a Transformer module to improve instance prediction and mask extraction. Finally, our network's predictions are refined using traditional algorithm-based postprocessing. For evaluation, we annotated a real-world dataset and corrected annotation errors in the existing RoofN3D dataset. Experimental results show that our method achieves state-of-the-art performance on our dataset, as well as both the original and reannotated RoofN3D datasets. Moreover, our model is not sensitive to plane boundary annotations during training, significantly reducing the annotation burden. Through comprehensive experiments, we also identified key factors influencing roof plane segmentation performance: in addition to roof types, variations in point cloud density, density uniformity, and 3D point precision have a considerable impact. These findings underscore the importance of incorporating data augmentation strategies that account for point cloud quality to enhance model robustness under diverse and challenging conditions.

cs.CV↗

Three-dimensional topological disclination in acoustic crystals

Topological disclinations, crystallographic defects that break rotation lattice symmetry, have attracted great interest and exhibited wide applications in cavities, waveguides, and lasers. However, topological disclinations have thus far been predominantly restricted to two-dimensional (2D) systems owing to the substantial challenges in constructing such defects in three-dimensional (3D) systems and characterizing their topological features. Here we report the theoretical proposal and experimental demonstration of a 3D topological disclination that exhibits fractional (1/2) charge and zero-dimensional (0D) topological bound states, realized by cutting-and-gluing a 3D acoustic topological crystalline insulator. Using acoustic pump-probe measurements, we directly observe 0D topological disclination states at the disclination core, consistent with the tight-binding model and full-wave simulation results. Our results extend the research frontier of topological disclinations and open a new paradigm for exploring the interplay between momentum-space band topology and the real-space defect topology in 3D and higher dimensions.

physics.app-ph↗

Identification of Stochastic Gravitational Wave Backgrounds from Cosmic String Using Machine Learning

Cosmic strings play a crucial role in enhancing our understanding of the fundamental structure and evolution of the universe, unifying our knowledge of cosmology, and potentially unveiling new physical laws and phenomena. The advent and operation of space-based detectors provide an important opportunity for detecting stochastic gravitational wave backgrounds (SGWB) generated by cosmic strings. However, the intricate nature of SGWB poses a formidable challenge in distinguishing its signal from the complex noise by some traditional methods. Therefore, we attempt to identify SGWB based on machine learning. Our findings show that the joint detection of LISA and Taiji significantly outperforms individual detectors, and even in the presence of numerous low signal-to-noise ratio(SNR) signals, the identification accuracy remains exceptionally high with 95%. Although our discussion is based solely on simulated data, the relevant methods can provide data-driven analytical capabilities for future observations of SGWB.

gr-qc↗

How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation

The massive amounts of web-mined parallel data contain large amounts of noise. Semantic misalignment, as the primary source of the noise, poses a challenge for training machine translation systems. In this paper, we first introduce a process for simulating misalignment controlled by semantic similarity, which closely resembles misaligned sentences in real-world web-crawled corpora. Under our simulated misalignment noise settings, we quantitatively analyze its impact on machine translation and demonstrate the limited effectiveness of widely used pre-filters for noise detection. This underscores the necessity of more fine-grained ways to handle hard-to-detect misalignment noise. With an observation of the increasing reliability of the model's self-knowledge for distinguishing misaligned and clean data at the token level, we propose self-correction, an approach that gradually increases trust in the model's self-knowledge to correct the training supervision. Comprehensive experiments show that our method significantly improves translation performance both in the presence of simulated misalignment noise and when applied to real-world, noisy web-mined datasets, across a range of translation tasks.

cs.CL↗

Model Inversion in Split Learning for Personalized LLMs: New Insights from Information Bottleneck Theory

Personalized Large Language Models (LLMs) have become increasingly prevalent, showcasing the impressive capabilities of models like GPT-4. This trend has also catalyzed extensive research on deploying LLMs on mobile devices. Feasible approaches for such edge-cloud deployment include using split learning. However, previous research has largely overlooked the privacy leakage associated with intermediate representations transmitted from devices to servers. This work is the first to identify model inversion attacks in the split learning framework for LLMs, emphasizing the necessity of secure defense. For the first time, we introduce mutual information entropy to understand the information propagation of Transformer-based LLMs and assess privacy attack performance for LLM blocks. To address the issue of representations being sparser and containing less information than embeddings, we propose a two-stage attack system in which the first part projects representations into the embedding space, and the second part uses a generative model to recover text from these embeddings. This design breaks down the complexity and achieves attack scores of 38%-75% in various scenarios, with an over 60% improvement over the SOTA. This work comprehensively highlights the potential privacy risks during the deployment of personalized LLMs on the edge side.

cs.LG↗