SearcharxivSearch

arXiv subjects

Fanyu Meng

Publications and source records attributed to Fanyu Meng.

36 records · Page 2Linked to original sources

Nodeless superconductivity in 4H$_{b}$-TaS$_{2}$ with broken time reversal symmetry

The transition metal dichalcogenide 4H$_{b}$-TaS$_{2}$ exhibits characteristics of topological edge modes and two-component superconductivity with time-reversal symmetry breaking (TRSB). The nature of the superconducting order parameter is a crucial issue that requires experimental investigation. Here, we report measurements of the magnetic penetration depth using a tunnel-diode-oscillator based technique, as well as the specific heat. Both the specific heat and the change in magnetic penetration depth ($\Delta$$\lambda$(T)) display an exponentially-activated temperature dependence, providing evidence for nodeless superconductivity in 4H$_{b}$-TaS$_{2}$. Moreover, the deduced superfluid density can be well described by a two-gap $s$-wave model, and such multigap superconductivity is consistent with there being multiple bands crossing the Fermi energy. These results constrain the possible pairing symmetries of 4H$_{b}$-TaS$_{2}$.

cond-mat.supr-con

Kill Two Birds with One Stone! Trajectory enabled Unified Online Detection of Adversarial Examples and Backdoor Attacks

The proposed UniGuard is the first unified online detection framework capable of simultaneously addressing adversarial examples and backdoor attacks. UniGuard builds upon two key insights: first, both AE and backdoor attacks have to compromise the inference phase, making it possible to tackle them simultaneously during run-time via online detection. Second, an adversarial input, whether a perturbed sample in AE attacks or a trigger-carrying sample in backdoor attacks, exhibits distinctive trajectory signatures from a benign sample as it propagates through the layers of a DL model in forward inference. The propagation trajectory of the adversarial sample must deviate from that of its benign counterpart; otherwise, the adversarial objective cannot be fulfilled. Detecting these trajectory signatures is inherently challenging due to their subtlety; UniGuard overcomes this by treating the propagation trajectory as a time-series signal, leveraging LSTM and spectrum transformation to amplify differences between adversarial and benign trajectories that are subtle in the time domain. UniGuard exceptional efficiency and effectiveness have been extensively validated across various modalities (image, text, and audio) and tasks (classification and regression), ranging from diverse model architectures against a wide range of AE attacks and backdoor attacks, including challenging partial backdoors and dynamic triggers. When compared to SOTA methods, including ContraNet (NDSS 22) specific for AE detection and TED (IEEE SP 24) specific for backdoor detection, UniGuard consistently demonstrates superior performance, even when matched against each method's strengths in addressing their respective threats-each SOTA fails to parts of attack strategies while UniGuard succeeds for all.

cs.CR

$PD^3F$: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models

Large Language Models (LLMs), due to substantial computational requirements, are vulnerable to resource consumption attacks, which can severely degrade server performance or even cause crashes, as demonstrated by denial-of-service (DoS) attacks designed for LLMs. However, existing works lack mitigation strategies against such threats, resulting in unresolved security risks for real-world LLM deployments. To this end, we propose the Pluggable and Dynamic DoS-Defense Framework ($PD^3F$), which employs a two-stage approach to defend against resource consumption attacks from both the input and output sides. On the input side, we propose the Resource Index to guide Dynamic Request Polling Scheduling, thereby reducing resource usage induced by malicious attacks under high-concurrency scenarios. On the output side, we introduce the Adaptive End-Based Suppression mechanism, which terminates excessive malicious generation early. Experiments across six models demonstrate that $PD^3F$ significantly mitigates resource consumption attacks, improving users' access capacity by up to 500% during adversarial load. $PD^3F$ represents a step toward the resilient and resource-aware deployment of LLMs against resource consumption attacks.

cs.CR

Implet: A Post-hoc Subsequence Explainer for Time Series Models

Explainability in time series models is crucial for fostering trust, facilitating debugging, and ensuring interpretability in real-world applications. In this work, we introduce Implet, a novel post-hoc explainer that generates accurate and concise subsequence-level explanations for time series models. Our approach identifies critical temporal segments that significantly contribute to the model's predictions, providing enhanced interpretability beyond traditional feature-attribution methods. Based on it, we propose a cohort-based (group-level) explanation framework designed to further improve the conciseness and interpretability of our explanations. We evaluate Implet on several standard time-series classification benchmarks, demonstrating its effectiveness in improving interpretability. The code is available at https://github.com/LbzSteven/implet

cs.LG

A Comprehensive Survey on Long Context Language Modeling

Efficient processing of long contexts has been a persistent pursuit in Natural Language Processing. With the growing number of long documents, dialogues, and other textual data, it is important to develop Long Context Language Models (LCLMs) that can process and analyze extensive inputs in an effective and efficient way. In this paper, we present a comprehensive survey on recent advances in long-context modeling for large language models. Our survey is structured around three key aspects: how to obtain effective and efficient LCLMs, how to train and deploy LCLMs efficiently, and how to evaluate and analyze LCLMs comprehensively. For the first aspect, we discuss data strategies, architectural designs, and workflow approaches oriented with long context processing. For the second aspect, we provide a detailed examination of the infrastructure required for LCLM training and inference. For the third aspect, we present evaluation paradigms for long-context comprehension and long-form generation, as well as behavioral analysis and mechanism interpretability of LCLMs. Beyond these three key aspects, we thoroughly explore the diverse application scenarios where existing LCLMs have been deployed and outline promising future development directions. This survey provides an up-to-date review of the literature on long-context LLMs, which we wish to serve as a valuable resource for both researchers and engineers. An associated GitHub repository collecting the latest papers and repos is available at: \href{https://github.com/LCLM-Horizon/A-Comprehensive-Survey-For-Long-Context-Language-Modeling}{\color[RGB]{175,36,67}{LCLM-Horizon}}.

cs.CL

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not taken into account the LLM's capability of identifying and handling unsafe information in detail. To address these issues, we propose a fine-grained benchmark SafeDialBench for evaluating the safety of LLMs across various jailbreak attacks in multi-turn dialogues. Specifically, we design a two-tier hierarchical safety taxonomy that considers 6 safety dimensions and generates more than 4000 multi-turn dialogues in both Chinese and English under 22 dialogue scenarios. We employ 7 jailbreak attack strategies, such as reference attack and purpose reverse, to enhance the dataset quality for dialogue generation. Notably, we construct an innovative assessment framework of LLMs, measuring capabilities in detecting, and handling unsafe information and maintaining consistency when facing jailbreak attacks. Experimental results across 17 LLMs reveal that Yi-34B-Chat and GLM4-9B-Chat demonstrate superior safety performance, while Llama3.1-8B-Instruct and o3-mini exhibit safety vulnerabilities.

cs.CL

Evidence for multiband gapless superconductivity in the topological superconductor candidate 4Hb-TaS2

We present the ultralow-temperature thermal conductivity measurements on single crystals of transition-metal dichalcogenide material 4Hb-TaS$_{2}$, which has recently been proposed as a topological superconductor candidate. In zero field, a small residual linear term $\kappa_{0}/T$ is observed, indicating the existence of a residual density of states in the superconducting state. The slow field dependence of $\kappa_{0}/T$ at low fields rules out the presence of nodes in the superconducting gap, and the S-shaped field dependence across the full field range suggests multiple superconducting gaps in 4Hb-TaS$_{2}$. Our results provide evidence for multiband gapless superconductivity in 4Hb-TaS$_{2}$, and the residual density of states come from certain gapless Fermi surfaces.

cond-mat.supr-con

CohEx: A Generalized Framework for Cohort Explanation

eXplainable Artificial Intelligence (XAI) has garnered significant attention for enhancing transparency and trust in machine learning models. However, the scopes of most existing explanation techniques focus either on offering a holistic view of the explainee model (global explanation) or on individual instances (local explanation), while the middle ground, i.e., cohort-based explanation, is less explored. Cohort explanations offer insights into the explainee's behavior on a specific group or cohort of instances, enabling a deeper understanding of model decisions within a defined context. In this paper, we discuss the unique challenges and opportunities associated with measuring cohort explanations, define their desired properties, and create a generalized framework for generating cohort explanations based on supervised clustering.

cs.LG

Interpreting Inflammation Prediction Model via Tag-based Cohort Explanation

Machine learning is revolutionizing nutrition science by enabling systems to learn from data and make intelligent decisions. However, the complexity of these models often leads to challenges in understanding their decision-making processes, necessitating the development of explainability techniques to foster trust and increase model transparency. An under-explored type of explanation is cohort explanation, which provides explanations to groups of instances with similar characteristics. Unlike traditional methods that focus on individual explanations or global model behavior, cohort explainability bridges the gap by providing unique insights at an intermediate granularity. We propose a novel framework for identifying cohorts within a dataset based on local feature importance scores, aiming to generate concise descriptions of the clusters via tags. We evaluate our framework on a food-based inflammation prediction model and demonstrated that the framework can generate reliable explanations that match domain knowledge.

cs.LG

Correlated electrons of the flat band in charge density wave state of 4Hb-TaSexS2-x

Many intriguing quantum states of matter, such as unconventional superconductivity, magnetic phases and fractional quantum Hall physics, emergent from the spatially-correlated localized electrons in the flat band of solid materials. By using scanning tunneling microscopy and spectroscopy (STM/STS), we report the real-space investigation of correlated electrons in the flat band of superlattice 4Hb-TaSexS2-x. In contrast with the pristine 4Hb-TaS2, the selenium (Se) substitutions significantly affect the interfacial transfer of correlated electrons between the CDW states of 1T- and 1H-TaS2 layers, and contribute a real-space fractional electron-filling configurations with the distributed electron-filled and -void SoD clusters of 1T-layer. The site-specific STS spectra directly reveal their respective prominent spectra weight above EF and symmetric Mott-like spectra. In addition, the spatial distributions of these electron-filled SoDs in the 1T-layer of 4Hb-TaSe0.7S1.3 demonstrate different local short-range patterning, clearly indicating the complex neighboring interactions among the localized electrons in the flat band of 1T-layer. Our results not only provide an in-depth insight of correlated electrons in the flat CDW band, and provide a simple platform to manipulate the electron-correlation-related quantum states.

cond-mat.mtrl-sci

Possible spin-polarized Cooper pairing in high temperature FeSe superconductor

Superconductivity and long-range ferromagnetism hardly coexist in a uniform manner. The counter-example has been observed, in uranium-based superconductors for instance, with a coexisting temperature limited to about 1 K. Here, we report the coexistence of high temperature superconductivity and itinerant ferromagnetism in lithium intercalated FeSe flakes. In superconducting samples with transition temperature around 40 K, we observe the anomalous Hall effect with a hysteresis loop in transverse resistivity and a butterfly-like pattern of magneto-resistance. Intriguingly, such ferromagnetism persists down to a temperature at which the zero-field resistance fully vanishes. Furthermore, the superconductivity is enhanced under an in-plane magnetic field, suggestive of the participation of spin-polarized Cooper pairs. The surprising finding underscores a uniform coexistence of the two antagonistic phenomena on a record-high energy scale.

cond-mat.supr-con

Extreme orbital ab-plane upper critical fields far beyond Pauli limit in 4Hb-Ta(S, Se)2 bulk crystals

Transition metal disulfides 4Hb-Ta(S, Se)2 with natural heterostructure of 1T- and 1H-Ta(S, Se)2 layers have became the focus of correlated materials their unique combinations of Mott physics and possible topological superconductivity. In this work, we study the upper critical fields mu0Hc2 of 4Hb-TaS2 and 4Hb-TaS1.99Se0.01 single crystals systematically. Transport measurements up to 35 T show that both of ab-plane and c-axis upper critical fields (mu0Hc2,ab and mu0Hc2,c) for 4Hb-TaS2 and 4Hb-TaS1.99Se0.01 exhibit a linear temperature dependent behavior down to 0.3 K, suggesting the three-dimensional superconductivity with dominant orbital depairing mechanism in bulk 4Hb-Ta(S, Se)2. However, the zero-temperature mu0Hc2,ab(0) for both crystals are far beyond the Pauli paramagnetic limit mu0HP. It could be explained by the effects of spin-momentum locking in 1H-Ta(S, Se)2 layers with local inversion symmetry broken and the relatively weak intersublattice interaction between 1H layers due to the existence of 1T layers.

cond-mat.supr-con

Highly Efficient Room-Temperature Nonvolatile Magnetic Switching by Current in Fe3GaTe2 Thin Flakes

Effectively tuning magnetic state by using current is essential for novel spintronic devices. Magnetic van der Waals (vdW) materials have shown superior properties for the applications of magnetic information storage based on the efficient spin torque effect. However, for most of known vdW ferromagnets, the ferromagnetic transition temperatures lower than room temperature strongly impede their applications and the room-temperature vdW spintronic device with low energy consumption is still a long-sought goal. Here, we realize the highly efficient room-temperature nonvolatile magnetic switching by current in a single-material device based on vdW ferromagnet Fe3GaTe2. Moreover, the switching current density and power dissipation are about 300 and 60000 times smaller than conventional spin-orbit-torque devices of magnet/heavy-metal heterostructures. These findings make an important progress on the applications of magnetic vdW materials in the fields of spintronics and magnetic information storage.

cond-mat.mtrl-sci

Nearly-room-temperature ferromagnetism and tunable anomalous Hall effect in atomically thin Fe4CoGeTe2

Itinerant ferromagnetism at room temperature is a key ingredient for spin transport and manipulation. Here, we report the realization of nearly-room-temperature itinerant ferromagnetism in Co doped Fe5GeTe2 thin flakes. The ferromagnetic transition temperature TC (~ 323 K - 337 K) is almost unchanged when thickness is down to 12 nm and is still about 284 K at 2 nm (bilayer thickness). Theoretical calculations further indicate that the ferromagnetism persists in monolayer Fe4CoGeTe2. In addition to the robust ferromagnetism down to the ultrathin limit, Fe4CoGeTe2 exhibits an unusual temperature- and thickness-dependent intrinsic anomalous Hall effect. We propose that it could be ascribed to the dependence of band structure on thickness that changes the Berry curvature near the Fermi energy level subtly. The nearly-room-temperature ferromagnetism and tunable anomalous Hall effect in atomically thin Fe4CoGeTe2 provide opportunities to understand the exotic transport properties of two-dimensional van der Waals magnetic materials and explore their potential applications in spintronics.

cond-mat.mtrl-sci

Causal Explanation for Reinforcement Learning: Quantifying State and Temporal Importance

Explainability plays an increasingly important role in machine learning. Furthermore, humans view the world through a causal lens and thus prefer causal explanations over associational ones. Therefore, in this paper, we develop a causal explanation mechanism that quantifies the causal importance of states on actions and such importance over time. We also demonstrate the advantages of our mechanism over state-of-the-art associational methods in terms of RL policy explanation through a series of simulation studies, including crop irrigation, Blackjack, collision avoidance, and lunar lander.

cs.AI

TODSum: Task-Oriented Dialogue Summarization with State Tracking

Previous dialogue summarization datasets mainly focus on open-domain chitchat dialogues, while summarization datasets for the broadly used task-oriented dialogue haven't been explored yet. Automatically summarizing such task-oriented dialogues can help a business collect and review needs to improve the service. Besides, previous datasets pay more attention to generate good summaries with higher ROUGE scores, but they hardly understand the structured information of dialogues and ignore the factuality of summaries. In this paper, we introduce a large-scale public Task-Oriented Dialogue Summarization dataset, TODSum, which aims to summarize the key points of the agent completing certain tasks with the user. Compared to existing work, TODSum suffers from severe scattered information issues and requires strict factual consistency, which makes it hard to directly apply recent dialogue summarization models. Therefore, we introduce additional dialogue state knowledge for TODSum to enhance the faithfulness of generated summaries. We hope a better understanding of conversational content helps summarization models generate concise and coherent summaries. Meanwhile, we establish a comprehensive benchmark for TODSum and propose a state-aware structured dialogue summarization model to integrate dialogue state information and dialogue history. Exhaustive experiments and qualitative analysis prove the effectiveness of dialogue structure guidance. Finally, we discuss the current issues of TODSum and potential development directions for future work.

cs.CL

A Hierarchical Multi-Objective Programming Approach to Planning Locations for Macro and Micro Fire Stations

Fire stations are among the most crucial emergency facilities in urban emergency control system in terms of their quick response to fires and other emergencies. Location plannings for fire stations have a significant influence on their effectiveness and capability of emergency responses trading off with the cost of constructions. To obtain efficient and practical siting plans for fire stations, various major requirements including effectiveness maximization, distance constraint and workload limitation are required to be considered in location models, especially for multi-level fire facility systems with macro and micro fire stations having specific aims and construction requirements. This paper proposes a novel hierarchical optimization approach taking all the major requirements for location planning into consideration and bonds functional connections between different levels of fire stations at the same time. A single-objective and a multi-objective optimization model are established to solve the location siting problems for macro and micro fire stations simultaneously. Genetic algorithm with elitist reservation and Pareto-based multi-objective evolutionary algorithm are adopted to solve the problem with NP-hard nature. The proposed hierarchical location model is further performed in a case study of Futian District in Shenzhen, and the siting result justifies the effectiveness and practicality of our novel approach.

math.OC

Open and Programmable 5G Network-in-a-Box: Technology Demonstration and Evaluation Results

The fifth-generation (5G) mobile/cellular technology is a game changer for industrial systems. Private 5G deployments are promising to address the challenges faced by industrial networks. Programmability and open-source are two key aspects which bring unprecedented flexibility and customizability to private 5G networks. Recent regulatory initiatives are removing barriers for industrial stakeholders to deploy their own local 5G networks with dedicated equipment. To this end, this demonstration showcases an open and programmable 5G network-in-a-box solution for private deployments. The network-in-a-box provides an integrated solution, based on open-source software stack and general-purpose hardware, for operation in 5G non-standalone (NSA) as well as 4G long-term evolution (LTE) modes. The demonstration also shows the capability of operation in different sub-6 GHz frequency bands, some of which are specifically available for private networks. Performance results, in terms of end-to-end latency and data rates, with a commercial off-the-shelf (COTS) 5G device are shown as well.

cs.NI