SearcharxivSearch

arXiv subjects

Chengyu Song

Publications and source records attributed to Chengyu Song.

At least 19 recordsLinked to original sources

Microflow: Microarchitectural Causal Observability for Deep Cross-Layer Analysis and Optimization

Existing architectural simulators expose aggregate metrics or raw traces, but fail to reveal complex interactions among microarchitectural events and their relationship to program execution. Consequently, architects observe performance symptoms but cannot systematically attribute them to root causes across abstraction layers. This paper introduces Microflow, an observability framework elevating causality to a first-class analytical object. Microflow transforms execution traces into the Microflow Intermediate Representation (MFIR), explicitly capturing dependencies across software semantics, instructions, microarchitectural events, and hardware resources. By unifying these elements, MFIR enables direct traversal from observed stalls to their underlying causes, paving the way for automated root-cause analysis. Microflow precisely attributes stalls, reveals unobservable phenomena, and enables exact critical-path decomposition through counterfactual analysis. These capabilities allow systematic reasoning about complex hardware-software interactions opaque to existing tools. Making causality queryable, Microflow provides a strong foundation for performance analysis and hardware-software co-design. We demonstrate it on two SPEC CPU 2017 benchmarks, uncovering bottlenecks invisible from aggregate symptoms: hidden misprediction costs in leela and cross-loop-iteration contention in mcf.

cs.AR

Natural Language based Specification and Verification

Recent frontier large language models (LLMs) have shown strong performance in identifying security vulnerabilities in large, mature open-source systems. As LLM-generated code becomes increasingly common, a natural goal is to prevent such models from producing vulnerable implementations in the first place. Formal verification offers a principled route to this objective, but existing verification pipelines typically require specifications written in rigid formal languages. Prior work has explored using LLMs to synthesize such specifications, with limited success. In this paper, we investigate a different approach: using LLMs both to generate specifications and to verify implementations compositionally when the specifications are expressed in natural language. Our preliminary results suggest that this approach is promising.

cs.SE

Topochemically-engineered coexistence of charge and spin orders in intercalated endotaxial heterostructures

Correlated electron systems that host multiple electronic orders offer routes to multifunctional quantum materials, but strong competition between these orders often prevents their coexistence. Here we show that nanoscale, metastable intercalated heterostructures can stabilize a rare combination of long-range magnetism and a commensurate charge density wave (C-CDW) order in a single material. We synthesize a two-dimensional (2D) metastable crystal, T/H-Fe$_x$TaS2, which comprises an endotaxial polytype heterostructure of 1T-TaS$_2$ and H-TaS$_2$ with Fe intercalated in the van der Waals interfaces. In T/H-Fe$_x$TaS2, Fe intercalants provide localized spins that support ferromagnetism, while 1T layers host a robust commensurate charge density wave (C-CDW) that persists to room temperature. In these intercalated heterostructures, Fe content simultaneously tunes ordering of spin and charge degrees of freedom, positioning topochemically-prepared intercalated endotaxial heterostructures as a route to stabilize and control competing quantum phases in 2D materials.

cond-mat.str-el

Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure

Laboratories are prone to severe injuries from minor unsafe actions, yet continuous safety monitoring -- beyond mandatory pre-lab safety training -- is limited by human availability. Vision language models (VLMs) offer promise for autonomous laboratory safety monitoring, but their effectiveness in realistic settings is unclear due to the lack of visual evaluation data, as most safety incidents are documented primarily as unstructured text. To address this gap, we first introduce a structured data generation pipeline that converts textual laboratory scenarios into aligned triples of (image, scene graph, ground truth), using large language models as scene graph architects and image generation models as renderers. Our experiments on the synthetic dataset of 1,207 samples across 362 unique scenarios and seven open- and closed-source models show that VLMs perform effectively given textual scene graph, but degrade substantially in visual-only settings indicating difficulty in extracting structured object relationships directly from pixels. To overcome this, we propose a post-training context-engineering approach, scene-graph-guided alignment, to bridge perceptual gaps in VLMs by translating visual inputs into structured scene graphs better aligned with VLM reasoning, improving hazard detection performance in visual only settings.

cs.CV

Nanocrystal Geometry Governs Phase Transformation Pathways in Palladium Hydride

Pathways and structural dynamics of phase transformations impact performance of materials in energy and information storage technologies. Palladium hydride ($\mathrm{PdH}_x$) nanocrystals are an ideal model system for studying solute-induced phase transformations, where elastic energy from lattice mismatch between $\alpha$-$\mathrm{PdH}_x$ and $\beta$-$\mathrm{PdH}_x$ phases is often considered a key to determining the transformation pathways. $\alpha/\beta$-$\mathrm{PdH}_x$ interfacial elastic energy is affected by the confined geometry of a nanocrystal. However, how nanocrystal geometry influences phase transformation pathways is largely unknown. Using in situ liquid phase transmission electron microscopy, we directly visualize hydrogenation in Pd nanocrystals with two geometries -- a nanocube and a hexagonal nanoplate. Both follow similar sequences of an initially curved nucleus, interface flattening, and reverse-stage nucleation; however, their evolving $\alpha/\beta$-$\mathrm{PdH}_x$ interfaces exhibit geometry-dependent crystallographic alignments. In nanocubes, $\{100\}$-aligned configurations conform to static elastic energy ordering, representing a pathway that maintains a local mechanical equilibrium, whereas nanoplates display both $\{110\}$- and $\{211\}$-aligned interfaces. Theoretical simulations show that geometry determines the accessibility of alternative phase transformation pathways as the system is driven far from equilibrium during hydrogenation. These findings identify geometry as a fundamental parameter for directing phase transformation pathways, offering design principles for accessing atypical configurations and improving properties of intercalation-based devices.

cond-mat.mes-hall

PBFuzz: Agentic Directed Fuzzing for PoV Generation

Proof-of-Vulnerability (PoV) input generation is a critical task in software security and supports downstream applications such as path generation and validation. Generating a PoV input requires solving two sets of constraints: (1) reachability constraints for reaching vulnerable code locations, and (2) triggering constraints for activating the target vulnerability. Existing approaches, including directed greybox fuzzing and LLM-assisted fuzzing, struggle to efficiently satisfy these constraints. This work presents an agentic method that mimics human experts. Human analysts iteratively study code to extract semantic reachability and triggering constraints, form hypotheses about PoV triggering strategies, encode them as test inputs, and refine their understanding using debugging feedback. We automate this process with an agentic directed fuzzing framework called PBFuzz. PBFuzz tackles four challenges in agentic PoV generation: autonomous code reasoning for semantic constraint extraction, custom program-analysis tools for targeted inference, persistent memory to avoid hypothesis drift, and property-based testing for efficient constraint solving while preserving input structure. Experiments on the Magma benchmark show strong results. PBFuzz triggered 57 vulnerabilities, surpassing all baselines, and uniquely triggered 17 vulnerabilities not exposed by existing fuzzers. PBFuzz achieved this within a 30-minute budget per target, while conventional approaches use 24 hours. Median time-to-exposure was 339 seconds for PBFuzz versus 8680 seconds for AFL++ with CmpLog, giving a 25.6x efficiency improvement with an API cost of 1.83 USD per vulnerability. In real-world application, PBFuzz reproduced three FFmpeg 1-day CVEs that had no public PoVs.

cs.CR

Critical Disconnect Between Structural and Electronic Recovery in Amorphous GaAs during Recrystallization

Understanding the evolution of structure and functionality through amorphous to crystalline phase transitions is critical for predicting and designing devices for application in extreme conditions. Here, we consider both aspects of recrystallization of irradiated GaAs. We find that structural evolution occurs in two stages, a low temperature regime characterized by slow, epitaxial front propagation and a high-temperature regime above dominated by rapid growth and formation of dense nanotwin networks. We link aspects of this structural evolution to local ordering, or paracrystallinity, within the amorphous phase. Critically, the electronic recovery of the materials is not commensurate with this structural evolution. The electronic properties of the recrystallized material deviate further from the pristine material than do those of the amorphous phase, highlighting the incongruence between structural and electronic recovery and the contrasting impact of loss of long range order versus localized defects on the functionality of semiconducting materials.

cond-mat.mtrl-sci

Insight-LLM: LLM-enhanced Multi-view Fusion in Insider Threat Detection

Insider threat detection (ITD) requires analyzing sparse, heterogeneous user behavior. Existing ITD methods predominantly rely on single-view modeling, resulting in limited coverage and missed anomalies. While multi-view learning has shown promise in other domains, its direct application to ITD introduces significant challenges: scalability bottlenecks from independently trained sub-models, semantic misalignment across disparate feature spaces, and view imbalance that causes high-signal modalities to overshadow weaker ones. In this work, we present Insight-LLM, the first modular multi-view fusion framework specifically tailored for insider threat detection. Insight-LLM employs frozen, pre-nes, achieving state-of-the-art detection with low latency and parameter overhead.

cs.CR

3D Strain Field Reconstruction by Inversion of Dynamical Scattering

Strain governs not only the mechanical response of materials but also their electronic, optical, and catalytic properties. For this reason, the measurement of the 3D strain field is crucial for a detailed understanding and for further developments of material properties through strain engineering. However, measuring strain variations along the electron beam direction has remained a major challenge for (scanning-) transmission electron microscopy (S/TEM). In this article, we present a method for 3D strain field determination using 4D-STEM. The method is based on the inversion of dynamical diffraction effects, which occur at strain field variations along the beam direction. We test the method against simulated data with a known ground truth and demonstrate its application to an experimental 4D-STEM dataset from an inclined pseudomorphically grown Al$_{0.47}$Ga$_{0.53}$N layer.

cond-mat.mtrl-sci

HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models

Large language models (LLMs) are increasingly being adopted as the cognitive core of embodied agents. However, inherited hallucinations, which stem from failures to ground user instructions in the observed physical environment, can lead to navigation errors, such as searching for a refrigerator that does not exist. In this paper, we present the first systematic study of hallucinations in LLM-based embodied agents performing long-horizon tasks under scene-task inconsistencies. Our goal is to understand to what extent hallucinations occur, what types of inconsistencies trigger them, and how current models respond. To achieve these goals, we construct a hallucination probing set by building on an existing benchmark, capable of inducing hallucination rates up to 40x higher than base prompts. Evaluating 12 models across two simulation environments, we find that while models exhibit reasoning, they fail to resolve scene-task inconsistencies-highlighting fundamental limitations in handling infeasible tasks. We also provide actionable insights on ideal model behavior for each scenario, offering guidance for developing more robust and reliable planning strategies.

cs.LG

Modeling shared micromobility as a label propagation process for detecting the overlapping communities

Shared micro-mobility such as e-scooters has gained significant popularity in many cities. However, existing methods for detecting community structures in mobility networks often overlook potential overlaps between communities. In this study, we conceptualize shared micro-mobility in urban spaces as a process of information exchange, where locations are connected through e-scooters, facilitating the interaction and propagation of community affiliations. As a result, similar locations are assigned the same label. Based on this concept, we developed a Geospatial Interaction Propagation model (GIP) by designing a Speaker-Listener Label Propagation Algorithm (SLPA) that accounts for geographic distance decay, incorporating anomaly detection to ensure the derived community structures reflect meaningful spatial patterns. We applied this model to detect overlapping communities within the e-scooter system in Washington, D.C. The results demonstrate that our algorithm outperforms existing model of overlapping community detection in both efficiency and modularity. However, existing methods for detecting community structures in mobility networks often overlook potential overlaps between communities. In this study, we conceptualize shared micro-mobility in urban spaces as a process of information exchange, where locations are connected through e-scooters, facilitating the interaction and propagation of community affiliations. As a result, similar locations are assigned the same label. Based on this concept, we developed a Geospatial Interaction Propagation model (GIP) by designing a Speaker-Listener Label Propagation Algorithm (SLPA) that accounts for geographic distance decay, incorporating anomaly detection to ensure the derived community structures reflect meaningful spatial patterns.

cs.SI

Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models

Vision-language models (VLMs) have improved significantly in their capabilities, but their complex architecture makes their safety alignment challenging. In this paper, we reveal an uneven distribution of harmful information across the intermediate layers of the image encoder and show that skipping a certain set of layers and exiting early can increase the chance of the VLM generating harmful responses. We call it as "Image enCoder Early-exiT" based vulnerability (ICET). Our experiments across three VLMs: LLaVA-1.5, LLaVA-NeXT, and Llama 3.2, show that performing early exits from the image encoder significantly increases the likelihood of generating harmful outputs. To tackle this, we propose a simple yet effective modification of the Clipped-Proximal Policy Optimization (Clip-PPO) algorithm for performing layer-wise multi-modal RLHF for VLMs. We term this as Layer-Wise PPO (L-PPO). We evaluate our L-PPO algorithm across three multimodal datasets and show that it consistently reduces the harmfulness caused by early exits.

cs.CL

In-situ Study of Understanding the Resistive Switching Mechanisms of Nitride-based Memristor Devices

Interface-type resistive switching (RS) devices with lower operation current and more reliable switching repeatability exhibits great potential in the applications for data storage devices and ultra-low-energy computing. However, the working mechanism of such interface-type RS devices are much less studied compared to that of the filament-type devices, which hinders the design and application of the novel interface-type devices. In this work, we fabricate a metal/TiOx/TiN/Si (001) thin film memristor by using a one-step pulsed laser deposition. In situ transmission electron microscopy (TEM) imaging and current-voltage (I-V) characteristic demonstrate that the device is switched between high resistive state (HRS) and low resistive state (LRS) in a bipolar fashion with sweeping the applied positive and negative voltages. In situ scanning transmission electron microscopy (STEM) experiments with electron energy loss spectroscopy (EELS) reveal that the charged defects (such as oxygen vacancies) can migrate along the intrinsic grain boundaries of TiOx insulating phase under electric field without forming obvious conductive filaments, resulting in the modulation of Schottky barriers at the metal/semiconductor interfaces. The fundamental insights gained from this study presents a novel perspective on RS processes and opens up new technological opportunities for fabricating ultra-low-energy nitride-based memristive devices.

physics.app-ph

Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection

Log-based insider threat detection (ITD) detects malicious user activities by auditing log entries. Recently, large language models (LLMs) with strong common sense knowledge have emerged in the domain of ITD. Nevertheless, diverse activity types and overlong log files pose a significant challenge for LLMs in directly discerning malicious ones within myriads of normal activities. Furthermore, the faithfulness hallucination issue from LLMs aggravates its application difficulty in ITD, as the generated conclusion may not align with user commands and activity context. In response to these challenges, we introduce Audit-LLM, a multi-agent log-based insider threat detection framework comprising three collaborative agents: (i) the Decomposer agent, breaking down the complex ITD task into manageable sub-tasks using Chain-of-Thought (COT) reasoning;(ii) the Tool Builder agent, creating reusable tools for sub-tasks to overcome context length limitations in LLMs; and (iii) the Executor agent, generating the final detection conclusion by invoking constructed tools. To enhance conclusion accuracy, we propose a pair-wise Evidence-based Multi-agent Debate (EMAD) mechanism, where two independent Executors iteratively refine their conclusions through reasoning exchange to reach a consensus. Comprehensive experiments conducted on three publicly available ITD datasets-CERT r4.2, CERT r5.2, and PicoDomain-demonstrate the superiority of our method over existing baselines and show that the proposed EMAD significantly improves the faithfulness of explanations generated by LLMs.

cs.CR

Tailored topotactic chemistry unlocks heterostructures of magnetic intercalation compounds

The construction of thin film heterostructures has been a widely successful archetype for fabricating materials with emergent physical properties. This strategy is of particular importance for the design of multilayer magnetic architectures in which direct interfacial spin--spin interactions between magnetic phases in dissimilar layers lead to emergent and controllable magnetic behavior. However, crystallographic incommensurability and atomic-scale interfacial disorder can severely limit the types of materials amenable to this strategy, as well as the performance of these systems. Here, we demonstrate a method for synthesizing heterostructures comprising magnetic intercalation compounds of transition metal dichalcogenides (TMDs), through directed topotactic reaction of the TMD with a metal oxide. The mechanism of the intercalation reaction enables thermally initiated intercalation of the TMD from lithographically patterned oxide films, giving access to a new family of multi-component magnetic architectures through the combination of deterministic van der Waals assembly and directed intercalation chemistry.

cond-mat.mtrl-sci

Cross-Modal Safety Alignment: Is textual unlearning all you need?

Recent studies reveal that integrating new modalities into Large Language Models (LLMs), such as Vision-Language Models (VLMs), creates a new attack surface that bypasses existing safety training techniques like Supervised Fine-tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF). While further SFT and RLHF-based safety training can be conducted in multi-modal settings, collecting multi-modal training datasets poses a significant challenge. Inspired by the structural design of recent multi-modal models, where, regardless of the combination of input modalities, all inputs are ultimately fused into the language space, we aim to explore whether unlearning solely in the textual domain can be effective for cross-modality safety alignment. Our evaluation across six datasets empirically demonstrates the transferability -- textual unlearning in VLMs significantly reduces the Attack Success Rate (ASR) to less than 8\% and in some cases, even as low as nearly 2\% for both text-based and vision-text-based attacks, alongside preserving the utility. Moreover, our experiments show that unlearning with a multi-modal dataset offers no potential benefits but incurs significantly increased computational demands, possibly up to 6 times higher.

cs.CL

Performance of Superconducting Resonators Suspended on SiN Membranes

Suspending devices on thin SiN membranes can limit their interaction with the bulk substrate and reduce parasitic capacitance to ground. While suspending devices on membranes is used in many fields including radiation detection using superconducting circuits, there has been less investigation into maximum membrane aspect ratios and achievable suspended device quality, metrics important to establish the applicable scope of the technique. Here, we investigate these metrics by fabricating superconducting coplanar waveguide resonators entirely atop thin ($\sim$110 nm) SiN membranes, where the membrane's shortest length to thickness yields an aspect ratio of approximately $7.4 \times 10^3$. We compare these membrane resonators to on-substrate resonators on the same chip, finding similar internal quality factors $\sim$$10^5$ at single photon levels. Furthermore, we confirm that these membranes do not adversely affect resonator thermalization and conduct further materials characterization. By achieving high quality superconducting circuit devices fully suspended on thin SiN membranes, our results help expand the technique's scope to potential uses including incorporating higher aspect ratio membranes for device suspension and creating larger footprint, high impedance, and high quality devices.

quant-ph

An Investigation of Patch Porting Practices of the Linux Kernel Ecosystem

Open-source software is increasingly reused, complicating the process of patching to repair bugs. In the case of Linux, a distinct ecosystem has formed, with Linux mainline serving as the upstream, stable or long-term-support (LTS) systems forked from mainline, and Linux distributions, such as Ubuntu and Android, as downstreams forked from stable or LTS systems for end-user use. Ideally, when a patch is committed in the Linux upstream, it should not introduce new bugs and be ported to all the applicable downstream branches in a timely fashion. However, several concerns have been expressed in prior work about the responsiveness of patch porting in this Linux ecosystem. In this paper, we mine the software repositories to investigate a range of Linux distributions in combination with Linux stable and LTS, and find diverse patch porting strategies and competence levels that help explain the phenomenon. Furthermore, we show concretely using three metrics, i.e., patch delay, patch rate, and bug inheritance ratio, that different porting strategies have different tradeoffs. We find that hinting tags(e.g., Cc stable tags and fixes tags) are significantly important to the prompt patch porting, but it is noteworthy that a substantial portion of patches remain devoid of these indicative tags. Finally, we offer recommendations based on our analysis of the general patch flow, e.g., interactions among various stakeholders in the ecosystem and automatic generation of hinting tags, as well as tailored suggestions for specific porting strategies.

cs.SE