SearcharxivSearch

arXiv subjects

Haifeng Lin

Publications and source records attributed to Haifeng Lin.

12 recordsLinked to original sources

Rise From The Ashes: LLM-based Static Analysis for Deep Learning Framework Bugs

Deep learning (DL) frameworks are critical AI infrastructures that often hide bugs with serious security implications. While dynamic approaches such as fuzzing are effective in uncovering these bugs, they require real test execution and incur high computational costs. Static analysis is a natural complement because it can detect bugs without runtime execution, offering fast and scalable testing. Unfortunately, there is still limited work targeting static analysis for DL frameworks due to their multilingual architectures and tensor-related program state. We present Phoenix, the first LLM-based static analysis technique for DL frameworks. Our key insight is that cross-language tensor flows in DL frameworks can be modeled, together with concrete code context, as a structured semantic bridge intermediate representation (SBIR) that LLMs can analyze for potential bugs in tensor semantic propagation. We implement this insight through a multi-agent workflow. A summarization agent first distills bug summaries from historical bug-fix patches and CWE rules. Guided by each summary, an extraction agent identifies bug-relevant repository symbols for code retrieval, and a generation agent synthesizes grounded SBIRs from the retrieved context. Finally, an analysis agent is leveraged to check SBIRs and report potential bugs. Our evaluation shows that Phoenix is a practical complement to dynamic DL framework testing for bug finding. To date, Phoenix has found 31 real new bugs in PyTorch for different heterogeneous hardware backends (Intel CPU, NVIDIA CUDA, and Apple MPS). Among them, 20 submitted bug-fixing patches have been merged into upstream.

cs.SE

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired with audio and systematically detailed textual descriptions. Specifically, every video and its audio is described across three hierarchical layers (factual, abstract, and intent), progressing from surface to depth. It contains over 22 million words, averaging ~225 words per sample. VideoMind's key distinction from existing datasets is its provision of intent expressions, which require contextual integration across the entire video and are not directly observable. These deep-cognitive expressions are generated using a Chain-of-Thought (COT) approach, prompting the mLLM through step-by-step reasoning. Each description includes annotations for subject, place, time, event, action, and intent, supporting downstream recognition tasks. Crucially, we establish a gold-standard benchmark with 3,000 manually validated samples for evaluating deep-cognitive video understanding. We design hybrid-cognitive retrieval experiments, scored by multi-level retrieval metrics, to appropriately assess deep video comprehension. Evaluation results for models (e.g., InternVideo, VAST, UMT-L) are released. VideoMind serves as a powerful benchmark for fine-grained cross-modal alignment and advances fields requiring in-depth video understanding, such as emotion and intent recognition. The data is publicly available on GitHub, HuggingFace, and OpenDataLab, https://github.com/cdx-cindy/VideoMind.

cs.CV

May the Feedback Be with You! Unlocking the Power of Feedback-Driven Deep Learning Framework Fuzzing via LLMs

Deep Learning (DL) frameworks have served as fundamental components in DL systems over the last decade. However, bugs in DL frameworks could lead to catastrophic consequences in critical scenarios. A simple yet effective way to find bugs in DL frameworks is fuzz testing (Fuzzing). Existing approaches focus on test generation, leaving execution results with high semantic value (e.g., coverage information, bug reports, and exception logs) in the wild, which can serve as multiple types of feedback. To fill this gap, we propose FUEL to effectively utilize the feedback information, which comprises two Large Language Models (LLMs): analysis LLM and generation LLM. Specifically, analysis LLM infers analysis summaries from feedback information, while the generation LLM creates tests guided by these summaries. Furthermore, based on multiple feedback guidance, we design two additional components: (i) a feedback-aware simulated annealing algorithm to select operators for test generation, enriching test diversity. (ii) a program self-repair strategy to automatically repair invalid tests, enhancing test validity. We evaluate FUEL on the two most popular DL frameworks, and experiment results show that FUEL can improve line code coverage of PyTorch and TensorFlow by 4.48% and 9.14% over four state-of-the-art baselines. By the time of submission, FUEL has detected 104 previously unknown bugs for PyTorch and TensorFlow, with 93 confirmed as new bugs, 53 already fixed. 14 vulnerabilities have been assigned CVE IDs, among which 7 are rated as high-severity with a CVSS score of "7.5 HIGH". Our artifact is available at https://github.com/NJU-iSE/FUEL

cs.SE

LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering

Aerospace manufacturing demands exceptionally high precision in technical parameters. The remarkable performance of Large Language Models (LLMs), such as GPT-4 and QWen, in Natural Language Processing has sparked industry interest in their application to tasks including process design, material selection, and tool information retrieval. However, LLMs are prone to generating "hallucinations" in specialized domains, producing inaccurate or false information that poses significant risks to the quality of aerospace products and flight safety. This paper introduces a set of evaluation metrics tailored for LLMs in aerospace manufacturing, aiming to assess their accuracy by analyzing their performance in answering questions grounded in professional knowledge. Firstly, key information is extracted through in-depth textual analysis of classic aerospace manufacturing textbooks and guidelines. Subsequently, utilizing LLM generation techniques, we meticulously construct multiple-choice questions with multiple correct answers of varying difficulty. Following this, different LLM models are employed to answer these questions, and their accuracy is recorded. Experimental results demonstrate that the capabilities of LLMs in aerospace professional knowledge are in urgent need of improvement. This study provides a theoretical foundation and practical guidance for the application of LLMs in aerospace manufacturing, addressing a critical gap in the field.

cs.CL

Copper-based disordered plasmonic system with dense nanoisland morphology

Dry synthesis is a highly versatile method for the fabrication of nanoporous metal films, since it enables easy and reproducible deposition of single or multi-layer(s) of nanostructured materials that can find intriguing applications in plasmonics, photochemistry and photocatalysis, to name a few. Here, we extend the use of this methodology to the preparation of copper nanoislands that represent an affordable and versatile example of disordered plasmonic substrate. We perform detailed characterizations of the system using several techniques such as spectroscopic ellipsometry, cathodoluminescence, electron energy loss spectroscopy, ultrafast pump-probe spectroscopy and second-harmonic generation with the aim to investigate the optical properties of these systems in an unprecedented systematic way. Our study represents the starting point for future applications of this new disordered plasmonic system ranging from sensing to photochemistry and photocatalysis.

physics.app-ph

Enhanced Safety in Autonomous Driving: Integrating Latent State Diffusion Model for End-to-End Navigation

With the advancement of autonomous driving, ensuring safety during motion planning and navigation is becoming more and more important. However, most end-to-end planning methods suffer from a lack of safety. This research addresses the safety issue in the control optimization problem of autonomous driving, formulated as Constrained Markov Decision Processes (CMDPs). We propose a novel, model-based approach for policy optimization, utilizing a conditional Value-at-Risk based Soft Actor Critic to manage constraints in complex, high-dimensional state spaces effectively. Our method introduces a worst-case actor to guide safe exploration, ensuring rigorous adherence to safety requirements even in unpredictable scenarios. The policy optimization employs the Augmented Lagrangian method and leverages latent diffusion models to predict and simulate future trajectories. This dual approach not only aids in navigating environments safely but also refines the policy's performance by integrating distribution modeling to account for environmental uncertainties. Empirical evaluations conducted in both simulated and real environment demonstrate that our approach outperforms existing methods in terms of safety, efficiency, and decision-making capabilities.

cs.AI

Structural and optical characterization of NiO polycrystalline thin films fabricated by spray-pyrolysis

Nickel (II) oxide, NiO, a wide band gap Mott insulator characterized by strong Coulomb repulsion between d-electrons and displaying antiferromagnetic order at room temperature, has gained attention in recent years as a very promising candidate for applications in a broad set of areas, including chemistry and metallurgy to spintronics and energy harvesting. Here, we report on the synthesis of polycrystalline NiO fabricated using spray-pyrolysis technique, which is a deposition technique able to produce quite uniform films of pure and crystalline materials without the need of high vacuum or inert atmospheres. We then characterized the composition and structure of our NiO thin films using X-ray diffraction, and atomic force and scanning electron microscopies, respectively. We completed our study by looking at the phononic and magnonic properties of our NiO thin films via Raman spectroscopy, and at the ultrafast electron dynamics by using optical pump probe spectroscopy. We found that our NiO samples display the same phononic and magnonic dispersion expected for single crystal NiO at room temperature, and that electron dynamics in our system is similar to those of previously reported NiO mono- and poli-crystalline systems synthesized with different techniques. These results prove that spray-pyrolysis can be used as affordable and large-scale fabrication technique to synthetize strongly correlated materials for a large set of applications.

cond-mat.mtrl-sci

Dry synthesis of bi-layer nanoporous metal films as plasmonic metamaterial

Nanoporous metals are a class of nanostructured materials finding extensive applications in multiple fields thanks to their unique properties attributed to their high surface area and interconnected nanoscale ligaments. They can be pre-pared following different strategies, but the deposition of an arbitrary pure porous metal is still challenging. Recently, a dry synthesis of nanoporous films based on the plasma treat-ment of metal thin layers deposited by physical vapour deposition has been demonstrated, as a general route to form pure nanoporous films from a large set of metals. An interest-ing aspect related to this approach is the possibility to apply the same methodology to deposit the porous films as a multilayer. In this way, it is possible to explore the properties of different porous metals in close contact. As demonstrated in this paper, interesting plasmonic properties emerge in a nanoporous Au-Ag bi-layer. The versatility of the method coupled with the possibility to include many different metals, provides an opportunity to tailor their optical resonances and to exploit the chemical and mechanical properties of compo-nents, which is of great interest to applications ranging from sensing, to photochemistry and photocatalysis.

physics.app-ph

Generation of ultrashort light pulses carrying orbital angular momentum using a vortex plate retarder-based approach

We use a vortex retarder-based approach to generate few optical cycles light pulses carrying orbital angular momentum (known also as twisted light or optical vortex) from a Yb:KGW oscillator pumping a noncollinear optical parametric amplifier generating sub-10 fs linearly polarized light pulses in the near infrared spectral range (central wavelength 850 nm). We characterize such vortices both spatially and temporally by using astigmatic imaging technique and second harmonic generation-based frequency resolved optical gating, respectively. The generation of optical vortices is analyzed, and its structure reconstructed by estimating the spatio-spectral field and Fourier transforming it into the temporal domain. As a proof of concept, we show that we can also generate sub-20 fs light pulses carrying orbital angular momentum and with arbitrary polarization on the first-order Poincar\'e sphere.

physics.optics

NP-RDMA: Using Commodity RDMA without Pinning Memory

Remote Direct Memory Access (RDMA) has been haunted by the need of pinning down memory regions. Pinning limits the memory utilization because it impedes on-demand paging and swapping. It also increases the initialization latency of large memory applications from seconds to minutes. To remove memory pining, existing approaches often require special hardware which supports page fault, and still have inferior performance. We propose NP-RDMA, which removes memory pinning during memory registration and enables dynamic page fault handling with commodity RDMA NICs. NP-RDMA does not require NICs to support page fault. Instead, by monitoring local memory paging and swapping with MMU-notifier, combining with IOMMU/SMMU-based address mapping, NP-RDMA efficiently detects and handles page fault in the software with near-zero additional latency to non-page-fault RDMA verbs. We implement an LD_PRELOAD library (with a modified kernel module), which is fully compatible with existing RDMA applications. Experiments show that NP-RDMA adds only 0.1{\sim}2 {\mu}s latency under non-page-fault scenarios. Moreover, NP-RDMA adds only 3.5{\sim}5.7 {\mu}s and 60 {\mu}s under minor or major page faults, respectively, which is 500x faster than ODP which uses advanced NICs that support page fault. With non-pinned memory, Spark initialization is 20x faster and the physical memory usage reduces by 86% with only 5.4% slowdown. Enterprise storage can expand to 5x capacity with SSDs while the average latency is only 10% higher. To the best of our knowledge, NP-RDMA is the first efficient and application-transparent software approach to remove memory pinning using commodity RDMA NICs.

cs.NI

Advances in ultrafast plasmonics

In the past twenty years, we have reached a broad understanding of many light-driven phenomena in nanoscale systems. The temporal dynamics of the excited states are instead quite challenging to explore, and, at the same time, crucial to study for understanding the origin of fundamental physical and chemical processes. In this review we examine the current state and prospects of ultrafast phenomena driven by plasmons both from a fundamental and applied point of view. This research area is referred to as ultrafast plasmonics and represents an outstanding playground to tailor and control fast optical and electronic processes at the nanoscale, such as ultrafast optical switching, single photon emission and strong coupling interactions to tailor photochemical reactions. Here, we provide an overview of the field, and describe the methodologies to monitor and control nanoscale phenomena with plasmons at ultrafast timescales in terms of both modeling and experimental characterization. Various directions are showcased, among others recent advances in ultrafast plasmon-driven chemistry and multi-functional plasmonics, in which charge, spin, and lattice degrees of freedom are exploited to provide active control of the optical and electronic properties of nanoscale materials. As the focus shifts to the development of practical devices, such as all-optical transistors, we also emphasize new materials and applications in ultrafast plasmonics and highlight recent development in the relativistic realm. The latter is a promising research field with potential applications in fusion research or particle and light sources providing properties such as attosecond duration.

physics.optics

Prediction of ICD Codes with Clinical BERT Embeddings and Text Augmentation with Label Balancing using MIMIC-III

This paper achieves state of the art results for the ICD code prediction task using the MIMIC-III dataset. This was achieved through the use of Clinical BERT (Alsentzer et al., 2019). embeddings and text augmentation and label balancing to improve F1 scores for both ICD Chapter as well as ICD disease codes. We attribute the improved performance mainly to the use of novel text augmentation to shuffle the order of sentences during training. In comparison to the Top-32 ICD code prediction (Keyang Xu, et. al.) with an F1 score of 0.76, we achieve a final F1 score of 0.75 but on a total of the top 50 ICD codes.

cs.CL