Searcharxiv⌕ Search

arXiv subjects

Zhuoyun Qian

Publications and source records attributed to Zhuoyun Qian.

4 recordsLinked to original sources

Behavior-Centric Malware Classification with Fine-Grained Malicious Logic Localization

Effective malware analysis requires understanding not only whether a program is malicious, but also which behaviors it exhibits and where those behaviors originate in the code. Existing machine-learning-based malware detectors largely operate as black boxes, providing limited insight into the malicious logic responsible for their decisions. This paper addresses malicious behavior localization and classification at the basic-block level. We propose a behavior-centric analysis framework that decomposes malware samples into behaviors and systematically links these behaviors to their originating code regions. Using context-sensitive backward slicing from security-relevant system API calls, we reconstruct control- and data-dependency chains and represent each behavior as a structured graph of related basic blocks. A Transformer-based model captures instruction-level semantics, while a Graph Neural Network models structural dependencies within behavior graphs. The resulting representations are fused to enable accurate and interpretable classification, with attention-based attribution identifying code regions responsible for malicious behaviors. We evaluate our approach using standard classification metrics and a behavior coverage metric that measures the detection of manually labeled malicious behaviors. Our results demonstrate that the proposed framework achieves accurate malware classification while providing fine-grained, behavior-aware localization of malicious logic.

cs.CR↗

Semantic Validation of Packer Identification Tools: Characterization, Repair, and Downstream Impact

Packer identification tools are a critical foundation of malware analysis, directly affecting unpacking, behavioral analysis, malware classification, and threat attribution. However, their semantic correctness is rarely validated. In practice, a tool may return a plausible packer label that is nevertheless semantically wrong, leading to failed unpacking and unreliable downstream analysis. This paper presents a semantic validation framework for testing and repairing packer identification tools. Our key idea is to use unpackers as executable semantic contracts. If a tool predicts a packer family, the corresponding unpacker should recover analyzable program content. This enables automatic test oracles without requiring manually labeled ground truth. Building on this idea, we develop a systematic pipeline for detecting, localizing, and repairing semantic faults in existing packer identification tools. We then conduct the first large-scale empirical study of semantic bugs in eleven open-source packer identification tools and six proprietary VirusTotal tools. Our results reveal that semantic bugs are widespread and recurring, largely due to incomplete signatures and unstable heuristic logic. After repair, packer identification coverage improves by up to 58.6%, and downstream malware classification performance improves by more than 13.6% on average. These findings show that semantic validation of packer identification tools is essential for building trustworthy malware analysis pipelines.

cs.CR↗

Feature-Oriented IoT Malware Analysis: Extraction, Classification, and Future Directions

As IoT devices continue to proliferate, their reliability is increasingly constrained by security concerns. In response, researchers have developed diverse malware analysis techniques to detect and classify IoT malware. These techniques typically rely on extracting features at different levels from IoT applications, giving rise to a wide range of feature extraction methods. However, current approaches still face significant challenges when applied in practice. This survey provides a comprehensive review of feature extraction techniques for IoT malware analysis from multiple perspectives. We first examine static and dynamic feature extraction methods, followed by hybrid approaches. We then explore feature representation strategies based on graph learning. Finally, we compare the strengths and limitations of existing techniques, highlight open challenges, and outline promising directions for future research.

cs.CR↗

Software Vulnerability Analysis Across Programming Language and Program Representation Landscapes: A Survey

Modern software systems are developed in diverse programming languages and often harbor critical vulnerabilities that attackers can exploit to compromise security. These vulnerabilities have been actively targeted in real-world attacks, causing substantial harm to users and cyberinfrastructure. Since many of these flaws originate from the code itself, a variety of techniques have been proposed to detect and mitigate them prior to software deployment. However, a comprehensive comparative study that spans different programming languages, program representations, bug types, and analysis techniques is still lacking. As a result, the relationships among programming languages, abstraction levels, vulnerability types, and detection approaches remain fragmented, and the limitations and research gaps across the landscape are not clearly understood. This article aims to bridge that gap by systematically examining widely used programming languages, levels of program representation, categories of vulnerabilities, and mainstream detection techniques. The survey provides a detailed understanding of current practices in vulnerability discovery, highlighting their strengths, limitations, and distinguishing characteristics. Furthermore, it identifies persistent challenges and outlines promising directions for future research in the field of software security.

cs.CR↗