SearcharxivSearch

arXiv subjects

Zhengkai Wang

Publications and source records attributed to Zhengkai Wang.

5 recordsLinked to original sources

SherAgent: Scaling Attack Investigation in the Wild via LLM-Empowered Iterative Query-Filter Backtracking

Provenance-based attack investigation enables viable automation by standardizing data and query logic; however, it is critically hindered in practice by dependency explosions and fragmented causal chains in the wild. Towards designing a robust and automated investigation tool, we collaborated with the SOC of a major Internet corporation serving billions of users. By engaging in real-world incident response, we are able to evaluate and refine their existing LLM-based investigation workflows, which processes tens of thousands of raw alerts daily, leaving thousands for manual triage, to find out the root causes of investigation failures and major challenges in their existing tools. Motivated by these findings, we propose SherAgent, an LLM-empowered automated investigation system. Operating on an iterative ``query-filter'' backtracking paradigm over provenance graphs, SherAgent leverages the semantic reasoning capabilities of LLMs to process unstructured data, such as investigation context and threat intelligence. To overcome fragmented causal chains caused by missing events, the system dynamically calibrates query conditions to broaden the search scope. Concurrently, it performs precision result filtering and strategic nodes selection for subsequent exploration, thereby mitigating dependency explosions. Extensive evaluations in the wild demonstrate that SherAgent improves the end-to-end investigation success rate by 31.1% and 63.7% compared to both legacy enterprise baselines and SOTA approaches, respectively. Furthermore, it operates with remarkable efficiency, incurring under $0.10 in API costs and requiring less than 4 minutes per investigation. Finally, our user study confirms that SherAgent provides accurate and clear insights, significantly reducing the analytical overhead for security experts.

cs.CR

Minos: A Multi-Agent Collaborative Framework for Provenance-Based Backward Tracking

Sophisticated cyber attacks, particularly Advanced Persistent Threats (APTs), require effective post-intrusion forensic analysis. Provenance-based backward tracking reconstructs attack scenarios by tracing causality from security alerts, but existing methods rely on low-level statistical features and rigid traversal strategies, limiting their ability to capture high-level adversarial intent and suffering from dependency explosion. We present Minos, a multi-agent framework that formulates backward tracking as an LLM-driven reasoning process. Minos adopts a two-tiered architecture: for event-level analysis, it combines hierarchical context management, retrieval-augmented reasoning with citation verification, and adversarial deliberation to improve reasoning quality; for graph exploration, it coordinates four specialized agents under a finite state machine (FSM), replacing exhaustive traversal with hypothesis-guided reasoning and count-first query protocols to efficiently prune the search space. Experiments on 14 attack scenarios across five public datasets show that Minos achieves an average recall of 0.92 and precision of 0.64, significantly outperforming state-of-the-art baselines while producing attack subgraphs that are 49% more compact. Moreover, Minos generates interpretable reasoning throughout the tracking process, facilitating forensic auditing and system refinement. These results demonstrate the effectiveness of LLM-driven reasoning for automated provenance-based backward tracking.

cs.CR

Quantifying the effect of graph structure on strong Feller property of SPDEs

This paper investigates how the structure of the underlying graph influences the behavior of stochastic partial differential equations (SPDEs) on finite tree graphs, where each edge is driven by space-time white noise. We first introduce a novel graph-based null decomposition approach to analyzing the strong Feller property of the Markov semigroup generated by SPDEs on tree graphs. By examining the positions of zero entries in eigenfunctions of the graph Laplacian operator, we establish a sharp upper bound on the number of noise-free edges that ensures both the strong Feller property and irreducibility. Interestingly, we find that the addition of noise to any single edge is sufficient for chain graphs, whereas for star graphs, at most one edge can remain noise-free without compromising the system's properties. Furthermore, under a dissipative condition, we prove the existence and exponential ergodicity of a unique invariant measure.

math.PR

From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models

Evolving attacker capabilities demand realistic and continuously updated cyberattack emulation for threat-informed defense and security benchmarking. Towards automated attack emulation, this paper defines modular attack actions and a linking model to organize and chain heterogeneous attack tools into causality-preserving cyberattacks. Building on this foundation, we introduce Aurora: an automated cyberattack emulation system powered by symbolic planning and large language models (LLMs). Aurora crafts actionable, causality-preserving attack chains tailored to Cyber Threat Intelligence (CTI) reports and target environments, and automatically executes these emulations. Using Aurora, we generated an extensive cyberattack emulation dataset from 250 attack reports, 15 times larger than the leading expert-crafted dataset. Our evaluation shows that Aurora significantly outperforms existing methods in creating actionable, diverse, and realistic attack chains. We release the dataset and use it to evaluate three state-of-the-art intrusion detection systems, whose performance differed notably from results on older datasets, highlighting the need for up-to-date, automated attack emulation.

cs.CR

Spatial Distributions of Sunspot Oscillation Modes at Different Temperatures

Three- and five-minute oscillations of sunspots have different spatial distributions in the solar atmospheric layers. The spatial distributions are crucial to reveal the physical origin of sunspot oscillations and to investigate their propagation. In this study, six sunspots observed by Solar Dynamics Observatory/Atmospheric Imaging Assembly were used to obtain the spatial distributions of three- and five-minute oscillations. The fast Fourier transform method is applied to represent the power spectra of oscillation modes. We find that, from the temperature minimum to the lower corona, the powers of the five-minute oscillation exhibit a circle-shape distribution around its umbra, and the shapes gradually expand with temperature increase. However, the circle-shape is disappeared and the powers of the oscillations appear to be very disordered in the higher corona. This indicates that the five-minute oscillation can be suppressed in the high-temperature region. For the three-minute oscillations, from the temperature minimum to the high corona, their powers mostly distribute within an umbra, and part of them locate at the coronal fan loop structures. Moreover, those relative higher powers are mostly concentrated in the position of coronal loop footpoints.

astro-ph.SR