SearcharxivSearch

arXiv subjects

Nathan Clarke

Publications and source records attributed to Nathan Clarke.

11 recordsLinked to original sources

Deep Recurrent Hidden Markov Learning Framework for Multi-Stage Advanced Persistent Threat Prediction

Advanced Persistent Threats (APTs) represent hidden, multi\-stage cyberattacks whose long term persistence and adaptive behavior challenge conventional intrusion detection systems (IDS). Although recent advances in machine learning and probabilistic modeling have improved APT detection performance, most existing approaches remain reactive and alert\-centric, providing limited capability for stage-aware prediction and principled inference under uncertainty, particularly when observations are sparse or incomplete. This paper proposes E\-HiDNet, a unified hybrid deep probabilistic learning framework that integrates convolutional and recurrent neural networks with a Hidden Markov Model (HMM) to allow accurate prediction of the progression of the APT campaign. The deep learning component extracts hierarchical spatio\-temporal representations from correlated alert sequences, while the HMM models latent attack stages and their stochastic transitions, allowing principled inference under uncertainty and partial observability. A modified Viterbi algorithm is introduced to handle incomplete observations, ensuring robust decoding under uncertainty. The framework is evaluated using a synthetically generated yet structurally realistic APT dataset (S\-DAPT\-2026). Simulation results show that E\-HiDNet achieves up to 98.8\-100\% accuracy in stage prediction and significantly outperforms standalone HMMs when four or more observations are available, even under reduced training data scenarios. These findings highlight that combining deep semantic feature learning with probabilistic state\-space modeling enhances predictive APT stage performance and situational awareness for proactive APT defense.

cs.CR

S-DAPT-2026: A Stage-Aware Synthetic Dataset for Advanced Persistent Threat Detection

The detection of advanced persistent threats (APTs) remains a crucial challenge due to their stealthy, multistage nature and the limited availability of realistic, labeled datasets for systematic evaluation. Synthetic dataset generation has emerged as a practical approach for modeling APT campaigns; however, existing methods often rely on computationally expensive alert correlation mechanisms that limit scalability. Motivated by these limitations, this paper presents a near realistic synthetic APT dataset and an efficient alert correlation framework. The proposed approach introduces a machine learning based correlation module that employs K Nearest Neighbors (KNN) clustering with a cosine similarity metric to group semantically related alerts within a temporal context. The dataset emulates multistage APT campaigns across campus and organizational network environments and captures a diverse set of fourteen distinct alert types, exceeding the coverage of commonly used synthetic APT datasets. In addition, explicit APT campaign states and alert to stage mappings are defined to enable flexible integration of new alert types and support stage aware analysis. A comprehensive statistical characterization of the dataset is provided to facilitate reproducibility and support APT stage predictions.

cs.CR

Optimizing Password Cracking for Digital Investigations

Efficient password cracking is a critical aspect of digital forensics, enabling investigators to decrypt protected content during criminal investigations. Traditional password cracking methods, including brute-force, dictionary and rule-based attacks face challenges in balancing efficiency with increasing computational complexity. This study explores rule based optimisation strategies to enhance the effectiveness of password cracking while minimising resource consumption. By analysing publicly available password datasets, we propose an optimised rule set that reduces computational iterations by approximately 40%, significantly improving the speed of password recovery. Additionally, the impact of national password recommendations were examined, specifically, the UK National Cyber Security Centre's three word password guideline on password security and forensic recovery. Through user generated password surveys, we evaluate the crackability of three word passwords using dictionaries of varying common word proportions. Results indicate that while three word passwords provide improved memorability and usability, they remain vulnerable when common word combinations are used, with up to 77.5% of passwords cracked using a 30% common word dictionary subset. The study underscores the importance of dynamic password cracking strategies that account for evolving user behaviours and policy driven password structures. Findings contribution to both forensic efficiency and cyber security awareness, highlight the dual impact of password policies on security and investigative capabilities. Future work will focus upon refining rule based cracking techniques and expanding research on password composition trends.

cs.CR

An Identity and Interaction Based Network Forensic Analysis

In todays landscape of increasing electronic crime, network forensics plays a pivotal role in digital investigations. It aids in understanding which systems to analyse and as a supplement to support evidence found through more traditional computer based investigations. However, the nature and functionality of the existing Network Forensic Analysis Tools (NFATs) fall short compared to File System Forensic Analysis Tools (FS FATs) in providing usable data. The analysis tends to focus upon IP addresses, which are not synonymous with user identities, a point of significant interest to investigators. This paper presents several experiments designed to create a novel NFAT approach that can identify users and understand how they are using network based applications whilst the traffic remains encrypted. The experiments build upon the prior art and investigate how effective this approach is in classifying users and their actions. Utilising an in-house dataset composed of 50 million packers, the experiments are formed of three incremental developments that assist in improving performance. Building upon the successful experiments, a proposed NFAT interface is presented to illustrate the ease at which investigators would be able to ask relevant questions of user interactions. The experiments profiled across 27 users, has yielded an average 93.3% True Positive Identification Rate (TPIR), with 41% of users experiencing 100% TPIR. Skype, Wikipedia and Hotmail services achieved a notably high level of recognition performance. The study has developed and evaluated an approach to analyse encrypted network traffic more effectively through the modelling of network traffic and to subsequently visualise these interactions through a novel network forensic analysis tool.

cs.CR

From Age Estimation to Age-Invariant Face Recognition: Generalized Age Feature Extraction Using Order-Enhanced Contrastive Learning

Generalized age feature extraction is crucial for age-related facial analysis tasks, such as age estimation and age-invariant face recognition (AIFR). Despite the recent successes of models in homogeneous-dataset experiments, their performance drops significantly in cross-dataset evaluations. Most of these models fail to extract generalized age features as they only attempt to map extracted features with training age labels directly without explicitly modeling the natural ordinal progression of aging. In this paper, we propose Order-Enhanced Contrastive Learning (OrdCon), a novel contrastive learning framework designed explicitly for ordinal attributes like age. Specifically, to extract generalized features, OrdCon aligns the direction vector of two features with either the natural aging direction or its reverse to model the ordinal process of aging. To further enhance generalizability, OrdCon leverages a novel soft proxy matching loss as a second contrastive objective, ensuring that features are positioned around the center of each age cluster with minimal intra-class variance and proportionally away from other clusters. By modeling the ageing process, the framework can enhance generalizability by improving the alignment of samples from the same class and reducing the divergence of direction vectors. We demonstrate that our proposed method achieves comparable results to state-of-the-art methods on various benchmark datasets in homogeneous-dataset evaluations for both age estimation and AIFR. In cross-dataset experiments, OrdCon outperforms other methods by reducing the mean absolute error by approximately 1.38 on average for the age estimation task and boosts the average accuracy for AIFR by 1.87%.

cs.CV

GPT-Enabled Cybersecurity Training: A Tailored Approach for Effective Awareness

This study explores the limitations of traditional Cybersecurity Awareness and Training (CSAT) programs and proposes an innovative solution using Generative Pre-Trained Transformers (GPT) to address these shortcomings. Traditional approaches lack personalization and adaptability to individual learning styles. To overcome these challenges, the study integrates GPT models to deliver highly tailored and dynamic cybersecurity learning expe-riences. Leveraging natural language processing capabilities, the proposed approach personalizes training modules based on individual trainee pro-files, helping to ensure engagement and effectiveness. An experiment using a GPT model to provide a real-time and adaptive CSAT experience through generating customized training content. The findings have demonstrated a significant improvement over traditional programs, addressing issues of en-gagement, dynamicity, and relevance. GPT-powered CSAT programs offer a scalable and effective solution to enhance cybersecurity awareness, provid-ing personalized training content that better prepares individuals to miti-gate cybersecurity risks in their specific roles within the organization.

cs.CR

A Unified Knowledge Graph to Permit Interoperability of Heterogeneous Digital Evidence

The modern digital world is highly heterogeneous, encompassing a wide variety of communications, devices, and services. This interconnectedness generates, synchronises, stores, and presents digital information in multidimensional, complex formats, often fragmented across multiple sources. When linked to misuse, this digital information becomes vital digital evidence. Integrating and harmonising these diverse formats into a unified system is crucial for comprehensively understanding evidence and its relationships. However, existing approaches to date have faced challenges limiting investigators' ability to query heterogeneous evidence across large datasets. This paper presents a novel approach in the form of a modern unified data graph. The proposed approach aims to seamlessly integrate, harmonise, and unify evidence data, enabling cross-platform interoperability, efficient data queries, and improved digital investigation performance. To demonstrate its efficacy, a case study is conducted, highlighting the benefits of the proposed approach and showcasing its effectiveness in enabling the interoperability required for advanced analytics in digital investigations.

cs.CR

A proactive malicious software identification approach for digital forensic examiners

Digital investigators often get involved with cases, which seemingly point the responsibility to the person to which the computer belongs, but after a thorough examination malware is proven to be the cause, causing loss of precious time. Whilst Anti-Virus (AV) software can assist the investigator in identifying the presence of malware, with the increase in zero-day attacks and errors that exist in AV tools, this is something that cannot be relied upon. The aim of this paper is to investigate the behaviour of malware upon various Windows operating system versions in order to determine and correlate the relationship between malicious software and OS artifacts. This will enable an investigator to be more efficient in identifying the presence of new malware and provide a starting point for further investigation.

cs.CR

The detection of an extremely bright fast radio burst in a phased array feed survey

We report the detection of an ultra-bright fast radio burst (FRB) from a modest, 3.4-day pilot survey with the Australian Square Kilometre Array Pathfinder. The survey was conducted in a wide-field fly's-eye configuration using the phased-array-feed technology deployed on the array to instantaneously observe an effective area of $160$ deg$^2$, and achieve an exposure totaling $13200$ deg$^2$ hr. We constrain the position of FRB 170107 to a region $8'\times8'$ in size (90% containment) and its fluence to be $58\pm6$ Jy ms. The spectrum of the burst shows a sharp cutoff above $1400$ MHz, which could be either due to scintillation or an intrinsic feature of the burst. This confirms the existence of an ultra-bright ($>20$ Jy ms) population of FRBs.

astro-ph.HE

A Multi-Beam Radio Transient Detector With Real-Time De-Dispersion Over a Wide DM Range

Isolated, short dispersed pulses of radio emission of unknown origin have been reported and there is strong interest in wide-field, sensitive searches for such events. To achieve high sensitivity, large collecting area is needed and dispersion due to the interstellar medium should be removed. To survey a large part of the sky in reasonable time, a telescope that forms multiple simultaneous beams is desirable. We have developed a novel FPGA-based transient search engine that is suitable for these circumstances. It accepts short-integration-time spectral power measurements from each beam of the telescope, performs incoherent de-dispersion simultaneously for each of a wide range of dispersion measure (DM) values, and automatically searches the de-dispersed time series for pulse-like events. If the telescope provides buffering of the raw voltage samples of each beam, then our system can provide trigger signals to allow data in those buffers to be saved when a tentative detection occurs; this can be done with a latency of tens of ms, and only the buffers for beams with detections need to be saved. In one version of our implementation, intended for the ASKAP array of 36 antennas (currently under construction in Australia), 36 beams are simultaneously de-dispersed for 448 different DMs with an integration time of 1.0 ms. In the absence of such a multi-beam telescope, we have built a second version that handles up to 6 beams at 0.1 ms integration time and 512 DMs. We have deployed and tested this at a 34-m antenna of the Deep Space Network in Goldstone, California. A third version that processes up to 6 beams at an integration time of 2.0 ms and 1,024 DMs has been built and deployed at the Murchison Widefield Array telescope.

astro-ph.IM

Performance of a novel fast transients detection system

We investigate the S/N of a new incoherent dedispersion algorithm optimized for FPGA-based architectures intended for deployment on ASKAP and other SKA precursors for fast transients surveys. Unlike conventional CPU- and GPU-optimized incoherent dedispersion algorithms, this algorithm has the freedom to maximize the S/N by way of programmable dispersion profiles that enable the inclusion of different numbers of time samples per spectral channel. This allows, for example, more samples to be summed at lower frequencies where intra-channel dispersion smearing is larger, or it could even be used to optimize the dedispersion sum for steep spectrum sources. Our analysis takes into account the intrinsic pulse width, scatter broadening, spectral index and dispersion measure of the signal, and the system's frequency range, spectral and temporal resolution, and number of trial dedispersions. We show that the system achieves better than 80% of the optimal S/N where the temporal resolution and the intra-channel smearing time are smaller than a quarter of the average width of the pulse across the system's frequency band (after including scatter smearing). Coarse temporal resolutions suffer a Delta_t^(-1/2) decay in S/N, and coarse spectral resolutions cause a Delta_nu^(-1/2) decay in S/N, where Delta_t and Delta_nu are the temporal and spectral resolutions of the system, respectively. We show how the system's S/N compares with that of matched filter and boxcar filter detectors. We further present a new algorithm for selecting trial dispersion measures for a survey that maintains a given minimum S/N performance across a range of dispersion measures.

astro-ph.IM