Searcharxiv⌕ Search

arXiv subjects

Negin Ayoughi

Publications and source records attributed to Negin Ayoughi.

3 recordsLinked to original sources

Synthesizing Behavioural Models of CPS Using Automata Learning and Statistical Machine Learning

Inferring behavioural models from system executions is essential for supporting formal verification and analysis of complex, heterogeneous cyber-physical systems (CPS). Automata learning provides an effective way to infer state machine models from system executions. However, CPS inputs and outputs often consist of numeric time-series data, while automata learning algorithms assume inputs over a finite symbolic alphabet. As a result, raw numeric data must first be abstracted into a finite set of symbols. In this article, we present MELA, a passive automata learning approach enhanced with machine learning to synthesize behavioural models from numeric time-series data generated by CPS. MELA systematically combines statistical machine learning with automata learning to automatically abstract raw numeric signals into interpretable intervals that are strongly correlated with system states. Specifically, MELA uses information-theoretic variable selection and decision-tree-based range abstraction to transform numeric traces into symbolic representations suitable for automata learning. We evaluate MELA on two CPS: a commercial network intrusion detection system developed by our industry partner, RabbitRun Technologies, and a publicly available industrial autopilot benchmark from the aerospace domain. Compared with expertise-based numeric data abstraction, MELA reduces the number of states and transitions in the learned state machines by 49.20% on average, while improving accuracy by 41.71% on average. Furthermore, the learned state machines support system-level requirement verification and help practitioners explore behaviours that are not explicit in the system requirements. We make our implementation and experimental data available online. Keywords: Automata learning, Cyber-physical systems, Behavioural model synthesis, Decision trees, Model checking, Intrusion detection, Simulink.

cs.SE↗

DSL or Code? Evaluating the Quality of LLM-Generated Algebraic Specifications: A Case Study in Optimization at Kinaxis

Model-driven engineering (MDE) provides abstraction and analytical rigour, but industrial adoption in many domains has been limited by the cost of developing and maintaining models. Large language models (LLMs) can help shift this cost balance by supporting direct generation of models from natural-language (NL) descriptions. For domain-specific languages (DSLs), however, LLM-generated models may be less accurate than LLM-generated code in mainstream languages such as Python, due to the latter's dominance in LLM training corpora. We investigate this issue in mathematical optimization, with AMPL, a DSL with established industrial use. We introduce EXEOS, an LLM-based approach that derives AMPL models and Python code from NL problem descriptions and iteratively refines them with solver feedback. Using a public optimization dataset and real-world supply-chain cases from our industrial partner Kinaxis, we evaluate generated AMPL models against Python code in terms of executability and correctness. An ablation study with two LLM families shows that AMPL is competitive with, and sometimes better than, Python, and that our design choices in EXEOS improve the quality of generated specifications.

cs.SE↗

Enhancing Automata Learning with Statistical Machine Learning: A Network Security Case Study

Intrusion detection systems are crucial for network security. Verification of these systems is complicated by various factors, including the heterogeneity of network platforms and the continuously changing landscape of cyber threats. In this paper, we use automata learning to derive state machines from network-traffic data with the objective of supporting behavioural verification of intrusion detection systems. The most innovative aspect of our work is addressing the inability to directly apply existing automata learning techniques to network-traffic data due to the numeric nature of such data. Specifically, we use interpretable machine learning (ML) to partition numeric ranges into intervals that strongly correlate with a system's decisions regarding intrusion detection. These intervals are subsequently used to abstract numeric ranges before automata learning. We apply our ML-enhanced automata learning approach to a commercial network intrusion detection system developed by our industry partner, RabbitRun Technologies. Our approach results in an average 67.5% reduction in the number of states and transitions of the learned state machines, while achieving an average 28% improvement in accuracy compared to using expertise-based numeric data abstraction. Furthermore, the resulting state machines help practitioners in verifying system-level security requirements and exploring previously unknown system behaviours through model checking and temporal query checking. We make our implementation and experimental data available online.

cs.CR↗