SearcharxivSearch

arXiv subjects

Jonathan Shek

Publications and source records attributed to Jonathan Shek.

5 recordsLinked to original sources

Wind Turbine Maintenance Log Labelling Framework: LLM-Driven Data Correction and Enrichment via Semantic Extraction of Reliability Intelligence

As wind turbine fleets age, data-driven reliability engineering and maintenance optimisation are essential to manage lifecycle expenditure and support asset life extension. Historical maintenance records offer a vital source of field evidence, yet their analytical use is impeded by inconsistent system codes, generic categorical fields, and unstructured technician text. This paper presents a topology-aware large language model (LLM) workflow for reviewing legacy labels, extracting candidate maintenance and failure-mode taxonomies, and assigning structured semantic fields at record level. The workflow processed 16,316 maintenance records from 280 turbines across 32 onshore wind farms, spanning 9.2 years of operational history. It combines system-specific batch synthesis with granular labelling, deterministic exclusions, structured outputs, record-level provenance, and explicit review routes. Of 2,984 records targeted by three system-code tasks, 2,178 proposed labels met the operational acceptance rule of a 'High' self-reported confidence tier and no human-review flag. Accepted maintenance-type and action labels were assigned to 14,251 and 13,179 records, respectively. Failure-mode evidence profiles were assigned to 11,662 records; 3,441 records were classified as containing insufficient information, and 1,213 records were excluded as 'Not applicable' by deterministic workflow rules. The resulting fields reveal changes in system and maintenance-type distributions, a broader component-level action vocabulary, and topology-specific candidate evidence profiles. The recorded API expenditure was $368.86, or $0.0226 per processed record, and the total wall-clock duration was 6.83 hours. The provenance-linked outputs constitute candidate semantic evidence for subsequent multi-source event reconstruction, exposure-based reliability analysis, and failure modes and effects analysis (FMEA).

cs.CL

Exploratory Semantic Reliability Analysis of Wind Turbine Maintenance Logs using Large Language Models

A wealth of operational intelligence is locked within the unstructured free-text of wind turbine maintenance logs, a resource largely inaccessible to traditional quantitative reliability analysis. While machine learning has been applied to this data, existing approaches typically stop at classification, categorising text into predefined labels. This paper addresses the gap in leveraging modern large language models (LLMs) for more complex reasoning tasks. We introduce an exploratory framework that uses LLMs to move beyond classification and perform deep semantic analysis. We apply this framework to a large industrial dataset to execute four analytical workflows: failure mode identification, causal chain inference, comparative site analysis, and data quality auditing. The results demonstrate that LLMs can function as powerful "reliability co-pilots," moving beyond labelling to synthesise textual information and generate actionable, expert-level hypotheses. This work contributes a novel and reproducible methodology for using LLMs as a reasoning tool, offering a new pathway to enhance operational intelligence in the wind energy sector by unlocking insights previously obscured in unstructured data.

cs.CL

Analysis and Control of Acoustic Emissions from Marine Energy Converters

Environmental licensing related to underwater acoustic emissions represents a critical bottleneck for the commercial deployment of marine renewable energy. This study presents a control engineering framework to mitigate acoustic risks from tidal current converters without compromising project viability. A MATLAB/Simulink model of a tidal current converter was utilised to evaluate two distinct mitigation tiers: (1) architectural modification, comparing a geared induction generator against a direct-drive permanent magnet synchronous generator, and (2) operational control, analysing the impact of switching frequencies and maximum power point tracking coefficient tuning. Results indicate that lowering switching frequencies is ineffective, increasing power electronic losses by over 2000% with negligible acoustic benefit. Conversely, the direct-drive permanent magnet synchronous generator architecture reduced sound pressure levels, effectively eliminating mechanical tonal noise. For existing geared systems, de-tuning the maximum power point tracking coefficient by a factor of 1.2 reduced the probability of exceeding temporary threshold shift limits for marine mammals, with a quantified energy yield reduction of 3.58%. These findings propose a hierarchical mitigation strategy: selecting direct-drive topologies for acoustically sensitive sites, and utilising maximum power point tracking coefficient based power curtailment as a transient operational mode during critical biological migration periods.

eess.SY

A Comparative Benchmark of Large Language Models for Labelling Wind Turbine Maintenance Logs

Effective Operation and Maintenance (O&M) is critical to reducing the Levelised Cost of Energy (LCOE) from wind power, yet the unstructured, free-text nature of turbine maintenance logs presents a significant barrier to automated analysis. Our paper addresses this by presenting a novel and reproducible framework for benchmarking Large Language Models (LLMs) on the task of classifying these complex industrial records. To promote transparency and encourage further research, this framework has been made publicly available as an open-source tool. We systematically evaluate a diverse suite of state-of-the-art proprietary and open-source LLMs, providing a foundational assessment of their trade-offs in reliability, operational efficiency, and model calibration. Our results quantify a clear performance hierarchy, identifying top models that exhibit high alignment with a benchmark standard and trustworthy, well-calibrated confidence scores. We also demonstrate that classification performance is highly dependent on the task's semantic ambiguity, with all models showing higher consensus on objective component identification than on interpretive maintenance actions. Given that no model achieves perfect accuracy and that calibration varies dramatically, we conclude that the most effective and responsible near-term application is a Human-in-the-Loop system, where LLMs act as a powerful assistant to accelerate and standardise data labelling for human experts, thereby enhancing O&M data quality and downstream reliability analysis.

cs.CL

Reconfigurable Power Electronics Topologies

This paper presents two novel topologies for automatically transforming power converter topology from three-phase 3-level cascaded H-bridge to three-phase 2-level converter design. These techniques are implemented by flicking specific switches to rearrange circuit connections. The switches can be controlled by signals in order to realize automation.

eess.SP