Searcharxiv⌕ Search

arXiv subjects

Roshan G. Ragel

Publications and source records attributed to Roshan G. Ragel.

8 recordsLinked to original sources

Architecture-aware Robustness Evaluation of Explainable Deep Learning for Breast Cancer Diagnosis

Explainable Artificial Intelligence (XAI) has become essential in medical image analysis to ensure transparency of deep learning (DL)-based diagnostic systems. However, selecting appropriate XAI techniques for breast cancer recognition remains largely ad hoc, with limited systematic evaluation across different DL architectures. This study presents a systematic architecture-aware evaluation protocol to assess the effectiveness of nine widely used XAI techniques across four categories of DL models: very deep, lightweight, transformer-based and hybrid neural networks. The evaluation is conducted on a breast ultrasound dataset comprising 780 images using clinically aligned spatial metrics, including Pointing Game, Intersection over Union and Mean Coverage, to quantify agreement between generated explanations and expert-annotated lesion regions. Results indicate that explanation quality is primarily influenced by the interaction between model architecture and XAI method, rather than any single technique consistently outperforming others. Hybrid architectures produce more spatially coherent explanations, while lightweight and transformer-based models exhibit greater variability across methods. The findings show that no single technique generalises across architectures and evaluation criteria, emphasising the need for joint selection of DL models and XAI techniques. Explainability depends on both model design and explanation strategy and should not be considered independently. This work provides a structured evaluation protocol and practical guidance for selecting XAI techniques in breast cancer diagnosis, supporting more transparent clinical decision-support systems. \textcolor{blue}{Code is publicly available at https://github.com/Nishan-Charlie/Explainable-AI}

cs.CV↗

UtVAA: Ultra-tiny Vision Transformer with Affix Attention for Mobile Image Classification

Vision Transformers (ViTs) have demonstrated strong representation capability in image classification. However, their quadratic self-attention complexity and large parameter counts limit deployment on resource-constrained mobile and edge devices. This paper introduces UtVAA, an ultra-tiny Vision Transformer architecture designed for efficient visual recognition under strict computational budgets. It incorporates a novel Affix Attention block that combines depthwise-pointwise local feature extraction, linear self-attention, coordinate attention for spatial dependency modelling, and a lightweight ternary fusion strategy to integrate local and global representations. In addition, Dilated Bottleneck blocks expand the receptive field using dilated depthwise separable convolutions while maintaining low FLOPs and stable optimisation through residual connections. UtVAA is implemented in scalable Tiny, Medium, and Large variants, with the smallest model containing 204.67K parameters and 53.95M FLOPs. Experimental results on CIFAR-10, CIFAR-100, PlantVillage-Tomato and SLIF-Tomato datasets show that UtVAA achieves competitive accuracy within a sub-million-parameter regime. Overall, the results demonstrate that transformer-based vision models can be redesigned into ultra-tiny architectures without significant loss in discriminative performance, making UtVAA suitable for mobile and edge deployment. Code is available at https://github.com/romiyal/UtVAA

cs.CV↗

U-FedTomAtt: Ultra-lightweight Federated Learning with Attention for Tomato Disease Recognition

Federated learning has emerged as a privacy-preserving and efficient approach for deploying intelligent agricultural solutions. Accurate edge-based diagnosis across geographically dispersed farms is crucial for recognising tomato diseases in sustainable farming. Traditional centralised training aggregates raw data on a central server, leading to communication overhead, privacy risks and latency. Meanwhile, edge devices require lightweight networks to operate effectively within limited resources. In this paper, we propose U-FedTomAtt, an ultra-lightweight federated learning framework with attention for tomato disease recognition in resource-constrained and distributed environments. The model comprises only 245.34K parameters and 71.41 MFLOPS. First, we propose an ultra-lightweight neural network with dilated bottleneck (DBNeck) modules and a linear transformer to minimise computational and memory overhead. To mitigate potential accuracy loss, a novel local-global residual attention (LoGRA) module is incorporated. Second, we propose the federated dual adaptive weight aggregation (FedDAWA) algorithm that enhances global model accuracy. Third, our framework is validated using three benchmark datasets for tomato diseases under simulated federated settings. Experimental results show that the proposed method achieves 0.9910% and 0.9915% Top-1 accuracy and 0.9923% and 0.9897% F1-scores on SLIF-Tomato and PlantVillage tomato datasets, respectively.

q-bio.QM↗

An optimized Parallel Failure-less Aho-Corasick algorithm for DNA sequence matching

The Aho-Corasick algorithm is multiple patterns searching algorithm running sequentially in various applications like network intrusion detection and bioinformatics for finding several input strings within a given large input string. The parallel version of the Aho-Corasick algorithm is called as Parallel Failure-less Aho-Corasick algorithm because it doesn't need failure links like in the original Aho-Corasick algorithm. In this research, we implemented an application specific parallel failureless Aho-Corasick algorithm to the general purpose graphics processing unit by applying several cache optimization techniques for matching DNA sequences. Our parallel Aho-Corasick algorithm shows better performance than the available parallel Aho-Corasick algorithm library due to its simplicity and optimized cache memory usage of graphics processing units for matching DNA sequences.

cs.DC↗

To Use or Not to Use: CPUs' Cache Optimization Techniques on GPGPUs

General Purpose Graphic Processing Unit(GPGPU) is used widely for achieving high performance or high throughput in parallel programming. This capability of GPGPUs is very famous in the new era and mostly used for scientific computing which requires more processing power than normal personal computers. Therefore, most of the programmers, researchers and industry use this new concept for their work. However, achieving high-performance or high-throughput using GPGPUs are not an easy task compared with conventional programming concepts in the CPU side. In this research, the CPU's cache memory optimization techniques have been adopted to the GPGPU's cache memory to identify rare performance improvement techniques compared to GPGPU's best practices. The cache optimization techniques of blocking, loop fusion, array merging and array transpose were tested on GPGPUs for finding suitability of these techniques. Finally, we identified that some of the CPU cache optimization techniques go well with the cache memory system of the GPGPU and shows performance improvements while some others show the opposite effect on the GPGPUs compared with the CPUs.

cs.DC↗

SecureD: A Secure Dual Core Embedded Processor

Security of embedded computing systems is becoming of paramount concern as these devices become more ubiquitous, contain personal information and are increasingly used for financial transactions. Security attacks targeting embedded systems illegally gain access to the information in these devices or destroy information. The two most common types of attacks embedded systems encounter are code-injection and power analysis attacks. In the past, a number of countermeasures, both hardware- and software-based, were proposed individually against these two types of attacks. However, no single system exists to counter both of these two prominent attacks in a processor based embedded system. Therefore, this paper, for the first time, proposes a hardware/software based countermeasure against both code-injection attacks and power analysis based side-channel attacks in a dual core embedded system. The proposed processor, named SecureD, has an area overhead of just 3.80% and an average runtime increase of 20.0% when compared to a standard dual processing system. The overhead were measured using a set of industry standard application benchmarks, with two encryption and five other programs.

cs.AR↗

Plagiarism Detection on Electronic Text based Assignments using Vector Space Model (ICIAfS14)

Plagiarism is known as illegal use of others' part of work or whole work as one's own in any field such as art, poetry, literature, cinema, research and other creative forms of study. Plagiarism is one of the important issues in academic and research fields and giving more concern in academic systems. The situation is even worse with the availability of ample resources on the web. This paper focuses on an effective plagiarism detection tool on identifying suitable intra-corpal plagiarism detection for text based assignments by comparing unigram, bigram, trigram of vector space model with cosine similarity measure. Manually evaluated, labelled dataset was tested using unigram, bigram and trigram vector. Even though trigram vector consumes comparatively more time, it shows better results with the labelled data. In addition, the selected trigram vector space model with cosine similarity measure is compared with tri-gram sequence matching technique with Jaccard measure. In the results, cosine similarity score shows slightly higher values than the other. Because, it focuses on giving more weight for terms that do not frequently exist in the dataset and cosine similarity measure using trigram technique is more preferable than the other. Therefore, we present our new tool and it could be used as an effective tool to evaluate text based electronic assignments and minimize the plagiarism among students.

cs.IR↗

Efficient Switch Architectures for Pre-configured Backup Protection with Sharing in Elastic Optical Networks (EON)

In this paper, we address the problem of providing survivability in elastic optical networks (EONs). EONs use fine granular frequency slots or flexible grids, when compared to the conventional fixed grid networks and therefore utilize the frequency spectrum efficiently. For providing survivability in EONs, we consider a recently proposed survivability method for conventional fixed grid networks, known as pre-configured backup protection with sharing (PBPS), because of its benefits over the traditional survivability approaches such as dedicated and shared protection. In PBPS, backup paths can be pre-configured and at the same time they can share resources. Therefore, both short recovery time and efficient resource usage can be achieved. We find that the existing switch architectures do not support both PBPS and EONs. Specifically, we identify and illustrate that, if a switch architecture is not carefully designed, several key problems/issues might arise in certain scenarios. Such problems include unnecessary resource consumption, inability of using existing free resources, and incapability of sharing backup paths. These problems appear when PBPS is adopted in EONs and they do not arise in fixed grid networks. In this paper, we propose new switch architectures which support both PBPS and EONs. Particularly, we illustrate that, our switch architectures avoid the specific problems/issues mentioned above. Therefore, our switch architectures support using resources more efficiently and reducing blocking of requests.

cs.NI↗