SearcharxivSearch

arXiv subjects

Muhammad Imran

Publications and source records attributed to Muhammad Imran.

At least 19 recordsLinked to original sources

Exact quantum splitting and the structure of finite algebras

Berlekamp's algorithm factors a squarefree polynomial $f\in\mathbb{F}_q[x]$ by deterministic linear algebra, reducing the problem to splitting an explicit commutative algebra $B\cong\mathbb{F}_q^r$ into its $r$ simple factors. For large odd $q$, the standard efficient splitting step is randomized, while known derandomizations are conditional on the Extended Riemann Hypothesis. We give an unconditional exact quantum implementation in a circuit model permitting single-qubit rotations through efficiently computable angles. The construction uses an unconditional counting argument. For a block containing $s\ge2$ irreducible factors, a quadratic-character test in odd characteristic and an absolute-trace test in characteristic $2$ yield a nonconstant test element with probability $p_{q,s}\ge\tfrac12$, known exactly in advance and depending only on $q$ and $s$, not on the unknown factorization. Exact amplitude amplification therefore converts each randomized test into a procedure succeeding with certainty after one amplification iteration. The resulting algorithm uses exactly $r-1$ quantum splitting rounds and $O(n^3\log q)$ quantum $\mathbb{F}_q$-operations and $O(n^3)$ classical operations, requiring no primitive root, quadratic non-residue, or distinct-degree preprocessing. The method also splits arbitrary finite-dimensional separable commutative $\\mathbb{F}_q$-algebras given by structure constants. Combined with R'onyai's classical structure theory, which computes the radical deterministically and reduces the remaining tasks deterministically to polynomial factorization, it yields the radical and the Wedderburn decomposition of $A/\mathrm{Rad}(A)$ into minimal two-sided ideals, with certainty, for any $n$-dimensional associative $\mathbb{F}_q$-algebra given by structure constants, using $O(n^4\log q)$ quantum $\mathbb{F}_q$-operations.

quant-ph

Filtering-out poor-quality images for data preparation

Filtering noise is a fundamental part of data preparation that enhances image quality for applications such as object segmentation, detection, and recognition. Various noise reduction techniques are proposed in the literature, including the use of median, Gaussian, and bilateral filters. Convolutional neural networks (CNNs) have gained popularity in image denoising owing to their ability to extract complex patterns and features from data. CNNs are highly adaptable, making them effective tools for various image-denoising tasks. One drawback of CNN-based techniques is that they require an appropriate training dataset and all images to be resized. Another notable drawback of all these filtering techniques is that they work for certain types of environmental and camera noises. To bridge this research gap, in this paper, for the first time, instead of denoising, we propose an approach that filters out poor-quality images for various environmental and camera impacts. In our approach, quality is assessed using an image quality assessment metric and an optimum threshold is used to filter out poor-quality images. We also ensure that a sufficient number of images remain to develop the deep learning (DL) model. The results produced using real and simulated traffic and object recognition data demonstrate the performance supremacy of the proposed approach compared with the state-of-the-art approaches. The average recognition accuracy for our proposed approach is 93.8% for the traffic sign recognition dataset and 84.9% for the object recognition dataset. This indicates our model's potential for real-life applications such as autonomous vehicles.

cs.CV

Static Metrics Are Insufficient: Predicting Java Method Energy Usage with Execution Time

The increasing energy demand of software systems is raising concerns about their environmental impact and associated costs. Reasoning on energy usage early in the development flow has the potential to significantly reduce the overall energy usage of a software system, as it allows developers to make informed design and refactoring decisions before inefficiencies propagate. However, assessing energy usage without repeated profiling and direct measurement is difficult, which limits early reasoning in practice. This study investigates the limits of method-level energy prediction in Java, examining whether static source code metrics complemented with method-level execution time can estimate the energy consumption of Java methods. We profile 2,786 Java methods to extract 33 static features and measure execution time and energy, then train and compare eleven regression models. Our findings show that static source code metrics alone yield poor predictive performance, with average R2 values close to zero. Incorporating execution time as a lightweight dynamic input significantly improves accuracy, raising R2 to as high as 0.46. Execution time, internal method calls, and cyclomatic complexity consistently emerge as the strongest predictors of energy consumption.

cs.SE

ErgoGlide: A Wearable Trackball Device for Ergonomic Text Entry in Virtual Reality

In virtual reality, it is challenging to achieve satisfactory text entry speed/accuracy, ergonomics, usability, and learnability. To address this issue, we developed ErgoGlide, a novel lightweight and compact wearable device that facilitates text entry tasks in virtual environments. The proposed ErgoGlide can be regarded as a small trackball that is wearable on a user's finger like a ring. By using ErgoGlide with a hive-like virtual keyboard, the user can rotate the ball for key selections, making text entry intuitive and accurate. We conducted three user studies to evaluate ErgoGlide and found that key confirmation techniques have significant effects on text entry speed and the hive-like keyboard design significantly reduced thumb movements. Furthermore, ErgoGlide can significantly improve typing accuracy, ergonomics, and usability over previous text entry methods. Experimental results also indicated that the typing speed of ErgoGlide can be notably improved after training.

cs.HC

Verification and Validation (V&V)-in-the-Loop for RISC-V Design: The Holistic Vision of BZL

The Barcelona Zetascale Lab (BZL) project aims to strengthening Europe's capacity in the design and manufacture of RISC-V based high-performance computing chips. In this context, we present a holistic pre-silicon verification and validation (V&V) methodology targeting highly robust RISC-V chip designs. This paper provides an overview of BZL's V&V approach, which integrates three complementary platforms: (1) a UVM-based verification environment to thoroughly validate RTL functionality; (2) an FPGA-based validation platform that enables system-level pre-silicon hardware-software RTL validation; and (3) a CI/CD flow that continuously automates build, deployment, and tests across these domains. By embedding these platforms into an industrial-grade V&V loop and exploiting large-scale CPU and FPGA hardware infrastructures, the BZL project enables continuous evolution of reliable hardware development and software integration. We believe that the BZL's V&V flow represents a robust and scalable foundation for ensuring the pre-silicon functional correctness and system level validation of RISC-V chip designs, and can serve as a key enabler for strategic initiatives in Europe, such as EPI and DARE, and beyond.

cs.AR

Dark solitons in nonlinear Su-Schrieffer-Heeger lattices

The introduction of nonlinearities into lattices with topological band structures has led to the discovery of various types of solitons. The Su-Schrieffer-Heeger (SSH) lattice, as the most fundamental topological model, has been extended into the nonlinear regime. In particular, nonlinear edge states and bulk solitons exhibiting intensity humps against a zero background have been extensively studied in nonlinear SSH lattices. In this paper, we systematically investigate dark solitons in nonlinear SSH lattices. These dark solitons maintain a nonzero and constant background, featuring intensity dips either in the bulk of the lattice or at its edges, and residing spectrally in the semi-infinite gap or the middle finite gap. Regardless of the specific type of dark soliton, the intensity dip remains wellpreserved and is not affected by the band structure of the original linear lattice. Although the dark solitons we have identified are generally dynamically unstable across a broad range of parameters, several types exhibit linear stability when the intracell coupling is much larger than the intercell coupling. Our findings may provide valuable insights for the exploration of novel types of solitons in nonlinear topological lattices.

nlin.PS

Symmetry-breaking bifurcation of coupled topological edge states

We propose that the symmetry-breaking bifurcation of coupled topological edge states (CTESs) can be used as a general principle for achieving spontaneous symmetry breaking (SSB) in a nonlinear topological lattice. Using an optical resonator array composed of two Su-Schrieffer-Heeger (SSH) chains as an example, we find that as the nonlinearity strength increases, the symmetric CTESs undergo a supercritical bifurcation. Beyond the critical threshold, the originally stable symmetric state becomes unstable, leading to the formation of a pair of stable asymmetric states. Both sides of the symmetric CTESs exhibit sublattice polarization, while the side of the asymmetric CTESs that is predominantly occupied demonstrates stronger sublattice polarization. We further find that as interchain coupling increases, the frequency range for stable CTESs expands, while the frequency range for stable asymmetric CTESs decreases. Our work provides a universal mechanism for realizing SSB in nonlinear topological lattices.

physics.optics

Architecting Trust: A Framework for Secure IoT Systems Through Trusted Execution and Semantic Middleware

The Internet of Things (IoT) security landscape requires the architectural solutions that can address the technical and operational challenges across the heterogeneous environments. The IoT systems operate in different conditions, and security issues continue to increase. This paper presents the comprehensive security framework for IoT that should integrate the Trusted Execution Environments (TEEs) with the semantic middleware and blockchain technologies. The work provides a systematic analysis of the architectural patterns based on more than twenty recent research works and the existing standards, and it proposes a layered security architecture. The architecture includes the hardware rooted trust at peripheral level, the zero trust principles at network level, and the semantic security mechanisms at application level. The framework focuses on practical implementation aspects such as the performance overhead, interoperability requirements, and the compliance with new regulations, which are very important for the real IoT deployments. The paper reports quantitative metrics which include the cryptographic performance on Cortex-M class microcontrollers with the detection accuracy rates and the energy consumption values. The proposed architecture shows that cross-layer security integration can provide defense in depth while it still satisfies the constraints of resource-limited IoT environments. The discussion highlights open challenges and the future research directions for the IoT security architectures that include the post-quantum migration, secure federated model exchange and the automated compliance verification.

cs.CR

DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in general-purpose domains, its suitability for the complex and safety-critical reasoning required in disaster response remains unclear. We introduce DisasterVQA, a benchmark dataset designed for perception and reasoning in crisis contexts. DisasterVQA consists of 1,395 real-world images and 4,405 expert-curated question-answer pairs spanning diverse events such as floods, wildfires, and earthquakes. Grounded in humanitarian frameworks including FEMA ESF and OCHA MIRA, the dataset includes binary, multiple-choice, and open-ended questions covering situational awareness and operational decision-making tasks. We benchmark seven state-of-the-art vision-language models and find performance variability across question types, disaster categories, regions, and humanitarian tasks. Although models achieve high accuracy on binary questions, they struggle with fine-grained quantitative reasoning, object counting, and context-sensitive interpretation, particularly for underrepresented disaster scenarios. DisasterVQA provides a challenging and practical benchmark to guide the development of more robust and operationally meaningful vision-language models for disaster response. The dataset is publicly available at https://doi.org/10.5281/zenodo.18267769.

cs.CV

MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models

We introduce MATEX (Multi-scale Attention and Text-guided Explainability), a novel framework that advances interpretability in medical vision-language models by incorporating anatomically informed spatial reasoning. MATEX synergistically combines multi-layer attention rollout, text-guided spatial priors, and layer consistency analysis to produce precise, stable, and clinically meaningful gradient attribution maps. By addressing key limitations of prior methods, such as spatial imprecision, lack of anatomical grounding, and limited attention granularity, MATEX enables more faithful and interpretable model explanations. Evaluated on the MS-CXR dataset, MATEX outperforms the state-of-the-art M2IB approach in both spatial precision and alignment with expert-annotated findings. These results highlight MATEX's potential to enhance trust and transparency in radiological AI applications.

cs.CV

Predicting When to Trust Vision-Language Models for Spatial Reasoning

Vision-Language Models (VLMs) demonstrate impressive capabilities across multimodal tasks, yet exhibit systematic spatial reasoning failures, achieving only 49% (CLIP) to 54% (BLIP-2) accuracy on basic directional relationships. For safe deployment in robotics and autonomous systems, we need to predict when to trust VLM spatial predictions rather than accepting all outputs. We propose a vision-based confidence estimation framework that validates VLM predictions through independent geometric verification using object detection. Unlike text-based approaches relying on self-assessment, our method fuses four signals via gradient boosting: geometric alignment between VLM claims and coordinates, spatial ambiguity from overlap, detection quality, and VLM internal uncertainty. We achieve 0.674 AUROC on BLIP-2 (34.0% improvement over text-based baselines) and 0.583 AUROC on CLIP (16.1% improvement), generalizing across generative and classification architectures. Our framework enables selective prediction: at 60% target accuracy, we achieve 61.9% coverage versus 27.6% baseline (2.2x improvement) on BLIP-2. Feature analysis reveals vision-based signals contribute 87.4% of model importance versus 12.7% from VLM confidence, validating that external geometric verification outperforms self-assessment. We demonstrate reliable scene graph construction where confidence-based pruning improves precision from 52.1% to 78.3% while retaining 68.2% of edges.

cs.CV

Exogenous Metal Cations in the Synthesis of CsPbBr3 Nanocrystals and their Interplay with Tertiary Amines

Current syntheses of CsPbBr3 halide perovskite nanocrystals (NCs) rely on over-stoichiometric amounts of Pb2+ precursors, resulting in unreacted lead ions at the end of the process. In our synthesis scheme of CsPbBr3 NCs we replaced excess Pb2+ with different exogenous metal cations (M) and investigated their effect on the synthesis products. These cations can be divided into two groups: group 1 delivers monodisperse CsPbBr3 cubes capped with oleate species (as for the case when Pb2+ is used in excess) and with photoluminescence quantum yield (PLQY) as high as 90% with some cations (for example with M= In3+); group 2 yields irregularly shaped CsPbBr3 NCs with broad size distributions. In both cases, the addition of a tertiary ammonium cation (didodecylmethyl ammonium, DDMA+) during the synthesis, after the nucleation of the NCs, reshapes the NCs to monodisperse truncated cubes. Such NCs feature a mixed oleate/DDMA+ surface termination with PLQY values up to 90%. For group 1 cations, this happens only if the ammonium cation is directly added as a salt (DDMA-Br) while for group 2 cations this happens even if the corresponding tertiary amine (DDMA) is added, instead of DDMA-Br. This is attributed to the fact that only group 2 cations can facilitate the protonation of DDMA by the excess oleic acid present in the reaction environment. In all cases studied, the incorporation of M cations is marginal and the reshaping of the NCs is only transient: if the reactions are run for a long time the truncated cubes evolve to cubes.

cond-mat.mtrl-sci

Halide Perovskite-Chalcohalide Nanocrystal Heterostructures as a Platform for the Synthesis and Investigation of the CsPbCl3-CsPbI3 Epitaxial Interface

Halide exchange in lead-based halide perovskites has been studied extensively. While mixed Cl-Br or Br-I alloy compositions can be formed with no miscibility gaps, this is precluded for mixed Cl-I compositions, due to the large difference in Cl and I ionic radii. Here, we exploit perovskite-chalcohalide CsPbCl3-Pb4S3Cl2 heterostructures to study the Cl-I exchange and isolate new types of intermediate structures. The epitaxial interface between the Pb4S3Cl2 chalcohalide and the CsPbCl3 perovskite domain significantly influences the intermediate stages of halide exchange in the perovskite domain, leading to coexisting CsPbCl3 and CsPbI3 domains, thereby delivering segmented CsPbI3-CsPbCl3-Pb4S3Cl2 energetically favorable heterostructures, with partial iodide alloying of the CsPbCl3 domain and at the perovskite-chalcohalide interface. The I:CsPbCl3 domain between CsPbI3 and Pb4S3Cl2 enables a gradual lattice expansion across the heterostructure. This design accommodates interfacial strain, with a 5.6% mismatch at the CsPbCl3-CsPbI3 interface and a 3.4% mismatch at the perovskite-chalcohalide interface. Full halide exchange leads to CsPbI3-Pb4S3Cl2 heterostructures. Both in intermediate and fully exchanged heterostructures, the CsPbI3 domain is emissive. In the intermediate structures, the band alignment between the two perovskite domains is type-I, with carriers photogenerated in the CsPbCl3 domain quickly transferring to the CsPbI3 domain, where they can recombine radiatively.

cond-mat.mtrl-sci

Rare-Earth Engineering of NaAlO3 Perovskites Unlocks Unified Optoelectronic, Thermoelectric, and Spintronic Functionalities

Perovskite oxides are promising for energy and quantum technologies, but wide-gap hosts such as NaAlO3 suffer from deep-UV absorption and limited carrier transport. Using first-principles GGA+U+SOC calculations, we investigate Eu3+-, Gd3+-, and Tb3+-doped NaAlO3 and evaluate their electronic, optical, elastic, and thermoelectric properties. Rare-earth substitution is thermodynamically favorable (formation energies 1.2-1.6 eV) and induces strong f-p hybridization, reducing the pristine band gap (about 6.2 eV) to about 3.1 eV for Tb. Spin-resolved band structures reveal Gd-driven half-metallicity, Eu-induced spin-selective metallicity, and Tb-stabilized p-type semiconducting behavior. The optical spectra show a red-shifted absorption edge (about 2.0-2.2 eV), a large static dielectric response (epsilon1(0) about 95 for Eu), and plasmonic resonances near 4 eV, enabling visible-light harvesting. Elastic analysis indicates mild lattice softening with preserved ductility (Pugh ratio B/G about 1.56-1.57). Thermoelectric performance is enhanced, with Seebeck coefficients greater than 210 uV/K for Eu and Tb and ZT about 0.45 at 500 K. These results identify rare-earth-doped NaAlO3 as a multifunctional perovskite platform for photovoltaics, photocatalysis, thermoelectrics, and spintronics.

cond-mat.mtrl-sci

GeoResponder: Towards Building Geospatial LLMs for Time-Critical Disaster Response

LLMs excel at linguistic tasks but lack the inner geospatial capabilities needed for time-critical disaster response, where reasoning about road networks, coordinates, and access to essential infrastructure such as hospitals, shelters, and pharmacies is vital. We introduce GeoResponder, a framework that instills robust spatial reasoning through a scaffolded instruction-tuning curriculum. By stratifying geospatial learning into different cognitive layers, we anchor semantic knowledge to the continuous coordinate manifold and enforce the internalization of spatial axioms. Extensive evaluations across four topologically distinct cities and diverse tasks demonstrate that GeoResponder significantly outperforms both state-of-the-art foundation models and domain-specific baselines. These results suggest that LLMs can begin to internalize and generalize geospatial structures, pointing toward the future development of language models capable of supporting disaster response needs.

cs.CL

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Recent advances in vision-language models have significantly expanded the frontiers of automated image analysis. However, applying these models in safety-critical contexts remains challenging due to the complex relationships between objects, subtle visual cues, and the heightened demand for transparency and reliability. This paper presents the Multi-Modal Explainable Learning (MMEL) framework, designed to enhance the interpretability of vision-language models while maintaining high performance. Building upon prior work in gradient-based explanations for transformer architectures (Grad-eclip), MMEL introduces a novel Hierarchical Semantic Relationship Module that enhances model interpretability through multi-scale feature processing, adaptive attention weighting, and cross-modal alignment. Our approach processes features at multiple semantic levels to capture relationships between image regions at different granularities, applying learnable layer-specific weights to balance contributions across the model's depth. This results in more comprehensive visual explanations that highlight both primary objects and their contextual relationships with improved precision. Through extensive experiments on standard datasets, we demonstrate that by incorporating semantic relationship information into gradient-based attribution maps, MMEL produces more focused and contextually aware visualizations that better reflect how vision-language models process complex scenes. The MMEL framework generalizes across various domains, offering valuable insights into model decisions for applications requiring high interpretability and reliability.

cs.CV

Evaluating Compositional Approaches for Focus and Sentiment Analysis

This paper summarizes the results of evaluating a compositional approach for Focus Analysis (FA) in Linguistics and Sentiment Analysis (SA) in Natural Language Processing (NLP). While quantitative evaluations of compositional and non-compositional approaches in SA exist in NLP, similar quantitative evaluations are very rare in FA in Linguistics that deal with linguistic expressions representing focus or emphasis such as "it was John who left". We fill this gap in research by arguing that compositional rules in SA also apply to FA because FA and SA are closely related meaning that SA is part of FA. Our compositional approach in SA exploits basic syntactic rules such as rules of modification, coordination, and negation represented in the formalism of Universal Dependencies (UDs) in English and applied to words representing sentiments from sentiment dictionaries. Some of the advantages of our compositional analysis method for SA in contrast to non-compositional analysis methods are interpretability and explainability. We test the accuracy of our compositional approach and compare it with a non-compositional approach VADER that uses simple heuristic rules to deal with negation, coordination and modification. In contrast to previous related work that evaluates compositionality in SA on long reviews, this study uses more appropriate datasets to evaluate compositionality. In addition, we generalize the results of compositional approaches in SA to compositional approaches in FA.

cs.CL

On the Classical Hardness of the Semidirect Discrete Logarithm Problem in Finite Groups

The semidirect discrete logarithm problem (SDLP) in finite groups was proposed as a foundation for post-quantum cryptographic protocols, based on the belief that its non-abelian structure would resist quantum attacks. However, recent results have shown that SDLP in finite groups admits efficient quantum algorithms, undermining its quantum resistance. This raises a fundamental question: does the SDLP offer any computational advantages over the standard discrete logarithm problem (DLP) against classical adversaries? In this work, we investigate the classical hardness of SDLP across different finite group platforms. We establish that the group-case SDLP can be reformulated as a generalized discrete logarithm problem, enabling adaptation of classical algorithms to study its complexity. We present a concrete adaptation of the Baby-Step Giant-Step algorithm for SDLP, achieving time and space complexity $O(\sqrt{r})$ where $r$ is the period of the underlying cycle structure. Through theoretical analysis and experimental validation in SageMath, we demonstrate that the classical hardness of SDLP is highly platform-dependent and does not uniformly exceed that of standard DLP. In finite fields $\mathbb{F}_p^*$, both problems exhibit comparable complexity. Surprisingly, in elliptic curves $E(\mathbb{F}_p)$, the SDLP becomes trivial due to the bounded automorphism group, while in elementary abelian groups $\mathbb{F}_p^n$, the SDLP can be harder than DLP, with complexity varying based on the eigenvalue structure of the automorphism. Our findings reveal that the non-abelian structure of semidirect products does not inherently guarantee increased classical hardness, suggesting that the search for classically hard problems for cryptographic applications requires more careful consideration of the underlying algebraic structures.

cs.CR