SearcharxivSearch

arXiv subjects

Ali Hassan

Publications and source records attributed to Ali Hassan.

At least 19 recordsLinked to original sources

Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression

Large language and sentence-embedding models provide rich semantic representations, but their high dimensionality poses a challenge for near-term quantum machine learning (QML), where quantum circuits can process only a limited number of input features. We investigate a hybrid quantum-classical pipeline that transforms high-dimensional sentence embeddings into compact representations for variational quantum classification. The workflow combines a pretrained sentence-embedding model, dimensionality reduction, angle encoding, a variational quantum circuit (VQC), and a classical decision layer. We systematically compare principal component analysis (PCA), neighborhood components analysis (NCA), and linear discriminant analysis (LDA), covering both unsupervised and supervised dimensionality reduction. Using the TREC question-classification dataset, we study the relationship between representation dimensionality, information retention, qubit count, and classification performance. Preliminary PCA experiments reveal a strong information bottleneck: reducing 768-dimensional embeddings to 3, 4, 5, and 8 dimensions retains about 8.2%, 10.2%, 11.9%, and 16.4% of the variance, with corresponding classification accuracies of 50.3%, 51.2%, 57.9%, and 63.4%. In contrast, supervised reduction is substantially more efficient. LDA reaches 85.3% accuracy and NCA reaches 83.1% using only 5 dimensions, under a leakage-free cross-validation protocol, compared with 85.1% for a full 384-dimensional classical baseline. These results indicate that supervised dimensionality reduction can preserve task-relevant information far more effectively than variance-based compression, making compact representations a promising route toward practical hybrid quantum-classical NLP models.

cs.LG

A Stackelberg-Bayesian Capacity-Market Game of Carbon Regulation and Second-Life Battery Investment under AI Data-Center Load Growth

Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technology-specific investors (followers) decide capacity and operation under incomplete information, yielding a Bayesian Nash equilibrium. The AI impact is captured parsimoniously as an additional load-growth factor on a greenfield-incremental expansion, isolating how much new capacity the growth pulls in and which technology fills it. Within this framework, we consider second-life battery (SLB) storage competing against new/first-life storage for capacity-market revenue. We quantify how a carbon tax, a renewable subsidy, and an SLB subsidy reshape the equilibrium investment mix, carbon emissions, and profit. Different scenarios are compared at the end based on cost-effectiveness and reduced carbon emissions.

eess.SY

Inflation with nondynamic distortion to leading order in slow roll

We study inflation in metric-affine gravity. We write an action that contains all the second order algebraic distortion terms, and all first order distortion terms with a single covariant derivative, coupled to a scalar field. We include the Einstein--Hilbert term with nonminimal coupling, a scalar field potential, and impose projective invariance. The distortion equation of motion is algebraic by construction, and the distortion is integrated out analytically. This yields a kinetic term sourced entirely by distortion, with a kinetic coupling function determined by the 13 free coupling constants of the starting action. We compute inflationary observables for three model classes with a monomial distortion coupling. For a monomial potential, the spectral index and tensor-to-scalar ratio depend only on the ratio of the exponents, with the starting coupling constants dropping out entirely; however, this model lies outside the Planck + BK18 $2\sigma$ contours. For a potential of the $\alpha$-attractor form, the observables are governed by a single parameter and approach the Starobinsky predictions as a limit. Including a nonminimal coupling to the Ricci scalar with a monomial potential can also yield an asymptotically flat effective potential with the same modified Starobinsky observables.

astro-ph.CO

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved content from overriding operator policy. Layer 3 audits model output using a policy rule engine and semantic drift detector before delivery. A continuous audit loop aggregates structured logs and supports retraining to adapt the classifier to emerging attack patterns. The framework is model-agnostic and deploys as middleware without modifying the underlying LLM. Evaluation on 5,080 samples across GPT-4o, Llama 3, and Mistral 7B shows that the framework reduces Attack Success Rate (ASR) from 71.4\% to 11.3\%, outperforming the best single-layer baseline by 27.3 percentage points and a published guardrail system by 23.8 percentage points, while maintaining a 4.8\% false positive rate and a median latency overhead of 61.2 ms. Ablation studies confirm that all three layers provide complementary protection and that their combined effect exceeds the sum of individual contributions.

cs.CR

Probing Singlet Vector-Like Top Quarks in the Hadronic tZ Channel at the HL-LHC using Machine and Deep Learning Architectures

In this work, we study the single production of a vector-like singlet top partner \( T \) at the 14 TeV HL-LHC in the channel \( pp \to T j \) with \( T \to t Z \), \( t \to b W \to b j j \), and \( Z \to \nu \bar{\nu} \). Signal and background samples are generated with MadGraph5\_aMC@NLO v3.5.11, showered with Pythia 8, and passed through Delphes. The dominant backgrounds are \( t \bar{t} \), \( t Z j \), \( ZZ j j \), and \( W Z j j \) (including charge conjugates). A hadronic pre-selection (\( N_j \geq 3 \), \( N_b \geq 1 \), \( N_\ell = 0 \)) is imposed as trigger, followed by optimized kinematic cuts. We perform multivariate classification with Extreme Gradient Boosting (XGBoost) and a Graph Neural Network (GNN) based on jet-level features. Sensitivities at 3000 fb\(^{-1}\) are quoted using the Asimov significance, \( S / \sqrt{S + B} \), and an Asimov variant with a 20\% background systematic. The model parameters \( g^* \) and \( R_L \) are defined in Sec.~2, and a single global working point is used to avoid per-mass tuning bias. In the \( (g^*, m_T) \) scan, we present 2\(\sigma\) exclusion and 5\(\sigma\) discovery contours for \( R_L = 0 \) and \( R_L = 0.5 \). For \( R_L = 0 \), 2\(\sigma\) exclusion corresponds to \( g^* \in [0.17, 0.49] \) (\( 0.16, 0.43 \)) over \( m_T \in [1.8, 2.7] \) TeV, while 5\(\sigma\) discovery corresponds to \( g^* \in [0.27, 0.44] \) (\( 0.26, 0.40 \)) over \( m_T \in [1.8, 2.2] \) TeV for XGBoost and GNN respectively. For \( R_L = 0.5 \), the 2\(\sigma\) reach is \( g^* \in [0.21, 0.48] \) (\( 0.20, 0.43 \)) over \( m_T \in [1.8, 2.5] \) TeV, and the 5\(\sigma\) reach is \( g^* \in [0.33, 0.43] \) (\( 0.31, 0.49 \)) over \( m_T \in [1.8, 2.2] \) TeV, with the GNN yielding slightly stronger and smoother limits across the scan.

hep-ph

Surrogate modeling for convection-dominated parametric problems based on error learning

Convection-dominated problems are known for their slow Kolmogorov $n$-width decays and are challenging for model order reduction (MOR). In this work, we propose a hybrid surrogate modeling approach and a non-intrusive variant that overcome some drawbacks of linear MOR methods. The proposed hybrid surrogate model is a projection-based reduced-order model (ROM), corrected by the error learned from a deep neural network. With the aid of deep learning, the model component of the surrogate model can be kept in a small reduced dimension. The neural network component and the model component are sequentially but separately built during the offline stage. At the online stage, they are easily coupled to output the solution predictions. Due to the intrusive nature of the hybrid-ROM, the numerically discretized operators of the original model must be available. For problems solved using black-box solvers, where the details of the numerical discretization are not accessible, we further propose a non-intrusive variant of the hybrid surrogate. Compared to the existing MOR methods with nonlinear manifolds, the proposed hybrid ROM is more easily built and is also easily assembled for online prediction. In contrast to the surrogate modeling approaches purely based on deep-learning, the proposed non-intrusive variant has a lighter neural network structure with much fewer parameters to be learned. We test the proposed methods on two nonlinear convection parametric problems. The first is the 1D inviscid Burgers' equation with one parameter, and the second is the 2D inviscid Burgers' equation with two parameters. Since both methods are based on error correction, their online predictions exhibit higher accuracy yet with largely reduced prediction time, compared to state-of-the-art methods.

math.DS

LCC-LLM: Leveraging Code-Centric Large Language Models for Malware Attribution

LLMs are increasingly explored for malware analysis; however, current LLM-based malware attribution remains limited by unsupported indicators and insufficient code-level grounding for identifying malicious and vulnerable code segments. To address these limitations, this research introduces LCC-LLM, a code-centric benchmark dataset and evidence-grounded framework for malware attribution and multi-task static malware analysis. The proposed LCCD dataset contains approximately 34K PE samples processed through a large-scale reverse-engineering pipeline and represented using decompiled C code, assembly code, CFG/FCG artifacts, hexadecimal data, PE metadata, suspicious API evidence, and structural features. Beyond dataset construction, LCC-LLM integrates LangGraph-orchestrated static analysis with multi-source cybersecurity knowledge to support evidence-grounded malware reasoning. The framework employs a seven-layer retrieval-augmented generation pipeline, CoVe for IoC validation, and a multi-dimensional quality gate to improve factual reliability and analyst-oriented decision support. Curriculum-ordered instruction data is used to fine-tune DeepSeek-R1-Distill-Qwen-14B and Qwen3-Coder-30B-A3B using QLoRA. Evaluation across 43 malware-analysis task types achieves an average semantic similarity of 0.634, with the highest task-level performance in structured report generation, IoC extraction, vulnerability assessment, malware configuration extraction, and malware class detection. In a real-world case study using MalwareBazaar samples, the grounded pipeline achieves a 10/10 structured analysis pass rate, producing CFG/FCG evidence, MITRE ATT&CK mappings, detection guidance, and analyst-ready reports. These results show that code-centric representations, retrieval grounding, and verification-guided reasoning improve the reliability and operational usefulness of LLM-assisted malware attribution.

cs.CR

Inflation with the Gauss-Bonnet term in the Palatini formulation

We consider the Gauss-Bonnet term coupled to the inflaton in the Palatini formulation of gravity. Unlike in the metric formulation, the Gauss-Bonnet term is not always a total derivative. We solve for the connection and insert it into the action, exactly for the spatially flat FLRW spacetime, and using the gradient approximation and order reduction for a general spacetime. We consider three cases: when the connection is unconstrained, and when non-metricity or torsion is put to zero. In all cases, the leading order change to the inflaton kinetic has the same form as that generated by the Chern-Simons term, but a negative sign. The modification of the gravitational wave sector also has the same form as in the Chern-Simons case but with a negative sign, except possibly for zero torsion, depending on the coupling and the potential. Within the range of validity of our approximations, differences from the metric formulation are small unless the kinetic term flips sign or is close to doing so.

astro-ph.CO

Grid Integration of AI Data Centers: A Critical Review of Energy Storage Solutions

Artificial intelligence (AI) is driving unprecedented growth in data center (DC) scale and power demand. AI workloads impose highly dynamic, difficult-to-forecast power profiles on the utility grid, creating reliability and stability challenges that conventional DC architectures are not designed to address. This paper provides a critical review of energy storage systems (ESSs) as the key enabling technology for reliable grid integration of AI DCs. We organize the review around a four-layer hierarchical taxonomy, namely chip-level buffering, rack/server-level ESSs, facility-level uninterruptible power supply (UPS) systems, and grid-scale battery energy storage systems (BESSs), supplemented by non-battery technologies including fuel cells (FCs) and thermal energy storage (TES). Each layer is analyzed with respect to response timescale, power and energy ratings, operational role, integration challenges, and coordination requirements. Key findings include: (i) AI DC load profiles differ fundamentally from traditional loads in their sub-second variability, making conventional ESS dispatch strategies insufficient; (ii) hierarchical, coordinated ESS deployment across all layers is necessary for effective load smoothing and grid support; and (iii) significant gaps remain in simulation tools, degradation modeling, load forecasting, and optimal multi-layer sizing. This review identifies open research challenges and future directions at the intersection of AI computing infrastructure and power system integration.

eess.SY

Advancing Industry 4.0: Multimodal Sensor Fusion for AI-Based Fault Detection in 3D Printing

Additive manufacturing, particularly fused deposition modeling, is transforming modern production by enabling rapid prototyping and complex part fabrication. However, its layer-by-layer process remains vulnerable to faults such as nozzle clogging, filament runout, and layer misalignment, which compromise print quality and reliability. Traditional inspection methods are costly, time-intensive, and often limited to post-process analysis, making them unsuitable for real-time intervention. In this current study, the authors developed a novel, low-cost, and portable faultdetection system that leverages multimodal sensor fusion and artificial intelligence for real-time monitoring in FDM-based 3D printing. The system integrates acoustic, vibration, and thermal sensing into a non-intrusive architecture, capturing complementary data streams that reflect both mechanical and process-related anomalies. Acoustic and thermal sensors operate in a fully contactless manner, while the vibration sensor requires minimal attachment such that it will not interfere with printer hardware, thereby preserving portability and ease of deployment. The multimodal signals are processed into spectrograms and time-frequency features, which are classified using convolutional neural networks for intelligent fault detection. The proposed system advances Industry 4.0 objectives by offering an affordable, scalable, and practical monitoring solution that improves faultdetection accuracy, reduces waste, and supports sustainable, adaptive manufacturing.

eess.SP

Towards eco friendly cybersecurity: machine learning based anomaly detection with carbon and energy metrics

The rising energy footprint of artificial intelligence has become a measurable component of US data center emissions, yet cybersecurity research seldom considers its environmental cost. This study introduces an eco aware anomaly detection framework that unifies machine learning based network monitoring with real time carbon and energy tracking. Using the publicly available Carbon Aware Cybersecurity Traffic Dataset comprising 2300 flow level observations, we benchmark Logistic Regression, Random Forest, Support Vector Machine, Isolation Forest, and XGBoost models across energy, carbon, and performance dimensions. Each experiment is executed in a controlled Colab environment instrumented with the CodeCarbon toolkit to quantify power draw and equivalent CO2 output during both training and inference. We construct an Eco Efficiency Index that expresses F1 score per kilowatt hour to capture the trade off between detection quality and environmental impact. Results reveal that optimized Random Forest and lightweight Logistic Regression models achieve the highest eco efficiency, reducing energy consumption by more than forty percent compared to XGBoost while sustaining competitive detection accuracy. Principal Component Analysis further decreases computational load with negligible loss in recall. Collectively, these findings establish that integrating carbon and energy metrics into cybersecurity workflows enables environmentally responsible machine learning without compromising operational protection. The proposed framework offers a reproducible path toward sustainable carbon accountable cybersecurity aligned with emerging US green computing and federal energy efficiency initiatives.

cs.CR

Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B

Benchmarks for large language models (LLMs) often rely on rubric-scented prompts that request visible reasoning and strict formatting, whereas real deployments demand terse, contract-bound answers. We investigate whether such "evaluation scent" inflates measured performance without commensurate capability gains. Using a single open-weights model (GPT-OSS-20B), we run six paired A/B scenarios that hold task content and decoding fixed while varying framing (evaluation-oriented vs. real-world) and reasoning depth (Medium/High): deterministic math, strict code-fix, citation generation, incentive flips (caution vs. competence), CoT visibility, and multilingual (Urdu) headers. Deterministic validators compute accuracy, answer-only compliance, hedging/refusals, chain-of-thought (CoT) length, and schema compliance, with pre-registered deltas and composite indices. Across scenarios, evaluation framing reliably inflates CoT (hundreds to >1000 characters) and reduces answer-only compliance, with limited or inconsistent accuracy gains. In structured outputs, it improves wrappers (e.g., fenced blocks, enumerated lists) but not regex-validated substance. Incentive wording reweights error composition: praising caution modestly improves accuracy at high reasoning and reduces wrong-but-confident errors, whereas praising competence yields terser but riskier outputs. Urdu rubric headers reproduce these signatures and can decrease accuracy at higher reasoning depth, indicating multilingual parity risks. We provide a reproducible A/B framework (prompt banks, validators, per-run scores, scripts; versioned DOI) and practical guidance: neutral phrasing or dual-framing checks, contract-aware grading, style-delta reporting, confidence governance, and multilingual dashboards to ensure that benchmark gains reflect deployable capability.

cs.CL

Impact of Medium and Heavy-Duty Electric Vehicle Electrification on Distribution System Stability

Medium and heavy-duty (MHD) commercial vehicles contribute significantly to carbon emissions, accounting for 21\% of the total emissions in the transportation sector. To curb this, U.S. government is increasingly focusing on achieving 100\% fleet electrification over the next decade. However, the integration of megawatt-scale charging stations designed for MHD vehicles poses challenges to the stability of secondary distribution systems. This study investigates the impact of megawatt-scale charging station loads on a benchmark IEEE 33-bus distribution system using real data from the HEVI-LOAD software for MHD electrification planning developed by Lawrence Berkeley National Laboratory (LBNL). The results reveal significant violations of per-unit (p.u.) voltage values at various nodes of the distribution system, indicating that substantial upgrades to the distribution infrastructure will be necessary to accommodate the projected MHDEV charging loads and meet electrification targets.

nlin.AO

Large Language Models for Solving Economic Dispatch Problem

This paper investigates the capability of off-the-shelf large language models (LLMs) to solve the economic dispatch (ED) problem. ED is a hard-constrained optimization problem solved on a day-ahead timescale by grid operators to minimize electricity generation costs while accounting for physical and engineering constraints. Numerous approaches have been proposed, but these typically require either mathematical formulations, face convergence issues, or depend on extensive labeled data and training time. This work implements LLMs enhanced with reasoning capabilities to address the classic lossless ED problem. The proposed approach avoids the need for explicit mathematical formulations, does not suffer from convergence challenges, and requires neither labeled data nor extensive training. A few-shot learning technique is utilized in two different prompting contexts. The IEEE 118-bus system with 19 generation units serves as the evaluation benchmark. Results demonstrate that various prompting strategies enable LLMs to effectively solve the ED problem, offering a convenient and efficient alternative. Consequently, this approach presents a promising future solution for ED tasks, particularly when foundational power system models are available.

eess.SY

Deep Reinforcement Learning-Based Optimization of Second-Life Battery Utilization in Electric Vehicles Charging Stations

The rapid rise in electric vehicle (EV) adoption presents significant challenges in managing the vast number of retired EV batteries. Research indicates that second-life batteries (SLBs) from EVs typically retain considerable residual capacity, offering extended utility. These batteries can be effectively repurposed for use in EV charging stations (EVCS), providing a cost-effective alternative to new batteries and reducing overall planning costs. Integrating battery energy storage systems (BESS) with SLBs into EVCS is a promising strategy to alleviate system overload. However, efficient operation of EVCS with integrated BESS is hindered by uncertainties such as fluctuating EV arrival and departure times and variable power prices from the grid. This paper presents a deep reinforcement learning-based (DRL) planning framework for EV charging stations with BESS, leveraging SLBs. We employ the advanced soft actor-critic (SAC) approach, training the model on a year's worth of data to account for seasonal variations, including weekdays and holidays. A tailored reward function enables effective offline training, allowing real-time optimization of EVCS operations under uncertainty.

eess.SY

V-CAS: A Realtime Vehicle Anti Collision System Using Vision Transformer on Multi-Camera Streams

This paper introduces a real-time Vehicle Collision Avoidance System (V-CAS) designed to enhance vehicle safety through adaptive braking based on environmental perception. V-CAS leverages the advanced vision-based transformer model RT-DETR, DeepSORT tracking, speed estimation, brake light detection, and an adaptive braking mechanism. It computes a composite collision risk score based on vehicles' relative accelerations, distances, and detected braking actions, using brake light signals and trajectory data from multiple camera streams to improve scene perception. Implemented on the Jetson Orin Nano, V-CAS enables real-time collision risk assessment and proactive mitigation through adaptive braking. A comprehensive training process was conducted on various datasets for comparative analysis, followed by fine-tuning the selected object detection model using transfer learning. The system's effectiveness was rigorously evaluated on the Car Crash Dataset (CCD) from YouTube and through real-time experiments, achieving over 98% accuracy with an average proactive alert time of 1.13 seconds. Results indicate significant improvements in object detection and tracking, enhancing collision avoidance compared to traditional single-camera methods. This research demonstrates the potential of low-cost, multi-camera embedded vision transformer systems to advance automotive safety through enhanced environmental perception and proactive collision avoidance mechanisms.

cs.RO

Inflation with the Chern-Simons term in the Palatini formulation

We consider the Chern--Simons term coupled to the inflaton in the Palatini formulation of general relativity. In contrast to the metric formulation, here the Chern--Simons term affects also the background evolution. We approximately solve for the connection, insert it back into the action, and reduce the order of the equations to obtain an effective theory in the gradient approximation. We consider three cases: when the connection is unconstrained, and when non-metricity or torsion is put to zero. In the first two cases, the inflaton kinetic term is modified with a term proportional to the square of the potential. For polynomial potentials dominated by the highest power of the field, the Chern--Simons term solves the problem that higher order corrections spoil the flatness of the potential. For Higgs inflation, the tensor-to-scalar ratio can be as large as the current observational bound, and the non-minimal coupling to the Ricci scalar can be as small as in the metric case. The Palatini contribution cures the known instability of the tensor modes due to the Chern--Simons term in the metric formulation.

astro-ph.CO

"hasSignification()": une nouvelle fonction de distance pour soutenir la détection de données personnelles

Today with Big Data and data lakes, we are faced of a mass of data that is very difficult to manage it manually. The protection of personal data in this context requires an automatic analysis for data discovery. Storing the names of attributes already analyzed in a knowledge base could optimize this automatic discovery. To have a better knowledge base, we should not store any attributes whose name does not make sense. In this article, to check if the name of an attribute has a meaning, we propose a solution that calculate the distances between this name and the words in a dictionary. Our studies on the distance functions like N-Gram, Jaro-Winkler and Levenshtein show limits to set an acceptance threshold for an attribute in the knowledge base. In order to overcome these limitations, our solution aims to strengthen the score calculation by using an exponential function based on the longest sequence. In addition, a double scan in dictionary is also proposed in order to process the attributes which have a compound name.

cs.CL