SearcharxivSearch

arXiv subjects

Rishabh Gupta

Publications and source records attributed to Rishabh Gupta.

At least 19 recordsLinked to original sources

Uncovering expert objectives in production planning via inverse optimization: An industrial case study

Production planning in the manufacturing industry often relies on the use of optimization models, but defining an appropriate objective function can be a challenge. In practice, planners must balance competing goals, manage uncertainty, and account for qualitative business preferences that are difficult to quantify. As a result, many optimization models fail to match expert behavior, limiting trust and adoption. In this work, we propose a data-driven inverse optimization framework to infer the objective function implicitly captured in expert planners' decisions. We formulate the production planning problem as a mixed-integer linear program, where the unknown objective function is represented as a weighted sum of hypothesized cost terms. A suboptimality-loss-based inverse optimization method is then applied to learn the objective weights from historical production plans. The proposed approach is applied to a real industrial case provided by Dow, where the inferred weights reveal that avoiding inventory shortages and maintaining consistent cycle lengths dominate the planners' decision-making. Time- and product-dependent extensions further improve predictive accuracy and uncover evolving priorities. Expert interviews confirm the practical validity of these insights. Overall, this study shows that inverse optimization can transform tacit human expertise into interpretable models, enabling more accurate and trusted decision-support tools for complex industrial systems.

math.OC

An inverse mixed-integer optimization framework for learning interpretable models of expert decision making

Understanding how experts make decisions and being able to transfer that knowledge is important, especially in complex engineering applications. It is highly valuable for training novices, improving the performance of human-machine systems, and potentially enabling fully autonomous systems that perform as well as human experts. However, an expert's decision-making strategy, developed through years of experience, is often not directly accessible, since the implicit preferences and decision rules involved can be difficult to specify explicitly. This has motivated the use of observed decisions made by the expert to learn an interpretable model that captures the expert's decision-making process. In this work, we develop an inverse optimization approach to jointly learn the decision-maker's preferences (or perceived costs) and the decision rules governing their choices. We demonstrate the general applicability of our approach using three case studies that consider a shift assignment problem, a production planning problem, and a real-world routing problem, respectively. Across these case studies, modeling both perceived costs and decision rules leads to better predictions, highlighting the value of the proposed framework and its greater flexibility in capturing and replicating expert decision making.

math.OC

MAGE-HEP: Monte Carlo Analysis and Graphical Environment for High-Energy Physics

Monte Carlo event generators are central to high-energy physics analysis. However, workflows based on handwritten scripts can be difficult to reuse, modify, and reproduce when multiple Monte Carlo models, tune variations, run variations, and output formats are involved. We present MAGE-HEP, short for Monte Carlo Analysis and Graphical Environment for High-Energy Physics, a Graphical User Interface (GUI) driven workflow environment for reproducible Monte Carlo-based analyses in high-energy physics. MAGE-HEP organizes analysis workflows through a project-study-run hierarchy. The project stores the workspace, the study stores the reusable analysis context, and each run represents a controlled execution of that context. The MAGE-HEP Node API provides the analysis-building layer for defining generator configurations, observables, selections, output rules, and generated C++/ROOT analysis code. A study context can be inspected, reused, or exported as a \texttt{.mcx} context bundle, while the project state can be exported as a portable \texttt{.mgp} bundle. The current beta implementation validates the core idea using a PYTHIA8 and ROOT workflow. It includes background execution, manifest-based run tracking, live ROOT inspection, and particle-table summaries for supported output layouts. This paper describes the architecture, workflow, and current beta implementation of MAGE-HEP.

hep-ph

Read, Extract, Classify: A Tool for Smarter Requirements Engineering

This paper presents the ReXCL tool, which automates the extraction and classification processes in requirements engineering, enhancing the software development life-cycle. The tool features two main modules: Extraction, which processes raw requirement documents into a predefined schema using heuristics and predictive modeling, and Classification, which assigns class labels to requirements using adaptive fine-tuning of encoder-based models. The final output can be exported to external requirement engineering tools. Performance evaluations indicate that ReXCL significantly improves efficiency and accuracy in managing requirements, marking a novel approach to automating the schematization of semi-structured requirement documents.

cs.SE

Inferring identified hadron production in $pp$ collisions with physics-informed machine learning at the LHC

Machine learning has become a powerful tool in high-energy collider experiments, which enables the studies based on data-driven approaches to complex reconstruction and regression tasks. The study of identified hadron spectra in pseudorapidity regions beyond detector acceptance, which is limited to mid-rapidity regions, carries important information about particle production, yet remains unmeasured. In this work, we develop a physics-informed neural network, trained on PYTHIA8 $pp$ collisions at $\sqrt{s}=13.6$ TeV, to infer $p_{\rm T}$ spectra of $\pi^{\pm}$, $K^{\pm}$, $p/\bar{p}$, $\Lambda/\bar{\Lambda}$, and $K^{0}_{\mathrm{s}}$ in different rapidity regions. Physics-motivated constraints, including particle yield ratios, spectral shape, and smoothness, are incorporated into the loss function. A staged hyperparameter optimization strategy is used to ensure stability. The model achieves yield uncertainties of ${\sim}1.5\%$, $1.8\%$, and $5.83\%$ in the training, interpolation, and extrapolation regimes, respectively, outperforming XGBoost and LightGBM. It further reproduces key observables such as particle yield ratios, the multiplicity dependence of $\langle p_{\rm T} \rangle$, and kinetic freeze-out parameters, indicating that the model captures the underlying physics and provides reliable predictions beyond the measured phase space.

hep-ph

EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition

How do you build a sign language recognizer that works on a phone? That question drove this work. We built EfficientSign, a lightweight model which takes EfficientNet-B0 and focuses on two attention modules (Squeeze-and-Excitation for channel focus, and a spatial attention layer that focuses on the hand gestures). We tested it against five other approaches on 12,637 images of Indian Sign Language alphabets, all 26 classes, using 5-fold cross-validation. EfficientSign achieves the accuracy of 99.94% (+/-0.05%), which matches the performance of ResNet18's 99.97% accuracy, but with 62% fewer parameters (4.2M vs 11.2M). We also experimented with feeding deep features (1,280-dimensional vectors pulled from EfficientNet-B0's pooling layer) into classical classifiers. SVM achieved the accuracy of 99.63%, Logistic Regression achieved the accuracy of 99.03% and KNN achieved accuracy of 96.33%. All of these blow past the 92% that SURF-based methods managed on a similar dataset back in 2015. Our results show that attention-enhanced learning model provides an efficient and deployable solution for ISL recognition without requiring a massive model or hand-tuned feature pipelines anymore.

cs.CV

Spectral methods: crucial for machine learning, natural for quantum computers?

This article presents an argument for why quantum computers could unlock new methods for machine learning. We argue that spectral methods, in particular those that learn, regularise, or otherwise manipulate the Fourier spectrum of a machine learning model, are often natural for quantum computers. For example, if a generative machine learning model is represented by a quantum state, the Quantum Fourier Transform allows us to manipulate the Fourier spectrum of the state using the entire toolbox of quantum routines, an operation that is usually prohibitive for classical models. At the same time, spectral methods are surprisingly fundamental to machine learning: A spectral bias has recently been hypothesised to be the core principle behind the success of deep learning; support vector machines have been known for decades to regularise in Fourier space, and convolutional neural nets build filters in the Fourier space of images. Could, then, quantum computing open fundamentally different, much more direct and resource-efficient ways to design the spectral properties of a model? We discuss this potential in detail here, hoping to stimulate a direction in quantum machine learning research that puts the question of ``why quantum?'' first.

quant-ph

Entropic signatures of market response under concentrated policy communication

The first 100 days of Donald Trump second presidential term (January 20th - April 30th, 2025) featured policy actions with potential market repercussions, constituting a well-suited case study of a concentrated policy scenario. Here, we provide a first look at this period, rooted in the information theory, by analyzing major stock indices across the Americas, Europe as well as Asia and Oceania. Our approach jointly examines dispersion (standard deviation) and information complexity (entropy), but also employs a sliding window cumulative entropy to localize extreme events. We find a notable decoupling between the first two measures, indicating that entropy is not merely a proxy for amplitude but reflects the diversity of populated outcomes. As such, they allow us to capture both market volatility and narrative constraints, signaling large and coherent moves driven by policy changes. In turn, the cumulative entropy is found to notably increase during regional episodes with high information density, providing effective signatures of such events. We argue that the obtained results indicate short-term globally coupled, yet regionally modulated, market impacts with clear connection to introduced policies. In what follows, the presented entropic framework emerges as an efficient complement to standard methods for characterizing markets under turbulent conditions, with potential to enhance forecasting strategies such as the stochastic modeling.

q-fin.ST

An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention

The rapid advancement of large language models (LLMs) has led to a significant increase in automated tools in the software engineering, capable of performing various code-related tasks such as code generation, completion, and translation. Despite these advancements, its effectiveness is constrained by fixed context lengths, limiting its ability to generalize across long, domain-specific code sequences. To address this challenge, we investigate zero-shot, inference-only methods aimed at improving position encodings and optimizing attention mechanisms. Our goal is to provide a thorough analysis of current approaches that facilitate context length extrapolation in code, particularly in the context of long code completion tasks.

cs.SE

Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models

Large Language Models (LLMs) have revolutionized both general natural language processing and domain-specific applications such as code synthesis, legal reasoning, and finance. However, while prior studies have explored individual model capabilities, a systematic cross-domain comparison that unifies linguistic, reasoning, and code understanding abilities remains underexplored. In this work, we present a comprehensive evaluation of five general-purpose and three code-specific state-of-the-art LLMs across six diverse benchmarks encompassing linguistic competence, mathematical reasoning, and trustworthiness. Additionally, we analyze model behavior on the CoNaLa dataset for code explanation, comparing natural language and code-specialized LLMs. Our findings reveal that models optimized for code (e.g., CodeLLaMA variants) exhibit strong reasoning and syntactic precision, that even for non-coding tasks can show measurable performance gains, in contrast to general-purpose models like Mistral-7B and Llama-3-8B.

cs.SE

LLM Based Long Code Translation using Identifier Replacement

In the domain of software development, LLMs have been utilized to automate tasks such as code translation, where source code from one programming language is translated to another while preserving its functionality. However, LLMs often struggle with long source codes that don't fit into the context window, which produces inaccurate translations. To address this, we propose a novel zero-shot code translation method that incorporates identifier replacement. By substituting user-given long identifiers with generalized placeholders during translation, our method allows the LLM to focus on the logical structure of the code, by reducing token count and memory usage, which improves the efficiency and cost-effectiveness of long code translation. Our empirical results demonstrate that our approach preserves syntactical and hierarchical information and produces translation results with reduced tokens.

cs.SE

ReXCL: A Tool for Requirement Document Extraction and Classification

This paper presents the ReXCL tool, which automates the extraction and classification processes in requirement engineering, enhancing the software development lifecycle. The tool features two main modules: Extraction, which processes raw requirement documents into a predefined schema using heuristics and predictive modeling, and Classification, which assigns class labels to requirements using adaptive fine-tuning of encoder-based models. The final output can be exported to external requirement engineering tools. Performance evaluations indicate that ReXCL significantly improves efficiency and accuracy in managing requirements, marking a novel approach to automating the schematization of semi-structured requirement documents.

cs.SE

Maximal Entropy Formalism and the Restricted Boltzmann Machine

The connection between the Maximum Entropy (MaxEnt) formalism and Restricted Boltzmann Machines (RBMs) is natural, as both give rise to a Boltzmann-like distribution with constraints enforced by Lagrange multipliers, which corresponds to RBM parameters. We integrate RBMs into quantum state tomography (QST) by using them as probabilistic models to approximate quantum states while satisfying MaxEnt constraints. Additionally, we employ polynomially efficient quantum sampling techniques to enhance RBM training, enabling scalable and high-fidelity quantum state reconstruction. This approach provides a computationally efficient framework for applying RBMs to MaxEnt-based quantum tomography. Furthermore, our method applies to the general and previously unaddressed case of reconstructing arbitrary mixed quantum states from incomplete and potentially non-commuting sets of expectations of observables while still ensuring maximal entropy.

quant-ph

Entropy-Assisted Quality Pattern Identification in Finance

Short-term patterns in financial time series form the cornerstone of many algorithmic trading strategies, yet extracting these patterns reliably from noisy market data remains a formidable challenge. In this paper, we propose an entropy-assisted framework for identifying high-quality, non-overlapping patterns that exhibit consistent behavior over time. We ground our approach in the premise that historical patterns, when accurately clustered and pruned, can yield substantial predictive power for short-term price movements. To achieve this, we incorporate an entropy-based measure as a proxy for information gain. Patterns that lead to high one-sided movements in historical data, yet retain low local entropy, are more informative in signaling future market direction. Compared to conventional clustering techniques such as K-means and Gaussian Mixture Models (GMM), which often yield biased or unbalanced groupings, our approach emphasizes balance over a forced visual boundary, ensuring that quality patterns are not lost due to over-segmentation. By emphasizing both predictive purity (low local entropy) and historical profitability, our method achieves a balanced representation of Buy and Sell patterns, making it better suited for short-term algorithmic trading strategies.

q-fin.TR

Learning Large Neighborhood Search for Maritime Inventory Routing Optimization

Maritime inventory routing optimization is an important yet challenging combinatorial optimization problem. We propose a machine learning-based local search approach for finding feasible solutions of large-scale maritime inventory routing optimization problems. Given the combinatorial complexity of the problems, we integrate a graph neural network-based neighborhood selection method to enhance local search efficiency. Our approach enables a structured exploration of different neighborhoods by imitating an optimization-based expert neighborhood selection policy, improving solution quality while maintaining computational efficiency. Through extensive computational experiments on realistic instances, we demonstrate that our method outperforms direct mixed-integer programming as well as benchmark local search approaches in solution time and solution quality.

math.OC

MAIDS: Malicious Agent Identification-based Data Security Model for Cloud Environments

With the vigorous development of cloud computing, most organizations have shifted their data and applications to the cloud environment for storage, computation, and sharing purposes. During storage and data sharing across the participating entities, a malicious agent may gain access to outsourced data from the cloud environment. A malicious agent is an entity that deliberately breaches the data. This information accessed might be misused or revealed to unauthorized parties. Therefore, data protection and prediction of malicious agents have become a demanding task that needs to be addressed appropriately. To deal with this crucial and challenging issue, this paper presents a Malicious Agent Identification-based Data Security (MAIDS) Model which utilizes XGBoost machine learning classification algorithm for securing data allocation and communication among different participating entities in the cloud system. The proposed model explores and computes intended multiple security parameters associated with online data communication or transactions. Correspondingly, a security-focused knowledge database is produced for developing the XGBoost Classifier-based Malicious Agent Prediction (XC-MAP) unit. Unlike the existing approaches, which only identify malicious agents after data leaks, MAIDS proactively identifies malicious agents by examining their eligibility for respective data access. In this way, the model provides a comprehensive solution to safeguard crucial data from both intentional and non-intentional breaches, by granting data to authorized agents only by evaluating the agents behavior and predicting the malicious agent before granting data.

cs.CR

FedMUP: Federated Learning driven Malicious User Prediction Model for Secure Data Distribution in Cloud Environments

Cloud computing is flourishing at a rapid pace. Significant consequences related to data security appear as a malicious user may get unauthorized access to sensitive data which may be misused, further. This raises an alarm-ringing situation to tackle the crucial issue related to data security and proactive malicious user prediction. This article proposes a Federated learning driven Malicious User Prediction Model for Secure Data Distribution in Cloud Environments (FedMUP). This approach firstly analyses user behavior to acquire multiple security risk parameters. Afterward, it employs the federated learning-driven malicious user prediction approach to reveal doubtful users, proactively. FedMUP trains the local model on their local dataset and transfers computed values rather than actual raw data to obtain an updated global model based on averaging various local versions. This updated model is shared repeatedly at regular intervals with the user for retraining to acquire a better, and more efficient model capable of predicting malicious users more precisely. Extensive experimental work and comparison of the proposed model with state-of-the-art approaches demonstrate the efficiency of the proposed work. Significant improvement is observed in the key performance indicators such as malicious user prediction accuracy, precision, recall, and f1-score up to 14.32%, 17.88%, 14.32%, and 18.35%, respectively.

cs.CR

Selective Shot Learning for Code Explanation

Code explanation plays a crucial role in the software engineering domain, aiding developers in grasping code functionality efficiently. Recent work shows that the performance of LLMs for code explanation improves in a few-shot setting, especially when the few-shot examples are selected intelligently. State-of-the-art approaches for such Selective Shot Learning (SSL) include token-based and embedding-based methods. However, these SSL approaches have been evaluated on proprietary LLMs, without much exploration on open-source Code-LLMs. Additionally, these methods lack consideration for programming language syntax. To bridge these gaps, we present a comparative study and propose a novel SSL method (SSL_ner) that utilizes entity information for few-shot example selection. We present several insights and show the effectiveness of SSL_ner approach over state-of-the-art methods across two datasets. To the best of our knowledge, this is the first systematic benchmarking of open-source Code-LLMs while assessing the performances of the various few-shot examples selection approaches for the code explanation task.

cs.SE