SearcharxivSearch

arXiv subjects

Nitin Sharma

Publications and source records attributed to Nitin Sharma.

8 recordsLinked to original sources

NEMO-4-PAYPAL: Leveraging NVIDIA's Nemo Framework for empowering PayPal's Commerce Agent

We present the development and optimization of PayPal's Commerce Agent, powered by NEMO-4-PAYPAL, a multi-agent system designed to revolutionize agentic commerce on the PayPal platform. Through our strategic partnership with NVIDIA, we leveraged the NeMo Framework for LLM model fine-tuning to enhance agent performance. Specifically, we optimized the Search and Discovery agent by replacing our base model with a fine-tuned Nemotron small language model (SLM). We conducted comprehensive experiments using the llama3.1-nemotron-nano-8B-v1 architecture, training LoRA-based models through systematic hyperparameter sweeps across learning rates, optimizers (Adam, AdamW), cosine annealing schedules, and LoRA ranks. Our contributions include: (1) the first application of NVIDIA's NeMo Framework to commerce-specific agent optimization, (2) LLM powered fine-tuning strategy for retrieval-focused commerce tasks, (3) demonstration of significant improvements in latency and cost while maintaining agent quality, and (4) a scalable framework for multi-agent system optimization in production e-commerce environments. Our results demonstrate that the fine-tuned Nemotron SLM effectively resolves the key performance issue in the retrieval component, which represents over 50\% of total agent response time, while maintaining or enhancing overall system performance.

cs.AI

From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise

Accurate domain-specific benchmarking of LLMs is essential, specifically in domains with direct implications for humans, such as law, healthcare, and education. However, existing benchmarks are documented to be contaminated and are based on multiple-choice questions, which suffer from inherent biases. To measure domain-specific knowledge in LLMs, we present a deterministic pipeline that transforms raw domain corpora into completion-style benchmarks without relying on other LLMs or costly human annotation. Our approach first extracts domain-specific keywords and related target vocabulary from an input corpus. It then constructs prompt-target pairs where domain-specific words serve as prediction targets. By measuring LLMs' ability to complete these prompts, we provide a direct assessment of domain knowledge at low computational cost. Our pipeline avoids benchmark contamination, enables automated updates with new domain data, and facilitates fair comparisons between base and instruction-tuned (chat) models. We validate our approach by showing that model performances on our benchmark significantly correlate with those on an expert-curated benchmark. We then demonstrate how our benchmark provides insights into knowledge acquisition in domain-adaptive, continual, and general pretraining. Finally, we examine the effects of instruction fine-tuning by comparing base and chat models within our unified evaluation framework. In conclusion, our pipeline enables scalable, domain-specific, LLM-independent, and unbiased evaluation of both base and chat models.

cs.CL

Koopman-Based Model Predictive Control of Functional Electrical Stimulation for Ankle Dorsiflexion and Plantarflexion Assistance

Functional Electrical Stimulation (FES) can be an effective tool to augment paretic muscle function and restore normal ankle function. Our approach incorporates a real-time, data-driven Model Predictive Control (MPC) scheme, built upon a Koopman operator theory (KOT) framework. This framework adeptly captures the complex nonlinear dynamics of ankle motion in a linearized form, enabling application of linear control approaches for highly nonlinear FES-actuated dynamics. Utilizing inertial measurement units (IMUs), our method accurately predicts the FES-induced ankle movements, while accounting for nonlinear muscle actuation dynamics, including the muscle activation for both plantarflexors, and dorsiflexors (Tibialis Anterior (TA)). The linear prediction model derived through KOT allowed us to formulate the MPC problem with linear state space dynamics, enhancing the real-time feasibility, precision and adaptability of the FES driven control. The effectiveness and applicability of our approach have been demonstrated through comprehensive simulations and experimental trials, including three participants with no disability and a participant with Multiple Sclerosis. Our findings highlight the potential of a KOT-based MPC approach for FES based gait assistance that offers effective and personalized assistance for individuals with gait impairment conditions.

eess.SY

Investigating Continual Pretraining in Large Language Models: Insights and Implications

Continual learning (CL) in large language models (LLMs) is an evolving domain that focuses on developing efficient and sustainable training strategies to adapt models to emerging knowledge and achieve robustness in dynamic environments. Our primary emphasis is on continual domain-adaptive pretraining, a process designed to equip LLMs with the ability to integrate new information from various domains while retaining previously learned knowledge. Since existing works concentrate mostly on continual fine-tuning for a limited selection of downstream tasks or training domains, we introduce a new benchmark designed to measure the adaptability of LLMs to changing pretraining data landscapes. We further examine the impact of model size on learning efficacy and forgetting, as well as how the progression and similarity of emerging domains affect the knowledge transfer within these models. Our findings uncover several key insights: (i) continual pretraining consistently improves <1.5B models studied in this work and is also superior to domain adaptation, (ii) larger models always achieve better perplexity than smaller ones when continually pretrained on the same corpus, (iii) smaller models are particularly sensitive to continual pretraining, showing the most significant rates of both learning and forgetting, (iv) continual pretraining boosts downstream task performance of GPT-2 family, (v) continual pretraining enables LLMs to specialize better when the sequence of domains shows semantic similarity while randomizing training domains leads to better transfer and final performance otherwise. We posit that our research establishes a new benchmark for CL in LLMs, providing a more realistic evaluation of knowledge retention and transfer across diverse domains.

cs.CL

Variational quantum algorithm for unconstrained black box binary optimization: Application to feature selection

We introduce a variational quantum algorithm to solve unconstrained black box binary optimization problems, i.e., problems in which the objective function is given as black box. This is in contrast to the typical setting of quantum algorithms for optimization where a classical objective function is provided as a given Quadratic Unconstrained Binary Optimization problem and mapped to a sum of Pauli operators. Furthermore, we provide theoretical justification for our method based on convergence guarantees of quantum imaginary time evolution. To investigate the performance of our algorithm and its potential advantages, we tackle a challenging real-world optimization problem: feature selection. This refers to the problem of selecting a subset of relevant features to use for constructing a predictive model such as fraud detection. Optimal feature selection -- when formulated in terms of a generic loss function -- offers little structure on which to build classical heuristics, thus resulting primarily in 'greedy methods'. This leaves room for (near-term) quantum algorithms to be competitive to classical state-of-the-art approaches. We apply our quantum-optimization-based feature selection algorithm, termed VarQFS, to build a predictive model for a credit risk data set with 20 and 59 input features (qubits) and train the model using quantum hardware and tensor-network-based numerical simulations, respectively. We show that the quantum method produces competitive and in certain aspects even better performance compared to traditional feature selection techniques used in today's industry.

quant-ph

A Supplementary Condition for the Convergence of the Control Policy during Adaptive Dynamic Programming

Reinforcement learning based adaptive/approximate dynamic programming (ADP) is a powerful technique to determine an approximate optimal controller for a dynamical system. These methods bypass the need to analytically solve the nonlinear Hamilton-Jacobi-Bellman equation, whose solution is often to difficult to determine but is needed to determine the optimal control policy. ADP methods usually employ a policy iteration algorithm that evaluates and improves a value function at every step to find the optimal control policy. Previous works in ADP have been lacking a stronger condition that ensures the convergence of the policy iteration algorithm. This paper provides a sufficient but not necessary condition that guarantees the convergence of an ADP algorithm. This condition may provide a more solid theoretical framework for ADP-based control algorithm design for nonlinear dynamical systems.

math.OC

A Theoretical Difficulty in Approximate Dynamic Programming with Input Constraints

Equipping approximate dynamic programming (ADP) with inputconstraints has a tremendous significance. This enables ADP to be applied tothe systems with actuator limitations, which is quite common for dynamicalsystems. In a conventional constrained ADP framework, the optimal control issearched via a policy iteration algorithm, where the value under a constrainedcontrol is solved from a Hamilton-Jacobi-Bellman (HJB) equation while theconstrained control policy is improved based on the current estimated value.This concise and applicable method has been widely-used. However, the con-vergence of the existing policy iteration algorithm may possesses a theoreticaldifficulty, which might be caused by forcibly evaluating the same trajectoryeven though the control policy has already changed. This problem will beexplored in this paper.

math.OC

Mathematical Modeling to Study the Dynamics of A Diatomic Molecule N2 in Water

In the present work an attempt has been made to study the dynamics of a diatomic molecule N2 in water. The proposed model consists of Langevin stochastic differential equation whose solution is obtained through Euler's method. The proposed work has been concluded by studying the behavior of statistical parameters like variance in position, variance in velocity and covariance between position and velocity. This model incorporates the important parameters like acceleration, intermolecular force, frictional force and random force.

cs.CE