SearcharxivSearch

arXiv subjects

Aswin Kannan

Publications and source records attributed to Aswin Kannan.

11 recordsLinked to original sources

Enhancing Model Based Derivative Free Optimization using Direct Search

We consider single and multiobjective simulation-based optimization problems. Simulation-based optimization has traditionally used both model-based and search-based methods, often in isolation. Model-based methods include trust region approaches and Bayesian optimization, while search methods include genetic algorithms and Direct Search-type techniques. In this work, we propose a switching framework that leverages Direct Search methods to enhance the performance of any model-based optimizer. Our contributions are twofold. First, in the single-objective setting, we analyze and prove the asymptotic convergence of the proposed switching approach. Second, motivated by applications in machine learning, we consider both classification and regression problems, where the objectives span accuracy, computational time, algorithmic bias, and sparsity. The models range from complex neural networks and decision trees to simpler KNN-type baselines. For machine learning in particular, we also introduce a warm-starting mechanism using weights from previous hyperparameter or architectural configurations to exploit problem structure and accelerate training. Finally, beyond ML tasks, we evaluate the method on the standard CUTEr test problems and compare its performance against classical Bayesian and trust region solvers. We observe consistently strong numerical performance, suggesting promise for the proposed switching-based approach.

math.OC

Pseudoconvex Problems in Operational Decision Systems: Algorithms for Joint Learning and Optimization

We consider joint optimization and learning problems arising in real-time decision systems. While most existing work focuses primarily on convex, revenue-based objectives, we extend this line of research to multi-objective formulations. In energy systems, for instance, we incorporate metrics such as renewable penetration and generation costs. Our key focus, however, is on a class of problems with a pseudoconvex structure - a natural relaxation of convexity. Representative examples include fractional objectives in energy management and logit-based revenue models in retail. The outer-level problem optimizes these pseudoconvex objectives, while the inner-level problem involves training a machine learning model using historical data. Our contributions are twofold. First, we propose a simultaneous learning-and-optimization framework that iteratively updates both inner- and outer-level variables. Second, we develop convergent algorithms for these problem classes under realistic mathematical assumptions. Using real-world datasets, we evaluate the computational performance of our methods and highlight an important observation: there exist clear trade-offs between inexact learning and computational time when assessing final solution quality.

math.OC

Adaptive Methods for Multiobjective Unit Commitment

This work considers a multiobjective version of the unit commitment problem that deals with finding the optimal generation schedule of a firm, over a period of time and a given electrical network. With growing importance of environmental impact, some objectives of interest include CO2 emission levels and renewable energy penetration, in addition to the standard generation costs. Some typical constraints include limits on generation levels and up/down times on generation units. This further entails solving a multiobjective mixed integer optimization problem. The related literature has predominantly focused on heuristics (like Genetic Algorithms) for solving larger problem instances. Our major intent in this work is to propose scalable versions of mathematical optimization based approaches that help in speeding up the process of estimating the underlying Pareto frontier. Our contributions are computational and rest on two key embodiments. First, we use the notion of both epsilon constraints and adaptive weights to solve a sequence of single objective optimization problems. Second, to ease the computational burden, we propose a Mccormick-type relaxation for quadratic type constraints that arise due to the resulting formulation types. We test the proposed computational framework on real network data from [1,50] and compare the same with standard solvers like Gurobi. Results show a significant reduction in complexity (computational time) when deploying the proposed framework.

math.OC

A Cournot-Nash Model for a Coupled Hydrogen and Electricity Market

We present a novel model of a coupled hydrogen and electricity market on the intraday time scale, where hydrogen gas is used as a storage device for the electric grid. Electricity is produced by renewable energy sources or by extracting hydrogen from a pipeline that is shared by non-cooperative agents. The resulting model is a generalized Nash equilibrium problem. Under certain mild assumptions, we prove that an equilibrium exists. Perspectives for future work are presented.

math.OC

Benefits of Multiobjective Learning in Solar Energy Prediction

While the space of renewable energy forecasting has received significant attention in the last decade, literature has primarily focused on machine learning models that train on only one objective at a time. A host of classification (and regression) tasks in energy markets lead to highly imbalanced training data. Say, to balance reserves, it is natural for market regulators to have a choice to be more/less averse to false negatives (can lead to poor operating efficiency and costs) than to false positives (can lead to market shortfall). Besides accuracy, other metrics like algorithmic bias, RMBE (in regression problems), inferencing time, and model sparsity are also very crucial. This paper is amongst the firsts in the field of renewable energy forecasting that attempts to present a Pareto frontier of solutions (tradeoffs), that answers the question on handling multiple objectives by means of using the XGBoost model (Gradient Boosted Trees). Our proposed algorithm relies on using a sequence of weighted (uniform meshes) single objective model training routines. Real world data examples from the Amherst (Massachusetts, United States) solar energy prediction panels with both triobjective (focus on accuracy) and biojective (focus on fairness/bias) classification instances are considered. Numerical experiments appear promising and clear advantages over single objective methods are seen by observing the spread and variety of solutions (model configurations).

math.OC

Document Structure aware Relational Graph Convolutional Networks for Ontology Population

Ontologies comprising of concepts, their attributes, and relationships are used in many knowledge based AI systems. While there have been efforts towards populating domain specific ontologies, we examine the role of document structure in learning ontological relationships between concepts in any document corpus. Inspired by ideas from hypernym discovery and explainability, our method performs about 15 points more accurate than a stand-alone R-GCN model for this task.

cs.AI

Optimal stochastic extragradient schemes for pseudomonotone stochastic variational inequality problems and their variants

We consider the stochastic variational inequality problem in which the map is expectation-valued in a component-wise sense. Much of the available convergence theory and rate statements for stochastic approximation schemes are limited to monotone maps. However, non-monotone stochastic variational inequality problems are not uncommon and are seen to arise from product pricing, fractional optimization problems, and subclasses of economic equilibrium problems. Motivated by the need to address a broader class of maps, we make the following contributions: (i) We present an extragradient-based stochastic approximation scheme and prove that the iterates converge to a solution of the original problem under either pseudomonotonicity requirements or a suitably defined acute angle condition. Such statements are shown to be generalizable to the stochastic mirror-prox framework; (ii) Under strong pseudomonotonicity, we show that the mean-squared error in the solution iterates produced by the extragradient SA scheme converges at the optimal rate of O(1/k) statements that were hitherto unavailable K in this regime. Notably, we optimize the initial steplength by obtaining an ε-infimum of a discontinuous nonconvex function. Similar statements are derived for mirror-prox generalizations and can accommodate monotone SVIs under a weak-sharpness requirement. Finally, both the asymptotics and the empirical rates of the schemes are studied on a set of variational problems where it is seen that the theoretically specified initial steplength leads to significant performance benefits.

math.OC

Document Structure Measure for Hypernym discovery

Hypernym discovery is the problem of finding terms that have is-a relationship with a given term. We introduce a new context type, and a relatedness measure to differentiate hypernyms from other types of semantic relationships. Our Document Structure measure is based on hierarchical position of terms in a document, and their presence or otherwise in definition text. This measure quantifies the document structure using multiple attributes, and classes of weighted distance functions.

cs.CL

Fine Grained Classification of Personal Data Entities

Entity Type Classification can be defined as the task of assigning category labels to entity mentions in documents. While neural networks have recently improved the classification of general entity mentions, pattern matching and other systems continue to be used for classifying personal data entities (e.g. classifying an organization as a media company or a government institution for GDPR, and HIPAA compliance). We propose a neural model to expand the class of personal data entities that can be classified at a fine grained level, using the output of existing pattern matching systems as additional contextual features. We introduce new resources, a personal data entities hierarchy with 134 types, and two datasets from the Wikipedia pages of elected representatives and Enron emails. We hope these resource will aid research in the area of personal data discovery, and to that effect, we provide baseline results on these datasets, and compare our method with state of the art models on OntoNotes dataset.

cs.CL

FSCNMF: Fusing Structure and Content via Non-negative Matrix Factorization for Embedding Information Networks

Analysis and visualization of an information network can be facilitated better using an appropriate embedding of the network. Network embedding learns a compact low-dimensional vector representation for each node of the network, and uses this lower dimensional representation for different network analysis tasks. Only the structure of the network is considered by a majority of the current embedding algorithms. However, some content is associated with each node, in most of the practical applications, which can help to understand the underlying semantics of the network. It is not straightforward to integrate the content of each node in the current state-of-the-art network embedding methods. In this paper, we propose a nonnegative matrix factorization based optimization framework, namely FSCNMF which considers both the network structure and the content of the nodes while learning a lower dimensional representation of each node in the network. Our approach systematically regularizes structure based on content and vice versa to exploit the consistency between the structure and content to the best possible extent. We further extend the basic FSCNMF to an advanced method, namely FSCNMF++ to capture the higher order proximities in the network. We conduct experiments on real world information networks for different types of machine learning applications such as node clustering, visualization, and multi-class classification. The results show that our method can represent the network significantly better than the state-of-the-art algorithms and improve the performance across all the applications that we consider.

cs.SI

Distributed Stochastic Optimization under Imperfect Information

We consider a stochastic convex optimization problem that requires minimizing a sum of misspecified agentspecific expectation-valued convex functions over the intersection of a collection of agent-specific convex sets. This misspecification is manifested in a parametric sense and may be resolved through solving a distinct stochastic convex learning problem. Our interest lies in the development of distributed algorithms in which every agent makes decisions based on the knowledge of its objective and feasibility set while learning the decisions of other agents by communicating with its local neighbors over a time-varying connectivity graph. While a significant body of research currently exists in the context of such problems, we believe that the misspecified generalization of this problem is both important and has seen little study, if at all. Accordingly, our focus lies on the simultaneous resolution of both problems through a joint set of schemes that combine three distinct steps: (i) An alignment step in which every agent updates its current belief by averaging over the beliefs of its neighbors; (ii) A projected (stochastic) gradient step in which every agent further updates this averaged estimate; and (iii) A learning step in which agents update their belief of the misspecified parameter by utilizing a stochastic gradient step. Under an assumption of mere convexity on agent objectives and strong convexity of the learning problems, we show that the sequences generated by this collection of update rules converge almost surely to the solution of the correctly specified stochastic convex optimization problem and the stochastic learning problem, respectively.

math.OC