Searcharxiv⌕ Search

arXiv subjects

Amitabh Basu

Publications and source records attributed to Amitabh Basu.

At least 37 records · Page 2Linked to original sources

Helly systems and certificates in optimization

Inspired by branch-and-bound and cutting plane proofs in mixed-integer optimization and proof complexity, we develop a general approach via Hoffman's Helly systems. This helps to distill the main ideas behind optimality and infeasibility certificates in optimization. The first part of the paper formalizes the notion of a certificate and its size in this general setting. The second part of the paper establishes lower and upper bounds on the sizes of these certificates in various different settings. We show that some important techniques existing in the literature are purely combinatorial in nature and do not depend on any underlying geometric notions.

math.OC↗

Two-halfspace closure

We define a new cutting plane closure for pure integer programs called the two-halfspace closure. It is a natural generalization of the well-known Chvátal-Gomory closure. We prove that the two-halfspace closure is polyhedral. We also study the corresponding $2$-halfpsace rank of any valid inequality and show that it is at most the split rank of the inequality. Moreover, while the split rank can be strictly larger than the two-halfspace rank, the split rank is at most twice the two-halfspace rank. A key step of our analysis shows that the split closure of a rational polyhedron can be obtained by considering the split closures of all $k$-dimensional (rational) projections of the polyhedron, for any fixed $k\geq 2$. This result may be of independent interest.

math.OC↗

Enumerating integer points in polytopes with bounded subdeterminants

We show that one can enumerate the vertices of the convex hull of integer points in polytopes whose constraint matrices have bounded and nonzero subdeterminants, in time polynomial in the dimension and encoding size of the polytope. This extends a previous result by Artmann et al. who showed that integer linear optimization in such polytopes can be done in polynomial time.

math.CO↗

Leveraging semantically similar queries for ranking via combining representations

In modern ranking problems, different and disparate representations of the items to be ranked are often available. It is sensible, then, to try to combine these representations to improve ranking. Indeed, learning to rank via combining representations is both principled and practical for learning a ranking function for a particular query. In extremely data-scarce settings, however, the amount of labeled data available for a particular query can lead to a highly variable and ineffective ranking function. One way to mitigate the effect of the small amount of data is to leverage information from semantically similar queries. Indeed, as we demonstrate in simulation settings and real data examples, when semantically similar queries are available it is possible to gainfully use them when ranking with respect to a particular query. We describe and explore this phenomenon in the context of the bias-variance trade off and apply it to the data-scarce settings of a Bing navigational graph and the Drosophila larva connectome.

cs.LG↗

Complexity of branch-and-bound and cutting planes in mixed-integer optimization -- II

We study the complexity of cutting planes and branching schemes from a theoretical point of view. We give some rigorous underpinnings to the empirically observed phenomenon that combining cutting planes and branching into a branch-and-cut framework can be orders of magnitude more efficient than employing these tools on their own. In particular, we give general conditions under which a cutting plane strategy and a branching scheme give a provably exponential advantage in efficiency when combined into branch-and-cut. The efficiency of these algorithms is evaluated using two concrete measures: number of iterations and sparsity of constraints used in the intermediate linear/convex programs. To the best of our knowledge, our results are the first mathematically rigorous demonstration of the superiority of branch-and-cut over pure cutting planes and pure branch-and-bound.

math.OC↗

Complexity of branch-and-bound and cutting planes in mixed-integer optimization

We investigate the theoretical complexity of branch-and-bound (BB) and cutting plane (CP) algorithms for mixed-integer optimization. In particular, we study the relative efficiency of BB and CP, when both are based on the same family of disjunctions. We extend a result of Dash to the nonlinear setting which shows that for convex 0/1 problems, CP does at least as well as BB, with variable disjunctions. We sharpen this by giving instances of the stable set problem where we can provably establish that CP does exponentially better than BB. We further show that if one moves away from 0/1 sets, this advantage of CP over BB disappears; there are examples where BB finishes in O(1) time, but CP takes infinitely long to prove optimality, and exponentially long to get to arbitrarily close to the optimal value (for variable disjunctions). We next show that if the dimension is considered a fixed constant, then the situation reverses and BB does at least as well as CP (up to a polynomial blow up), no matter which family of disjunctions is used. This is also complemented by examples where this gap is exponential (in the size of the input data).

math.OC↗

Split cuts in the plane

We provide a polynomial time cutting plane algorithm based on split cuts to solve integer programs in the plane. We also prove that the split closure of a polyhedron in the plane has polynomial size.

math.OC↗

Admissibility of solution estimators for stochastic optimization

We look at stochastic optimization problems through the lens of statistical decision theory. In particular, we address admissibility, in the statistical decision theory sense, of the natural sample average estimator for a stochastic optimization problem (which is also known as the empirical risk minimization (ERM) rule in learning literature). It is well known that for some simple stochastic optimization problems, the sample average estimator may not be admissible. This is known as {\em Stein's paradox} in the statistics literature. We show in this paper that for optimizing stochastic linear functions over compact sets, the sample average estimator is admissible. Moreover, we study problems with convex quadratic objectives subject to box constraints. Stein's paradox holds when there are no constraints and the dimension of the problem is at least three. We show that in the presence of box constraints, admissibility is recovered for dimensions 3 and 4.

math.OC↗

Optimal Probabilistic Catalogue Matching for Radio Sources

Cross-matching catalogues from radio surveys to catalogues of sources at other wavelengths is extremely hard, because radio sources are often extended, often consist of several spatially separated components, and often no radio component is coincident with the optical/infrared host galaxy. Traditionally, the cross-matching is done by eye, but this does not scale to the millions of radio sources expected from the next generation of radio surveys. We present an innovative automated procedure, using Bayesian hypothesis testing, that models trial radio-source morphologies with putative positions of the host galaxy. This new algorithm differs from an earlier version by allowing more complex radio source morphologies, and performing a simultaneous fit over a large field. We show that this technique performs well in an unsupervised mode.

astro-ph.IM↗

Robust Registration of Astronomy Catalogs with Applications to the Hubble Space Telescope

Astrometric calibration of images with a small field of view is often inferior to the internal accuracy of the source detections due to the small number of accessible guide stars. One important experiment with such challenges is the Hubble Space Telescope (HST). A possible solution is to cross-calibrate overlapping fields instead of just relying on standard stars. Following the approach of \citet{2012ApJ...761..188B}, we use infinitesimal 3D rotations for fine-tuning the calibration but devise a better objective that is robust to a large number of false candidates in the initial set of associations. Using Bayesian statistics, we accommodate bad data by explicitly modeling the quality, which yields a formalism essentially identical to an $M$-estimation in robust statistics. Our results on simulated and real catalogs show great potentials for improving the HST calibration, and those with similar challenges.

astro-ph.IM↗

Non-unique lifting of integer variables in minimal inequalities

We explore the lifting question in the context of cut-generating functions. Most of the prior literature on this question focuses on cut-generating functions that have the unique lifting property. We develop a general theory for understanding the lifting question for cut-generating functions that do not necessarily have the unique lifting property.

math.OC↗

Can cut generating functions be good and efficient?

Making cut generating functions (CGFs) computationally viable is a central question in modern integer programming research. One would like to find CGFs that are simultaneously good, i.e., there are good guarantees for the cutting planes they generate, and efficient, meaning that the values of the CGFs can be computed cheaply (with procedures that have some hope of being implemented in current solvers). We investigate in this paper to what extent this balance can be struck. We propose a family of CGFs which, in a sense, achieves this harmony between good and efficient. In particular, we provide a parameterized family of $b+\Z^n$ free sets to derive CGFs from and show that our proposed CGFs give a good approximation of the closure given by CGFs obtained from all maximal $b+\Z^n$ free sets and their so-called trivial liftings, and simultaneously, show that these CGFs can be computed with explicit, efficient procedures. We provide a constructive framework to identify these sets as well as computing their trivial lifting. We follow it up with computational experiments to demonstrate this and to evaluate their practical use. Our proposed family of cuts seem to give some tangible improvement on randomly generated instances compared to GMI cuts; however, in MIPLIB 3.0 instances, and vertex cover and stable problems on random graph instances, their performance is poor.

math.OC↗

An exposition of special relativity without appeal to "constancy of speed of light" hypotheses

We present the theory of special relativity here through the lens of differential geometry. In particular, we explicitly avoid any reference to hypotheses of the form "The laws of physics take the same form in all inertial reference frames" and "The speed of light is constant in all inertial reference frames", or to any other electrodynamic phenomenon. For the author, the clearest understanding of relativity comes about when developing the theory out of just the primitive concept of time (which is also a concept inherent in any standard exposition) and the basic tenets of differential geometry. Perhaps surprisingly, once the theory is framed in this way, one can predict existence of a "universal velocity" which stays the same in all "inertial reference frames". This prediction can be made by performing much more basic time measurement physical experiments that we outline in these notes, rather than experiments of an electrodynamic nature. Thus, had these physical experiments been performed prior to Michelson-Morley type experiments (which, in principle, could have been done in any period with precise enough time keeping instruments), the Michelson-Morley experiments would simply give us an example of a physical entity, i.e., light, which enjoys this special "universal" status.

physics.hist-ph↗

Probabilistic Cross-identification of Multiple Catalogs in Crowded Fields

Matching astronomical catalogs in crowded regions of the sky is challenging both statistically and computationally due to the many possible alternative associations. Budavári and Basu (2016) modeled the two-catalog situation as an Assignment Problem and used the famous Hungarian algorithm to solve it. Here we treat cross-identification of multiple catalogs by introducing a different approach based on integer linear programming. We first test this new method on problems with two catalogs and compare with the previous results. We then test the efficacy of the new approach on problems with three catalogs. The performance and scalability of the new approach is discussed in the context of large surveys.

astro-ph.IM↗

Mixed-integer bilevel representability

We study the representability of sets that admit extended formulations using mixed-integer bilevel programs. We show that feasible regions modeled by continuous bilevel constraints (with no integer variables), complementarity constraints, and polyhedral reverse convex constraints are all finite unions of polyhedra. Conversely, any finite union of polyhedra can be represented using any one of these three paradigms. We then prove that the feasible region of bilevel problems with integer constraints exclusively in the upper level is a finite union of sets representable by mixed-integer programs and vice versa. Further, we prove that, up to topological closures, we do not get additional modeling power by allowing integer variables in the lower level as well. To establish the last statement, we prove that the family of sets that are finite unions of mixed-integer representable sets forms an algebra of sets (up to topological closures).

math.OC↗

The structure of the infinite models in integer programming

The infinite models in integer programming can be described as the convex hull of some points or as the intersection of halfspaces derived from valid functions. In this paper we study the relationships between these two descriptions. Our results have implications for corner polyhedra. One consequence is that nonnegative, continuous valid functions suffice to describe corner polyhedra (with or without rational data).

math.OC↗

Optimal cutting planes from the group relaxations

We study quantitative criteria for evaluating the strength of valid inequalities for Gomory and Johnson's finite and infinite group models and we describe the valid inequalities that are optimal for these criteria. We justify and focus on the criterion of maximizing the volume of the nonnegative orthant cut off by a valid inequality. For the finite group model of prime order, we show that the unique maximizer is an automorphism of the {\em Gomory Mixed-Integer (GMI) cut} for a possibly {\em different} finite group problem of the same order. We extend the notion of volume of a simplex to the infinite dimensional case. This is used to show that in the infinite group model, the GMI cut maximizes the volume of the nonnegative orthant cut off by an inequality.

math.OC↗

Understanding Deep Neural Networks with Rectified Linear Units

In this paper we investigate the family of functions representable by deep neural networks (DNN) with rectified linear units (ReLU). We give an algorithm to train a ReLU DNN with one hidden layer to *global optimality* with runtime polynomial in the data size albeit exponential in the input dimension. Further, we improve on the known lower bounds on size (from exponential to super exponential) for approximating a ReLU deep net function by a shallower ReLU net. Our gap theorems hold for smoothly parametrized families of "hard" functions, contrary to countable, discrete families known in the literature. An example consequence of our gap theorems is the following: for every natural number $k$ there exists a function representable by a ReLU DNN with $k^2$ hidden layers and total size $k^3$, such that any ReLU DNN with at most $k$ hidden layers will require at least $\frac{1}{2}k^{k+1}-1$ total nodes. Finally, for the family of $\mathbb{R}^n\to \mathbb{R}$ DNNs with ReLU activations, we show a new lowerbound on the number of affine pieces, which is larger than previous constructions in certain regimes of the network architecture and most distinctively our lowerbound is demonstrated by an explicit construction of a *smoothly parameterized* family of functions attaining this scaling. Our construction utilizes the theory of zonotopes from polyhedral theory.

cs.LG↗