SearcharxivSearch

arXiv subjects

Michael Forbes

Publications and source records attributed to Michael Forbes.

13 recordsLinked to original sources

Acceleration Techniques for Learning Optimal Classification Trees with Integer Programming

Decision trees are a popular machine learning model which are traditionally trained by heuristic methods. Massive improvements in computing power and optimisation techniques has led to renewed interest in learning globally optimal decision trees. Empirical evidence shows that optimal classification trees (OCTs) have better out-of-sample performance than heuristic methods. The dominant optimisation paradigms for training OCTs are mixed-integer programming (MIP) and dynamic programming (DP). MIP formulations offer flexibility in the objectives and constraints that are modelled, but suffer from poor scaling in the size of the training dataset and the maximum tree depth. DP models represent the state of the art in scaling for OCTs, but lack some of the flexibility of MIP models. In this paper we present progress on using advanced integer programming methods to integrate ideas from DP models into MIP formulations to begin bridging the scaling gap. Using the existing BendOCT model from the literature as a base model, we introduce valid inequalities, cutting planes, and a primal heuristic to improve the scaling of MIP formulations. We show that these techniques significantly improve the ability of BendOCT to find provably optimal solutions over a wide range of datasets.

math.OC

Hybrid restricted master problem for Boolean matrix factorisation

We present bfact, a Python package for performing accurate low-rank Boolean matrix factorisation (BMF). bfact uses a hybrid combinatorial optimisation approach based on a priori candidate factors generated from clustering algorithms. It selects the best disjoint factors before performing either a second combinatorial or heuristic algorithm to recover the BMF. We show that bfact does particularly well at estimating the true rank of matrices in simulated settings. In real benchmarks, using a collation of single-cell RNA-sequencing datasets from the Human Lung Cell Atlas, we show that bfact achieves strong signal recovery, with a much lower rank.

q-bio.QM

DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit Differentiation

While differentiable control has emerged as a powerful paradigm combining model-free flexibility with model-based efficiency, the iterative Linear Quadratic Regulator (iLQR) remains underexplored as a differentiable component. The scalability of differentiating through extended iterations and horizons poses significant challenges, hindering iLQR from being an effective differentiable controller. This paper introduces DiLQR, a framework that facilitates differentiation through iLQR, allowing it to serve as a trainable and differentiable module, either as or within a neural network. A novel aspect of this framework is the analytical solution that it provides for the gradient of an iLQR controller through implicit differentiation, which ensures a constant backward cost regardless of iteration, while producing an accurate gradient. We evaluate our framework on imitation tasks on famous control benchmarks. Our analytical method demonstrates superior computational performance, achieving up to 128x speedup and a minimum of 21x speedup compared to automatic differentiation. Our method also demonstrates superior learning performance ($10^6$x) compared to traditional neural network policies and better model loss with differentiable controllers that lack exact analytical gradients. Furthermore, we integrate our module into a larger network with visual inputs to demonstrate the capacity of our method for high-dimensional, fully end-to-end tasks. Codes can be found on the project homepage https://sites.google.com/view/dilqr/.

cs.RO

Enhancements of Fragment Based Algorithms for Vehicle Routing Problems

The method of fragments was recently proposed, and its effectiveness has been empirically shown for three specialised pickup and delivery problems. We propose an enhanced fragment algorithm that for the first time, effectively solves the Pickup and Delivery Problem with Time Windows. Additionally, we describe the approach in general terms to exemplify its theoretical applicability to vehicle routing problems without pickup and delivery requirements. We then apply it to the Truck-Based Drone Delivery Routing Problem Problem with Time Windows. The algorithm uses a fragment formulation rather than a route one. The definition of a fragment is problem specific, but generally, they can be thought of as enumerable segments of routes with a particular structure. A resource expanded network is constructed from the fragments and is iteratively updated via dynamic discretization discovery. Additionally, we introduce two new concepts called formulation leveraging and column enumeration for row elimination that are crucial for solving difficult problems. These use the strong linear relaxation of the route formulation to strengthen the fragment formulation. We test our algorithm on instances of the Pickup and Delivery Problem with Time Windows and the Truck-Based Drone Delivery Routing Problem with Time Windows. Our approach is competitive with, or outperforms the state-of-the-art algorithm for both.

math.OC

Active Set methods for solving large sample average approximations of chance constrained optimisation problems

This article describes a novel approach to chance-constrained programming based on the sample average approximation (SAA) method. Recent work focuses on heuristic approximations to the SAA problem and we introduce a novel approach which improves on some existing methods. Our Active Set method allows one to solve SAAs of chance-constrained programs with very large numbers of scenarios quickly. We demonstrate that increasing the number of scenarios is more important than improving accuracy with small numbers of scenarios. We use an example of the portfolio selection problem to demonstrate the relative performance of previous and new methods. Extending the Active Set method to an integer-programming model further highlights its applicability and further improves over previous approaches.

math.OC

Combining Optimisation and Simulation Using Logic-Based Benders Decomposition

Operations research practitioners frequently want to model complicated functions that are are difficult to encode in their underlying optimisation framework. A common approach is to solve an approximate model, and to use a simulation to evaluate the true objective value of one or more solutions. We propose a new approach to integrating simulation into the optimisation model itself. The idea is to run the simulation at each incumbent solution to the master problem. The simulation data is then used to guide the trajectory of the optimisation model itself using logic-based Benders cuts. We test the approach on a class of stochastic resource allocation problems with monotonic performance measures. We derive strong, novel Benders cuts that are provably valid for all problems of the given form. We consider two concrete examples: a nursing home shift scheduling problem, and an airport check in counter allocation problem. While previous papers on these applications could only approximately solve realistic instances, we are able to solve them exactly within a reasonable amount of time. Moreover, while those papers account for the inherent variance of the problem by including estimates of the underlying random variables as model parameters, we are able to compute sample average approximations to optimality with up to 100 scenarios.

math.OC

Disaggregated Benders decomposition and lazy constraints for solving the budget-constrained dynamic uncapacitated facility location and network design problem

We present an approach for solving to optimality the budget-constrained Dynamic Uncapacitated Facility Location and Network Design problem (DUFLNDP). This is a problem where a network must be constructed or expanded and facilities placed in the network, subject to a budget, in order to satisfy a number of demands. With the demands satisfied, the objective is to minimise the running cost of the network and the cost of moving demands to facilities. The problem can be disaggregated over two different sets simultaneously, leading to many smaller models which can be solved more easily. Using disaggregated Benders decomposition and lazy constraints, we solve many instances to optimality that have not previously been solved. We use an analytic procedure to generate Benders optimality cuts which are provably Pareto-optimal.

math.OC

Small hitting-sets for tiny arithmetic circuits or: How to turn bad designs into good

We show that if we can design poly($s$)-time hitting-sets for $\Sigma\wedge^a\Sigma\Pi^{O(\log s)}$ circuits of size $s$, where $a=\omega(1)$ is arbitrarily small and the number of variables, or arity $n$, is $O(\log s)$, then we can derandomize blackbox PIT for general circuits in quasipolynomial time. This also establishes that either E$\not\subseteq$\#P/poly or that VP$\ne$VNP. In fact, we show that one only needs a poly($s$)-time hitting-set against individual-degree $a'=\omega(1)$ polynomials that are computable by a size-$s$ arity-$(\log s)$ $\Sigma\Pi\Sigma$ circuit (note: $\Pi$ fanin may be $s$). Alternatively, we claim that, to understand VP one only needs to find hitting-sets, for depth-$3$, that have a small parameterized complexity. Another tiny family of interest is when we restrict the arity $n=\omega(1)$ to be arbitrarily small. We show that if we can design poly($s,\mu(n)$)-time hitting-sets for size-$s$ arity-$n$ $\Sigma\Pi\Sigma\wedge$ circuits (resp.~$\Sigma\wedge^a\Sigma\Pi$), where function $\mu$ is arbitrary, then we can solve PIT for VP in quasipoly-time, and prove the corresponding lower bounds. Our methods are strong enough to prove a surprising {\em arity reduction} for PIT-- to solve the general problem completely it suffices to find a blackbox PIT with time-complexity $sd2^{O(n)}$. We give several examples of ($\log s$)-variate circuits where a new measure (called cone-size) helps in devising poly-time hitting-sets, but the same question for their $s$-variate versions is open till date: For eg., diagonal depth-$3$ circuits, and in general, models that have a {\em small} partial derivative space. We also introduce a new concept, called cone-closed basis isolation, and provide example models where it occurs, or can be achieved by a small shift.

cs.CC

Disaggregated Benders Decomposition for solving a Network Maintenance Scheduling Problem

We consider a problem concerning a network and a set of maintenance requests to be undertaken. We wish to schedule the maintenance in such a way as to minimise the impact on the total throughput of the network. We apply disaggregated Benders cuts and lazy constraints to solve the problem to optimality, as well as exploring the strengths and weaknesses of the technique. We prove that our Benders cuts are pareto optimal. Solutions to the LP relaxation also provide further valid inequalities to reduce total solve time. We implement these techniques on simulated data presented in previous papers, and compare our solution technique to previous methods and a direct MIP formulation. We prove optimality in many problem instances that have not previously been proven.

math.OC

Column Generation and Lazy Constraints for solving the Liner Ship Fleet Repositioning Problem with cargo flows

We consider an important problem in the shipping industry known as the liner shipping fleet repositioning problem (LSFRP). We examine a public data set for this problem including many instances which have not previously been solved to optimality. We present several improvements on a previous mathematical formulation, however the largest instances still result in models too difficult to solve in reasonable time. The implementation of column generation reduces the model size significantly, allowing all instances to be solved, with some taking two to three hours. A novel application of lazy constraints further reduces the size of the model, and results in all instances being solved to optimality in under four minutes.

math.OC

On the Locality of Codeword Symbols in Non-Linear Codes

Consider a possibly non-linear (n,K,d)_q code. Coordinate i has locality r if its value is determined by some r other coordinates. A recent line of work obtained an optimal trade-off between information locality of codes and their redundancy. Further, for linear codes meeting this trade-off, structure theorems were derived. In this work we give a new proof of the locality / redundancy trade-off and generalize structure theorems to non-linear codes.

cs.IT

Square root Bound on the Least Power Non-residue using a Sylvester-Vandermonde Determinant

We give a new elementary proof of the fact that the value of the least $k^{th}$ power non-residue in an arithmetic progression $\{bn+c\}_{n=0,1...}$, over a prime field $\F_p$, is bounded by $7/\sqrt{5} \cdot b \cdot \sqrt{p/k} + 4b + c$. Our proof is inspired by the so called \emph{Stepanov method}, which involves bounding the size of the solution set of a system of equations by constructing a non-zero low degree auxiliary polynomial that vanishes with high multiplicity on the solution set. The proof uses basic algebra and number theory along with a determinant identity that generalizes both the Sylvester and the Vandermonde determinant.

math.NT

Tensor Rank: Some Lower and Upper Bounds

The results of Strassen and Raz show that good enough tensor rank lower bounds have implications for algebraic circuit/formula lower bounds. We explore tensor rank lower and upper bounds, focusing on explicit tensors. For odd d, we construct field-independent explicit 0/1 tensors T:[n]^d->F with rank at least 2n^(floor(d/2))+n-Theta(d log n). This matches (over F_2) or improves (all other fields) known lower bounds for d=3 and improves (over any field) for odd d>3. We also explore a generalization of permutation matrices, which we denote permutation tensors. We show, by counting, that there exists an order-3 permutation tensor with super-linear rank. We also explore a natural class of permutation tensors, which we call group tensors. For any group G, we define the group tensor T_G^d:G^d->F, by T_G^d(g_1,...,g_d)=1 iff g_1 x ... x g_d=1_G. We give two upper bounds for the rank of these tensors. The first uses representation theory and works over large fields F, showing (among other things) that rank_F(T_G^d)<= |G|^(d/2). We also show that if this upper bound is tight, then super-linear tensor rank lower bounds would follow. The second upper bound uses interpolation and only works for abelian G, showing that over any field F that rank_F(T_G^d)<= O(|G|^(1+log d)log^(d-1)|G|). In either case, this shows that many permutation tensors have far from maximal rank, which is very different from the matrix case and thus eliminates many natural candidates for high tensor rank. We also explore monotone tensor rank. We give explicit 0/1 tensors T:[n]^d->F that have tensor rank at most dn but have monotone tensor rank exactly n^(d-1). This is a nearly optimal separation.

cs.CC