Searcharxiv⌕ Search

arXiv subjects

Daniel Levy

Publications and source records attributed to Daniel Levy.

At least 37 records · Page 2Linked to original sources

Using Multiple Vector Channels Improves E(n)-Equivariant Graph Neural Networks

We present a natural extension to E(n)-equivariant graph neural networks that uses multiple equivariant vectors per node. We formulate the extension and show that it improves performance across different physical systems benchmark tasks, with minimal differences in runtime or number of parameters. The proposed multichannel EGNN outperforms the standard singlechannel EGNN on N-body charged particle dynamics, molecular property predictions, and predicting the trajectories of solar system bodies. Given the additional benefits and minimal additional cost of multi-channel EGNN, we suggest that this extension may be of practical use to researchers working in machine learning for the physical sciences

cs.LG↗

Retail Pricing Format and Rigidity of Regular Prices

We study the price rigidity of regular and sale prices, and how it is affected by pricing formats (pricing strategies). We use data from three large Canadian stores with different pricing formats (Every-Day-Low-Price, Hi-Lo, and Hybrid) that are located within a 1 km radius of each other. Our data contains both the actual transaction prices and actual regular prices as displayed on the store shelves. We combine these data with two generated regular price series (filtered prices and reference prices) and study their rigidity. Regular price rigidity varies with store formats because different format stores treat sale prices differently, and consequently define regular prices differently. Correspondingly, the meanings of price cuts and sale prices vary across store formats. To interpret the findings, we consider the store pricing format distribution across the US.

econ.GN↗

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI's GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit "breakthrough" behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.

cs.CL↗

3 Lessons from Hyperinflationary Periods

Inflation is painful, for firms, customers, employees, and society. But careful study of periods of hyperinflation point to ways that firms can adapt. In particular, companies need to think about how to change prices regularly and cheaply, because constant price changes can ultimately be very, very expensive. And they should consider how to communicate those price changes to customers. Providing clarity and predictability can increase consumer trust and help firms in the long run.

econ.GN↗

Zero-Ending Prices, Cognitive Convenience, and Price Rigidity

We assess the role of cognitive convenience in the popularity and rigidity of 0 ending prices in convenience settings. Studies show that 0 ending prices are common at convenience stores because of the transaction convenience that 0 ending prices offer. Using a large store level retail CPI data, we find that 0 ending prices are popular and rigid at convenience stores even when they offer little transaction convenience. We corroborate these findings with two large retail scanner price datasets from Dominicks and Nielsen. In the Dominicks data, we find that there are more 0 endings in the prices of the items in the front end candies category than in any other category, even though these prices have no effect on the convenience of the consumers check out transaction. In addition, in both Dominicks and Nielsens datasets, we find that 0 ending prices have a positive effect on demand. Ruling out consumer antagonism and retailers use of heuristics in pricing, we conclude that 0 ending prices are popular and rigid, and that they increase demand at convenience settings, not only for their transaction convenience, but also for the cognitive convenience they offer.

econ.GN↗

Potterian Economics

Recent studies in psychology and neuroscience offer systematic evidence that fictional works exert a surprisingly strong influence on readers and have the power to shape their opinions and worldviews. Building on these findings, we study what we term Potterian economics, the economic ideas, insights, and structure, found in Harry Potter books, to assess how the books might affect economic literacy. A conservative estimate suggests that more than 7.3 percent of the world population has read the Harry Potter books, and millions more have seen their movie adaptations. These extraordinary figures underscore the importance of the messages the books convey. We explore the Potterian economic model and compare it to professional economic models to assess the consistency of the Potterian economic principles with the existing economic models. We find that some of the principles of Potterian economics are consistent with economists models. Many other principles, however, are distorted and contain numerous inaccuracies, contradicting professional economists views and insights. We conclude that Potterian economics can teach us about the formation and dissemination of folk economics, the intuitive notions of naive individuals who see market transactions as a zero-sum game, who care about distribution but fail to understand incentives and efficiency, and who think of prices as allocating wealth but not resources or their efficient use.

econ.GN↗

Learning with User-Level Privacy

We propose and analyze algorithms to solve a range of learning tasks under user-level differential privacy constraints. Rather than guaranteeing only the privacy of individual samples, user-level DP protects a user's entire contribution ($m \ge 1$ samples), providing more stringent but more realistic protection against information leaks. We show that for high-dimensional mean estimation, empirical risk minimization with smooth losses, stochastic convex optimization, and learning hypothesis classes with finite metric entropy, the privacy cost decreases as $O(1/\sqrt{m})$ as users provide more samples. In contrast, when increasing the number of users $n$, the privacy cost decreases at a faster $O(1/n)$ rate. We complement these results with lower bounds showing the minimax optimality of our algorithms for mean estimation and stochastic convex optimization. Our algorithms rely on novel techniques for private mean estimation in arbitrary dimension with error scaling as the concentration radius $τ$ of the distribution rather than the entire range.

cs.LG↗

Distributionally Robust Multilingual Machine Translation

Multilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memory-efficiency of deployed models. However, the heavy data imbalance between languages hinders the model from performing uniformly across language pairs. In this paper, we propose a new learning objective for MNMT based on distributionally robust optimization, which minimizes the worst-case expected loss over the set of language pairs. We further show how to practically optimize this objective for large translation corpora using an iterated best response scheme, which is both effective and incurs negligible additional computational cost compared to standard empirical risk minimization. We perform extensive experiments on three sets of languages from two datasets and show that our method consistently outperforms strong baseline methods in terms of average and per-language performance under both many-to-one and one-to-many translation settings.

cs.CL↗

Adapting to Function Difficulty and Growth Conditions in Private Optimization

We develop algorithms for private stochastic convex optimization that adapt to the hardness of the specific function we wish to optimize. While previous work provide worst-case bounds for arbitrary convex functions, it is often the case that the function at hand belongs to a smaller class that enjoys faster rates. Concretely, we show that for functions exhibiting $κ$-growth around the optimum, i.e., $f(x) \ge f(x^*) + λκ^{-1} \|x-x^*\|_2^κ$ for $κ> 1$, our algorithms improve upon the standard ${\sqrt{d}}/{n\varepsilon}$ privacy rate to the faster $({\sqrt{d}}/{n\varepsilon})^{\tfracκ{κ- 1}}$. Crucially, they achieve these rates without knowledge of the growth constant $κ$ of the function. Our algorithms build upon the inverse sensitivity mechanism, which adapts to instance difficulty (Asi & Duchi, 2020), and recent localization techniques in private optimization (Feldman et al., 2020). We complement our algorithms with matching lower bounds for these function classes and demonstrate that our adaptive algorithm is \emph{simultaneously} (minimax) optimal over all $κ\ge 1+c$ whenever $c = Θ(1)$.

cs.LG↗

Large-Scale Methods for Distributionally Robust Optimization

We propose and analyze algorithms for distributionally robust optimization of convex losses with conditional value at risk (CVaR) and $χ^2$ divergence uncertainty sets. We prove that our algorithms require a number of gradient evaluations independent of training set size and number of parameters, making them suitable for large-scale applications. For $χ^2$ uncertainty sets these are the first such guarantees in the literature, and for CVaR our guarantees scale linearly in the uncertainty level rather than quadratically as in previous work. We also provide lower bounds proving the worst-case optimality of our algorithms for CVaR and a penalized version of the $χ^2$ problem. Our primary technical contributions are novel bounds on the bias of batch robust risk estimation and the variance of a multilevel Monte Carlo gradient estimator due to [Blanchet & Glynn, 2015]. Experiments on MNIST and ImageNet confirm the theoretical scaling of our algorithms, which are 9--36 times more efficient than full-batch methods.

math.OC↗

Faster CryptoNets: Leveraging Sparsity for Real-World Encrypted Inference

Homomorphic encryption enables arbitrary computation over data while it remains encrypted. This privacy-preserving feature is attractive for machine learning, but requires significant computational time due to the large overhead of the encryption scheme. We present Faster CryptoNets, a method for efficient encrypted inference using neural networks. We develop a pruning and quantization approach that leverages sparse representations in the underlying cryptosystem to accelerate inference. We derive an optimal approximation for popular activation functions that achieves maximally-sparse encodings and minimizes approximation error. We also show how privacy-safe training techniques can be used to reduce the overhead of encrypted inference for real-world datasets by leveraging transfer learning and differential privacy. Our experiments show that our method maintains competitive accuracy and achieves a significant speedup over previous methods. This work increases the viability of deep learning systems that use homomorphic encryption to protect user privacy.

cs.CR↗

Generalizing Hamiltonian Monte Carlo with Neural Networks

We present a general-purpose method to train Markov chain Monte Carlo kernels, parameterized by deep neural networks, that converge and mix quickly to their target distribution. Our method generalizes Hamiltonian Monte Carlo and is trained to maximize expected squared jumped distance, a proxy for mixing speed. We demonstrate large empirical gains on a collection of simple but challenging distributions, for instance achieving a 106x improvement in effective sample size in one case, and mixing when standard HMC makes no measurable progress in a second. Finally, we show quantitative and qualitative gains on a real-world task: latent-variable generative modeling. We release an open source TensorFlow implementation of the algorithm.

stat.ML↗

Deterministic Policy Optimization by Combining Pathwise and Score Function Estimators for Discrete Action Spaces

Policy optimization methods have shown great promise in solving complex reinforcement and imitation learning tasks. While model-free methods are broadly applicable, they often require many samples to optimize complex policies. Model-based methods greatly improve sample-efficiency but at the cost of poor generalization, requiring a carefully handcrafted model of the system dynamics for each task. Recently, hybrid methods have been successful in trading off applicability for improved sample-complexity. However, these have been limited to continuous action spaces. In this work, we present a new hybrid method based on an approximation of the dynamics as an expectation over the next state under the current policy. This relaxation allows us to derive a novel hybrid policy gradient estimator, combining score function and pathwise derivative estimators, that is applicable to discrete action spaces. We show significant gains in sample complexity, ranging between $1.7$ and $25\times$, when learning parameterized policies on Cart Pole, Acrobot, Mountain Car and Hand Mass. Our method is applicable to both discrete and continuous action spaces, when competing pathwise methods are limited to the latter.

cs.AI↗

Fast Amortized Inference and Learning in Log-linear Models with Randomly Perturbed Nearest Neighbor Search

Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amortized cost. Our idea hinges on using Gumbel random variable perturbations and a pre-computed Maximum Inner Product Search data structure to access the most-likely elements in sublinear amortized time. Our method yields provable runtime and accuracy guarantees. Further, we present empirical experiments on ImageNet and Word Embeddings showing significant speedups for sampling, inference, and learning in log-linear models.

cs.LG↗

Constraining the expansion history of the universe from the red shift evolution of cosmic shear

We present a quantitative analysis of the constraints on the total equation of state parameter that can be obtained from measuring the red shift evolution of the cosmic shear. We compare the constraints that can be obtained from measurements of the spin two angular multipole moments of the cosmic shear to those resulting from the two dimensional and three dimensional power spectra of the cosmic shear. We find that if the multipole moments of the cosmic shear are measured accurately enough for a few red shifts the constraints on the dark energy equation of state parameter improve significantly compared to those that can be obtained from other measurements.

astro-ph.CO↗

Expressing the equation of state parameter in terms of the three dimensional cosmic shear

We study the functional dependence of the spin-weighted angular moments of the two-point correlation function of the three dimensional cosmic shear on the expansion history of the universe. We first express the redshift dependent total equation of state parameter in terms of the growing mode of the gauge invariant metric perturbation in the conformal-Newtonian gauge for the case of adiabatic perturbations. We then express the redshift dependent angular moments of the shear two-point correlation function as an integral in terms of the metric perturbation. We present the final explicit expression for the case of a Harrison-Zeldovich spectrum of primordial perturbations. Our analysis is restricted to the linear regime. We use our results to make a preliminary study of the required sensitivity that will allow cosmic shear observations to add significant information about the expansion history of the universe.

astro-ph↗

Smectic Polymer Vesicles

Polymer vesicles are stable robust vesicles made from block copolymer amphiphiles. Recent progress in the chemical design of block copolymers opens up the exciting possibility of creating a wide variety of polymer vesicles with varying fine structure, functionality and geometry. Polymer vesicles not only constitute useful systems for drug delivery and micro/nano-reactors but also provide an invaluable arena for exploring the ordering of matter on curved surfaces embedded in three dimensions. By choosing suitable liquid-crystalline polymers for one of the copolymer components one can create vesicles with smectic stripes. Smectic order on shapes of spherical topology inevitably possesses topological defects (disclinations) that are themselves distinguished regions for potential chemical functionalization and nucleators of vesicle budding. Here we report on glassy striped polymer vesicles formed from amphiphilic block copolymers in which the hydrophobic block is a smectic liquid crystal polymer containing cholesteryl-based mesogens. The vesicles exhibit two-dimensional smectic order and are ellipsoidal in shape with defects, or possible additional budding into isotropic vesicles, at the poles.

cond-mat.soft↗