SearcharxivSearch

arXiv subjects

B. N. Kausik

Publications and source records attributed to B. N. Kausik.

9 recordsLinked to original sources

Resolving the automation paradox: falling labor share, rising wages

A central socioeconomic concern about Artificial Intelligence is that it will lower wages by depressing the labor share - the fraction of economic output paid to labor. We show that declining labor share is more likely to raise wages. In a competitive economy with constant returns to scale, we prove that the wage-maximizing labor share depends only on the capital-to-labor ratio, implying a non-monotonic relationship between labor share and wages. When labor share exceeds this wage-maximizing level, further automation increases wages even while reducing labor's output share. Using data from the United States and eleven other industrialized countries, we estimate that labor share is too high in all twelve, implying that further automation should raise wages. Moreover, we find that falling labor share accounted for 16\% of U.S. real wage growth between 1954 and 2019. These wage gains notwithstanding, automation-driven shifts in labor share are likely to pose significant social and political challenges.

econ.GN

Scaling Efficient LLMs

Recent LLMs have hundreds of billions of parameters consuming vast resources. Furthermore, the so called "AI scaling law" for transformers suggests that the number of parameters must scale linearly with the size of the data. In response, we inquire into efficient LLMs, i.e. those with the fewest parameters that achieve the desired accuracy on a training corpus. Specifically, by comparing theoretical and empirical estimates of the Kullback-Leibler divergence, we derive a natural AI scaling law that the number of parameters in an efficient LLM scales as $D^γ$ where $D$ is the size of the training data and $ γ\in [0.44, 0.72]$, suggesting the existence of more efficient architectures. Against this backdrop, we propose recurrent transformers, combining the efficacy of transformers with the efficiency of recurrent networks, progressively applying a single transformer layer to a fixed-width sliding window across the input sequence. Recurrent transformers (a) run in linear time in the sequence length, (b) are memory-efficient and amenable to parallel processing in large batches, (c) learn to forget history for language tasks, or accumulate history for long range tasks like copy and selective copy, and (d) are amenable to curriculum training to overcome vanishing gradients. In our experiments, we find that recurrent transformers perform favorably on benchmark tests.

cs.CL

Occam Gradient Descent

Deep learning neural network models must be large enough to adapt to their problem domain, while small enough to avoid overfitting training data during gradient descent. To balance these competing demands, over-provisioned deep learning models such as transformers are trained for a single epoch on large data sets, and hence inefficient with both computing resources and training data. In response to these inefficiencies, we derive a provably good algorithm that can combine any training and pruning methods to simultaneously optimize efficiency and accuracy, identifying conditions that resist overfitting and reduce model size while outperforming the underlying training algorithm. We then use the algorithm to combine gradient descent with magnitude pruning into "Occam Gradient Descent." With respect to loss, compute and model size (a) on image classification benchmarks, linear and convolutional neural networks trained with Occam Gradient Descent outperform traditional gradient descent with or without post-train pruning; (b) on a range of tabular data classification tasks, neural networks trained with Occam Gradient Descent outperform traditional gradient descent, as well as Random Forests; (c) on natural language transformers, Occam Gradient Descent outperforms traditional gradient descent.

cs.LG

Equity Premium in Efficient Markets

Equity premium, the surplus returns of stocks over bonds, has been an enduring puzzle. While numerous prior works approach the problem assuming the utility of money is invariant across contexts, our approach implies that in efficient markets the utility of money is polymorphic, with risk aversion dependent on the information available in each context, i.e. the discount on each future cash flow depends on all information available on that cash flow. Specifically, we prove that in efficient markets, informed investors maximize return on volatility by being risk-neutral with riskless bonds, and risk-averse with equities, thereby resolving the puzzle. We validate our results on historical data with surprising consistency. JEL Classification: C58, G00, G12, G17

econ.GN

Cognitive Aging and Labor Share

Labor share, the fraction of economic output accrued as wages, is inexplicably declining in industrialized countries. Whilst numerous prior works attempt to explain the decline via economic factors, our novel approach links the decline to biological factors. Specifically, we propose a theoretical macroeconomic model where labor share reflects a dynamic equilibrium between the workforce automating existing outputs, and consumers demanding new output variants that require human labor. Industrialization leads to an aging population, and while cognitive performance is stable in the working years it drops sharply thereafter. Consequently, the declining cognitive performance of aging consumers reduces the demand for new output variants, leading to a decline in labor share. Our model expresses labor share as an algebraic function of median age, and is validated with surprising accuracy on historical data across industrialized economies via non-linear stochastic regression.

econ.GN

Long Tails, Automation and Labor

A central question in economics is whether automation will displace human labor and diminish standards of living. Whilst prior works typically frame this question as a competition between human labor and machines, we frame it as a competition between human consumers and human suppliers. Specifically, we observe that human needs favor long tail distributions, i.e., a long list of niche items that are substantial in aggregate demand. In turn, the long tails are reflected in the goods and services that fulfill those needs. With this background, we propose a theoretical model of economic activity on a long tail distribution, where innovation in demand for new niche outputs competes with innovation in supply automation for mature outputs. Our model yields analytic expressions and asymptotes for the shares of automation and labor in terms of just four parameters: the rates of innovation in supply and demand, the exponent of the long tail distribution and an initial value. We validate the model via non-linear stochastic regression on historical US economic data with surprising accuracy.

econ.GN

Psychophysical Machine Learning

The Weber Fechner Law of psychophysics observes that human perception is logarithmic in the stimulus. We present an algorithm for incorporating the Weber Fechner law into loss functions for machine learning, and use the algorithm to enhance the performance of deep learning networks.

cs.LG

Accelerating Machine Learning via the Weber-Fechner Law

The Weber-Fechner Law observes that human perception scales as the logarithm of the stimulus. We argue that learning algorithms for human concepts could benefit from the Weber-Fechner Law. Specifically, we impose Weber-Fechner on simple neural networks, with or without convolution, via the logarithmic power series of their sorted output. Our experiments show surprising performance and accuracy on the MNIST data set within a few training iterations and limited computational resources, suggesting that Weber-Fechner can accelerate machine learning of human concepts.

cs.LG

Income Inequality, Cause and Cure

We argue that the recent growth in income inequality is driven by disparate growth in investment income rather than by disparate growth in wages. Specifically, we present evidence that real wages are flat across a range of professions, doctors, software engineers, auto mechanics and cashiers, while stock ownership favors higher education and income levels. Artificial Intelligence and automation allocate an increased share of job tasks towards capital and away from labor. The rewards of automation accrue to capital, and are reflected in the growth of the stock market with several companies now valued in the trillions. We propose a Deferred Investment Payroll plan to enable all workers to participate in the rewards of automation and analyze the performance of such a plan. JEL Classification: J31, J33, O33

econ.GN