Searcharxiv⌕ Search

arXiv subjects

H. Oliver Gao

Publications and source records attributed to H. Oliver Gao.

12 recordsLinked to original sources

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center (TMC) tasks such as anomaly detection, incident reporting, and traveler information. Deploying multiple such models across TMC functions raises a portfolio question: which model should serve each function, in which deployment mode, and under what shared hardware budget? We formulate this as the Foundation Model Deployment Portfolio (FMDP) problem, a mixed-integer program minimizing total cost of ownership (TCO) subject to per-function quality, latency, and safety constraints over shared GPU capacity. We prove the problem NP-hard by reduction from the 0-1 knapsack problem and propose a polynomial-time greedy heuristic. In an illustrative case study with five TMC functions and 19 candidate (model, mode) pairs, FMDP identifies a mixed portfolio costing $34/mo (97% below the cheapest feasible all-closed-API baseline) by routing four functions to open-source APIs and the one function whose quality floor no open-source model meets to a closed API. Break-even analysis shows that on-premise GPU investment becomes reasonable only above approximately 309 vision queries/hour or if API prices double.

cs.AI↗

Artificial collectives of specialists and generalists excel at different tasks

Collective artificial intelligence, where multiple agents work on shared tasks, holds potential to solve expansive problems in fields from medicine to collective governance. But while prescriptive engineering solutions abound, we lack descriptive scientific understanding of artificial collectives, and therefore principles for how to design resource efficient multi-agent systems. Through systematic experiments with optimizing agents, we characterize how agent interpretive abilities, rationality bounds, and task qualities interact to shape collective performance. Agents range from specialists, with narrow interpretive abilities, to generalists, with broad ones. Collectives of specialists correspond to sparse, centralized networks, while collectives of generalists correspond to dense, decentralized ones. We show that interpretive network properties have small performance effects on average (0.07 standard deviations of performance). However, for specific task qualities, these effects are 4.5 times larger (0.33 sd) and can reach much higher for certain task qualities (1.84 sd). This leads collectives of generalists to perform better on tasks that involve generating, choosing, and coordinating, while collectives of specialists with a few generalist mediators perform better on tasks that involve negotiating. Rationality bounds then moderate these relationships. At loose bounds, specialists outperform generalists through more effective sampling of high-dimensional decision spaces. At tight bounds, generalists outperform specialists through better gradient estimation. A fundamental trade-off between performance and convergence speed emerges at moderate bounds. These findings suggest that multi-agent design could benefit from matching interpretive networks to both task demands and agents' computational limits, with implications for the efficiency and energy costs of multi-agent systems.

cs.MA↗

Emissions-Robust Portfolios

We study portfolio choice when firm-level emissions intensities are measured with error. We introduce a scope-specific penalty operator that rescales asset payoffs as a smooth function of revenue-normalized emissions intensity. Under payoff homogeneity, unit-scale invariance, mixture linearity, and a curvature semigroup axiom, the operator is unique and has the closed form $P^{(m)}_j(r,λ)=\bigl(1-λ/λ_{\max,j}\bigr)^m r$. Combining this operator with norm- and moment-constrained ambiguity sets yields robust mean-variance and CVaR programs with exact linear and second-order cone reformulations and economically interpretable dual variables. In a U.S. large-cap equity universe with monthly rebalancing and uniform transaction costs, the resulting strategy reduces average Scope~1 emissions intensity by roughly 92\% relative to equal weight while exhibiting no statistically detectable reduction in the Sharpe ratio under block-bootstrap inference and no statistically detectable change in average returns under HAC inference. We report the return-emissions Pareto frontier, sensitivity to robustness and turnover constraints, and uncertainty propagation from multiple imputation of emissions disclosures.

q-fin.MF↗

AI-Driven Adaptive Air Transit Network with Modular Aerial Pods

This paper presents an adaptive air transit network leveraging modular aerial pods and artificial intelligence (AI) to address urban mobility challenges. Passenger demand, forecasted from AI models, serves as input parameters for a Mixed-Integer Nonlinear Programming (MINLP) optimization model that dynamically adjusts pod dispatch schedules and train lengths in response to demand variations. The results reveal a complex interplay of factors, including demand levels, headway bounds, train configurations, and fleet sizes, which collectively influence network performance and service quality. The proposed system demonstrates the importance of dynamic adjustments, where modularity mitigates capacity bottlenecks and improves operational efficiency. Additionally, the framework enhances energy efficiency and optimizes resource utilization through flexible and adaptive scheduling. This framework provides a foundation for a responsive and sustainable urban air mobility solution, supporting the shift from static planning to agile, data-driven operations.

math.OC↗

Global Geolocated Realtime Data of Interfleet Urban Transit Bus Idling

Urban transit bus idling is a contributor to ecological stress, economic inefficiency, and medically hazardous health outcomes due to emissions. The global accumulation of this frequent pattern of undesirable driving behavior is enormous. In order to measure its scale, we propose GRD-TRT-BUF-4I (Ground Truth Buffer for Idling) an extensible, realtime detection system that records the geolocation and idling duration of urban transit bus fleets internationally. Using live vehicle locations from General Transit Feed Specification (GTFS) Realtime, the system detects approximately 200,000 idling events per day from over 50 cities across North America, Europe, Oceania, and Asia. This realtime data was created dynamically to serve operational decision-making and fleet management to reduce the frequency and duration of idling events as they occur, as well as to capture its accumulative effects. Civil and Transportation Engineers, Urban Planners, Epidemiologists, Policymakers, and other stakeholders might find this useful for emissions modeling, traffic management, route planning, and other urban sustainability efforts at a variety of geographic and temporal scales.

eess.SY↗

Twenty-five years of random asset exchange modeling

The last twenty-five years have seen the development of a significant literature within the subfield of econophysics which attempts to model economic inequality as an emergent property of stochastic interactions among ensembles of agents. In this article, the literature surrounding this approach to the study of wealth and income distributions, henceforth the "random asset exchange" literature following the terminology of Sinha (2003), is thoroughly reviewed for the first time. The foundational papers of Dragulescu and Yakovenko (2000), Chakraborti and Chakrabarti (2000), and Bouchaud and Mezard (2000) are discussed in detail, and principal canonical models within the random asset exchange literature are established. The most common variations upon these canonical models are enumerated, and significant papers within each kind of modification are introduced. The successes of such models, as well as the limitations of their underlying assumptions, are discussed, and it is argued that the literature should move in the direction of more explicit representations of economic structure and processes to acquire greater explanatory power.

cond-mat.stat-mech↗

A formulation of the relaxation phenomenon for lane changing dynamics in an arbitrary car following model

Lane changing dynamics are an important part of traffic microsimulation and are vital for modeling weaving sections and merge bottlenecks. However, there is often much more emphasis placed on car following and gap acceptance models, whereas lane changing dynamics such as tactical, cooperation, and relaxation models receive comparatively little attention. This paper develops a general relaxation model which can be applied to an arbitrary parametric or nonparametric microsimulation model. The relaxation model modifies car following dynamics after a lane change, when vehicles can be far from equilibrium. Relaxation prevents car following models from reacting too strongly to the changes in space headway caused by lane changing, leading to more accurate and realistic simulated trajectories. We also show that relaxation is necessary for correctly simulating traffic breakdown with realistic values of capacity drop.

eess.SY↗

Variance Reduction for Score Functions Using Optimal Baselines

Many problems involve the use of models which learn probability distributions or incorporate randomness in some way. In such problems, because computing the true expected gradient may be intractable, a gradient estimator is used to update the model parameters. When the model parameters directly affect a probability distribution, the gradient estimator will involve score function terms. This paper studies baselines, a variance reduction technique for score functions. Motivated primarily by reinforcement learning, we derive for the first time an expression for the optimal state-dependent baseline, the baseline which results in a gradient estimator with minimum variance. Although we show that there exist examples where the optimal baseline may be arbitrarily better than a value function baseline, we find that the value function baseline usually performs similarly to an optimal baseline in terms of variance reduction. Moreover, the value function can also be used for bootstrapping estimators of the return, leading to additional variance reduction. Our results give new insight and justification for why value function baselines and the generalized advantage estimator (GAE) work well in practice.

cs.LG↗

Particulate Matter Exposure at a Densely Populated Urban Traffic Intersection and Crosswalk

Exposure to elevated particulate matter pollution is of great concern to both the general public and air quality management agencies. At urban traffic intersections, for example, pedestrians are often at a higher risk of exposure to near-source PM pollution from traffic while waiting on the roadside or while walking in the crosswalk. This study offers an in-depth investigation of pedestrian exposure to PM pollution at an urban traffic intersection. Fixed-site measurements near an urban intersection were conducted to examine the variations in particles of various sizes through traffic signal cycles. This process aids in the identification of major PM dispersion patterns on the roadside. In addition, mobile measurements of pedestrian exposure to PM were conducted across six time intervals that correspond to different segments of a pedestrian's journey when passing through the intersection. Measurement results are used to estimate and compare the cumulative deposited doses of PM by size categories and journey segments for pedestrians at an intersection. Furthermore, comparisons of pedestrian exposure to PM on a sunny day and a cloudy day were analyzed. The results indicate the importance of reducing PM pollution at intersections and provide policymakers with a foundation for possible measures to reduce pedestrian PM exposure at urban traffic intersections.

physics.soc-ph↗

Fast Calibration of Car Following models to Trajectory data using the Adjoint Method

Before a car-following model can be applied in practice, it must first be validated against real data in a process known as calibration. This paper discusses the formulation of calibration as an optimization problem, and compares different algorithms for its solution. The optimization consists of an arbitrary car following model, posed as either an ordinary or delay differential equation, being calibrated to an arbitrary source of trajectory data which may include lane changes. Typically, the calibration problem is solved using gradient free optimization. In this work, the gradient of the optimization problem is derived analytically using the adjoint method. The computational cost of the adjoint method does not scale with the number of model parameters, which makes it more efficient than evaluating the gradient numerically using finite differences. Numerical results are presented which show that quasi-newton algorithms using the adjoint method are significantly faster than a genetic algorithm, and also achieve slightly better accuracy of the calibrated model.

eess.SY↗

Analyzing 50 years of major fog events across the central coastal plain of Israel

This report presents an analysis of 152 major fog events that have been occurring for five decades (1967-2017) across the central coastal plain of Israel. Analysis of the meteorological data shows that fog events in the experimental area predominantly occur under two sets of synoptic conditions - Red Sea Trough (44%) and Ridge (41%), while the incidence of fog events peaks between March and June. In particular, the results obtained indicate a decreasing trend in the number of fog events and their duration over time where the frequency of radiation fog has decreased over time when compared to the incidence of advection fog. This note provides a long-term analysis of data in a region that lacks reliable time series of this length, and highlight important insights for future research.

physics.ao-ph↗

Increasing Traffic Throughput by Controlling Autonomous Vehicles at Low Penetration Rates

Human drivers may behave in an imprecise/unstable manner, leading to traffic oscillations which are harmful to traffic throughput. Recent field experiments have shown that the control of a single autonomous vehicle (AV) can increase traffic throughput on a circular test track, as well as reduce traffic oscillations on straight roads. We consider a mixed traffic environment consisting of humans and autonomous vehicles, where the goal is to find a control policy for the autonomous vehicles which maximizes traffic throughput by preventing oscillations in speed. We formulate this problem as an optimization problem which can be solved using gradient based optimization. Numerical experiments on a circular road show that the optimized control policy improves traffic throughput by 28%.

eess.SY↗