SearcharxivSearch

arXiv subjects

Ruolin Li

Publications and source records attributed to Ruolin Li.

17 recordsLinked to original sources

CatchBench: When Can an Agent Failure Be Caught?

When can an agent failure be caught? An audit is usually limited by the record rather than by the method. CatchBench therefore puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST). Prior benchmarks fix one of these states or vary the telemetry; to our knowledge none scores all three under one task-method interface. Each state admits different questions, so seven task contracts carry their own labels and metrics rather than one leaderboard. Four are evidential; three are Gold-derived mechanism diagnostics. The release scores 72 entrants, from rule scanners and structural models to eleven LLM judges across nine model families (GPT, Claude, Gemini, Gemma, Llama, Qwen, DeepSeek, Mistral, Nova), over 1187 declared configurations and 1162 recorded runs. Most of the arena does not order: 56 of 138 registered contrasts separate, and the rest are published unresolved rather than ranked. The two sharpest results cut against our own data. One rule ignores every name and permission; it flags each capability declared after the first. On one of six configuration sources it reaches a perfect F1, so a score there measures how the corpus was built rather than how well a method reasons. Our admissibility bar then rejected one injected substrate and withheld evidential status from the other. A benchmark number is therefore not interpretable until the process behind its labels is published and tested for the shortcut it may leave. We report both, and regenerate every ordering from released predictions with no model call.

cs.LG

When Altruism Meets Autonomy: Managing Bottleneck Congestion with Strategic Autonomous Vehicles

Weaving ramps are critical bottlenecks in highway networks due to conflicting traffic flows and complex interactions among heterogeneous vehicle types. In mixed-autonomy settings, the presence of controllable autonomous vehicles (AVs) introduces new opportunities to influence system-level outcomes, yet the structural impact of such control remains poorly understood. This paper develops a unified equilibrium framework to capture, predict, and optimize aggregate lane-choice behavior in weaving ramps with heterogeneous vehicle populations. We first formulate a Wardrop-based model capturing the selfish behavior of human-driven vehicles (HDVs) and establish existence, uniqueness, and validity of the resulting equilibrium. We then introduce a Stackelberg--Wardrop formulation in which AVs act as strategic leaders optimizing system performance, while HDVs respond through equilibrium adaptation. The framework is further generalized to incorporate heterogeneous behavioral preferences of HDVs and AVs via a Social Value Orientation (SVO) model. Our analysis reveals a fundamental structural property of mixed-autonomy traffic systems: under selfish HDV behavior, the impact of AV penetration is inherently non-increasing, exhibiting plateau regions where performance remains unchanged and improves only at critical thresholds. These results provide principled guidance for the design of AV control and incentive mechanisms in the presence of selfish human behavior, and demonstrate how strategically controlled autonomous agents can be deployed to induce system-level efficiency gains in mixed-autonomy transportation networks.

eess.SY

Traffic Equilibrium in Mixed-Autonomy Network with Capped Customer Waiting

This paper develops a unified modeling framework to capture the equilibrium-state interactions among ride-hailing companies, travelers, and traffic of mixed-autonomy transportation networks. Our framework integrates four interrelated sub-modules: (i) the operational behavior of representative ride-hailing Mixed-Fleet Traffic Network Companies (MiFleet TNCs) managing autonomous vehicle (AV) and human-driven vehicle (HV) fleets, (ii) traveler mode-choice decisions taking into account travel costs and waiting time, (iii) capped customer waiting times to reflect the option available to travelers not to wait for TNCs' service beyond his/her patience and to resort to existing travel modes, and (iv) a flow-dependent traffic congestion model for travel times. A key modeling feature distinguishes AVs and HVs across the pickup and service (customer-on-board) stages: AVs follow Wardrop pickup routes but may deviate during service under company coordination, whereas HVs operate in the reverse manner. The overall framework is formulated as a Nonlinear Complementarity Problem (NCP), which is equivalent to a Variational Inequality(VI) formulation based on which the existence of a variational equilibrium solution to the traffic model is established. Numerical experiments examine how AV penetration and Wardrop relaxation factors, which bound route deviation, affect company, traveler, and system performance to various degrees. The results provide actionable insights for policymakers on regulating AV adoption and company vehicle deviation behavior in modern-day traffic systems that are fast changing due to the advances in technology and information accessibility.

eess.SY

A Dual-stage Prompt-driven Privacy-preserving Paradigm for Person Re-Identification

With growing concerns over data privacy, researchers have started using virtual data as an alternative to sensitive real-world images for training person re-identification (Re-ID) models. However, existing virtual datasets produced by game engines still face challenges such as complex construction and poor domain generalization, making them difficult to apply in real scenarios. To address these challenges, we propose a Dual-stage Prompt-driven Privacy-preserving Paradigm (DPPP). In the first stage, we generate rich prompts incorporating multi-dimensional attributes such as pedestrian appearance, illumination, and viewpoint that drive the diffusion model to synthesize diverse data end-to-end, building a large-scale virtual dataset named GenePerson with 130,519 images of 6,641 identities. In the second stage, we propose a Prompt-driven Disentanglement Mechanism (PDM) to learn domain-invariant generalization features. With the aid of contrastive learning, we employ two textual inversion networks to map images into pseudo-words representing style and content, respectively, thereby constructing style-disentangled content prompts to guide the model in learning domain-invariant content features at the image level. Experiments demonstrate that models trained on GenePerson with PDM achieve state-of-the-art generalization performance, surpassing those on popular real and virtual Re-ID datasets.

cs.CV

To Stay or to Bypass: Unraveling Mainline Vehicles' Aggregate Strategic Decision-Making at Highway Weaving Ramps

The weaving ramp scenario is a critical bottleneck in highway networks due to conflicting flows and complex interactions among merging, exiting, and through vehicles. In this work, we propose a game-theoretic model to capture and predict the aggregate lane choice behavior of mainline through vehicles as they approach the weaving zone. Faced with potential conflicts from merging and exiting vehicles, mainline vehicles can either bypass the conflict zone by changing to an adjacent lane or stay steadfast in their current lane. Our model effectively captures these strategic choices using a small set of parameters, requiring only limited traffic measurements for calibration. The model's validity is demonstrated through SUMO simulations, achieving high predictive accuracy. The simplicity and flexibility of the proposed framework make it a practical tool for analyzing bottleneck weaving scenarios and informing traffic management strategies.

eess.SY

Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning

Graphs are data structures used to represent irregular networks and are prevalent in numerous real-world applications. Previous methods directly model graph structures and achieve significant success. However, these methods encounter bottlenecks due to the inherent irregularity of graphs. An innovative solution is converting graphs into textual representations, thereby harnessing the powerful capabilities of Large Language Models (LLMs) to process and comprehend graphs. In this paper, we present a comprehensive review of methodologies for applying LLMs to graphs, termed LLM4graph. The core of LLM4graph lies in transforming graphs into texts for LLMs to understand and analyze. Thus, we propose a novel taxonomy of LLM4graph methods in the view of the transformation. Specifically, existing methods can be divided into two paradigms: Graph2text and Graph2token, which transform graphs into texts or tokens as the input of LLMs, respectively. We point out four challenges during the transformation to systematically present existing methods in a problem-oriented perspective. For practical concerns, we provide a guideline for researchers on selecting appropriate models and LLMs for different graphs and hardware constraints. We also identify five future research directions for LLM4graph.

cs.LG

Performance assessment of ADAS in a representative subset of critical traffic situations

As a variety of automated collision prevention systems gain presence within personal vehicles, rating and differentiating the automated safety performance of car models has become increasingly important for consumers, manufacturers, and insurers. In 2023, Swiss Re and partners initiated an eight-month long vehicle testing campaign conducted on a recognized UNECE type approval authority and Euro NCAP accredited proving ground in Germany. The campaign exposed twelve mass-produced vehicle models and one prototype vehicle fitted with collision prevention systems to a selection of safety-critical traffic scenarios representative of United States and European Union accident landscape. In this paper, we compare and evaluate the relative safety performance of these thirteen collision prevention systems (hardware and software stack) as demonstrated by this testing campaign. We first introduce a new scoring system which represents a test system's predicted impact on overall real-world collision frequency and reduction of collision impact energy, weighted based on the real-world relevance of the test scenario. Next, we introduce a novel metric that quantifies the realism of the protocol and confirm that our test protocol is a plausible representation of real-world driving. Finally, we find that the prototype system in its pre-release state outperforms the mass-produced (post-consumer-release) vehicles in the majority of the tested scenarios on the test track.

cs.RO

Realistic Extreme Behavior Generation for Improved AV Testing

This work introduces a framework to diagnose the strengths and shortcomings of Autonomous Vehicle (AV) collision avoidance technology with synthetic yet realistic potential collision scenarios adapted from real-world, collision-free data. Our framework generates counterfactual collisions with diverse crash properties, e.g., crash angle and velocity, between an adversary and a target vehicle by adding perturbations to the adversary's predicted trajectory from a learned AV behavior model. Our main contribution is to ground these adversarial perturbations in realistic behavior as defined through the lens of data-alignment in the behavior model's parameter space. Then, we cluster these synthetic counterfactuals to identify plausible and representative collision scenarios to form the basis of a test suite for downstream AV system evaluation. We demonstrate our framework using two state-of-the-art behavior prediction models as sources of realistic adversarial perturbations, and show that our scenario clustering evokes interpretable failure modes from a baseline AV policy under evaluation.

math.OC

Micro-Expression Recognition by Motion Feature Extraction based on Pre-training

Micro-expressions (MEs) are spontaneous, unconscious facial expressions that have promising applications in various fields such as psychotherapy and national security. Thus, micro-expression recognition (MER) has attracted more and more attention from researchers. Although various MER methods have emerged especially with the development of deep learning techniques, the task still faces several challenges, e.g. subtle motion and limited training data. To address these problems, we propose a novel motion extraction strategy (MoExt) for the MER task and use additional macro-expression data in the pre-training process. We primarily pretrain the feature separator and motion extractor using the contrastive loss, thus enabling them to extract representative motion features. In MoExt, shape features and texture features are first extracted separately from onset and apex frames, and then motion features related to MEs are extracted based on the shape features of both frames. To enable the model to more effectively separate features, we utilize the extracted motion features and the texture features from the onset frame to reconstruct the apex frame. Through pre-training, the module is enabled to extract inter-frame motion features of facial expressions while excluding irrelevant information. The feature separator and motion extractor are ultimately integrated into the MER network, which is then fine-tuned using the target ME data. The effectiveness of proposed method is validated on three commonly used datasets, i.e., CASME II, SMIC, SAMM, and CAS(ME)3 dataset. The results show that our method performs favorably against state-of-the-art methods.

cs.CV

Large Language Models for Mobility Analysis in Transportation Systems: A Survey on Forecasting Tasks

Mobility analysis is a crucial element in the research area of transportation systems. Forecasting traffic information offers a viable solution to address the conflict between increasing transportation demands and the limitations of transportation infrastructure. Predicting human travel is significant in aiding various transportation and urban management tasks, such as taxi dispatch and urban planning. Machine learning and deep learning methods are favored for their flexibility and accuracy. Nowadays, with the advent of large language models (LLMs), many researchers have combined these models with previous techniques or applied LLMs to directly predict future traffic information and human travel behaviors. However, there is a lack of comprehensive studies on how LLMs can contribute to this field. This survey explores existing approaches using LLMs for time series forecasting problems for mobility in transportation systems. We provide a literature review concerning the forecasting applications within transportation systems, elucidating how researchers utilize LLMs, showcasing recent state-of-the-art advancements, and identifying the challenges that must be overcome to fully leverage LLMs in this domain.

cs.LG

A Unified Toll Lane Framework for Autonomous and High-Occupancy Vehicles in Interactive Mixed Autonomy

In this study, we introduce a toll lane framework that optimizes the mixed flow of autonomous and high-occupancy vehicles on freeways, where human-driven and autonomous vehicles of varying commuter occupancy share a segment. Autonomous vehicles, with their ability to maintain shorter headways, boost traffic throughput. Our framework designates a toll lane for autonomous vehicles with high occupancy to use free of charge, while others pay a toll. We explore the lane choice equilibria when all vehicles minimize travel costs, and characterize the equilibria by ranking vehicles by their mobility enhancement potential, a concept we term the mobility degree. Through numerical examples, we demonstrate the framework's utility in addressing design challenges such as setting optimal tolls, determining occupancy thresholds, and designing lane policies, showing how it facilitates the integration of high-occupancy and autonomous vehicles. We also propose an algorithm for assigning rational tolls to decrease total commuter delay and examine the effects of toll non-compliance. Our findings suggest that self-interest-driven behavior mitigates moderate non-compliance impacts, highlighting the framework's resilience. This work presents a pioneering comprehensive analysis of a toll lane framework that emphasizes the coexistence of autonomous and high-occupancy vehicles, offering insights for traffic management improvements and the integration of autonomous vehicles into existing transportation infrastructures.

eess.SY

A Simple Structure For Building A Robust Model

As deep learning applications, especially programs of computer vision, are increasingly deployed in our lives, we have to think more urgently about the security of these applications.One effective way to improve the security of deep learning models is to perform adversarial training, which allows the model to be compatible with samples that are deliberately created for use in attacking the model.Based on this, we propose a simple architecture to build a model with a certain degree of robustness, which improves the robustness of the trained network by adding an adversarial sample detection network for cooperative training. At the same time, we design a new data sampling strategy that incorporates multiple existing attacks, allowing the model to adapt to many different adversarial attacks with a single training.We conducted some experiments to test the effectiveness of this design based on Cifar10 dataset, and the results indicate that it has some degree of positive effect on the robustness of the model.Our code could be found at https://github.com/dowdyboy/simple_structure_for_robust_model .

cs.CV

Employing Altruistic Vehicles at On-ramps to Improve the Social Traffic Conditions

Highway on-ramps are regarded as typical bottlenecks in transportation networks. In previous work, mainline vehicles' selfish lane choice behavior at on-ramps is studied and regarded as one cause leading to on-ramp inefficiency. When on-ramp vehicles plan to merge into the mainline of the highway, mainline vehicles choose to either stay steadfast on the current lane or bypass the merging area by switching to a neighboring lane farther from the on-ramp. Selfish vehicles make the decisions to minimize their own travel delay, which compromises the efficiency of the whole on-ramp. Results in previous work have shown that, if we can encourage a proper portion of mainline vehicles to bypass rather than to stay steadfast, the social traffic conditions can be improved. In this work, we consider employing a proportion of altruistic vehicles among the selfish mainline vehicles to improve the efficiency of the on-ramps. The altruistic vehicles are individual optimizers, making decisions whether to stay steadfast or bypass to minimize their own altruistic cost, which is a weighted sum of the travel delay and their negative impact on other vehicles. We first consider the ideal case that altruistic costs can be perfectly measured by altruistic vehicles. We give the conditions for the proportion of altruistic vehicles and the weight configuration of the altruistic costs, under which the social delay can be decreased or reach the optimal. Subsequently, we consider the impact of uncertainty in the measurement of altruistic costs and we give the optimal weight configuration for altruistic vehicles which minimizes the worst-case social delay under such uncertainty.

eess.SY

A Highway Toll Lane Framework that Unites Autonomous Vehicles and High-occupancy Vehicles

We consider the scenario where human-driven/autonomous vehicles with low/high occupancy are sharing a segment of highway and autonomous vehicles are capable of increasing the traffic throughput by preserving a shorter headway than human-driven vehicles. We propose a toll lane framework where a lane on the highway is reserved freely for autonomous vehicles with high occupancy, which have the greatest capability to increase social mobility, and the other three classes of vehicles can choose to use the toll lane with a toll or use the other regular lanes freely. All vehicles are assumed to be only interested in minimizing their own travel costs. We explore the resulting lane choice equilibria under the framework and establish desirable properties of the equilibria, which implicitly compare high-occupancy vehicles with autonomous vehicles in terms of their capabilities to increase social mobility. We further use numerical examples in the optimal toll design, the occupancy threshold design, and the policy design problems to clarify the various potential applications of this toll lane framework that unites high-occupancy vehicles and autonomous vehicles. To our best knowledge, this is the first work that systematically studies a toll lane framework that unites autonomous vehicles and high-occupancy vehicles on the roads.

eess.SY

Improving Urban Traffic Throughput with Vehicle Platooning: Theory and Experiments

In this paper we present a model-predictive control (MPC) based approach for vehicle platooning in an urban traffic setting. Our primary goal is to demonstrate that vehicle platooning has the potential to significantly increase throughput at intersections, which can create bottlenecks in the traffic flow. To do so, our approach relies on vehicle connectivity: vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication. In particular, we introduce a customized V2V message set which features a velocity forecast, i.e. a prediction on the future velocity trajectory, which enables platooning vehicles to accurately maintain short following distances, thereby increasing throughput. Furthermore, V2I communication allows platoons to react immediately to changes in the state of nearby traffic lights, e.g. when the traffic phase becomes green, enabling additional gains in traffic efficiency. We present our design of the vehicle platooning system, and then evaluate performance by estimating the potential gains in terms of throughput using our results from simulation, as well as experiments conducted with real test vehicles on a closed track. Lastly, we briefly overview our demonstration of vehicle platooning on public roadways in Arcadia, CA.

eess.SY

An Extended Game-Theoretic Model for Aggregate Lane Choice Behavior of Vehicles at Traffic Diverges with a Bifurcating Lane

Road network junctions, such as merges and diverges, often act as bottlenecks that initiate and exacerbate congestion. More complex junction configurations lead to more complex driver behaviors, resulting in aggregate congestion patterns that are more difficult to predict and mitigate. In this paper, we discuss diverge configurations where vehicles on some lanes can enter only one of the downstream roads, but vehicles on other lanes can enter one of several downstream roads. Counterintuitively, these bifurcating lanes, rather than relieving congestion (by acting as a versatile resource that can serve either downstream road as the demand changes), often cause enormous congestion due to lane changing. We develop an aggregate lane--changing model for this situation that is expressive enough to model drivers' choices and the resultant congestion, but simple enough to easily analyze. We use a game-theoretic framework to model the aggregate lane choice behavior of selfish vehicles as a Wardrop equilibrium (an aggregate type of Nash equilibrium). We then establish the existence and uniqueness of this equilibrium. We explain how our model can be easily calibrated using simulation data or real data, and we present results showing that our model successfully predicts the aggregate behavior that emerges from widely-used behavioral lane-changing models. Our model's expressiveness, ease of calibration, and accuracy may make it a useful tool for mitigating congestion at these complex diverges.

cs.GT

A Game Theoretic Macroscopic Model of Bypassing at Traffic Diverges with Applications to Mixed Autonomy Networks

Vehicle bypassing is known to negatively affect delays at traffic diverges. However, due to the complexities of this phenomenon, accurate and yet simple models of such lane change maneuvers are hard to develop. In this work, we present a macroscopic model for predicting the number of vehicles that bypass at a traffic diverge. We take into account the selfishness of vehicles in selecting their lanes; every vehicle selects lanes such that its own cost is minimized. We discuss how we model the costs experienced by the vehicles. Then, taking into account the selfish behavior of the vehicles, we model the lane choice of vehicles at a traffic diverge as a Wardrop equilibrium. We state and prove the properties of Wardrop equilibrium in our model. We show that there always exists an equilibrium for our model. Moreover, unlike most nonlinear asymmetrical routing games, we prove that the equilibrium is unique under mild assumptions. We discuss how our model can be easily calibrated by running a simple optimization problem. Using our calibrated model, we validate it through simulation studies and demonstrate that our model successfully predicts the aggregate lane change maneuvers that are performed by vehicles for bypassing at a traffic diverge. We further discuss how our model can be employed to obtain the optimal lane choice behavior of the vehicles, where the social or total cost of vehicles is minimized. Finally, we demonstrate how our model can be utilized in scenarios where a central authority can dictate the lane choice and trajectory of certain vehicles so as to increase the overall vehicle mobility at a traffic diverge. Examples of such scenarios include the case when both human driven and autonomous vehicles coexist in the network. We show how certain decisions of the central authority can affect the total delays in such scenarios via an example.

cs.GT