SearcharxivSearch

arXiv subjects

Jin Yao

Publications and source records attributed to Jin Yao.

14 recordsLinked to original sources

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that geometry, represented by point clouds, offers a more natural state space for driving because it explicitly captures spatial structure and the rigid and non-rigid transformations that govern scene evolution while directly aligning with the space in which driving actions are executed. Building on this insight, we introduce \textbf{GeoWAM}, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego trajectories. Extensive open-loop and closed-loop evaluations show that visual geometry world modeling yields substantially stronger driving policies than image-based alternatives, establishing future-geometry prediction as an effective pretraining objective for autonomous driving.

cs.CV

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler and more efficient alternative by representing each video with a single searchable vector. However, general-purpose models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events, such as turning left versus right or accelerating versus decelerating. In this work, we study how to adapt a general-purpose multimodal embedding model to driving-video retrieval. We first fine-tune Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective. While this stage substantially improves overall retrieval, caption supervision alone remains insufficient for fine-grained motion understanding. We therefore introduce TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization. Trajectories serve only as privileged training supervision; retrieval still operates on single-vector video embeddings without ego poses, expert rules, or auxiliary perception outputs. We further construct a driving-video retrieval benchmark from nuReasoning. Experiments show that TraVEL improves motion-centric retrieval across model scales: relative to SFT, it raises longitudinal and lateral mAP by 9.8 and 4.7 points at 2B, with corresponding gains of 7.2 and 1.5 points at 8B. TraVEL thus combines physically grounded supervision with efficient embedding-based search.

cs.CV

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\,m average) and 3-second collision rate (0.18\%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.

cs.CV

LabelAny3D: Label Any Object 3D in the Wild

Detecting objects in 3D space from monocular input is crucial for applications ranging from robotics to scene understanding. Despite advanced performance in the indoor and autonomous driving domains, existing monocular 3D detection models struggle with in-the-wild images due to the lack of 3D in-the-wild datasets and the challenges of 3D annotation. We introduce LabelAny3D, an \emph{analysis-by-synthesis} framework that reconstructs holistic 3D scenes from 2D images to efficiently produce high-quality 3D bounding box annotations. Built on this pipeline, we present COCO3D, a new benchmark for open-vocabulary monocular 3D detection, derived from the MS-COCO dataset and covering a wide range of object categories absent from existing 3D datasets. Experiments show that annotations generated by LabelAny3D improve monocular 3D detection performance across multiple benchmarks, outperforming prior auto-labeling approaches in quality. These results demonstrate the promise of foundation-model-driven annotation for scaling up 3D recognition in realistic, open-world settings.

cs.CV

Open Vocabulary Monocular 3D Object Detection

We propose and study open-vocabulary monocular 3D detection, a novel task that aims to detect objects of any categores in metric 3D space from a single RGB image. Existing 3D object detectors either rely on costly sensors such as LiDAR or multi-view setups, or remain confined to closed vocabularies settings with limited categories, restricting their applicability. We identify two key challenges in this new setting. First, the scarcity of 3D bounding box annotations limits the ability to train generalizable models. To reduce dependence on 3D supervision, we propose a framework that effectively integrates pretrained 2D and 3D vision foundation models. Second, missing labels and semantic ambiguities (\eg, table vs. desk) in existing datasets hinder reliable evaluation. To address this, we design a novel metric that captures model performance while mitigating annotation issues. Our approach achieves state-of-the-art results in zero-shot 3D detection of novel categories as well as in-domain detection on seen classes. We hope our method provides a strong baseline and our evaluation protocol establishes a reliable benchmark for future research.

cs.CV

Structural evolution of iron oxides melts at Earth's outer-core pressures

Oxygen and other light elements comprise up to 5 wt% of the Earth's outer-core, and may significantly influence its physical properties and the operation of the geodynamo. Here we report in situ x-ray diffraction measurements of Fe, Fe + 4.5 FeO (atomic proportion), and Fe2O3 melts at 177-438 GPa, achieved using laser-driven shock compression at an x-ray free-electron laser. The melts exhibit Fe-O coordination numbers between 4.0(0.4) and 4.5(0.4), indicating predominantly four-fold coordination environments. These coordination states are significantly smaller than those of Fe-bearing lower-mantle phases such as bridgmanite and ferropericlase. Shorter Fe-Fe interatomic distances in compressed iron oxide melts drive the denser packing relative to ambient melts, while the structural differences between Fe + 4.5 FeO and Fe2O3 melts under shock indicate that the oxidation state modulates oxygen solubility in liquid Fe. At around 177 GPa (380 km below the core-mantle boundary), Fe2O3 melts exhibit higher Fe-O coordination, suggesting that local variations in oxygen content could contribute to the stratification in the uppermost outer-core inferred from seismological and geomagnetic observations.

cond-mat.mtrl-sci

Breaking the Cycle of Incarceration With Targeted Mental Health Outreach: A Case Study in Machine Learning for Public Policy

Many incarcerated individuals face significant and complex challenges, including mental illness, substance dependence, and homelessness, yet jails and prisons are often poorly equipped to address these needs. With little support from the existing criminal justice system, these needs can remain untreated and worsen, often leading to further offenses and a cycle of incarceration with adverse outcomes both for the individual and for public safety, with particularly large impacts on communities of color that continue to widen the already extensive racial disparities in criminal justice outcomes. Responding to these failures, a growing number of criminal justice stakeholders are seeking to break this cycle through innovative approaches such as community-driven and alternative approaches to policing, mentoring, community building, restorative justice, pretrial diversion, holistic defense, and social service connections. Here we report on a collaboration between Johnson County, Kansas, and Carnegie Mellon University to perform targeted, proactive mental health outreach in an effort to reduce reincarceration rates. This paper describes the data used, our predictive modeling approach and results, as well as the design and analysis of a field trial conducted to confirm our model's predictive power, evaluate the impact of this targeted outreach, and understand at what level of reincarceration risk outreach might be most effective. Through this trial, we find that our model is highly predictive of new jail bookings, with more than half of individuals in the trial's highest-risk group returning to jail in the following year. Outreach was most effective among these highest-risk individuals, with impacts on mental health utilization, EMS dispatches, and criminal justice involvement.

cs.LG

Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models

As Large Language Models (LLMs) become increasingly powerful and accessible to human users, ensuring fairness across diverse demographic groups, i.e., group fairness, is a critical ethical concern. However, current fairness and bias research in LLMs is limited in two aspects. First, compared to traditional group fairness in machine learning classification, it requires that the non-sensitive attributes, in this case, the prompt questions, be the same across different groups. In many practical scenarios, different groups, however, may prefer different prompt questions and this requirement becomes impractical. Second, it evaluates group fairness only for the LLM's final output without identifying the source of possible bias. Namely, the bias in LLM's output can result from both the pretraining and the finetuning. For finetuning, the bias can result from both the RLHF procedure and the learned reward model. Arguably, evaluating the group fairness of each component in the LLM pipeline could help develop better methods to mitigate the possible bias. Recognizing those two limitations, this work benchmarks the group fairness of learned reward models. By using expert-written text from arXiv, we are able to benchmark the group fairness of reward models without requiring the same prompt questions across different demographic groups. Surprisingly, our results demonstrate that all the evaluated reward models (e.g., Nemotron-4-340B-Reward, ArmoRM-Llama3-8B-v0.1, and GRM-llama3-8B-sftreg) exhibit statistically significant group unfairness. We also observed that top-performing reward models (w.r.t. canonical performance metrics) tend to demonstrate better group fairness.

cs.CL

TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning Agents

Reinforcement learning (RL) trains an agent from experiences interacting with the environment. In scenarios where online interactions are impractical, offline RL, which trains the agent using pre-collected datasets, has become popular. While this new paradigm presents remarkable effectiveness across various real-world domains, like healthcare and energy management, there is a growing demand to enable agents to rapidly and completely eliminate the influence of specific trajectories from both the training dataset and the trained agents. To meet this problem, this paper advocates Trajdeleter, the first practical approach to trajectory unlearning for offline RL agents. The key idea of Trajdeleter is to guide the agent to demonstrate deteriorating performance when it encounters states associated with unlearning trajectories. Simultaneously, it ensures the agent maintains its original performance level when facing other remaining trajectories. Additionally, we introduce Trajauditor, a simple yet efficient method to evaluate whether Trajdeleter successfully eliminates the specific trajectories of influence from the offline RL agent. Extensive experiments conducted on six offline RL algorithms and three tasks demonstrate that Trajdeleter requires only about 1.5% of the time needed for retraining from scratch. It effectively unlearns an average of 94.8% of the targeted trajectories yet still performs well in actual environment interactions after unlearning. The replication package and agent parameters are available online.

cs.LG

Machine Unlearning of Pre-trained Large Language Models

This study investigates the concept of the `right to be forgotten' within the context of large language models (LLMs). We explore machine unlearning as a pivotal solution, with a focus on pre-trained models--a notably under-researched area. Our research delineates a comprehensive framework for machine unlearning in pre-trained LLMs, encompassing a critical analysis of seven diverse unlearning methods. Through rigorous evaluation using curated datasets from arXiv, books, and GitHub, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over $10^5$ times more computationally efficient than retraining. Our results show that integrating gradient ascent with gradient descent on in-distribution data improves hyperparameter robustness. We also provide detailed guidelines for efficient hyperparameter tuning in the unlearning process. Our findings advance the discourse on ethical AI practices, offering substantive insights into the mechanics of machine unlearning for pre-trained LLMs and underscoring the potential for responsible AI development.

cs.CL

How Does Risk Hedging Impact Operations? Insights from a Price-Setting Newsvendor Model

If a financial asset's price movement impacts a firm's product demand, the firm can respond to the impact by adjusting its operational decisions. For example, in the automotive industry, car makers decrease the selling prices of fuel-inefficient cars when the oil price rises. Meanwhile, the firm can implement a risk-hedging strategy using the financial asset jointly with its operational decisions. Motivated by this, we develop and solve a general risk-management model integrating risk hedging into a price-setting newsvendor. The optimal hedging strategy is calculated analytically, which leads to an explicit objective function for optimizing price and ``virtual production quantity'' (VPQ). (The latter determines the service level, i.e., the demand fulfillment probability.) We find that hedging generally reduces the optimal price {when the firm sets the target mean return as its production-only maximum expected profit. With the same condition on the target mean return}, hedging also reduces the optimal VPQ when the asset price trend positively impacts product demand; meanwhile, it may increase the VPQ by a small margin when the impact is negative. We construct the return-risk efficient frontier that characterizes the optimal return-risk trade-off. Our numerical study using data from a prominent automotive manufacturer shows that the markdowns in price and reduction in VPQ are small under our model and that the hedging strategy substantially reduces risk without materially reducing operational profit.

q-fin.RM

Weakly Supervised Lesion Localization With Probabilistic-CAM Pooling

Localizing thoracic diseases on chest X-ray plays a critical role in clinical practices such as diagnosis and treatment planning. However, current deep learning based approaches often require strong supervision, e.g. annotated bounding boxes, for training such systems, which is infeasible to harvest in large-scale. We present Probabilistic Class Activation Map (PCAM) pooling, a novel global pooling operation for lesion localization with only image-level supervision. PCAM pooling explicitly leverages the excellent localization ability of CAM during training in a probabilistic fashion. Experiments on the ChestX-ray14 dataset show a ResNet-34 model trained with PCAM pooling outperforms state-of-the-art baselines on both the classification task and the localization task. Visual examination on the probability maps generated by PCAM pooling shows clear and sharp boundaries around lesion regions compared to the localization heatmaps generated by CAM. PCAM pooling is open sourced at https://github.com/jfhealthcare/Chexpert.

cs.CV

Doubly Enhanced Third Harmonic Generation in Metal-Based Silicon Nanodisks

Doubly enhanced third harmonic generation (THG) is realized by the electric dipole resonance (EDR) in silicon nanodisks placed onto a metallic film at near-infrared. By introducing the metal substrate, the perfect electric conductor (PEC) surface effect can be formed to not only enhance the electric field in silicon, but also improve the near-field distribution. Meanwhile, in consideration of the periodic nanostructure, the silicon nanodisk, acting as a high-refractive-index dielectric grating, can generate a propagating surface plasmon resonance (PSPR) on the metal surface. By flexibly manipulating the array period, PSPR can effectively couple with EDR, which produces further enhancement and novel properties for EDR. On account of the dual enhancements from these two effects, it is demonstrated that the total THG conversion efficiency of the proposed nanostructure is raised by more than eight orders of magnitude as compared with that of the all-dielectric nanostructure. In addition, the influence of silicon Kerr effect on the coupling and THG is investigated as well, manifesting an unprecedent THG efficiency around 10^-2 with input internsity 2 GW/cm^2. This work will pave a new way for improving the multipolar mode in metal-dielectric nanosturctures and facilitate its engineered applications in nonlinear optics, e.g. frequency conversion, spectroscopic and biochemical sensing.

physics.optics

n-Type Chalcogenides by Ion Implantation

Carrier-type reversal to enable the formation of semiconductor p-n junctions is a prerequisite for many electronic applications. Chalcogenide glasses are p-type semiconductors and their applications have been limited by the extraordinary difficulty in obtaining n-type conductivity. The ability to form chalcogenide glass p-n junctions could improve the performance of phase-change memory and thermoelectric devices and allow the direct electronic control of nonlinear optical devices. Previously, carrier-type reversal has been restricted to the GeCh (Ch=S, Se, Te) family of glasses, with very high Bi or Pb doping concentrations (5 to 11 at.%) incorporated during high-temperature glass melting. Here we report the first n-type doping of chalcogenide glasses by ion implantation of Bi into GeTe and GaLaSO amorphous films, demonstrating rectification and photocurrent in a Bi-implanted GaLaSO device. The electrical doping effect of Bi is observed at a 100 times lower concentration than for Bi melt-doped GeCh glasses.

cond-mat.mtrl-sci