SearcharxivSearch

arXiv subjects

Kuan-Hui Lee

Publications and source records attributed to Kuan-Hui Lee.

At least 19 recordsLinked to original sources

Pluriclosed 3-folds with vanishing Bismut Ricci form: General theory in the quasi-regular case

We study compact complex $3$-dimensional non-Kähler Bismut Ricci flat pluriclosed Hermitian manifolds (BHE) via their dimensional reduction to a special Kähler geometry in complex dimension $2$, recently obtained by Barbaro, Streets and the first and third authors. We show that in the quasi-regular case, the reduced geometry satisfies a 6th order non-linear PDE which has infinite dimensional momentum map interpretation, similar to the much studied Kähler metrics of constant scalar curvature (cscK). We use this to associate to the reduced manifold or orbifold Mabuchi and Calabi functionals, as well as to obtain obstructions for the existence of solutions in terms of the authomorphism group, paralleling results by Futaki and Calabi-Lichnerowicz-Matsushima in the cscK case. This is used to characterize the Samelson locally homogeneous BHE geometries in complex dimension 3 as the only non-Kähler BHE $3$-folds with $2$-dimensional Bott-Chern $(1,1)$-cohomology group, for which the reduced space is a smooth Kähler surface. We also discuss explicit solutions of the PDE on orthotoric Kähler orbifold surfaces, extending examples found by Couzens-Gauntlett-Martelli-Sparks in the framework of supersymmetric ${\rm AdS}_3 \times Y_7$ type IIB supergravity. Our construction yields infinitely many non-Kähler BHE structures on $S^3\times S^3$ and $S^1\times S^2 \times S^3$, which are not locally isometric to a Samelson geometry. These appear to be the first such examples.

math.DG

Rigidity results for non-Kähler Calabi-Yau geometries on threefolds

We derive a canonical symmetry reduction associated to a compact non-Kähler Bismut-Hermitian-Einstein manifold. In real dimension $6$, the transverse geometry is conformally Kähler, and we give a complete description in terms of a single scalar PDE for the underlying Kähler structure. In the case when the soliton potential is constant, we show that that the Bott-Chern number $h^{1,1}_{BC} \geq 2$, and that equality holds if and only if the metric is Bismut-flat, and hence a quotient of either $\SU(2) \times \mathbb R \times \mathbb C$ or $\SU(2) \times \SU(2)$.

math.DG

Stability of Hyperkähler Flow

In this work, we discuss the stability of Donaldson's flow of surfaces in a hyperkähler 4-manifold. In \cite{WT2}, Wang and Tsai proved a uniqueness theorem and $C^1$ dynamic stability theorem of the mean curvature flow for minimal surface. We extend their results and obtain a similar dynamic stability theorem of the hyperkähler flow.

math.DG

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Robot manipulation has seen tremendous progress in recent years, with imitation learning policies enabling successful performance of dexterous and hard-to-model tasks. Concurrently, scaling data and model size has led to the development of capable language and vision foundation models, motivating large-scale efforts to create general-purpose robot foundation models. While these models have garnered significant enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting both the pace of development and inhibiting a nuanced understanding of current capabilities. In this paper, we rigorously evaluate multitask robot manipulation policies, referred to as Large Behavior Models (LBMs), by extending the Diffusion Policy paradigm across a corpus of simulated and real-world robot data. We propose and validate an evaluation pipeline to rigorously analyze the capabilities of these models with statistical confidence. We compare against single-task baselines through blind, randomized trials in a controlled setting, using both simulation and real-world experiments. We find that multi-task pretraining makes the policies more successful and robust, and enables teaching complex new tasks more quickly, using a fraction of the data when compared to single-task baselines. Moreover, performance predictably increases as pretraining scale and diversity grows. Project page: https://toyotaresearchinstitute.github.io/lbm1/

cs.RO

Dynamical stability of Pluriclosed and Generalized Ricci solitons

In this work, we discuss the stability of the pluriclosed flow and generalized Ricci flow. We proved that if the second variation of generalized Einstein--Hilbert functional is nonpositive and the infinitesimal deformations are integrable, the flow is dynamically stable. Moreover, we prove that the pluriclosed steady solitons are dynamically stable when the first Chern class vanishes.

math.DG

The stability of non-Kähler Calabi-Yau metrics

Non Kähler Calabi Yau theory is a newly developed subject and it arises naturally in mathematical physics and generalized geometry. The relevant geometrics are pluriclosed metrics which are critical points of the generalized Einstein Hilbert action. In this work, we study the critical points of the generalized Einstein Hilbert action and discuss the stability of critical points which are defined as pluriclosed steady solitons. We proved that all Bismut Hermitian Einstein manifolds are linearly stable which generalizes the work from Tian, Zhu, Hall, Murphy and Koiso In addition, all Bismut flat pluriclosed steady solitons with positive Ricci curvature are linearly strictly stable.

math.DG

Stability and moduli space of generalized Ricci solitons

The generalized Einstein Hilbert action is an extension of the classic scalar curvature energy and Perelman F functional which incorporates a closed three-form. The critical points are known as generalized Ricci solitons, which arise naturally in mathematical physics, complex geometry, and generalized geometry. Through a delicate analysis of the group of generalized gauge transformations, and implementing a novel connection, we give a simple formula for the second variation of this energy which generalizes the Lichnerowicz operator in the Einstein case. As an application, we show that all Bismut flat manifolds are linearly stable critical points, and admit nontrivial deformations arising from Lie theory. Furthermore, this leads to extensions of classic results of Koiso and Podesta, Spiro, Kröncke to the moduli space of generalized Ricci solitons. To finish we classify deformations of the Bismut-flat structure on S3 and show that some are integrable while others are not.

math.DG

The Stability of Generalized Ricci Solitons

In this paper, I computed the second variation formula of the generalized Einstein-Hilbert functional and prove that a Bismut-flat, Einstein manifold is linearly stable under some curvature assumption. In the last part of the paper, I prove that dynamical stability and linear stability are equivalent on a steady gradient generalized Ricci soliton $(g, H,f)$ which generalizes the result done by Kröncke, Haslhofer, Sesum, Raffero, and Vezzoni.

math.DG

Group Distributionally Robust Reinforcement Learning with Hierarchical Latent Variables

One key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task indicators. Robust RL has been applied to deal with task ambiguity, but may result in over-conservative policies. To balance the worst-case (robustness) and average performance, we propose Group Distributionally Robust Markov Decision Process (GDR-MDP), a flexible hierarchical MDP formulation that encodes task groups via a latent mixture model. GDR-MDP identifies the optimal policy that maximizes the expected return under the worst-possible qualified belief over task groups within an ambiguity set. We rigorously show that GDR-MDP's hierarchical structure improves distributional robustness by adding regularization to the worst possible outcomes. We then develop deep RL algorithms for GDR-MDP for both value-based and policy-based RL methods. Extensive experiments on Box2D control tasks, MuJoCo benchmarks, and Google football platforms show that our algorithms outperform classic robust training algorithms across diverse environments in terms of robustness under belief uncertainties. Demos are available on our project page (\url{https://sites.google.com/view/gdr-rl/home}).

cs.LG

Learning Optical Flow, Depth, and Scene Flow without Real-World Labels

Self-supervised monocular depth estimation enables robots to learn 3D perception from raw video streams. This scalable approach leverages projective geometry and ego-motion to learn via view synthesis, assuming the world is mostly static. Dynamic scenes, which are common in autonomous driving and human-robot interaction, violate this assumption. Therefore, they require modeling dynamic objects explicitly, for instance via estimating pixel-wise 3D motion, i.e. scene flow. However, the simultaneous self-supervised learning of depth and scene flow is ill-posed, as there are infinitely many combinations that result in the same 3D point. In this paper we propose DRAFT, a new method capable of jointly learning depth, optical flow, and scene flow by combining synthetic data with geometric self-supervision. Building upon the RAFT architecture, we learn optical flow as an intermediate task to bootstrap depth and scene flow learning via triangulation. Our algorithm also leverages temporal and geometric consistency losses across tasks to improve multi-task learning. Our DRAFT architecture simultaneously establishes a new state of the art in all three tasks in the self-supervised monocular setting on the standard KITTI benchmark. Project page: https://sites.google.com/tri.global/draft.

cs.CV

Heterogeneous-Agent Trajectory Forecasting Incorporating Class Uncertainty

Reasoning about the future behavior of other agents is critical to safe robot navigation. The multiplicity of plausible futures is further amplified by the uncertainty inherent to agent state estimation from data, including positions, velocities, and semantic class. Forecasting methods, however, typically neglect class uncertainty, conditioning instead only on the agent's most likely class, even though perception models often return full class distributions. To exploit this information, we present HAICU, a method for heterogeneous-agent trajectory forecasting that explicitly incorporates agents' class probabilities. We additionally present PUP, a new challenging real-world autonomous driving dataset, to investigate the impact of Perceptual Uncertainty in Prediction. It contains challenging crowded scenes with unfiltered agent class probabilities that reflect the long-tail of current state-of-the-art perception systems. We demonstrate that incorporating class probabilities in trajectory forecasting significantly improves performance in the face of uncertainty, and enables new forecasting capabilities such as counterfactual predictions.

cs.CV

CoCon: Cooperative-Contrastive Learning

Labeling videos at scale is impractical. Consequently, self-supervised visual representation learning is key for efficient video analysis. Recent success in learning image representations suggests contrastive learning is a promising framework to tackle this challenge. However, when applied to real-world videos, contrastive learning may unknowingly lead to the separation of instances that contain semantically similar events. In our work, we introduce a cooperative variant of contrastive learning to utilize complementary information across views and address this issue. We use data-driven sampling to leverage implicit relationships between multiple input video views, whether observed (e.g. RGB) or inferred (e.g. flow, segmentation masks, poses). We are one of the firsts to explore exploiting inter-instance relationships to drive learning. We experimentally evaluate our representations on the downstream task of action recognition. Our method achieves competitive performance on standard benchmarks (UCF101, HMDB51, Kinetics400). Furthermore, qualitative experiments illustrate that our models can capture higher-order class relationships.

cs.CV

An Interaction-aware Evaluation Method for Highly Automated Vehicles

It is important to build a rigorous verification and validation (V&V) process to evaluate the safety of highly automated vehicles (HAVs) before their wide deployment on public roads. In this paper, we propose an interaction-aware framework for HAV safety evaluation which is suitable for some highly-interactive driving scenarios including highway merging, roundabout entering, etc. Contrary to existing approaches where the primary other vehicle (POV) takes predetermined maneuvers, we model the POV as a game-theoretic agent. To capture a wide variety of interactions between the POV and the vehicle under test (VUT), we characterize the interactive behavior using level-k game theory and social value orientation and train a diverse set of POVs using reinforcement learning. Moreover, we propose an adaptive test case sampling scheme based on the Gaussian process regression technique to generate customized and diverse challenging cases. The highway merging is used as the example scenario. We found the proposed method is able to capture a wide range of POV behaviors and achieve better coverage of the failure modes of the VUT compared with other evaluation approaches.

cs.RO

Discovering Avoidable Planner Failures of Autonomous Vehicles using Counterfactual Analysis in Behaviorally Diverse Simulation

Automated Vehicles require exhaustive testing in simulation to detect as many safety-critical failures as possible before deployment on public roads. In this work, we focus on the core decision-making component of autonomous robots: their planning algorithm. We introduce a planner testing framework that leverages recent progress in simulating behaviorally diverse traffic participants. Using large scale search, we generate, detect, and characterize dynamic scenarios leading to collisions. In particular, we propose methods to distinguish between unavoidable and avoidable accidents, focusing especially on automatically finding planner-specific defects that must be corrected before deployment. Through experiments in complex multi-agent intersection scenarios, we show that our method can indeed find a wide range of critical planner failures.

cs.LG

Behaviorally Diverse Traffic Simulation via Reinforcement Learning

Traffic simulators are important tools in autonomous driving development. While continuous progress has been made to provide developers more options for modeling various traffic participants, tuning these models to increase their behavioral diversity while maintaining quality is often very challenging. This paper introduces an easily-tunable policy generation algorithm for autonomous driving agents. The proposed algorithm balances diversity and driving skills by leveraging the representation and exploration abilities of deep reinforcement learning via a distinct policy set selector. Moreover, we present an algorithm utilizing intrinsic rewards to widen behavioral differences in the training. To provide quantitative assessments, we develop two trajectory-based evaluation metrics which measure the differences among policies and behavioral coverage. We experimentally show the effectiveness of our methods on several challenging intersection scenes.

cs.LG

PillarFlow: End-to-end Birds-eye-view Flow Estimation for Autonomous Driving

In autonomous driving, accurately estimating the state of surrounding obstacles is critical for safe and robust path planning. However, this perception task is difficult, particularly for generic obstacles/objects, due to appearance and occlusion changes. To tackle this problem, we propose an end-to-end deep learning framework for LIDAR-based flow estimation in bird's eye view (BeV). Our method takes consecutive point cloud pairs as input and produces a 2-D BeV flow grid describing the dynamic state of each cell. The experimental results show that the proposed method not only estimates 2-D BeV flow accurately but also improves tracking performance of both dynamic and static objects.

cs.CV

It Is Not the Journey but the Destination: Endpoint Conditioned Trajectory Prediction

Human trajectory forecasting with multiple socially interacting agents is of critical importance for autonomous navigation in human environments, e.g., for self-driving cars and social robots. In this work, we present Predicted Endpoint Conditioned Network (PECNet) for flexible human trajectory prediction. PECNet infers distant trajectory endpoints to assist in long-range multi-modal trajectory prediction. A novel non-local social pooling layer enables PECNet to infer diverse yet socially compliant trajectories. Additionally, we present a simple "truncation-trick" for improving few-shot multi-modal trajectory prediction performance. We show that PECNet improves state-of-the-art performance on the Stanford Drone trajectory prediction benchmark by ~20.9% and on the ETH/UCY benchmark by ~40.8%. Project homepage: https://karttikeya.github.io/publication/htf/

cs.CV

Disentangling Human Dynamics for Pedestrian Locomotion Forecasting with Noisy Supervision

We tackle the problem of Human Locomotion Forecasting, a task for jointly predicting the spatial positions of several keypoints on the human body in the near future under an egocentric setting. In contrast to the previous work that aims to solve either the task of pose prediction or trajectory forecasting in isolation, we propose a framework to unify the two problems and address the practically useful task of pedestrian locomotion prediction in the wild. Among the major challenges in solving this task is the scarcity of annotated egocentric video datasets with dense annotations for pose, depth, or egomotion. To surmount this difficulty, we use state-of-the-art models to generate (noisy) annotations and propose robust forecasting models that can learn from this noisy supervision. We present a method to disentangle the overall pedestrian motion into easier to learn subparts by utilizing a pose completion and a decomposition module. The completion module fills in the missing key-point annotations and the decomposition module breaks the cleaned locomotion down to global (trajectory) and local (pose keypoint movements). Further, with Quasi RNN as our backbone, we propose a novel hierarchical trajectory forecasting network that utilizes low-level vision domain specific signals like egomotion and depth to predict the global trajectory. Our method leads to state-of-the-art results for the prediction of human locomotion in the egocentric view. Project pade: https://karttikeya.github.io/publication/plf/

cs.CV