Searcharxiv⌕ Search

arXiv subjects

Yurong You

Publications and source records attributed to Yurong You.

34 records · Page 2Linked to original sources

Ithaca365: Dataset and Driving Perception under Repeated and Challenging Weather Conditions

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under nice weather conditions. Yet, to achieve the high safety requirement, these perceptual systems must operate robustly under a wide variety of weather conditions including snow and rain. In this paper, we present a new dataset to enable robust autonomous driving via a novel data collection process - data is repeatedly recorded along a 15 km route under diverse scene (urban, highway, rural, campus), weather (snow, rain, sun), time (day/night), and traffic conditions (pedestrians, cyclists and cars). The dataset includes images and point clouds from cameras and LiDAR sensors, along with high-precision GPS/INS to establish correspondence across routes. The dataset includes road and object annotations using amodal masks to capture partial occlusions and 3D bounding boxes. We demonstrate the uniqueness of this dataset by analyzing the performance of baselines in amodal segmentation of road and objects, depth estimation, and 3D object detection. The repeated routes opens new research directions in object discovery, continual learning, and anomaly detection. Link to Ithaca365: https://ithaca365.mae.cornell.edu/

cs.CV↗

Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection

Self-driving cars must detect other vehicles and pedestrians in 3D to plan safe routes and avoid collisions. State-of-the-art 3D object detectors, based on deep learning, have shown promising accuracy but are prone to over-fit to domain idiosyncrasies, making them fail in new environments -- a serious problem if autonomous vehicles are meant to operate freely. In this paper, we propose a novel learning approach that drastically reduces this gap by fine-tuning the detector on pseudo-labels in the target domain, which our method generates while the vehicle is parked, based on replays of previously recorded driving sequences. In these replays, objects are tracked over time, and detections are interpolated and extrapolated -- crucially, leveraging future information to catch hard cases. We show, on five autonomous driving datasets, that fine-tuning the object detector on these pseudo-labels substantially reduces the domain gap to new driving environments, yielding drastic improvements in accuracy and detection reliability.

cs.CV↗

R4D: Utilizing Reference Objects for Long-Range Distance Estimation

Estimating the distance of objects is a safety-critical task for autonomous driving. Focusing on short-range objects, existing methods and datasets neglect the equally important long-range objects. In this paper, we introduce a challenging and under-explored task, which we refer to as Long-Range Distance Estimation, as well as two datasets to validate new methods developed for this task. We then proposeR4D, the first framework to accurately estimate the distance of long-range objects by using references with known distances in the scene. Drawing inspiration from human perception, R4D builds a graph by connecting a target object to all references. An edge in the graph encodes the relative distance information between a pair of target and reference objects. An attention module is then used to weigh the importance of reference objects and combine them into one target object distance prediction. Experiments on the two proposed datasets demonstrate the effectiveness and robustness of R4D by showing significant improvements compared to existing baselines. We are looking to make the proposed dataset, Waymo OpenDataset - Long-Range Labels, available publicly at waymo.com/open/download.

cs.CV↗

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior performance when compared to LiDAR-based techniques. Through systematic analysis, we identified that per-object depth estimation accuracy is a major factor bounding the performance. Motivated by this observation, we propose a multi-level fusion method that combines different representations (RGB and pseudo-LiDAR) and temporal information across multiple frames for objects (tracklets) to enhance per-object depth estimation. Our proposed fusion method achieves the state-of-the-art performance of per-object depth estimation on the Waymo Open Dataset, the KITTI detection dataset, and the KITTI MOT dataset. We further demonstrate that by simply replacing estimated depth with fusion-enhanced depth, we can achieve significant improvements in monocular 3D perception tasks, including detection and tracking.

cs.CV↗

Learning to Detect Mobile Objects from LiDAR Scans Without Labels

Current 3D object detectors for autonomous driving are almost entirely trained on human-annotated data. Although of high quality, the generation of such data is laborious and costly, restricting them to a few specific locations and object types. This paper proposes an alternative approach entirely based on unlabeled data, which can be collected cheaply and in abundance almost everywhere on earth. Our approach leverages several simple common sense heuristics to create an initial set of approximate seed labels. For example, relevant traffic participants are generally not persistent across multiple traversals of the same route, do not fly, and are never under ground. We demonstrate that these seed labels are highly effective to bootstrap a surprisingly accurate detector through repeated self-training without a single human annotated label.

cs.CV↗

Hindsight is 20/20: Leveraging Past Traversals to Aid 3D Perception

Self-driving cars must detect vehicles, pedestrians, and other traffic participants accurately to operate safely. Small, far-away, or highly occluded objects are particularly challenging because there is limited information in the LiDAR point clouds for detecting them. To address this challenge, we leverage valuable information from the past: in particular, data collected in past traversals of the same scene. We posit that these past data, which are typically discarded, provide rich contextual information for disambiguating the above-mentioned challenging cases. To this end, we propose a novel, end-to-end trainable Hindsight framework to extract this contextual information from past traversals and store it in an easy-to-query data structure, which can then be leveraged to aid future 3D object detection of the same scene. We show that this framework is compatible with most modern 3D detection architectures and can substantially improve their average precision on multiple autonomous driving datasets, most notably by more than 300% on the challenging cases.

cs.CV↗

Train in Germany, Test in The USA: Making 3D Object Detectors Generalize

In the domain of autonomous driving, deep learning has substantially improved the 3D object detection accuracy for LiDAR and stereo camera data alike. While deep networks are great at generalization, they are also notorious to over-fit to all kinds of spurious artifacts, such as brightness, car sizes and models, that may appear consistently throughout the data. In fact, most datasets for autonomous driving are collected within a narrow subset of cities within one country, typically under similar weather conditions. In this paper we consider the task of adapting 3D object detectors from one dataset to another. We observe that naively, this appears to be a very challenging task, resulting in drastic drops in accuracy levels. We provide extensive experiments to investigate the true adaptation challenges and arrive at a surprising conclusion: the primary adaptation hurdle to overcome are differences in car sizes across geographic areas. A simple correction based on the average car size yields a strong correction of the adaptation gap. Our proposed method is simple and easily incorporated into most 3D object detection frameworks. It provides a first baseline for 3D object detection adaptation across countries, and gives hope that the underlying problem may be more within grasp than one may have hoped to believe. Our code is available at https://github.com/cxy1997/3D_adapt_auto_driving.

cs.CV↗

End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection

Reliable and accurate 3D object detection is a necessity for safe autonomous driving. Although LiDAR sensors can provide accurate 3D point cloud estimates of the environment, they are also prohibitively expensive for many settings. Recently, the introduction of pseudo-LiDAR (PL) has led to a drastic reduction in the accuracy gap between methods based on LiDAR sensors and those based on cheap stereo cameras. PL combines state-of-the-art deep neural networks for 3D depth estimation with those for 3D object detection by converting 2D depth map outputs to 3D point cloud inputs. However, so far these two networks have to be trained separately. In this paper, we introduce a new framework based on differentiable Change of Representation (CoR) modules that allow the entire PL pipeline to be trained end-to-end. The resulting framework is compatible with most state-of-the-art networks for both tasks and in combination with PointRCNN improves over PL consistently across all benchmarks -- yielding the highest entry on the KITTI image-based 3D object detection leaderboard at the time of submission. Our code will be made available at https://github.com/mileyan/pseudo-LiDAR_e2e.

cs.CV↗

Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving

Detecting objects such as cars and pedestrians in 3D plays an indispensable role in autonomous driving. Existing approaches largely rely on expensive LiDAR sensors for accurate depth information. While recently pseudo-LiDAR has been introduced as a promising alternative, at a much lower cost based solely on stereo images, there is still a notable performance gap. In this paper we provide substantial advances to the pseudo-LiDAR framework through improvements in stereo depth estimation. Concretely, we adapt the stereo network architecture and loss function to be more aligned with accurate depth estimation of faraway objects --- currently the primary weakness of pseudo-LiDAR. Further, we explore the idea to leverage cheaper but extremely sparse LiDAR sensors, which alone provide insufficient information for 3D detection, to de-bias our depth estimation. We propose a depth-propagation algorithm, guided by the initial depth estimates, to diffuse these few exact measurements across the entire depth map. We show on the KITTI object detection benchmark that our combined approach yields substantial improvements in depth estimation and stereo-based 3D object detection --- outperforming the previous state-of-the-art detection accuracy for faraway objects by 40%. Our code is available at https://github.com/mileyan/Pseudo_Lidar_V2.

cs.CV↗

Design of reversible low-field magnetocaloric effect at room temperature in hexagonal MnMX ferromagnets

Giant magnetocaloric effect is widely achieved in hexagonal MnMX-based (M = Co or Ni, X = Si or Ge) ferromagnets at their first-order magnetostructural transition. However, the thermal hysteresis and the low sensitivity of the magnetostructural transition to the magnetic field inevitably lead to a sizeable irreversibility of the low-field magnetocaloric effect. In this work, we show an alternative way to realize a reversible low-field magnetocaloric effect in MnMX-based alloys by taking advantage of the second-order phase transition. With introducing Cu into Co in MnCoGe alloy, the martensitic transition is stabilized at high temperature, while the Curie temperature of the orthorhombic phase is reduced to room temperature. As a result, a second-order magnetic transition with negligible thermal hysteresis and a large magnetization change can be observed, enabling a large reversible magnetocaloric effect. By both calorimetric and direct measurements, a reversible adiabatic temperature change of about 1 K is obtained under a field change of 0-1 T at 304 K, which is larger than that obtained in a first-order magnetostructural transition. To get a better insight into the origin of these experimental results, first-principles calculations are carried out to characterize the chemical bonds and the magnetic exchange interaction. Our work provides a new understanding of the MnCoGe alloy and indicates a feasible route to improve the reversibility of the low-field magnetocaloric effect in the MnMX system.

cond-mat.mtrl-sci↗

Simple Black-box Adversarial Attacks

We propose an intriguingly simple method for the construction of adversarial images in the black-box setting. In constrast to the white-box scenario, constructing black-box adversarial images has the additional constraint on query budget, and efficient attacks remain an open problem to date. With only the mild assumption of continuous-valued confidence scores, our highly query-efficient algorithm utilizes the following simple iterative principle: we randomly sample a vector from a predefined orthonormal basis and either add or subtract it to the target image. Despite its simplicity, the proposed method can be used for both untargeted and targeted attacks -- resulting in previously unprecedented query efficiency in both settings. We demonstrate the efficacy and efficiency of our algorithm on several real world settings including the Google Cloud Vision API. We argue that our proposed algorithm should serve as a strong baseline for future black-box attacks, in particular because it is extremely fast and its implementation requires less than 20 lines of PyTorch code.

cs.LG↗

Planar topological Hall effect in a uniaxial van der Waals ferromagnet Fe3GeTe2

In this work, we reported the observation of a novel planar topological Hall effect (PTHE) in single crystal of Fe3GeTe2, a paradigmatic two-dimensional ferromagnet with strong uniaxial anisotropy. The Hall effect and magnetoresistance varied periodically when the external magnetic field rotated in the ac (or bc) plane, while the PTHE emerged and maintained robust with field swept across the hard-magnetized ab plane. The PTHE covers the whole temperature region below Tc (~150 K) and a comparatively large value is observed at 100 K. Emergence of an internal gauge field was proposed to explain the origin of this large PTHE, which is either generated by the possible topological domain structure of uniaxial Fe3GeTe2 or the non-coplanar spin structure formed during the in-plane magnetization. Our results promisingly provide an alternative detection method to the in-plane skyrmion formation and may bring brand-new prospective to magneto-transport studies in condensed matter physics.

cond-mat.mtrl-sci↗

Resource Aware Person Re-identification across Multiple Resolutions

Not all people are equally easy to identify: color statistics might be enough for some cases while others might require careful reasoning about high- and low-level details. However, prevailing person re-identification(re-ID) methods use one-size-fits-all high-level embeddings from deep convolutional networks for all cases. This might limit their accuracy on difficult examples or makes them needlessly expensive for the easy ones. To remedy this, we present a new person re-ID model that combines effective embeddings built on multiple convolutional network layers, trained with deep-supervision. On traditional re-ID benchmarks, our method improves substantially over the previous state-of-the-art results on all five datasets that we evaluate on. We then propose two new formulations of the person re-ID problem under resource-constraints, and show how our model can be used to effectively trade off accuracy and computation in the presence of resource constraints. Code and pre-trained models are available at https://github.com/mileyan/DARENet.

cs.CV↗

Tunable magnetic and transport properties of Mn3Ga thin films on Ta/Ru seedlayer

Hexagonal D019-type Mn3Z alloys that possess large anomalous and topological-like Hall effects have attracted much attention due to their great potential in the antiferromagnetic spintronic devices. Here, we report the preparation of Mn3Ga film in both tetragonal and hexagonal phases with a tuned Ta/Ru seed layer on the thermally oxidized Si substrate. A large coercivity together with a large anomalous Hall resistivity is found in the Ta-only sample with mixed tetragonal phase. By increasing the thickness of Ru layer, the tetragonal phase gradually disappears and a relatively pure hexagonal phase is obtained in the Ta(5)/Ru(30) buffered sample. Further magnetic and transport measurements revealed that the anomalous Hall conductivity nearly vanishes in the pure hexagonal sample, while an abnormal asymmetric hump structure emerges in the low field region. The extracted additional Hall term is robust in a large temperature range and presents a sign reversal above 200K. The abnormal Hall properties are proposed to be closely related with the frustrated spin structure of D019 Mn3Ga.

cond-mat.mtrl-sci↗

Virtual to Real Reinforcement Learning for Autonomous Driving

Reinforcement learning is considered as a promising direction for driving policy learning. However, training autonomous driving vehicle with reinforcement learning in real environment involves non-affordable trial-and-error. It is more desirable to first train in a virtual environment and then transfer to the real environment. In this paper, we propose a novel realistic translation network to make model trained in virtual environment be workable in real world. The proposed network can convert non-realistic virtual image input into a realistic one with similar scene structure. Given realistic frames as input, driving policy trained by reinforcement learning can nicely adapt to real world driving. Experiments show that our proposed virtual to real (VR) reinforcement learning (RL) works pretty well. To our knowledge, this is the first successful case of driving policy trained by reinforcement learning that can adapt to real world driving data.

cs.AI↗

Designing compensated magnetic states in tetragonal Mn3Ge-based alloys

Magnetic compensated state attracted much interests due to the observed large exchange bias and large coercivity, and its potential applications in the antiferromagnetic spintronics with merit of no stray field. In this work, by ab initio calculations with KKR-CPA for the treatment of random substitution, we obtain the complete compensated states in the Ni (Pd, Pt) doped Mn3Ge-based D022-type tetragonal Heusler alloys. We find the total moment change is asymmetric across the compensation point (at ~ x = 0.3) in Mn3-xYxGe (Y = Ni, Pd, Pt), which is highly conforming to that experimentally observed in Mn3Ga. In addition, an uncommon discontinuous jump is observed across the critical zero-moment point, indicating that some non-trivial properties can emerge at this point. Further electronic analysis for the three compensation compositions reveals large spin polarizations, together with the high Curie temperature of the host Mn3Ge, making them promising candidates for spin transfer torque applications.

cond-mat.mtrl-sci↗