SearcharxivSearch

arXiv subjects

Jeongsoo Park

Publications and source records attributed to Jeongsoo Park.

12 recordsLinked to original sources

Layer-polarized Transport via Gate-defined 1D and 0D PN Junctions in Double Bilayer Graphene

We fabricate twisted double bilayer graphene devices with zero twist angle and a set of local top and bottom gates aligned perpendicularly to each other. A 1D PN junction can be electrostatically defined when the gate voltages applied to the top gates are the same but different on the bottom gates. Resistance peaks are observed at finite doping instead of at the charge neutrality points, exhibiting an unconventional broken-cross shape that arises from layer polarization of the P and N region, which can be further enhanced with finite magnetic fields. A 0D point junction (PJ) can be electrostatically defined by applying different gate voltages to the top and bottom gates, such that the P and N sides of the device are connected at a single point in the center of the device. As finite magnetic field B increases, the quantum Hall (QH) states are selectively brought into contact or away from each other depending on their layer polarization, leading to unconventional quantum oscillations which characterize the layer-polarized band-crossing. Our work provides new insights into understanding band-structure evolution and layer polarization in twisted bilayers and paves the way for new device functionality based on manipulating layer-polarized electronic states.

cond-mat.mes-hall

Ferroelectric Quantum Point Contact in Twisted Transition Metal Dichalcogenides

In twisted transition metal dichalcogenides (tTMDs), atomic reconstruction gives rise to moiré domains with alternating ferroelectric polarization, whose domain size and overall electric dipole moment are tunable by an out-of-plane electric field. Previous transport measurements in Hall bar devices have successfully demonstrated the overall ferroelectric behavior of tTMDs from a collective ensemble of ferroelectric moiré domains. To locally probe a single ferroelectric moiré domain, we fabricate and study mesoscopic quantum transport via a gate-defined twisted molybdenum disulfide (tMoS2) quantum point contact (QPC). The local property of a single moiré domain is invulnerable to long-range disorder and twist-angle inhomogeneity, resulting in an unusually long conductance plateau with large electrical hysteresis. The comparison between local and global measurements confirms that antiferroelectricity can emerge from alternating polarization of individual ferroelectric domains. Using a QPC as a single charge sensor, we characterize the nature and time scale of different domain evolution mechanisms with single atomic dipole resolution. Our findings shed new light on the microscopic ferroelectric behavior and dynamics within a single tTMD moiré domain, paving the way toward more advanced ferroelectric quantum devices with tunable local Hamiltonian, such as ferroelectric tTMD quantum dots (QDs).

cond-mat.mes-hall

Nonreciprocal Transport in chiral Mo3Al2C Near the Superconducting to Normal Transition

We investigate nonreciprocal electrical transport in bulk single-crystalline Mo3Al2C, a material known to host crystallographic chirality, a polar charge-density-wave instability, and a superconducting transition near 8 K. Using AC transport measurements to analyze the first-harmonic and second-harmonic resistance responses, we observe a distinct nonreciprocal second-harmonic signal that is significantly enhanced near the boundary of the normal and superconducting phases. Phenomenologically, this response arises from direction-dependent coupling between the external magnetic field and the current-induced intrinsic magnetization within the chiral lattice. Furthermore, a persistent nonreciprocal response observed under perpendicular magnetic fields suggests a toroidal-induced effect linked to the electric polarization emerging from the charge-density-wave phase. These results demonstrate that bulk Mo3Al2C serves as an intrinsic platform for tunable nonreciprocal transport rooted in the interplay of chirality, polarity, and superconductivity.

cond-mat.supr-con

Community Forensics: Using Thousands of Generators to Train Fake Image Detectors

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this problem, and we propose a new dataset that is significantly larger and more diverse than prior work. As part of creating this dataset, we systematically download thousands of text-to-image latent diffusion models and sample images from them. We also collect images from dozens of popular open source and commercial models. The resulting dataset contains 2.7M images that have been sampled from 4803 different models. These images collectively capture a wide range of scene content, generator architectures, and image processing settings. Using this dataset, we study the generalization abilities of fake image detectors. Our experiments suggest that detection performance improves as the number of models in the training set increases, even when these models have similar architectures. We also find that detection performance improves as the diversity of the models increases, and that our trained detectors generalize better than those trained on other datasets. The dataset can be found in https://jespark.net/projects/2024/community_forensics

cs.CV

High-speed control and navigation for quadrupedal robots on complex and discrete terrain

High-speed legged navigation in discrete and geometrically complex environments is a challenging task because of the high-degree-of-freedom dynamics and long-horizon, nonconvex nature of the optimization problem. In this work, we propose a hierarchical navigation pipeline for legged robots that can traverse such environments at high speed. The proposed pipeline consists of a planner and tracker module. The planner module finds physically feasible foothold plans by sampling-based optimization with fast sequential filtering using heuristics and a neural network. Subsequently, rollouts are performed in a physics simulation to identify the best foothold plan regarding the engineered cost function and to confirm its physical consistency. This hierarchical planning module is computationally efficient and physically accurate at the same time. The tracker aims to accurately step on the target footholds from the planning module. During the training stage, the foothold target distribution is given by a generative model that is trained competitively with the tracker. This process ensures that the tracker is trained in an environment with the desired difficulty. The resulting tracker can overcome terrains that are more difficult than what the previous methods could manage. We demonstrated our approach using Raibo, our in-house dynamic quadruped robot. The results were dynamic and agile motions: Raibo is capable of running on vertical walls, jumping a 1.3-meter gap, running over stepping stones at 4 meters per second, and autonomously navigating on terrains full of 30° ramps, stairs, and boxes of various sizes.

cs.RO

Learning Semantic Traversability with Egocentric Video and Automated Annotation Strategy

For reliable autonomous robot navigation in urban settings, the robot must have the ability to identify semantically traversable terrains in the image based on the semantic understanding of the scene. This reasoning ability is based on semantic traversability, which is frequently achieved using semantic segmentation models fine-tuned on the testing domain. This fine-tuning process often involves manual data collection with the target robot and annotation by human labelers which is prohibitively expensive and unscalable. In this work, we present an effective methodology for training a semantic traversability estimator using egocentric videos and an automated annotation process. Egocentric videos are collected from a camera mounted on a pedestrian's chest. The dataset for training the semantic traversability estimator is then automatically generated by extracting semantically traversable regions in each video frame using a recent foundation model in image segmentation and its prompting technique. Extensive experiments with videos taken across several countries and cities, covering diverse urban scenarios, demonstrate the high scalability and generalizability of the proposed annotation method. Furthermore, performance analysis and real-world deployment for autonomous robot navigation showcase that the trained semantic traversability estimator is highly accurate, able to handle diverse camera viewpoints, computationally light, and real-world applicable. The summary video is available at https://youtu.be/EUVoH-wA-lA.

cs.RO

Cross-domain Sound Recognition for Efficient Underwater Data Analysis

This paper presents a novel deep learning approach for analyzing massive underwater acoustic data by leveraging a model trained on a broad spectrum of non-underwater (aerial) sounds. Recognizing the challenge in labeling vast amounts of underwater data, we propose a two-fold methodology to accelerate this labor-intensive procedure. The first part of our approach involves PCA and UMAP visualization of the underwater data using the feature vectors of an aerial sound recognition model. This enables us to cluster the data in a two dimensional space and listen to points within these clusters to understand their defining characteristics. This innovative method simplifies the process of selecting candidate labels for further training. In the second part, we train a neural network model using both the selected underwater data and the non-underwater dataset. We conducted a quantitative analysis to measure the precision, recall, and F1 score of our model for recognizing airgun sounds, a common type of underwater sound. The F1 score achieved by our model exceeded 84.3%, demonstrating the effectiveness of our approach in analyzing underwater acoustic data. The methodology presented in this paper holds significant potential to reduce the amount of labor required in underwater data analysis and opens up new possibilities for further research in the field of cross-domain data analysis.

cs.SD

RGB no more: Minimally-decoded JPEG Vision Transformers

Most neural networks for computer vision are designed to infer using RGB images. However, these RGB images are commonly encoded in JPEG before saving to disk; decoding them imposes an unavoidable overhead for RGB networks. Instead, our work focuses on training Vision Transformers (ViT) directly from the encoded features of JPEG. This way, we can avoid most of the decoding overhead, accelerating data load. Existing works have studied this aspect but they focus on CNNs. Due to how these encoded features are structured, CNNs require heavy modification to their architecture to accept such data. Here, we show that this is not the case for ViTs. In addition, we tackle data augmentation directly on these encoded features, which to our knowledge, has not been explored in-depth for training in this setting. With these two improvements -- ViT and data augmentation -- we show that our ViT-Ti model achieves up to 39.2% faster training and 17.9% faster inference with no accuracy loss compared to the RGB counterpart.

cs.CV

CochlScene: Acquisition of acoustic scene data using crowdsourcing

This paper describes a pipeline for collecting acoustic scene data by using crowdsourcing. The detailed process of crowdsourcing is explained, including planning, validation criteria, and actual user interfaces. As a result of data collection, we present CochlScene, a novel dataset for acoustic scene classification. Our dataset consists of 76k samples collected from 831 participants in 13 acoustic scenes. We also propose a manual data split of training, validation, and test sets to increase the reliability of the evaluation results. Finally, we provide a baseline system for future research.

eess.AS

Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint with a fast maximum inner-product search. To this end, we present a contrastive learning framework that derives from the segment-level search objective. Each update in training uses a batch consisting of a set of pseudo labels, randomly selected original samples, and their augmented replicas. These replicas can simulate the degrading effects on original audio signals by applying small time offsets and various types of distortions, such as background noise and room/microphone impulse responses. In the segment-level search task, where the conventional audio fingerprinting systems used to fail, our system using 10x smaller storage has shown promising results. Our code and dataset are available at \url{https://mimbres.github.io/neural-audio-fp/}.

cs.SD

Enhancing Music Features by Knowledge Transfer from User-item Log Data

In this paper, we propose a novel method that exploits music listening log data for general-purpose music feature extraction. Despite the wealth of information available in the log data of user-item interactions, it has been mostly used for collaborative filtering to find similar items or users and was not fully investigated for content-based music applications. We resolve this problem by extending intra-domain knowledge distillation to cross-domain: i.e., by transferring knowledge obtained from the user-item domain to the music content domain. The proposed system first trains the model that estimates log information from the audio contents; then it uses the model to improve other task-specific models. The experiments on various music classification and regression tasks show that the proposed method successfully improves the performances of the task-specific models.

cs.SD

Separation of Instrument Sounds using Non-negative Matrix Factorization with Spectral Envelope Constraints

Spectral envelope is one of the most important features that characterize the timbre of an instrument sound. However, it is difficult to use spectral information in the framework of conventional spectrogram decomposition methods. We overcome this problem by suggesting a simple way to provide a constraint on the spectral envelope calculated by linear prediction. In the first part of this study, we use a pre-trained spectral envelope of known instruments as the constraint. Then we apply the same idea to a blind scenario in which the instruments are unknown. The experimental results reveal that the proposed method outperforms the conventional methods.

cs.SD