SearcharxivSearch

arXiv subjects

Dong-Wook Kim

Publications and source records attributed to Dong-Wook Kim.

11 recordsLinked to original sources

CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation

While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.

cs.RO

How to Relieve Distribution Shifts in Semantic Segmentation for Off-Road Environments

Semantic segmentation is crucial for autonomous navigation in off-road environments, enabling precise classification of surroundings to identify traversable regions. However, distinctive factors inherent to off-road conditions, such as source-target domain discrepancies and sensor corruption from rough terrain, can result in distribution shifts that alter the data differently from the trained conditions. This often leads to inaccurate semantic label predictions and subsequent failures in navigation tasks. To address this, we propose ST-Seg, a novel framework that expands the source distribution through style expansion (SE) and texture regularization (TR). Unlike prior methods that implicitly apply generalization within a fixed source distribution, ST-Seg offers an intuitive approach for distribution shift. Specifically, SE broadens domain coverage by generating diverse realistic styles, augmenting the limited style information of the source domain. TR stabilizes local texture representation affected by style-augmented learning through a deep texture manifold. Experiments across various distribution-shifted target domains demonstrate the effectiveness of ST-Seg, with substantial improvements over existing methods. These results highlight the robustness of ST-Seg, enhancing the real-world applicability of semantic segmentation for off-road navigation.

cs.RO

From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision foundation models (VFMs) via semantic segmentation supervision. However, this paradigm faces three fundamental challenges that undermine its reliability: the task-agnostic design of VFMs, the ambiguity of traversability annotations, and the discrepancy between semantic labels and physical safety. We propose Vision-to-Traversability Adaptation (ViTA), a framework that adapts VFMs for reliable traversability estimation, instantiated on SAM2. ViTA injects task-specific knowledge through learnable traversability prompts while preserving the VFM's cross-domain generalization. To handle annotation ambiguity, we introduce Perspective-Diversified Training, which estimates semantic uncertainty to suppress confident predictions at ambiguous boundaries. To bridge the semantic-traversability discrepancy, we distill geometric knowledge during training, enabling slope and elevation reasoning from RGB images alone at inference. The semantic and geometric outputs are fused into a continuous traversability score that reflects both semantic uncertainty and geometric risk. Evaluations across diverse domains, including challenging real-world off-road datasets, demonstrate that ViTA achieves state-of-the-art IoU and Precision with substantial false-positive reduction and strong cross-domain generalization.

cs.CV

Temporal coupled mode theory for high-$Q$ resonances in dielectric metasurfaces

In this work, we propose a coupled mode theory for resonant response from quasi-guided modes in periodic dielectric metasurfaces. First, we derived a generic set of constraints imposed onto the parameters of the temporal coupled mode theory by energy conservation and time-reversal symmetry in an invariant form that allows for asymmetry between the coupling and decoupling coefficients. The proposed approach is applied to the problem of Fano resonances induced by isolated quasi-guided modes in the regime of specular reflection. Our central result is a generic formula for the line-shape of the Fano resonance in transmittance for the lossless metasurfaces in the framework of 2D electrodynamics. We consider all possible symmetries of the metasurface elementary cell and uncover the effects that the symmetry incurs on the profile of the Fano resonance induced by an isolated high-$Q$ mode. It is shown that the proposed approach correctly describes the presence of robust reflection and transmission zeros in the spectra as well as the spectral signatures of bound states in the continuum. The approach is applied to uniderictionally guided resonant modes in metasurfaces with an asymmetric elementary cell. It is found that the existence of such modes and the transmittance in their spectral vicinity are consistent with the theoretical predictions. Furthermore, the theory predicts that a uniderictionally guided resonant mode is dual to a counter-propagating mode of a peculiar type which is coupled with the outgoing wave on both sides of the metasurface but, nonetheless, exhibits only a single-sided coupling with incident waves.

physics.optics

Attend-and-Refine: Interactive keypoint estimation and quantitative cervical vertebrae analysis for bone age assessment

In pediatric orthodontics, accurate estimation of growth potential is essential for developing effective treatment strategies. Our research aims to predict this potential by identifying the growth peak and analyzing cervical vertebra morphology solely through lateral cephalometric radiographs. We accomplish this by comprehensively analyzing cervical vertebral maturation (CVM) features from these radiographs. This methodology provides clinicians with a reliable and efficient tool to determine the optimal timings for orthodontic interventions, ultimately enhancing patient outcomes. A crucial aspect of this approach is the meticulous annotation of keypoints on the cervical vertebrae, a task often challenged by its labor-intensive nature. To mitigate this, we introduce Attend-and-Refine Network (ARNet), a user-interactive, deep learning-based model designed to streamline the annotation process. ARNet features Interaction-guided recalibration network, which adaptively recalibrates image features in response to user feedback, coupled with a morphology-aware loss function that preserves the structural consistency of keypoints. This novel approach substantially reduces manual effort in keypoint identification, thereby enhancing the efficiency and accuracy of the process. Extensively validated across various datasets, ARNet demonstrates remarkable performance and exhibits wide-ranging applicability in medical imaging. In conclusion, our research offers an effective AI-assisted diagnostic tool for assessing growth potential in pediatric orthodontics, marking a significant advancement in the field.

cs.CV

Morphology-Aware Interactive Keypoint Estimation

Diagnosis based on medical images, such as X-ray images, often involves manual annotation of anatomical keypoints. However, this process involves significant human efforts and can thus be a bottleneck in the diagnostic process. To fully automate this procedure, deep-learning-based methods have been widely proposed and have achieved high performance in detecting keypoints in medical images. However, these methods still have clinical limitations: accuracy cannot be guaranteed for all cases, and it is necessary for doctors to double-check all predictions of models. In response, we propose a novel deep neural network that, given an X-ray image, automatically detects and refines the anatomical keypoints through a user-interactive system in which doctors can fix mispredicted keypoints with fewer clicks than needed during manual revision. Using our own collected data and the publicly available AASCE dataset, we demonstrate the effectiveness of the proposed method in reducing the annotation costs via extensive quantitative and qualitative results. A demo video of our approach is available on our project webpage.

cs.CV

RZSR: Reference-based Zero-Shot Super-Resolution with Depth Guided Self-Exemplars

Recent methods for single image super-resolution (SISR) have demonstrated outstanding performance in generating high-resolution (HR) images from low-resolution (LR) images. However, most of these methods show their superiority using synthetically generated LR images, and their generalizability to real-world images is often not satisfactory. In this paper, we pay attention to two well-known strategies developed for robust super-resolution (SR), i.e., reference-based SR (RefSR) and zero-shot SR (ZSSR), and propose an integrated solution, called reference-based zero-shot SR (RZSR). Following the principle of ZSSR, we train an image-specific SR network at test time using training samples extracted only from the input image itself. To advance ZSSR, we obtain reference image patches with rich textures and high-frequency details which are also extracted only from the input image using cross-scale matching. To this end, we construct an internal reference dataset and retrieve reference image patches from the dataset using depth information. Using LR patches and their corresponding HR reference patches, we train a RefSR network that is embodied with a non-local attention module. Experimental results demonstrate the superiority of the proposed RZSR compared to the previous ZSSR methods and robustness to unseen images compared to other fully supervised SISR methods.

cs.CV

GRDN:Grouped Residual Dense Network for Real Image Denoising and GAN-based Real-world Noise Modeling

Recent research on image denoising has progressed with the development of deep learning architectures, especially convolutional neural networks. However, real-world image denoising is still very challenging because it is not possible to obtain ideal pairs of ground-truth images and real-world noisy images. Owing to the recent release of benchmark datasets, the interest of the image denoising community is now moving toward the real-world denoising problem. In this paper, we propose a grouped residual dense network (GRDN), which is an extended and generalized architecture of the state-of-the-art residual dense network (RDN). The core part of RDN is defined as grouped residual dense block (GRDB) and used as a building module of GRDN. We experimentally show that the image denoising performance can be significantly improved by cascading GRDBs. In addition to the network architecture design, we also develop a new generative adversarial network-based real-world noise modeling method. We demonstrate the superiority of the proposed methods by achieving the highest score in terms of both the peak signal-to-noise ratio and the structural similarity in the NTIRE2019 Real Image Denoising Challenge - Track 2:sRGB.

eess.IV

Large-Scale Conformal Growth of Atomic-Thick MoS2 for Highly Efficient Photocurrent Generation

Controlling the interconnection of neighboring seeds (nanoflakes) to full coverage of the textured substrate is the main challenge for the large-scale conformal growth of atomic-thick transition metal dichalcogenides by chemical vapor deposition. Herein, we report on a controllable method for the conformal growth of monolayer MoS2 on not only planar but also micro- and nano-rugged SiO2/Si substrates via metal-organic chemical vapor deposition. The continuity of monolayer MoS2 on the rugged surface is evidenced by scanning electron microscopy, cross-section high-resolution transmission electron microscopy, photoluminescence (PL) mapping, and Raman mapping. Interestingly, the photo-responsivity (~254.5 mA/W) of as-grown MoS2 on the nano-rugged substrate exhibits 59 times higher than that of the planar sample (4.3 mA/W) under a small applied bias of 0.1 V. This value is record high when compared with all previous MoS2-based photocurrent generation under low or zero bias. Such a large enhancement in the photo-responsivity arises from a large active area for light-matter interaction and local strain for PL quenching, where the latter effect is the key factor and unique in the conformally grown monolayer on the nano-rugged surface. The result is a step toward the batch fabrication of modern atomic-thick optoelectronic devices.

physics.app-ph

Influence of Perfluorinated Ionomer in PEDOT:PSS on the Rectification and Degradation of Organic Photovoltaic Cells

Poly(3,4-ethylenedioxythiophene):poly(styrenesulfonate) (PEDOT:PSS) is widely used to build optoelectronic devices. However, as a hygroscopic water-based acidic material, it brings major concerns for stability and degradation, resulting in an intense effort to replace it in organic photovoltaic (OPV) devices. In this work, we focus on the perfluorinated ionomer (PFI) polymeric additive to PEDOT:PSS. We demonstrate that it can reduce the relative amplitude of OPV device burn-in, and find two distinct regimes of influence. At low concentrations there is a subtle effect on wetting and work function, for instance, with a detrimental impact on the device characteristics, and above a threshold it changes the electronic and device properties. The abrupt threshold in the conducting polymer occurs for PFI concentrations greater than or equal to the PSS concentration and was revealed by monitoring variations in transmission, topography, work-function, wettability and OPV device characteristics. Below this PFI concentration threshold, the power conversion efficiency (PCE) of OPVs based on poly(3-hexylthiophene-2,5-diyl):[6,6]-phenyl-C61-butyric acid methyl ester (P3HT:PCBM) are impaired largely by low fill-factors due to poor charge extraction. Above the PFI concentration threshold, we recover the PCE before it is improved beyond the pristine PEDOT:PSS layer based OPV devices. Supplementary to the performance enhancement, PFI improves OPV device stability and lifetime. Our degradation study leads to the conclusion that PFI prevents water from diffusing to and from the hygrosopic PEDOT:PSS layer, which slows down the deterioration of the PEDOT:PSS layer and the aluminum electrode. These findings reveal mechanisms and opportunities that should be taken into consideration when developing components to inhibit OPV degradation.

physics.app-ph

Influence of gas ambient on charge writing at the LaAlO3/SrTiO3 heterointerface

We investigated the influences charge writing on the surface work function and resistance of the LaAlO3/SrTiO3 (LAO/STO) heterointerface in several gas environments (air, O2, N2, and H2/N2). Charge writing decreased the surface work function and resistance of the LAO/STO sample quite a lot in air but slightly in O2.The interface carrier density was extracted from the measured sheet resistance and compared with that obtained from the proposed charge-writing mechanisms, such as carrier transfer via surface adsorbates and surface redox. Such quantitative analyses suggested that additional processes (e.g., electronic state modification and electrochemical surface reaction) were required to explain charge writing on the LAO/STO interface.

cond-mat.mtrl-sci