Searcharxiv⌕ Search

arXiv subjects

Eunji Kim

Publications and source records attributed to Eunji Kim.

30 records · Page 2Linked to original sources

Weakly Supervised Semantic Segmentation using Out-of-Distribution Data

Weakly supervised semantic segmentation (WSSS) methods are often built on pixel-level localization maps obtained from a classifier. However, training on class labels only, classifiers suffer from the spurious correlation between foreground and background cues (e.g. train and rail), fundamentally bounding the performance of WSSS. There have been previous endeavors to address this issue with additional supervision. We propose a novel source of information to distinguish foreground from the background: Out-of-Distribution (OoD) data, or images devoid of foreground object classes. In particular, we utilize the hard OoDs that the classifier is likely to make false-positive predictions. These samples typically carry key visual features on the background (e.g. rail) that the classifiers often confuse as foreground (e.g. train), so these cues let classifiers correctly suppress spurious background cues. Acquiring such hard OoDs does not require an extensive amount of annotation efforts; it only incurs a few additional image-level labeling costs on top of the original efforts to collect class labels. We propose a method, W-OoD, for utilizing the hard OoDs. W-OoD achieves state-of-the-art performance on Pascal VOC 2012.

cs.CV↗

Robust End-to-End Focal Liver Lesion Detection using Unregistered Multiphase Computed Tomography Images

The computer-aided diagnosis of focal liver lesions (FLLs) can help improve workflow and enable correct diagnoses; FLL detection is the first step in such a computer-aided diagnosis. Despite the recent success of deep-learning-based approaches in detecting FLLs, current methods are not sufficiently robust for assessing misaligned multiphase data. By introducing an attention-guided multiphase alignment in feature space, this study presents a fully automated, end-to-end learning framework for detecting FLLs from multiphase computed tomography (CT) images. Our method is robust to misaligned multiphase images owing to its complete learning-based approach, which reduces the sensitivity of the model's performance to the quality of registration and enables a standalone deployment of the model in clinical practice. Evaluation on a large-scale dataset with 280 patients confirmed that our method outperformed previous state-of-the-art methods and significantly reduced the performance degradation for detecting FLLs using misaligned multiphase CT images. The robustness of the proposed method can enhance the clinical adoption of the deep-learning-based computer-aided detection system.

eess.IV↗

Information-Theoretic Visual Explanation for Black-Box Classifiers

In this work, we attempt to explain the prediction of any black-box classifier from an information-theoretic perspective. For each input feature, we compare the classifier outputs with and without that feature using two information-theoretic metrics. Accordingly, we obtain two attribution maps--an information gain (IG) map and a point-wise mutual information (PMI) map. IG map provides a class-independent answer to "How informative is each pixel?", and PMI map offers a class-specific explanation of "How much does each pixel support a specific class?" Compared to existing methods, our method improves the correctness of the attribution maps in terms of a quantitative metric. We also provide a detailed analysis of an ImageNet classifier using the proposed method, and the code is available online.

cs.CV↗

The Effect of Robo-taxi User Experience on User Acceptance: Field Test Data Analysis

With the advancement of self-driving technology, the commercialization of Robo-taxi services is just a matter of time. However, there is some skepticism regarding whether such taxi services will be successfully accepted by real customers due to perceived safety-related concerns; therefore, studies focused on user experience have become more crucial. Although many studies statistically analyze user experience data obtained by surveying individuals' perceptions of Robo-taxi or indirectly through simulators, there is a lack of research that statistically analyzes data obtained directly from actual Robo-taxi service experiences. Accordingly, based on the user experience data obtained by implementing a Robo-taxi service in the downtown of Seoul and Daejeon in South Korea, this study quantitatively analyzes the effect of user experience on user acceptance through structural equation modeling and path analysis. We also obtained balanced and highly valid insights by reanalyzing meaningful causal relationships obtained through statistical models based on in-depth interview results. Results revealed that the experience of the traveling stage had the greatest effect on user acceptance, and the cutting edge of the service and apprehension of technology were emotions that had a great effect on user acceptance. Based on these findings, we suggest guidelines for the design and marketing of future Robo-taxi services.

cs.HC↗

XProtoNet: Diagnosis in Chest Radiography with Global and Local Explanations

Automated diagnosis using deep neural networks in chest radiography can help radiologists detect life-threatening diseases. However, existing methods only provide predictions without accurate explanations, undermining the trustworthiness of the diagnostic methods. Here, we present XProtoNet, a globally and locally interpretable diagnosis framework for chest radiography. XProtoNet learns representative patterns of each disease from X-ray images, which are prototypes, and makes a diagnosis on a given X-ray image based on the patterns. It predicts the area where a sign of the disease is likely to appear and compares the features in the predicted area with the prototypes. It can provide a global explanation, the prototype, and a local explanation, how the prototype contributes to the prediction of a single image. Despite the constraint for interpretability, XProtoNet achieves state-of-the-art classification performance on the public NIH chest X-ray dataset.

cs.CV↗

Anti-Adversarially Manipulated Attributions for Weakly and Semi-Supervised Semantic Segmentation

Weakly supervised semantic segmentation produces a pixel-level localization from a classifier, but it is likely to restrict its focus to a small discriminative region of the target object. AdvCAM is an attribution map of an image that is manipulated to increase the classification score. This manipulation is realized in an anti-adversarial manner, which perturbs the images along pixel gradients in the opposite direction from those used in an adversarial attack. It forces regions initially considered not to be discriminative to become involved in subsequent classifications, and produces attribution maps that successively identify more regions of the target object. In addition, we introduce a new regularization procedure that inhibits the incorrect attribution of regions unrelated to the target object and limits the attributions of the regions that already have high scores. On PASCAL VOC 2012 test images, we achieve mIoUs of 68.0 and 76.9 for weakly and semi-supervised semantic segmentation respectively, which represent a new state-of-the-art.

cs.CV↗

Interpretation of NLP models through input marginalization

To demystify the "black box" property of deep neural networks for natural language processing (NLP), several methods have been proposed to interpret their predictions by measuring the change in prediction probability after erasing each token of an input. Since existing methods replace each token with a predefined value (i.e., zero), the resulting sentence lies out of the training data distribution, yielding misleading interpretations. In this study, we raise the out-of-distribution problem induced by the existing interpretation methods and present a remedy; we propose to marginalize each token out. We interpret various NLP models trained for sentiment analysis and natural language inference using the proposed method.

cs.CL↗

A Study on Anxiety about Using Robo-taxis: HMI Design for Anxiety Factor Analysis and Anxiety Relief Based on Field Tests

Despite the approaching commercialization of robo-taxis, various anxiety factors concerning the safety of autonomous vehicles are expected to form a large barrier against consumers' use of robo-taxi services. The purpose of this study is to derive the various internal and external factors that contribute to the anxieties of robo-taxi passengers, and to propose a human-machine interface (HMI) concept to resolve such factors, by testing robo-taxi services on real, complex urban roads. In addition, a remote system for safely testing a robo-taxi in complex downtown areas was constructed, by adopting the Wizard of Oz (WOZ) methodology. From the results of our tests - conducted upon 28 subjects in the central area of Seoul - 19 major anxiety factors arising from autonomous driving were identified, and seven HMI functions to resolve such factors were designed. The functions were evaluated and their anxiety reduction effects verified. In addition, the various design insights required to increase the reliability of robo-taxis were provided through quantitative and qualitative analysis of the user experience surveys and interviews.

cs.HC↗

Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic Segmentation

When a deep neural network is trained on data with only image-level labeling, the regions activated in each image tend to identify only a small region of the target object. We propose a method of using videos automatically harvested from the web to identify a larger region of the target object by using temporal information, which is not present in the static image. The temporal variations in a video allow different regions of the target object to be activated. We obtain an activated region in each frame of a video, and then aggregate the regions from successive frames into a single image, using a warping technique based on optical flow. The resulting localization maps cover more of the target object, and can then be used as proxy ground-truth to train a segmentation network. This simple approach outperforms existing methods under the same level of supervision, and even approaches relying on extra annotations. Based on VGG-16 and ResNet 101 backbones, our method achieves the mIoU of 65.0 and 67.4, respectively, on PASCAL VOC 2012 test images, which represents a new state-of-the-art.

cs.CV↗

FickleNet: Weakly and Semi-supervised Semantic Image Segmentation using Stochastic Inference

The main obstacle to weakly supervised semantic image segmentation is the difficulty of obtaining pixel-level information from coarse image-level annotations. Most methods based on image-level annotations use localization maps obtained from the classifier, but these only focus on the small discriminative parts of objects and do not capture precise boundaries. FickleNet explores diverse combinations of locations on feature maps created by generic deep neural networks. It selects hidden units randomly and then uses them to obtain activation scores for image classification. FickleNet implicitly learns the coherence of each location in the feature maps, resulting in a localization map which identifies both discriminative and other parts of objects. The ensemble effects are obtained from a single network by selecting random hidden unit pairs, which means that a variety of localization maps are generated from a single image. Our approach does not require any additional training steps and only adds a simple layer to a standard convolutional neural network; nevertheless it outperforms recent comparable techniques on the Pascal VOC 2012 benchmark in both weakly and semi-supervised settings.

cs.CV↗

Control of InGaAs facets using metal modulation epitaxy (MME)

Control of faceting during epitaxy is critical for nanoscale devices. This work identifies the origins of gaps and different facets during regrowth of InGaAs adjacent to patterned features. Molecular beam epitaxy (MBE) near SiO2 or SiNx led to gaps, roughness, or polycrystalline growth, but metal modulated epitaxy (MME) produced smooth and gap-free "rising tide" (001) growth filling up to the mask. The resulting self-aligned FETs were dominated by FET channel resistance rather than source-drain access resistance. Higher As fluxes led first to conformal growth, then pronounced {111} facets sloping up away from the mask.

cond-mat.mtrl-sci↗

A variant of the Brillouin-Wigner perturbation theory with Epstein-Nesbet partitioning

We present an elementary pedagogical derivation of the Brillouin-Wigner and the Rayleigh-Schrödinger perturbation theories with Epstein-Nesbet partitioning. A variant of the Brillouin-Wigner perturbation theory is also introduced, which can be easily extended to the quasi-degenerate case. A main advantage of the new theory is that the computing time required for obtaining the successive higher-order results is negligible after the third-order calculation. We illustrate the accuracy of the new perturbation theory for some simple model systems like the perturbed harmonic oscillator and the particle in a box.

quant-ph↗