Searcharxiv⌕ Search

arXiv subjects

Chanho Kim

Publications and source records attributed to Chanho Kim.

23 records · Page 2Linked to original sources

Maximal Cliques on Multi-Frame Proposal Graph for Unsupervised Video Object Segmentation

Unsupervised Video Object Segmentation (UVOS) aims at discovering objects and tracking them through videos. For accurate UVOS, we observe if one can locate precise segment proposals on key frames, subsequent processes are much simpler. Hence, we propose to reason about key frame proposals using a graph built with the object probability masks initially generated from multiple frames around the key frame and then propagated to the key frame. On this graph, we compute maximal cliques, with each clique representing one candidate object. By making multiple proposals in the clique to vote for the key frame proposal, we obtain refined key frame proposals that could be better than any of the single-frame proposals. A semi-supervised VOS algorithm subsequently tracks these key frame proposals to the entire video. Our algorithm is modular and hence can be used with any instance segmentation and semi-supervised VOS algorithm. We achieve state-of-the-art performance on the DAVIS-2017 validation and test-dev dataset. On the related problem of video instance segmentation, our method shows competitive performance with the previous best algorithm that requires joint training with the VOS algorithm.

cs.CV↗

X-ray studies of the pulsar PSR J1420-6048 and its TeV pulsar wind nebula in the Kookaburra region

We present a detailed analysis of broadband X-ray observations of the pulsar PSR J1420-6048 and its wind nebula (PWN) in the Kookaburra region with Chandra, XMM-Newton, and NuSTAR. Using the archival XMM-Newton and new NuSTAR data, we detected 68 ms pulsations of the pulsar and characterized its X-ray pulse profile which exhibits a sharp spike and a broad bump separated by ~0.5 in phase. A high-resolution Chandra image revealed a complex morphology of the PWN: a torus-jet structure, a few knots around the torus, one long (~7') and two short tails extending in the northwest direction, and a bright diffuse emission region to the south. Spatially integrated Chandra and NuSTAR spectra of the PWN out to 2.5' are well described by a power law model with a photon index $Γ {\approx}$ 2. A spatially resolved spectroscopic study, as well as NuSTAR radial profiles of the 3--7 keV and 7--20 keV brightness, showed a hint of spectral softening with increasing distance from the pulsar. A multi-wavelength spectral energy distribution (SED) of the source was then obtained by supplementing our X-ray measurements with published radio, Fermi-LAT, and H.E.S.S. data. The SED and radial variations of the X-ray spectrum were fit with a leptonic multi-zone emission model. Our detailed study of the PWN may be suggestive of (1) particle transport dominated by advection, (2) a low magnetic-field strength (B ~ 5$μ$G), and (3) electron acceleration to ~PeV energies.

astro-ph.HE↗

Space Time Recurrent Memory Network

Transformers have recently been popular for learning and inference in the spatial-temporal domain. However, their performance relies on storing and applying attention to the feature tensor of each frame in video. Hence, their space and time complexity increase linearly as the length of video grows, which could be very costly for long videos. We propose a novel visual memory network architecture for the learning and inference problem in the spatial-temporal domain. We maintain a fixed set of memory slots in our memory network and propose an algorithm based on Gumbel-Softmax to learn an adaptive strategy to update this memory. Finally, this architecture is benchmarked on the video object segmentation (VOS) and video prediction problems. We demonstrate that our memory architecture achieves state-of-the-art results, outperforming transformer-based methods on VOS and other recent methods on video prediction while maintaining constant memory capacity independent of the sequence length.

cs.CV↗

Discriminative Appearance Modeling with Multi-track Pooling for Real-time Multi-object Tracking

In multi-object tracking, the tracker maintains in its memory the appearance and motion information for each object in the scene. This memory is utilized for finding matches between tracks and detections and is updated based on the matching result. Many approaches model each target in isolation and lack the ability to use all the targets in the scene to jointly update the memory. This can be problematic when there are similar looking objects in the scene. In this paper, we solve the problem of simultaneously considering all tracks during memory updating, with only a small spatial overhead, via a novel multi-track pooling module. We additionally propose a training strategy adapted to multi-track pooling which generates hard tracking episodes online. We show that the combination of these innovations results in a strong discriminative appearance model, enabling the use of greedy data association to achieve online tracking performance. Our experiments demonstrate real-time, state-of-the-art performance on public multi-object tracking (MOT) datasets.

cs.CV↗

2018 Robotic Scene Segmentation Challenge

In 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models. However, the limited background variation and simple motion rendered the dataset uninformative in learning about which techniques would be suitable for segmentation in real surgery. In 2017, at the same workshop in Quebec we introduced the robotic instrument segmentation dataset with 10 teams participating in the challenge to perform binary, articulating parts and type segmentation of da Vinci instruments. This challenge included realistic instrument motion and more complex porcine tissue as background and was widely addressed with modifications on U-Nets and other popular CNN architectures. In 2018 we added to the complexity by introducing a set of anatomical objects and medical devices to the segmented classes. To avoid over-complicating the challenge, we continued with porcine data which is dramatically simpler than human tissue due to the lack of fatty tissue occluding many organs.

cs.CV↗