SearcharxivSearch

arXiv subjects

Dong Hwan Kim

Publications and source records attributed to Dong Hwan Kim.

11 recordsLinked to original sources

Spectral-Adaptive Modulation Networks for Visual Perception

Recent studies have shown that 2D convolution and self-attention exhibit distinct spectral behaviors, and optimizing their spectral properties can enhance vision model performance. However, theoretical analyses remain limited in explaining why 2D convolution is more effective in high-pass filtering than self-attention and why larger kernels favor shape bias, akin to self-attention. In this paper, we employ graph spectral analysis to theoretically simulate and compare the frequency responses of 2D convolution and self-attention within a unified framework. Our results corroborate previous empirical findings and reveal that node connectivity, modulated by window size, is a key factor in shaping spectral functions. Leveraging this insight, we introduce a \textit{spectral-adaptive modulation} (SPAM) mixer, which processes visual features in a spectral-adaptive manner using multi-scale convolutional kernels and a spectral re-scaling mechanism to refine spectral components. Based on SPAM, we develop SPANetV2 as a novel vision backbone. Extensive experiments demonstrate that SPANetV2 outperforms state-of-the-art models across multiple vision tasks, including ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation.

cs.CV

H2G: Hierarchy-Aware Hyperbolic Grouping for 3D Scenes

Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on semantic labels or a fixed vocabulary. The main challenge is to transform 2D foundation-model cues into coherent hierarchy supervision and embed that hierarchy in a 3D representation. We propose H2G, a hyperbolic affinity field for hierarchical 3D grouping. Our method derives semantically organized tree supervision by interpreting foundation-model affinities through Dasgupta's objective for similarity-based hierarchical clustering. This supervision is distilled into a single Lorentz hyperbolic feature field, whose geometry is well suited for tree-like branching structures. A hierarchy-aware objective aligns the field with fine-level assignments, coarse object structure, compact feature clusters, and LCA (Lowest Common Ancestor) ordering. This formulation represents multiple grouping levels in one feature space, enabling semantic hierarchical grouping grounded in 2D foundation-model knowledge.

cs.CV

Initiation of Interaction Detection Framework using a Nonverbal Cue for Human-Robot Interaction

This paper describes an initiation of interaction(IoI) detection framework without keywords for human-robot interaction(HRI) based on audio and vision sensor fusion in a domestic environment. In the proposed framework, the robot has its own audio and vision sensors, and can employ external vision sensor for stable human detection and tracking. When the user starts to speak while looking at the robot, the robot can localize his or her position by its sound source localization together with human tracking information. Then the robot can detect the IoI if it perceives the face of the speaker faces the robot. In case that the user does not speak directly, the robot can also detect the IoI if he or she looks at the robot for more than predefined periods of time. A state transition model for the proposed IoI detection framework is designed and verified by experiments with a mobile robot. In order to implement and associate our model in a robot architecture, all the components are implemented and integrated in the Robot Operating System(ROS) environment.

cs.CV

Bound for Gaussian-state Quantum illumination using direct photon measurement

It is important to find feasible measurement bounds for quantum information protocols. We present analytic bounds for quantum illumination with Gaussian states when using an on-off detection or a photon number resolving (PNR) detection, where its performance is evaluated with signal-to-noise ratio. First, for coincidence counting measurement, the best performance is given by the two-mode squeezed vacuum (TMSV) state which outperforms the coherent state and the classically correlated thermal (CCT) state. However, the coherent state can beat the TMSV state with increasing signal mean photon number in the case of the on-off detection. Second, the performance is enhanced by taking Fisher information approach of all counting probabilities including non-detection events. In the Fisher information approach, the TMSV state still presents the best performance but the CCT state can beat the TMSV state with increasing signal mean photon number in the case of the on-off detection. Furthermore, we show that it is useful to take the PNR detection on the signal mode and the on-off detection on the idler mode, which reaches similar performance of using PNR detections on both modes.

quant-ph

SPANet: Frequency-balancing Token Mixer using Spectral Pooling Aggregation Modulation

Recent studies show that self-attentions behave like low-pass filters (as opposed to convolutions) and enhancing their high-pass filtering capability improves model performance. Contrary to this idea, we investigate existing convolution-based models with spectral analysis and observe that improving the low-pass filtering in convolution operations also leads to performance improvement. To account for this observation, we hypothesize that utilizing optimal token mixers that capture balanced representations of both high- and low-frequency components can enhance the performance of models. We verify this by decomposing visual features into the frequency domain and combining them in a balanced manner. To handle this, we replace the balancing problem with a mask filtering problem in the frequency domain. Then, we introduce a novel token-mixer named SPAM and leverage it to derive a MetaFormer model termed as SPANet. Experimental results show that the proposed method provides a way to achieve this balance, and the balanced representations of both high- and low-frequency components can improve the performance of models on multiple computer vision tasks. Our code is available at $\href{https://doranlyong.github.io/projects/spanet/}{\text{https://doranlyong.github.io/projects/spanet/}}$.

cs.CV

Gaussian Quantum Illumination via Monotone Metrics

Quantum illumination is to discern the presence or absence of a low reflectivity target, where the error probability decays exponentially in the number of copies used. When the target reflectivity is small so that it is hard to distinguish target presence or absence, the exponential decay constant falls into a class of objects called monotone metrics. We evaluate monotone metrics restricted to Gaussian states in terms of first-order moments and covariance matrix. Under the assumption of a low reflectivity target, we explicitly derive analytic formulae for decay constant of an arbitrary Gaussian input state. Especially, in the limit of large background noise and low reflectivity, there is no need of symplectic diagonalization which usually complicates the computation of decay constants. First, we show that two-mode squeezed vacuum (TMSV) states are the optimal probe among pure Gaussian states with fixed signal mean photon number. Second, as an alternative to preparing TMSV states with high mean photon number, we show that preparing a TMSV state with low mean photon number and displacing the signal mode is a more experimentally feasible setup without degrading the performance that much. Third, we show that it is of utmost importance to prepare an efficient idler memory to beat coherent states and provide analytic bounds on the idler memory transmittivity in terms of signal power, background noise, and idler memory noise. Finally, we identify the region of physically possible correlations between the signal and idler modes that can beat coherent states.

quant-ph

Squeezing Limit of the Josephson Ring Modulator as a Non-Degenerate Parametric Amplifier

Two-mode squeezed vacuum states are a crucial component of quantum technologies. In the microwave domain, they can be produced by Josephson ring modulator which acts as a three-wave mixing non-degenerate parametric amplifier. Here, we solve the master equation of three bosonic modes describing the Josephson ring modulator with a novel numerical method to compute squeezing of output fields and gain at low signal power. We show that the third-order interaction from the three-wave mixing process intrinsically limits squeezing and reduces gain. Since our results are related to other general cavity-based three-wave mixing processes, these imply that any non-degenerate parametric amplifier will have an intrinsic squeezing limit in the output fields.

quant-ph

Observable bound for Gaussian illumination

We propose observable bounds for Gaussian illumination to maximize the signal-to-noise ratio, which minimizes the discrimination error between the presence and absence of a low-reflectivity target using Gaussian states. The observable bounds are achieved with mode-by-mode measurements. In the quantum regime using a two-mode squeezed vacuum state, our observable receiver outperforms the other feasible receivers whereas it cannot approach the quantum Chernoff bound. The corresponding observable cannot be implemented with heterodyne detections due to the additional vacuum noise. In the classical regime using a thermal state, a receiver implemented with a photon number difference measurement approaches its bound regardless of the signal mean photon number, while it asymptotically approaches the classical bound in the limit of a huge idler mean photon number.

quant-ph

CycleMorph: Cycle Consistent Unsupervised Deformable Image Registration

Image registration is a fundamental task in medical image analysis. Recently, deep learning based image registration methods have been extensively investigated due to their excellent performance despite the ultra-fast computational time. However, the existing deep learning methods still have limitation in the preservation of original topology during the deformation with registration vector fields. To address this issues, here we present a cycle-consistent deformable image registration. The cycle consistency enhances image registration performance by providing an implicit regularization to preserve topology during the deformation. The proposed method is so flexible that can be applied for both 2D and 3D registration problems for various applications, and can be easily extended to multi-scale implementation to deal with the memory issues in large volume registration. Experimental results on various datasets from medical and non-medical applications demonstrate that the proposed method provides effective and accurate registration on diverse image pairs within a few seconds. Qualitative and quantitative evaluations on deformation fields also verify the effectiveness of the cycle consistency of the proposed method.

cs.CV

Planning for target retrieval using a robotic manipulator in cluttered and occluded environments

This paper presents planning algorithms for a robotic manipulator with a fixed base in order to grasp a target object in cluttered environments. We consider a configuration of objects in a confined space with a high density so no collision-free path to the target exists. The robot must relocate some objects to retrieve the target while avoiding collisions. For fast completion of the retrieval task, the robot needs to compute a plan optimizing an appropriate objective value directly related to the execution time of the relocation plan. We propose planning algorithms that aim to minimize the number of objects to be relocated. Our objective value is appropriate for the object retrieval task because grasping and releasing objects often dominate the total running time. In addition to the algorithm working in fully known and static environments, we propose algorithms that can deal with uncertain and dynamic situations incurred by occluded views. The proposed algorithms are shown to be complete and run in polynomial time. Our methods reduce the total running time significantly compared to a baseline method (e.g., 25.1% of reduction in a known static environment with 10 objects

cs.RO

Unsupervised Deformable Image Registration Using Cycle-Consistent CNN

Medical image registration is one of the key processing steps for biomedical image analysis such as cancer diagnosis. Recently, deep learning based supervised and unsupervised image registration methods have been extensively studied due to its excellent performance in spite of ultra-fast computational time compared to the classical approaches. In this paper, we present a novel unsupervised medical image registration method that trains deep neural network for deformable registration of 3D volumes using a cycle-consistency. Thanks to the cycle consistency, the proposed deep neural networks can take diverse pair of image data with severe deformation for accurate registration. Experimental results using multiphase liver CT images demonstrate that our method provides very precise 3D image registration within a few seconds, resulting in more accurate cancer size estimation.

cs.CV