SearcharxivSearch

arXiv subjects

Kanghee Lee

Publications and source records attributed to Kanghee Lee.

9 recordsLinked to original sources

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility

Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and spatial reasoning directly from multi-view images, without requiring explicit 3D reconstructions. Nevertheless, key challenges that frequently arise in real-world environments, such as partial visibility, occlusion, and low-overlap conditions that require reasoning from fragmented visual cues, remain under-explored. To address these limitations, we propose a scalable multi-view data generation and annotation pipeline that constructs realistic spatial reasoning QAs, resulting in SpatialMosaic, a comprehensive instruction-tuning dataset with 2M QA pairs. We further introduce SpatialMosaic-Bench, a challenging benchmark for evaluating multi-view spatial reasoning under complex and diverse scenarios, consisting of 1M QA pairs across 11 tasks with both multiple-choice and numerical-answer formats. Our dataset spans both indoor and outdoor scenes, enabling comprehensive evaluation across diverse real-world scenarios. In addition, we provide a practical baseline for multi-view settings by integrating geometry encoders into VLMs for improved cross-view consistency and spatial grounding. Extensive experiments demonstrate that our dataset effectively enhances spatial reasoning under challenging multi-view conditions, validating the effectiveness of our data generation pipeline in constructing realistic and challenging QAs.

cs.CV

Group-wise Scaling and Orthogonal Decomposition for Domain-Invariant Feature Extraction in Face Anti-Spoofing

Domain Generalizable Face Anti-Spoofing (DGFAS) methods effectively capture domain-invariant features by aligning the directions (weights) of local decision boundaries across domains. However, the bias terms associated with these boundaries remain misaligned, leading to inconsistent classification thresholds and degraded performance on unseen target domains. To address this issue, we propose a novel DGFAS framework that jointly aligns weights and biases through Feature Orthogonal Decomposition (FOD) and Group-wise Scaling Risk Minimization (GS-RM). Specifically, GS-RM facilitates bias alignment by balancing group-wise losses across multiple domains. FOD employs the Gram-Schmidt orthogonalization process to decompose the feature space explicitly into domain-invariant and domain-specific subspaces. By enforcing orthogonality between domain-specific and domain-invariant features during training using domain labels, FOD ensures effective weight alignment across domains without negatively impacting bias alignment. Additionally, we introduce Expected Calibration Error (ECE) as a novel evaluation metric for quantitatively assessing the effectiveness of our method in aligning bias terms across domains. Extensive experiments on benchmark datasets demonstrate that our approach achieves state-of-the-art performance, consistently improving accuracy, reducing bias misalignment, and enhancing generalization stability on unseen target domains.

cs.CV

BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse Scenes

Recent advances in deep learning-based point cloud registration have improved generalization, yet most methods still require retraining or manual parameter tuning for each new environment. In this paper, we identify three key factors limiting generalization: (a) reliance on environment-specific voxel size and search radius, (b) poor out-of-domain robustness of learning-based keypoint detectors, and (c) raw coordinate usage, which exacerbates scale discrepancies. To address these issues, we present a zero-shot registration pipeline called BUFFER-X by (a) adaptively determining voxel size/search radii, (b) using farthest point sampling to bypass learned detectors, and (c) leveraging patch-wise scale normalization for consistent coordinate bounds. In particular, we present a multi-scale patch-based descriptor generation and a hierarchical inlier search across scales to improve robustness in diverse scenes. We also propose a novel generalizability benchmark using 11 datasets that cover various indoor/outdoor scenarios and sensor modalities, demonstrating that BUFFER-X achieves substantial generalization without prior information or manual parameter tuning for the test datasets. Our code is available at https://github.com/MIT-SPARK/BUFFER-X.

cs.CV

3D Geometric Shape Assembly via Efficient Point Cloud Matching

Learning to assemble geometric shapes into a larger target structure is a pivotal task in various practical applications. In this work, we tackle this problem by establishing local correspondences between point clouds of part shapes in both coarse- and fine-levels. To this end, we introduce Proxy Match Transform (PMT), an approximate high-order feature transform layer that enables reliable matching between mating surfaces of parts while incurring low costs in memory and computation. Building upon PMT, we introduce a new framework, dubbed Proxy Match TransformeR (PMTR), for the geometric assembly task. We evaluate the proposed PMTR on the large-scale 3D geometric shape assembly benchmark dataset of Breaking Bad and demonstrate its superior performance and efficiency compared to state-of-the-art methods. Project page: https://nahyuklee.github.io/pmtr.

cs.CV

Learning to Register Unbalanced Point Pairs

Point cloud registration methods can effectively handle large-scale, partially overlapping point cloud pairs. Despite its practicality, matching the unbalanced pairs in terms of spatial extent and density has been overlooked and rarely studied. We present a novel method, dubbed UPPNet, for Unbalanced Point cloud Pair registration. We propose to incorporate a hierarchical framework that effectively finds inlier correspondences by gradually reducing search space. The proposed method first predicts subregions within target point cloud that are likely to be overlapped with query. Then following super-point matching and fine-grained refinement modules predict accurate inlier correspondences between the target and query. Additional geometric constraints are applied to refine the correspondences that satisfy spatial compatibility. The proposed network can be trained in an end-to-end manner, predicting the accurate rigid transformation with a single forward pass. To validate the efficacy of the proposed method, we create a carefully designed benchmark, named KITTI-UPP dataset, by augmenting the KITTI odometry dataset. Extensive experiments reveal that the proposed method not only outperforms state-of-the-art point cloud registration methods by large margins on KITTI-UPP benchmark, but also achieves competitive results on the standard pairwise registration benchmark including 3DMatch, 3DLoMatch, ScanNet, and KITTI, thus showing the applicability of our method on various datasets. The source code and dataset will be publicly released.

cs.CV

Non-Hermitian chiral degeneracy of gated graphene metasurfaces

Non-Hermitian degeneracies, also known as exceptional points (EPs), have been the focus of much attention due to their singular eigenvalue surface structure. Nevertheless, as pertaining to a non-Hermitian metasurface platform, the reduction of an eigenspace dimensionality at the EP has been investigated mostly in a passive repetitive manner. Here, we propose an electrical and spectral way of resolving chiral EPs and clarifying the consequences of chiral mode collapsing of a non-Hermitian gated graphene metasurface. More specifically, the measured non-Hermitian Jones matrix in parameter space enables the quantification of nonorthogonality of polarisation eigenstates and half-integer topological charges associated with a chiral EP. Interestingly, the output polarisation state can be made orthogonal to the coalesced polarisation eigenstate of the metasurface, revealing the missing dimension at the chiral EP. In addition, the maximal nonorthogonality at the chiral EP leads to a blocking of one of the cross-polarised transmission pathways and, consequently, the observation of enhanced asymmetric polarisation conversion. We anticipate that electrically controllable non-Hermitian metasurface platforms can serve as an interesting framework for the investigation of rich non-Hermitian polarisation dynamics around chiral EPs.

physics.optics

Revealing non-Hermitian band structures of photonic Floquet media

Periodically driven systems, characterised by their inherent non-equilibrium dynamics, are ubiquitously found in both classical and quantum regimes. In the field of photonics, these Floquet systems have begun to provide insight into how time periodicity can extend the concept of spatially periodic photonic crystals and metamaterials to the time domain. However, despite the necessity arising from the presence of non-reciprocal coupling between states in a photonic Floquet medium, a unified non-Hermitian band structure description remains elusive. Here, we experimentally reveal the unique Bloch-Floquet and non-Bloch band structures of a photonic Floquet medium emulated in the microwave regime with a one-dimensional array of time-periodically driven resonators. Specifically, these non-Hermitian band structures are shown to be two measurable distinct subsets of complex eigenfrequency surfaces of the photonic Floquet medium defined in complex momentum space. In the Bloch-Floquet band structure, the driving-induced non-reciprocal coupling between oppositely signed frequency states leads to opening of momentum gaps along the real momentum axis, at the edges of which exceptional phase transitions occur. More interestingly, we show that the non-Bloch band structure defined in the complex Brillouin zone supplements the information on the morphology of complex eigenfrequency surfaces of the photonic Floquet medium. Our work paves the way for a comprehensive understanding of photonic Floquet media in complex energy-momentum space and could provide general guidelines for the study of non-equilibrium photonic phases of matter.

physics.app-ph

Resonance-enhanced spectral funneling in Fabry-Perot resonators with a temporal boundary mirror

A temporal boundary refers to a specific time at which the properties of an optical medium are abruptly changed. When light interacts with the temporal boundary, its spectral content can be redistributed due to the breaking of continuous time-translational symmetry of the medium where light resides. In this work, we use this principle to demonstrate, at terahertz (THz) frequencies, the resonance-enhanced spectral funneling of light coupled to a Fabry-Perot resonator with a temporal boundary mirror. To produce a temporal boundary effect, we abruptly increase the reflectance of a mirror constituting the Fabry-Perot resonator and, correspondingly, its quality factor in a step-like manner. The abrupt increase in the mirror reflectance leads to a trimming of the coupled THz pulse that causes the pulse to broaden in the spectral domain. Through this dynamic resonant process, the spectral contents of the input THz pulse are redistributed into the modal frequencies of the high-Q Fabry-Perot resonator formed after the temporal boundary. An energy conversion efficiency of up to 33% was recorded for funneling into the fundamental mode with a Fabry-Perot resonator exhibiting a sudden Q-factor change from 4.8 to 48. We anticipate that the proposed resonance-enhanced spectral funneling technique could be further utilized in the development of efficient mechanically tunable narrowband terahertz sources for diverse applications.

physics.optics

Linear frequency conversion via sudden merging of resonances in time-variant metasurfaces

Energy conversion in a physical system requires time-translation invariance breaking according to Noether's theorem. Closely associated with this symmetry-conservation relation, the frequencies of electromagnetic waves are found to be converted as the waves propagate through a temporally varying medium. Thus, effective temporal control of the medium, be it artificial or natural, through which the waves are propagating, lies at the heart of linear optical frequency conversion. Here, we propose rapidly time-variant metasurfaces as a frequency-conversion platform and experimentally demonstrate their efficacy at THz frequencies. The proposed metasurface is designed for the sudden merging of two distinct resonances into a single resonance upon ultrafast optical excitation. From this spectrally-engineered temporal boundary onward, the merged-resonance frequency component is radiated. In addition, temporal coherence of the two original resonating modes with respect to the abrupt temporal boundary is found to be strongly related to the amount of frequency conversion as well as the phase of the converted wave. Due to their design flexibility, time-variant metasurfaces may become on-demand frequency synthesizers for various frequency ranges.

physics.optics