SearcharxivSearch

arXiv subjects

Hyunwoong Kim

Publications and source records attributed to Hyunwoong Kim.

6 recordsLinked to original sources

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing reveals a mismatch in scale: each finding occupies only a small region of the image, yet supervision is provided only at the global image-report level. This poses a central challenge: prior approaches spread weight densely across all patches rather than concentrating on the sparse subset relevant to a given query. To address this, we present GLINT (Gated Language-Image alignmeNT), a framework that explicitly models this sparse correspondence. On the alignment side, we introduce Sparsely Gated Alignment, a novel architecture in which a sigmoid gate over a separate gate embedding space activates only the patches relevant to each textual query, enforcing explicit sparsity. On the representation side, we add Dense Feature Regularization, which anchors the trainable encoder's intermediate features to a frozen self-supervised learning (SSL) teacher, preserving the fine-grained patch features that the gate relies on. The same recipe applies to both 2D chest X-ray (CXR) and 3D chest computed tomography (CT), built with DINOv3 and V-JEPA 2.1, respectively. GLINT enables zero-shot classification, grounding, and segmentation from free-text queries, and to our knowledge is the first to demonstrate zero-shot segmentation on 3D CT volumes without mask supervision. Notably, the most pronounced gains arise on zero-shot grounding and segmentation, where sparse, query-specific localization is required, consistent with our design intent. In downstream evaluation, GLINT outperforms both SSL encoders and medical VLMs on classification, report generation, and segmentation.

cs.CV

SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation

Group Relative Policy Optimization (GRPO) and its variants, originally developed for Large Language Models (LLMs), have recently been applied to Multimodal LLMs and produced strong results. However, their coarse-grained holistic credit assignment from a single scalar advantage underfits vision-language (VL) tasks, where outputs are often long-form responses grounded in semantically rich images. To address this limitation, we exploit a structured signal that single-scalar formulations discard: the natural segmentation of long-form VL outputs. Concretely, we propose Segment-Decomposed GRPO (SD-GRPO), which z-normalizes verifiable per-segment rewards across the rollout group, yielding a vector of per-segment advantages in place of a single scalar. We evaluate SD-GRPO across three settings spanning controlled and real-world long-form VL generation, organized by increasing semantic entanglement across segments. On a controlled multi-panel dense-captioning task constructed from DOCCI, where segments are semantically independent, SD-GRPO consistently outperforms the GRPO baseline, with larger gains at higher segment counts. Extending to a controlled multi-chart long-form VQA task constructed from MultiChartQA, we show both theoretically and empirically that rollout-level rewards suffer from cross-segment credit misattribution that scales with output length. On a real-world scientific figure captioning task on the MMSci dataset, where subfigure captions share context across the figure, blending holistic and per-segment rewards further improves on both, suggesting per-segment normalization alone is insufficient when segments are semantically entangled. Finally, by integrating SD-GRPO into Dr. GRPO, we confirm that it can be applied to any GRPO framework with minimal implementation overhead to enhance long-form VL generation.

cs.CV

Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion

Image-to-Image translation models can help mitigate various challenges inherent to medical image acquisition. Latent diffusion models (LDMs) leverage efficient learning in compressed latent space and constitute the core of state-of-the-art generative image models. However, this efficiency comes with a trade-off, potentially compromising crucial pixel-level detail essential for high-fidelity medical images. This limitation becomes particularly critical when generating clinically significant structures, such as lesions, which often occupy only a small portion of the image. Failure to accurately reconstruct these regions can severely impact diagnostic reliability and clinical decision-making. To overcome this limitation, we propose a novel post-training framework for LDMs in medical image-to-image translation by incorporating lesion-aware medical pixel space objectives. This approach is essential, as it not only enhances overall image quality but also improves the precision of lesion delineation. We evaluate our framework on brain CT-to-MRI translation in acute ischemic stroke patients, where early and accurate diagnosis is critical for optimal treatment selection and improved patient outcomes. While diffusion MRI is the gold standard for stroke diagnosis, its clinical utility is often constrained by high costs and low accessibility. Using a dataset of 817 patients, we demonstrate that our framework improves overall image quality and enhances lesion delineation when synthesizing DWI and ADC images from CT perfusion scans, outperforming existing image-to-image translation models. Furthermore, our post-training strategy is easily adaptable to pre-trained LDMs and exhibits substantial potential for broader applications across diverse medical image translation tasks.

cs.CV

Extraction of higher-order nonlinear electronic response to strong field excitation in solids using high harmonic generation

State-of-the-art experiments employ strong ultrafast optical fields to study the nonlinear response of electrons in solids on an attosecond time-scale. Notably, a recent experiment retrieved a 3rd order nonlinear susceptibility by comparing the nonlinear response induced by a strong laser field to a linear response induced by the otherwise identical weak field. In parallel, experiments have demonstrated high harmonic generation (HHG) in solids, a highly nonlinear process that until recently had only been observed in gases. The highly nonlinear nature of HHG has the potential to extract even higher order nonlinear susceptibility terms, and thereby characterize the entire response of the electronic system to strong field excitation. However, up till now, such characterization has been elusive due to a lack of direct correspondence between high harmonics and nonlinear susceptibilities. Here, we demonstrate a regime where such correspondence can be clearly made, extracting nonlinear susceptibilities (7th, 9th, and 11th) from sapphire of the same order as the measured high harmonics. The extracted high order susceptibilities show angular-resolved periodicities arising from variation in the band structure with crystal orientation. Nonlinear susceptibilities are key to ultrafast lightwave driven optoelectronics, allowing petahertz scaling manipulation of the signal. Our results open a door to multi-channel signal processing, controlled by laser polarization.

physics.optics

Spectral Interference in High Harmonic Generation from Solids

Various interference effects are known to exist in the process of high harmonic generation (HHG) both at the single atom and macroscopic levels. In particular, the quantum path difference between the long and short trajectories of electron excursion causes the HHG yield to experience interference-based temporal and spectral modulations. In solids, due to additional phenomena such as multi-band superposition and crystal symmetry dependency, the HHG mechanism appears to be more complicated than in gaseous atoms in identifying accompanying interference phenomena. Here, we first report experimental data showing intensity-dependent spectral modulation and broadening of high harmonics observed from bulk sapphire. Then, by adopting theoretical simulation, the extraordinary observation is interpreted as a result of the quantum path interference between the long and short electron/hole trajectories. Specifically, the long trajectory undergoes an intensity-dependent redshift, which coherently combines with the short trajectory to exhibit spectral splitting in an anomalous way of inverse proportion to the driving laser intensity. This quantum interference may be extended to higher harmonics with increasing the laser intensity, underpinning the potential for precise control of the phase matching and modulation even in the extreme ultraviolet and soft X-ray regime. Further, this approach may act as a novel tool for probing arbitrary crystals so as to adjust the electron dynamics of higher harmonics for attosecond spectroscopy.

physics.optics

Surface Plasmon Assisted Gentle Ablation of Nanostructures by Femtosecond Oscillator

We experimentally demonstrate the use of subwavelength optical nanoantennae to assist the gentle ablation of nanostructures directly using ultralow fluence from a Ti: sapphire oscillator through the excitation of surface plasmon waves. We show that this ablation mechanism is the same for metal and dielectric. The analytical solutions of ablation threshold are in excellent agreement with the experiment estimations. Surface plasmon assisted locally enhanced ablation at nanoscale provides a method for nanomachining, manipulation and modification the nanostructures without collateral thermal damage to the materials. It is also shown that this ablation can deposit low-density high quality thin nano film.

physics.optics