SearcharxivSearch

arXiv subjects

Haejun Chung

Publications and source records attributed to Haejun Chung.

At least 19 recordsLinked to original sources

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.

cs.LG

Inverse-Designed Lithium Niobate Wavelength Demultiplexer via Birefringent Effective Index Approximation

Inverse design of thin-film lithium niobate (TFLN) photonic devices is computationally demanding because optical birefringence and fabrication-induced slanted sidewalls generally require three-dimensional electromagnetic models. We introduce a birefringent effective-index (BEI) method to reduce this problem to two dimensions while retaining polarization-dependent slab confinement and a representative cross section of the etched geometry. The method is integrated with adjoint topology optimization and fabrication constraints to design a 30 x 10 um demultiplexer that routes 1550 and 775 nm light to separate output ports. Quantitative comparisons with three-dimensional finite-difference time-domain simulations establish the accuracy and etch-depth dependence of the reduced model. The fabricated device provides mean signal-to-crosstalk ratios of 13.9 dB across 1540-1560 nm and 13.3 dB across 770-780 nm. A two-stage cascaded configuration increases the output extinction ratio to 26.8 dB in the telecom band and 20.3 dB in the near-visible band. This fabrication-aware reduced-dimensional approach enables optimizations of the multifunctional photonic devices for nonlinear optical and quantum applications on the TFLN platform.

physics.optics

Scaling Limits of Multichannel Spectral Routers for Snapshot Imaging

Inverse-designed spectral routers can enable compact snapshot spectral imaging by directing different wavelengths to designated detector sub-pixels without mechanical scanning or absorptive filters. Here, we examine how routing performance varies as the number of spectral channels increases while the visible bandwidth remains fixed. A delay-bandwidth analysis relates channel count, worst-channel efficiency, and device thickness, explaining the growing optical capacity required to generate many spatially distinct spectral outputs. We then use adjoint-based topology optimization and three-dimensional finite-difference time-domain simulations to design TiO2/SiO2 routers with 9, 16, 25, and 36 channels under matched material, bandwidth, and thickness conditions. The average routing efficiency increases with sub-pixel size and approaches a plateau, while the maximum near-plateau efficiency decreases monotonically from 97.0% for 9 channels to 82.3% for 36 channels. At a fixed channel count, compressing the wavelength spacing to 8 nm or rearranging the wavelength-to-sub-pixel assignment changes the efficiency by less than one percentage point. These results indicate that the observed efficiency penalty is governed mainly by the number of wavelength-dependent routing constraints rather than by wavelength spacing or local wavelength arrangement. The findings quantify the trade-off between spectral sampling and optical throughput in compact snapshot spectral imagers.

physics.optics

Nyquist-Sampled Time-Domain Adjoint FDTD for Memory-Efficient Broadband Nanophotonic Inverse Design

Adjoint optimization is a cornerstone of broadband nanophotonic inverse design, but conventional time-domain implementations face a severe memory bottleneck because they retain forward-field histories at every finite-difference time-domain (FDTD) time step. Here, we show that this full time-step storage is unnecessary for broadband design objectives because the underlying fields are band-limited. By storing forward fields only at Nyquist intervals and using the resulting sparse fields during the adjoint pass, the proposed method enables on-the-fly gradient accumulation without retaining full forward-field histories. This Nyquist-sampled adjoint FDTD framework preserves the two-simulation scaling of time-domain adjoint optimization while substantially reducing the dominant field-storage. Because the broadband gradient is evaluated directly in the time domain, with no spectral discretization, the per-iteration cost is independent of the number of frequencies---in contrast to frequency-sampled adjoint formulations, whose cost grows with spectral sampling density. Gradient verification confirms that Nyquist sampling reproduces conventional full-storage adjoint gradients with negligible error, whereas undersampling beyond the Nyquist limit produces aliasing-induced gradient degradation. Across four two-dimensional broadband nanophotonic benchmarks and a fully three-dimensional metalens, the method maintains gradient fidelity and optimized device performance while reducing dominant field-storage memory by more than $100\times$ relative to full-history storage in a prototypical example. These results suggest that the principal memory barrier in broadband time-domain adjoint FDTD is not an intrinsic requirement of gradient evaluation but rather a consequence of redundant temporal field storage, thereby opening a practical route to large-scale three-dimensional nanophotonic inverse design.

physics.optics

Exoplanet Detection Using Adaptive Quantum-Optimal Measurement

Detecting terrestrial exoplanets in the habitable zones of nearby stars remains a critical challenge. Such planets can be \(10^8\) to \(10^{10}\) times fainter than their host stars and lie at diffraction-limited angular separations, where starlight strongly obscures the companion signal. Here we present an adaptive quantum measurement method for estimating the number, positions, and brightnesses of mutually incoherent point sources in the sub-Rayleigh, ultra-high-contrast regime, operating at contrasts down to \(10^{-8}\) -- five orders of magnitude beyond previous quantum imaging approaches to exoplanet detection. The method adopts a spatial-mode basis that is updated to maximize the quantum Fisher information per detected photon. Estimation is performed by maximum likelihood in log-brightness coordinates, and the source count is determined by Bayesian-information-criterion (BIC) model selection directly from photon-count statistics, without a tunable detection threshold. For point sources within sub-Rayleigh separations and with brightness ratios spanning eight orders of magnitude, the method reconstructs complete scenes with a mean success rate of \(72.5\%\). Furthermore, it is robust to misalignment, maintaining a \(71.3\%\) success rate under offsets of up to six pixels. These results demonstrate that terrestrial exoplanets can be detected below the Rayleigh limit, a regime previously inaccessible to direct imaging.

physics.optics

VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attractive alternative but operate under strict physical efficiency limits. We propose CODA, a co-design framework that optimizes a continuous-density meta-optic front-end for frozen-model recognition using differentiable image formation and adjoint-gradient updates of Maxwell-based simulations. CODA directly optimizes the cross-entropy loss of a fixed zero-shot CLIP classifier without learned reconstruction, image signal processing, or image-fidelity auxiliary objectives. In a two-dimensional simulated imaging benchmark on ImageNet-100, CODA improves CLIP ViT-L/14 zero-shot accuracy from 53.75 $\pm$ 3.57$\%$ with a focal-concentration baseline to 65.41 $\pm$ 3.99$\%$. The optimized optics further transfer without re-optimization across CLIP, SigLIP, and DINOv2 on ImageNet-100, CIFAR-100, and Food-101. These results demonstrate that, under constrained meta-optic imaging, downstream recognition can be improved by aligning optical design with frozen vision-model objectives rather than conventional image-formation criteria.

cs.CV

CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images

Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer-based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a general-purpose backbone for 2D and 3D thin-structure segmentation. It employs cross-shaped stripe self-attention to model long-range principal-axis context and incorporates cyclic shifts to enhance information exchange across stripes. To better preserve fine-grained details, we further introduce a detail-enhanced multi-scale self-attention module that aggregates contextual features from multi-resolution representations. In addition, we propose sparse-control dynamic snake convolution, which reconstructs reliable dense curvilinear kernels from sparsely predicted control points to better follow tortuous geometry. Extensive experiments on four benchmarks across ophthalmology, neurovascular imaging, and dermatology demonstrate that CSWinUNETR consistently outperforms state-of-the-art methods without task-specific post-processing or topology-aware losses. The code is available at https://github.com/labhai/CSWinUNETR.

cs.CV

When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often require costly expert adjudication, even though structured clinical variables are routinely available in tabular form. Self-supervised learning can leverage these unlabeled tables, and recent binning-based pretexts offer a promising inductive bias, but existing objectives fix a single global quantile discretization and apply feature-agnostic supervision. We propose Adaptive Binning, a training-adaptive discretization pretext for tabular SSL that couples discretization to learning through a feature-wise coarse-to-fine curriculum. Motivated by the spectral bias of neural networks and the principles of curriculum learning, our method progressively refines discretization per feature upon plateau detection and selects representation-aware splits to jointly improve value-space concentration and representation-space coherence. A heterogeneity-aware objective unifies categorical reconstruction with ordinal supervision for numerical features, and experiments on public medical tabular datasets under unified evaluation protocols show consistent gains for linear probing and fine-tuning without dataset-specific discretization tuning. We further introduce a medical tabular SSL benchmark with standardized protocols to support reproducible progress in this underexplored domain. Our code is available at https://github.com/labhai/Adaptive-Binning.

cs.LG

OTCHA: Optimal Transport-driven Confidence-aware Latent Hub Alignment for Multi-View Medical Image Classification

Multi-view imaging, such as mammography and chest radiography, is a standard component of clinical practice. However, medical images are often unregistered and contain view-specific artifacts or irrelevant background cues that can obscure diagnostically relevant findings. Many existing methods directly fuse per-view representations, allowing such irrelevant content to contaminate the fused embedding and reducing robustness under varying view configurations. We propose OTCHA, a confidence-aware latent hub token alignment module based on optimal transport (OT) that refines patch tokens before fusion for multi-view classification. OTCHA introduces a set of learnable latent hub tokens shared across views. For each view, we compute an OT plan between patch tokens and hub tokens that jointly considers feature similarity and geometry, and augment the OT formulation with token-conditional dustbins to enable partial matching and discard irrelevant tokens. The resulting transport plan provides token-wise matching confidence, which gates hub-mediated message passing and weights a novel optimal-transport-based representation alignment loss to stabilize refinement. Experiments on three multi-view medical image datasets demonstrate consistent improvements over competing baselines across diverse anatomies and view configurations. Our code is available at https://github.com/labhai/OTCHA.

cs.CV

Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling

Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human comparisons ($O(n^2)$). While sorting-based methods have reduced this burden to $O(n\log n)$, they still require expensive human judgment for every single comparison. To further improve annotation efficiency, we propose leveraging a Vision-Language Model (VLM) not as an annotator replacement, but as a \emph{question prioritizer} to identify which comparisons genuinely require human judgment. The proposed \textbf{Surprise-Guided MergeSort (SGS)} framework achieves this through three integrated components: (1) a bottom-up MergeSort scheduler that structures comparisons and exploits transitivity, (2) a composite Surprise Scorer -- combining position-bias-cancelled VLM confidence, Elo gap, and vote entropy -- to quantify comparison ambiguity, and (3) an adaptive budget allocator that routes high-surprise pairs to humans while automating low-surprise pairs via transitivity inference. Validation was conducted on six diverse benchmarks spanning text similarity (STS-B, BIOSSES, SICKR-STS) and image quality assessment (KonIQ-10k, TID2013, LIVE Challenge). SGS effectively identified and skipped up to 535 non-informative comparisons per session. Consequently, it achieved Kendall's $\tau{\times}100$ improvements of $+6$ to $+12$ over Active Elo under the same total budget. These results demonstrate that combining VLM-guided surprise metrics with algorithmic sorting provides a generally consistent accuracy-efficiency trade-off across diverse domains.

cs.LG

MetaRanker: Human-in-the-loop Active Ranking for Metalens Image Quality

Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses offer a path toward substantial miniaturization of optical modules, but practical designs often exhibit pronounced chromatic and field-dependent aberrations that necessitate computational reconstruction. In current metalens pipelines, reconstruction models are commonly trained and selected using distortion-based fidelity objectives, such as PSNR, yet these proxies can be weakly correlated with human preference and downstream utility, reflecting the well-known perception--distortion trade-off. We introduce MetaRanker, a human-in-the-loop active ranking framework that formalizes metalens image quality in terms of semantic interpretability, defined as the degree to which humans can reliably recognize objects and structures in the presence of optical artifacts. MetaRanker combines a probabilistic preference model with uncertainty-aware query selection, and leverages vision--language models to provide lightweight semantic priors. Importantly, these priors are used only to guide the sampling of informative comparisons; human judgments remain the primary supervision signal throughout. Across real-world and synthetic metalens datasets with distinct degradation profiles, MetaRanker produces rankings that align most closely with human assessments, while reducing the number of pairwise annotations required by approximately 80% relative to exhaustive pairwise evaluation. Finally, we show that standard image quality assessment metrics exhibit limited alignment with human interpretability in the metalens domain, positioning MetaRanker as a practical step toward perceptually grounded metalens evaluation and co-design.

cs.CV

Near-unity efficiency optical vortex generation in van der Waals materials

Optical spin-orbit coupling provides a promising, fabrication-free route for developing ultra-compact optical vortex generators. However, the conversion efficiency has been theoretically limited to 0.5. Here, we demonstrate enhanced vortex generation efficiency by employing a Bessel beam as the input and propagating it through van der Waals (vdW) crystals. The large birefringence of vdW crystals and the single transverse wave vector of a Bessel beam allow a near unity spin-orbit conversion efficiency and a topological charge transition of $\ell \rightarrow \ell + 2$. Through combined analytical and experimental investigations, we demonstrate a conversion efficiency of up to 0.82 in hexagonal boron nitride (hBN) crystals with a thickness of $27.4\,\mu\mathrm{m}$. The higher efficiency of Bessel input beams over Gaussian beams is attributed to their distinct transverse wave vector distribution of constituent plane wave components. Furthermore, we demonstrate the dependence of conversion efficiency on the numerical aperture (NA) of the objective lens, which is in good alignment with theoretical predictions. These demonstrations provide a fabrication-free route to highly efficient optical vortex generation via microscale vdW materials platforms.

physics.optics

Neural Adjoint Method for Meta-optics: Accelerating Volumetric Inverse Design via Fourier Neural Operators

Meta-optics promises compact, high-performance imaging and color routing. However, designing high-performance structures is a high-dimensional optimization problem: mapping a desired optical output back to a physical 3D structure requires solving computationally expensive Maxwell's equations iteratively. Even with adjoint optimization, broadband design can require thousands of Maxwell solves, making industrial-scale optimization slow and costly. To overcome this challenge, we propose the Neural Adjoint Method, a solver-supervised surrogate that predicts 3D adjoint gradient fields from a voxelized permittivity volume using a Fourier Neural Operator (FNO). By learning the dense, per-voxel sensitivity field that drives gradient-based updates, our method can replace per-iteration adjoint solves with fast predictions, greatly reducing the computational cost of full-wave simulations required during iterative refinement. To better preserve sensitivity peaks, we introduce a stage-wise FNO that progressively refines residual errors with increasing emphasis on higher-frequency components. We curate a meta-optics dataset from paired forward/adjoint FDTD simulations and evaluate it across three tasks: spectral sorting (color routers), achromatic focusing (metalenses), and waveguide mode conversion. Our method reduces design time from hours to seconds. These results suggest a practical route toward fast, large-scale volumetric meta-optical design enabled by AI-accelerated scientific computing.

cs.LG

Inverse-Designed Metasurfaces for Compact Optical Skyrmion Generation with High Topological Fidelity

Optical skyrmions are structured vector fields with nontrivial polarization topology and subwavelength-scale features. One common approach to generating optical skyrmions is the superposition of a zeroth-order Bessel beam and a higher-order Bessel beam carrying orbital angular momentum, with each beam possessing an orthogonal circular polarization state. However, creating such complex beams typically requires bulky free-space optical setups; therefore, recent efforts have focused on compact optical skyrmion generators based on metasurfaces. Nevertheless, achieving the degrees of freedom required for simultaneous phase and polarization control remains challenging because of the limited design flexibility of conventional meta-atoms. Here, we address this challenge by employing an inverse-design approach and demonstrate a single-layer metasurface that generates high-fidelity optical skyrmions. We employ an adjoint-based topology-optimization method to design a silicon metasurface that converts an incident beam into an optical skyrmion without the need for additional optical components. The optimized metasurface generates an optical skyrmion with skyrmion number $(N_\mathrm{sk}) = 0.970$. This work demonstrates that inverse design can be a promising route to compact skyrmion generators, and our approach provides a basis for near-field particle manipulation and the generation of independent topological bits in dense photonic integration.

physics.optics

Dodgersort: Uncertainty-Aware VLM-Guided Human-in-the-Loop Pairwise Ranking

Pairwise comparison labeling is emerging as it yields higher inter-rater reliability than conventional classification labeling, but exhaustive comparisons require quadratic cost. We propose Dodgersort, which leverages CLIP-based hierarchical pre-ordering, a neural ranking head and probabilistic ensemble (Elo, BTL, GP), epistemic--aleatoric uncertainty decomposition, and information-theoretic pair selection. It reduces human comparisons while improving the reliability of the rankings. In visual ranking tasks in medical imaging, historical dating, and aesthetics, Dodgersort achieves a 11--16\% annotation reduction while improving inter-rater reliability. Cross-domain ablations across four datasets show that neural adaptation and ensemble uncertainty are key to this gain. In FG-NET with ground-truth ages, the framework extracts 5--20$\times$ more ranking information per comparison than baselines, yielding Pareto-optimal accuracy--efficiency trade-offs.

cs.CV

EZ-Sort: Efficient Pairwise Comparison via Zero-Shot CLIP-Based Pre-Ordering and Human-in-the-Loop Sorting

Pairwise comparison is often favored over absolute rating or ordinal classification in subjective or difficult annotation tasks due to its improved reliability. However, exhaustive comparisons require a massive number of annotations (O(n^2)). Recent work has greatly reduced the annotation burden (O(n log n)) by actively sampling pairwise comparisons using a sorting algorithm. We further improve annotation efficiency by (1) roughly pre-ordering items using the Contrastive Language-Image Pre-training (CLIP) model hierarchically without training, and (2) replacing easy, obvious human comparisons with automated comparisons. The proposed EZ-Sort first produces a CLIP-based zero-shot pre-ordering, then initializes bucket-aware Elo scores, and finally runs an uncertainty-guided human-in-the-loop MergeSort. Validation was conducted using various datasets: face-age estimation (FGNET), historical image chronology (DHCI), and retinal image quality assessment (EyePACS). It showed that EZ-Sort reduced human annotation cost by 90.5% compared to exhaustive pairwise comparisons and by 19.8% compared to prior work (when n = 100), while improving or maintaining inter-rater reliability. These results demonstrate that combining CLIP-based priors with uncertainty-aware sampling yields an efficient and scalable solution for pairwise ranking.

cs.CV

Inverse-Designed Metasurfaces for Wavefront Restoration in Under-Display Camera Systems

Under-display camera (UDC) systems enable full-screen displays in smartphones by embedding the camera beneath the display panel, eliminating the need for notches or punch holes. However, the periodic pixel structures of display panels introduce significant optical diffraction effects, leading to imaging artifacts and degraded visual quality. Conventional approaches to mitigate these distortions, such as deep learning-based image reconstruction, are often computationally expensive and unsuitable for real-time applications in consumer electronics. This work introduces an inverse-designed metasurface for wavefront restoration, addressing diffraction-induced distortions without relying on external software processing. The proposed metasurface effectively suppresses higher-order diffraction modes caused by the metallic pixel structures, restores the optical wavefront, and enhances imaging quality across multiple wavelengths. By eliminating the need for software-based post-processing, our approach establishes a scalable, real-time optical solution for diffraction management in UDC systems. This advancement paves the way to achieve software-free real-time image restoration frameworks for many industrial applications.

physics.optics

Physics-guided and fabrication-aware inverse design of photonic devices using diffusion models

Designing free-form photonic devices is fundamentally challenging due to the vast number of possible geometries and the complex requirements of fabrication constraints. Traditional inverse-design approaches--whether driven by human intuition, global optimization, or adjoint-based gradient methods--often involve intricate binarization and filtering steps, while recent deep learning strategies demand prohibitively large numbers of simulations (10^5 to 10^6). To overcome these limitations, we present AdjointDiffusion, a physics-guided framework that integrates adjoint sensitivity gradients into the sampling process of diffusion models. AdjointDiffusion begins by training a diffusion network on a synthetic, fabrication-aware dataset of binary masks. During inference, we compute the adjoint gradient of a candidate structure and inject this physics-based guidance at each denoising step, steering the generative process toward high figure-of-merit (FoM) solutions without additional post-processing. We demonstrate our method on two canonical photonic design problems--a bent waveguide and a CMOS image sensor color router--and show that our method consistently outperforms state-of-the-art nonlinear optimizers (such as MMA and SLSQP) in both efficiency and manufacturability, while using orders of magnitude fewer simulations (approximately 2 x 10^2) than pure deep learning approaches (approximately 10^5 to 10^6). By eliminating complex binarization schedules and minimizing simulation overhead, AdjointDiffusion offers a streamlined, simulation-efficient, and fabrication-aware pipeline for next-generation photonic device design. Our open-source implementation is available at https://github.com/dongjin-seo2020/AdjointDiffusion.

physics.optics