SearcharxivSearch

arXiv subjects

Wonseok Shin

Publications and source records attributed to Wonseok Shin.

8 recordsLinked to original sources

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Diffusion multimodal large language models (dMLLMs) have recently emerged as a new decoding paradigm for multimodal generation. Starting from a fully masked sequence, dMLLMs progressively decode the sequence by unmasking a subset of the remaining masked positions at each step. Since the selected tokens serve as the prediction context for subsequent steps, deciding which tokens to decode is crucial to the quality of the final output. The most common strategy prioritizes tokens based on a certainty measure that tends to favor tokens frequently observed in the training data. Recent approaches instead order tokens according to their influence on subsequent predictions, but do not explicitly account for the input image. We propose the Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens. We further impose a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded subset. Extensive experiments on 7 captioning and VQA benchmarks with 3 open-source dMLLMs demonstrate the effectiveness of VIG-Sampler, which outperforms the Info-Gain Sampler by an average of 19.3 CIDEr points across the captioning benchmarks and surpasses it on COCO Caption while using only half as many decoding steps.

cs.CV

Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders

Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition. This opacity hinders the alignment of retrieval processes with human intent, as the entangled representations are difficult to analyze or control. In this work, we propose a method to disentangle the dense representations of sentence transformers (e.g., E5) into human-interpretable concepts using Top-k Sparse Autoencoders (SAEs). We demonstrate that these disentangled features align with specific semantic, syntactic, and pragmatic categories. Furthermore, we introduce an activation steering mechanism that allows for precise intervention in the retrieval process. By clamping specific latent features, we show that it is possible to re-rank search results to better align with user constraints without retraining the backbone model. Our findings suggest that SAE-based decomposition offers a viable path toward transparent and steerable neural information retrieval.

cs.IR

Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

Sparse Autoencoders (SAEs) have demonstrated significant success in interpreting Large Language Models (LLMs) by decomposing dense representations into sparse, semantic components. However, their potential for analyzing Vision Transformers (ViTs) remains largely under-explored. In this work, we present the first application of SAEs to the ViT [CLS] token for out-of-distribution (OOD) detection, addressing the limitation of existing methods that rely on entangled feature representations. We propose a novel framework utilizing a Top-k SAE to disentangle the dense [CLS] features into a structured latent space. Through this analysis, we reveal that in-distribution (ID) data exhibits consistent, class-specific activation patterns, which we formalize as Class Activation Profiles (CAPs). Our study uncovers a key structural invariant: while ID samples preserve a stable pattern within CAPs, OOD samples systematically disrupt this structure. Leveraging this insight, we introduce a scoring function based on the divergence of core energy profiles to quantify the deviation from ideal activation profiles. Our method achieves strong results on the FPR95 metric, critical for safety-sensitive applications across multiple benchmarks, while also achieving competitive AUROC. Overall, our findings demonstrate that the sparse, disentangled features revealed by SAEs can serve as a powerful, interpretable tool for robust OOD detection in vision models.

cs.CV

An Integrated Ultralow Noise Spiral Interferometric Laser

Photonic integration offers the potential to bring complex high-performance optical systems to the form factor of a compact semiconductor chip. However, the range of system functions accessible critically depends on the extent to which free-space and fiber components can be made integrable. The ultralow-expansion cavity-stabilized laser$-$often used in precision metrology, high-resolution sensors, and advanced systems in atomic physics$-$is one component that currently has no direct parallel on chip. Lasers stabilized to photonically-integrated resonators exist, but exhibit considerably higher frequency noise and are accompanied by large levels of frequency drift. We demonstrate here a new architecture for an ultranarrow linewidth integrated laser based on stabilization to a sinusoidal fringe of an interferometer having a long 25-m unbalanced delay line. Our interferometric laser not only advances the state-of-the-art for on-chip lasers, but we in addition introduce an amplitude locking scheme that greatly suppresses the laser's long-term frequency wander. We achieve a record on-chip fractional frequency noise of $5.6 \times 10^{-14}$, corresponding to a linewidth of 12 Hz centered at 1348 nm. To showcase the utility of this laser, we divide the optical carrier to microwave frequencies, demonstrating the ability to outperform state-of-the-art quartz crystal oscillators by 15 dB or more.

physics.optics

Efficient Hypergraph Pattern Matching via Match-and-Filter and Intersection Constraint

A hypergraph is a generalization of a graph, in which a hyperedge can connect multiple vertices, modeling complex relationships involving multiple vertices simultaneously. Hypergraph pattern matching, which is to find all isomorphic embeddings of a query hypergraph in a data hypergraph, is one of the fundamental problems. In this paper, we present a novel algorithm for hypergraph pattern matching by introducing (1) the intersection constraint, a necessary and sufficient condition for valid embeddings, which significantly speeds up the verification process, (2) the candidate hyperedge space, a data structure that stores potential mappings between hyperedges in the query hypergraph and the data hypergraph, and (3) the Match-and-Filter framework, which interleaves matching and filtering operations to maintain only compatible candidates in the candidate hyperedge space during backtracking. Experimental results on real-world datasets demonstrate that our algorithm significantly outperforms the state-of-the-art algorithms, by up to orders of magnitude in terms of query processing time.

cs.DB

Cardinality Estimation of Subgraph Matching: A Filtering-Sampling Approach

Subgraph counting is a fundamental problem in understanding and analyzing graph structured data, yet computationally challenging. This calls for an accurate and efficient algorithm for Subgraph Cardinality Estimation, which is to estimate the number of all isomorphic embeddings of a query graph in a data graph. We present FaSTest, a novel algorithm that combines (1) a powerful filtering technique to significantly reduce the sample space, (2) an adaptive tree sampling algorithm for accurate and efficient estimation, and (3) a worst-case optimal stratified graph sampling algorithm for difficult instances. Extensive experiments on real-world datasets show that FaSTest outperforms state-of-the-art sampling-based methods by up to two orders of magnitude and GNN-based methods by up to three orders of magnitude in terms of accuracy.

cs.DB

Triangular Contrastive Learning on Molecular Graphs

Recent contrastive learning methods have shown to be effective in various tasks, learning generalizable representations invariant to data augmentation thereby leading to state of the art performances. Regarding the multifaceted nature of large unlabeled data used in self-supervised learning while majority of real-word downstream tasks use single format of data, a multimodal framework that can train single modality to learn diverse perspectives from other modalities is an important challenge. In this paper, we propose TriCL (Triangular Contrastive Learning), a universal framework for trimodal contrastive learning. TriCL takes advantage of Triangular Area Loss, a novel intermodal contrastive loss that learns the angular geometry of the embedding space through simultaneously contrasting the area of positive and negative triplets. Systematic observation on embedding space in terms of alignment and uniformity showed that Triangular Area Loss can address the line-collapsing problem by discriminating modalities by angle. Our experimental results also demonstrate the outperformance of TriCL on downstream task of molecular property prediction which implies that the advantages of the embedding space indeed benefits the performance on downstream tasks.

cs.LG

Inverse design of large-area metasurfaces

We present a computational framework for efficient optimization-based "inverse design" of large-area "metasurfaces" (subwavelength-patterned surfaces) for applications such as multi-wavelength and multi-angle optimizations, and demultiplexers. To optimize surfaces that can be thousands of wavelengths in diameter, with thousands (or millions) of parameters, the key is a fast approximate solver for the scattered field. We employ a "locally periodic" approximation in which the scattering problem is approximated by a composition of periodic scattering problems from each unit cell of the surface, and validate it against brute-force Maxwell solutions. This is an extension of ideas in previous metasurface designs, but with greatly increased flexibility, e.g. to automatically balance tradeoffs between multiple frequencies, or to optimize a photonic device given only partial information about the desired field. Our approach even extends beyond the metasurface regime to non-subwavelength structures where additional diffracted orders must be included (but the period is not large enough to apply scalar diffraction theory).

physics.optics