SearcharxivSearch

arXiv subjects

Byeongchan Lee

Publications and source records attributed to Byeongchan Lee.

8 recordsLinked to original sources

CLIP Meets Diffusion: A Synergistic Approach to Anomaly Detection

Anomaly detection is a complex problem due to the ambiguity in defining anomalies, the diversity of anomaly types (e.g., local and global defect), and the scarcity of training data. As such, it necessitates a comprehensive model capable of capturing both low-level and high-level features, even with limited data. To address this, we propose CLIPFUSION, a method that leverages both discriminative and generative foundation models. Specifically, the CLIP-based discriminative model excels at capturing global features, while the diffusion-based generative model effectively captures local details, creating a synergistic and complementary approach. Notably, we introduce a methodology for utilizing cross-attention maps and feature maps extracted from diffusion models specifically for anomaly detection. Experimental results on benchmark datasets (MVTec-AD, VisA) demonstrate that CLIPFUSION consistently outperforms baseline methods, achieving outstanding performance in both anomaly segmentation and classification. We believe that our method underscores the effectiveness of multi-modal and multi-model fusion in tackling the multifaceted challenges of anomaly detection, providing a scalable solution for real-world applications.

cs.CV

Understanding Self-supervised Contrastive Learning through Supervised Objectives

Self-supervised representation learning has achieved impressive empirical success, yet its theoretical understanding remains limited. In this work, we provide a theoretical perspective by formulating self-supervised representation learning as an approximation to supervised representation learning objectives. Based on this formulation, we derive a loss function closely related to popular contrastive losses such as InfoNCE, offering insight into their underlying principles. Our derivation naturally introduces the concepts of prototype representation bias and a balanced contrastive loss, which help explain and improve the behavior of self-supervised learning algorithms. We further show how components of our theoretical framework correspond to established practices in contrastive learning. Finally, we empirically validate the effect of balancing positive and negative pair interactions. All theoretical proofs are provided in the appendix, and our code is included in the supplementary material.

cs.LG

Efficient LLM Collaboration via Planning

Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achieve remarkable results across diverse tasks, they often incur substantial monetary inference cost, making frequent use impractical for many applications. In contrast, small models are often freely available and easy to deploy locally, but their performance on complex tasks remains limited. This trade-off raises a natural question: how can small and large models efficiently collaborate to combine their complementary strengths? To bridge this trade-off, we propose COPE, a test-time collaboration framework. A planner model first generates a plan that serves as a lightweight intermediate that guides a downstream executor model. Small and large models take turns acting as planner and executor, exchanging plans in a multi-stage cascade to collaboratively solve tasks. Through comprehensive experiments on benchmarks spanning mathematical reasoning, code generation, open-ended tasks, and agent tasks, we demonstrate that COPE achieves performance comparable to large proprietary models, while drastically reducing the inference API cost. These results highlight planning as an effective prior for cost-efficient inference.

cs.AI

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Selecting a layer normalization (LN) strategy that stabilizes training and speeds convergence in Transformers remains difficult, even for today's large language models (LLM). We present a comprehensive analytical foundation for understanding how different LN strategies influence training dynamics in large-scale Transformers. Until recently, Pre-LN and Post-LN have long dominated practices despite their limitations in large-scale training. However, several open-source models have recently begun silently adopting a third strategy without much explanation. This strategy places normalization layer peripherally around sublayers, a design we term Peri-LN. While Peri-LN has demonstrated promising performance, its precise mechanisms and benefits remain almost unexplored. Our in-depth analysis delineates the distinct behaviors of LN strategies, showing how each placement shapes activation variance and gradient propagation. To validate our theoretical insight, we conduct extensive experiments on Transformers up to $3.2$B parameters, showing that Peri-LN consistently achieves more balanced variance growth, steadier gradient flow, and convergence stability. Our results suggest that Peri-LN warrants broader consideration for large-scale Transformer architectures, providing renewed insights into the optimal placement of LN.

cs.LG

Implicit Contrastive Representation Learning with Guided Stop-gradient

In self-supervised representation learning, Siamese networks are a natural architecture for learning transformation-invariance by bringing representations of positive pairs closer together. But it is prone to collapse into a degenerate solution. To address the issue, in contrastive learning, a contrastive loss is used to prevent collapse by moving representations of negative pairs away from each other. But it is known that algorithms with negative sampling are not robust to a reduction in the number of negative samples. So, on the other hand, there are algorithms that do not use negative pairs. Many positive-only algorithms adopt asymmetric network architecture consisting of source and target encoders as a key factor in coping with collapse. By exploiting the asymmetric architecture, we introduce a methodology to implicitly incorporate the idea of contrastive learning. As its implementation, we present a novel method guided stop-gradient. We apply our method to benchmark algorithms SimSiam and BYOL and show that our method stabilizes training and boosts performance. We also show that the algorithms with our method work well with small batch sizes and do not collapse even when there is no predictor. The code is available at https://github.com/bych-lee/gsg.

cs.LG

The Tersoff potential for extreme environment

A novel modification of the Tersoff potential for Si is presented. The modification improves the transferability of the Tersoff potential for liquid states without the change of original parameters and with no alteration of bulk properties. Also, the modification introduces a correction term for high-pressure states. The modification is meaningful considering that by high energy irradiations local liquid structures and unstable high-pressure manifolds may occur, therefore an interatomic potential must have an acceptable reliability on high thermal/pressure situations to simulate such phenomenon. Particularly, in the modification, a new screening function replaces the radial cutoff function and the bond order function is slightly changed. Also, a repulsive energy function is replaced by a correction function within a specific pair distance.

cond-mat.mtrl-sci

First-principles calculation of mechanical properties of Si <001> nanowires and comparison to nanomechanical theory

We report the results of first-principles density functional theory calculations of the Young's modulus and other mechanical properties of hydrogen-passivated Si <001> nanowires. The nanowires are taken to have predominantly {100} surfaces, with small {110} facets according to the Wulff shape. The Young's modulus, the equilibrium length and the constrained residual stress of a series of prismatic beams of differing sizes are found to have size dependences that scale like the surface area to volume ratio for all but the smallest beam. The results are compared with a continuum model and the results of classical atomistic calculations based on an empirical potential. We attribute the size dependence to specific physical structures and interactions. In particular, the hydrogen interactions on the surface and the charge density variations within the beam are quantified and used both to parameterize the continuum model and to account for the discrepancies between the two models and the first-principles results.

cond-mat.mtrl-sci

First-principles study of the Young's modulus of Si <001> nanowires

We report the results of first-principles density functional theory calculations of the Young's modulus and other mechanical properties of hydrogen-passivated Si <001> nanowires. The nanowires are taken to have predominantly {100} surfaces, with small {110} facets. The Young's modulus, the equilibrium length and the residual stress of a series of prismatic wires are found to have a size dependence that scales like the surface area to volume ratio for all but the smallest wires. We analyze the physical origin of the size dependence, and compare the results to two existing models.

cond-mat.mtrl-sci