SearcharxivSearch

arXiv subjects

Sangyeon Cho

Publications and source records attributed to Sangyeon Cho.

10 recordsLinked to original sources

Half-wave plasmonic nanolasers near the localization limit

Miniaturized lasers with sub-micron dimensions are of broad interest for optical science, on-chip communication, sensing, and biomedical barcoding. Recently, lowest-order half-wave lasing was demonstrated in semiconductor-on-metal cavities by operating away from the highly dispersive and absorptive surface-plasmon resonance. Here, we demonstrate half-wave-mode lasing near the surface-plasmon resonance at ~630 nm using high-gain indium phosphide (InP) nanoparticles on ultrasmooth gold substrates. The smallest lasing particle, estimated from simulated dispersion curves to have a length of ~115 nm and height of ~100 nm, emitted at 730 nm in air, representing one of the smallest reported active laser cavities. Linewidth and threshold pump fluence generally decreased as the lasing wavelength shifted farther from the plasmon resonance. In larger particles with lengths of 280~480 nm, we observed lasing attributable to second- and third-order plasmonic modes with progressively narrower linewidths. These results extend half-wave dipolar lasing toward near-infrared and visible wavelengths and further push laser miniaturization toward the plasmonic localization limit.

physics.optics

Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation

While LLMs are increasingly used in commercial services, they pose privacy risks such as leakage of sensitive personally identifiable information (PII). For LLMs trained on multilingual corpora, Multilingual Machine Unlearning (MMU) aims to remove information across multiple languages. However, prior MMU evaluations fail to capture such cross-linguistic distribution of information, being largely limited to direct extensions of per-language evaluation protocols. To this end, we propose two metrics to evaluate the information spread across languages: the Knowledge Separability Score (KSS) and the Knowledge Persistence Score (KPS). KSS measures the overall unlearning quality across multiple languages, while KPS more specifically aims to assess consistent removal of information among different language pairs. We evaluated various unlearning methods in the multilingual setting with these metrics and conducted comprehensive analyses. Through our investigation, we provide insights into unique phenomena exclusive to MMU and offer a new perspective on MMU evaluation.

cs.CL

Air-Stable Room-Temperature Quasi-2D Tin Iodide Perovskite Microlasers

Quasi-2D tin iodide perovskites (TIPs) are promising lead-free alternatives for optoelectronic applications, but achieving stable lasing remains challenging due to their limited environmental stability. Here, we report air-stable, room-temperature lasing from quasi-2D TIP microcrystals as small as 4 {\mu}m. Incorporation of the organic spacer 5IPA3 significantly enhanced the stability of these materials compared to previously reported TIPs. Lasing was observed from both dielectric (n=4) and plasmonic (n=3 and n=4) TIP microlasers. Under picosecond pumping, lasing was sustained for over 10^8 pump pulses in ambient conditions. These results represent a significant step toward practical photonic applications of tin-based perovskites.

physics.optics

Synergy-CLIP: Extending CLIP with Multi-modal Integration for Robust Representation Learning

Multi-modal representation learning has become a pivotal area in artificial intelligence, enabling the integration of diverse modalities such as vision, text, and audio to solve complex problems. However, existing approaches predominantly focus on bimodal interactions, such as image-text pairs, which limits their ability to fully exploit the richness of multi-modal data. Furthermore, the integration of modalities in equal-scale environments remains underexplored due to the challenges of constructing large-scale, balanced datasets. In this study, we propose Synergy-CLIP, a novel framework that extends the contrastive language-image pre-training (CLIP) architecture to enhance multi-modal representation learning by integrating visual, textual, and audio modalities. Unlike existing methods that focus on adapting individual modalities to vanilla-CLIP, Synergy-CLIP aligns and captures latent information across three modalities equally. To address the high cost of constructing large-scale multi-modal datasets, we introduce VGG-sound+, a triple-modal dataset designed to provide equal-scale representation of visual, textual, and audio data. Synergy-CLIP is validated on various downstream tasks, including zero-shot classification, where it outperforms existing baselines. Additionally, we introduce a missing modality reconstruction task, demonstrating Synergy-CLIP's ability to extract synergy among modalities in realistic application scenarios. These contributions provide a robust foundation for advancing multi-modal representation learning and exploring new research directions.

cs.LG

BioBridge: Unified Bio-Embedding with Bridging Modality in Code-Switched EMR

Pediatric Emergency Department (PED) overcrowding presents a significant global challenge, prompting the need for efficient solutions. This paper introduces the BioBridge framework, a novel approach that applies Natural Language Processing (NLP) to Electronic Medical Records (EMRs) in written free-text form to enhance decision-making in PED. In non-English speaking countries, such as South Korea, EMR data is often written in a Code-Switching (CS) format that mixes the native language with English, with most code-switched English words having clinical significance. The BioBridge framework consists of two core modules: "bridging modality in context" and "unified bio-embedding." The "bridging modality in context" module improves the contextual understanding of bilingual and code-switched EMRs. In the "unified bio-embedding" module, the knowledge of the model trained in the medical domain is injected into the encoder-based model to bridge the gap between the medical and general domains. Experimental results demonstrate that the proposed BioBridge significantly performance traditional machine learning and pre-trained encoder-based models on several metrics, including F1 score, area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), and Brier score. Specifically, BioBridge-XLM achieved enhancements of 0.85% in F1 score, 0.75% in AUROC, and 0.76% in AUPRC, along with a notable 3.04% decrease in the Brier score, demonstrating marked improvements in accuracy, reliability, and prediction calibration over the baseline XLM model. The source code will be made publicly available.

cs.CL

DSG-KD: Knowledge Distillation from Domain-Specific to General Language Models

The use of pre-trained language models fine-tuned to address specific downstream tasks is a common approach in natural language processing (NLP). However, acquiring domain-specific knowledge via fine-tuning is challenging. Traditional methods involve pretraining language models using vast amounts of domain-specific data before fine-tuning for particular tasks. This study investigates emergency/non-emergency classification tasks based on electronic medical record (EMR) data obtained from pediatric emergency departments (PEDs) in Korea. Our findings reveal that existing domain-specific pre-trained language models underperform compared to general language models in handling N-lingual free-text data characteristics of non-English-speaking regions. To address these limitations, we propose a domain knowledge transfer methodology that leverages knowledge distillation to infuse general language models with domain-specific knowledge via fine-tuning. This study demonstrates the effective transfer of specialized knowledge between models by defining a general language model as the student model and a domain-specific pre-trained model as the teacher model. In particular, we address the complexities of EMR data obtained from PEDs in non-English-speaking regions, such as Korea, and demonstrate that the proposed method enhances classification performance in such contexts. The proposed methodology not only outperforms baseline models on Korean PED EMR data, but also promises broader applicability in various professional and technical domains. In future works, we intend to extend this methodology to include diverse non-English-speaking regions and address additional downstream tasks, with the aim of developing advanced model architectures using state-of-the-art KD techniques. The code is available in https://github.com/JoSangYeon/DSG-KD.

cs.CL

ConCSE: Unified Contrastive Learning and Augmentation for Code-Switched Embeddings

This paper examines the Code-Switching (CS) phenomenon where two languages intertwine within a single utterance. There exists a noticeable need for research on the CS between English and Korean. We highlight that the current Equivalence Constraint (EC) theory for CS in other languages may only partially capture English-Korean CS complexities due to the intrinsic grammatical differences between the languages. We introduce a novel Koglish dataset tailored for English-Korean CS scenarios to mitigate such challenges. First, we constructed the Koglish-GLUE dataset to demonstrate the importance and need for CS datasets in various tasks. We found the differential outcomes of various foundation multilingual language models when trained on a monolingual versus a CS dataset. Motivated by this, we hypothesized that SimCSE, which has shown strengths in monolingual sentence embedding, would have limitations in CS scenarios. We construct a novel Koglish-NLI (Natural Language Inference) dataset using a CS augmentation-based approach to verify this. From this CS-augmented dataset Koglish-NLI, we propose a unified contrastive learning and augmentation method for code-switched embeddings, ConCSE, highlighting the semantics of CS sentences. Experimental results validate the proposed ConCSE with an average performance enhancement of 1.77\% on the Koglish-STS(Semantic Textual Similarity) tasks.

cs.CL

Half-Wave Dipolar Metal-Semiconductor Laser

Nano-scale lasers harnessing metallic plasmons hold promise across physical sciences and industrial applications. Plasmons are categorized as surface plasmon polaritons (SPP) and localized surface plasmons (LSP). While SPP has gained popularity for nano-lasers by fitting a few cycles of SPP waves into resonators, achieving LSP lasing in single nanoparticles remains an elusive goal. Here, we highlight the equivalence of LSP and SPP within resonant systems and present lasers oscillating in the lowest-order LSP or, equivalently, half-cycle SPP. This diffraction-limited dipolar emitter is realized through strong coupling of plasmonic oscillation in gold and dielectric resonance in high-gain III-V semiconductor in the near infrared away from surface plasmon frequencies. The resulting single-mode stimulated emission peak exhibits linewidth Q factors over 50 at room temperature, with wide tunability spanning from 1190 to 1460 nm determined by resonator sizes ranging from 190 to 280 nm. A semiconductor laser model elucidates the temporal and spectral buildup dynamics under optical pumping. Notably, linewidth Q values surpassing 250 are attained from higher-order, isolated laser particles within live biological cells. These results offer fresh perspectives in nanophotonics and indicate promising opportunities for multiplexed biological applications.

physics.optics

Ultrasmall InGa(As)P dielectric and plasmonic nanolasers

Nanolasers have great potential as both on-chip light sources and optical barcoding particles. We demonstrate ultrasmall InGaP and InGaAsP disk lasers with diameters down to 360 nm (198 nm in height) in the red spectral range. Optically pumped, room-temperature, single-mode lasing was achieved from both disk-on-pillar and isolated particles. When isolated disks were placed on gold, plasmon polariton lasing was obtained with Purcell-enhanced stimulated emission. UV lithography and plasma ashing enabled the fabrication of nanodisks on a wafer-scale, with intended random size variation. Silica-coated nanodisk particles generated stable sub-nanometer spectra from within biological cells across an 80 nm bandwidth from 635 to 715 nm.

physics.optics

Sub-micron single-particle perovskite plasmonic nanolasers at room temperature

Plasmonic nanolasers have received a substantial interest for their promising applications in integrated photonics, optical sensing, and biomedical imaging. To date, a room-temperature plasmonic nanolaser, submicron in all dimensions, remains elusive in the visible regime due to high metallic losses. Here, we demonstrate single-particle lasing around 2.3 eV with full-submicron, cesium lead bromide perovskite (CsPbBr3) crystals atop polymer-coated gold substrates at room temperature. With a large number (~100) of devices in total, we systematically study the lasing action of plasmonic test and photonic control groups. The achieved smallest plasmonic laser was 0.56 micrometer x 0.58 micrometer x 0.32 micrometer in size, ten-fold smaller than that of our smallest photonic laser. Key elements to efficient plasmonic lasing are identified as enhanced optical gain by the Purcell effect, long carrier diffusivity, a large spontaneous emission factor, and a high group index. Our results shed light on three-dimensional miniaturization of plasmonic lasers.

physics.optics