SearcharxivSearch

arXiv subjects

Sang Park

Publications and source records attributed to Sang Park.

7 recordsLinked to original sources

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.

cs.CL

Making Qwen3 Think in Korean with Reinforcement Learning

We present a two-stage fine-tuning approach to make the large language model Qwen3 14B "think" natively in Korean. In the first stage, supervised fine-tuning (SFT) on a high-quality Korean reasoning dataset establishes a strong foundation in Korean logical reasoning, yielding notable improvements in Korean-language tasks and even some gains in general reasoning ability. In the second stage, we employ reinforcement learning with a customized Group Relative Policy Optimization (GRPO) algorithm to further enhance both Korean reasoning alignment and overall problem-solving performance. We address critical stability challenges in GRPO training - such as reward hacking and policy collapse - by introducing an oracle judge model that calibrates the reward signal. Our approach achieves stable learning (avoiding the collapse observed in naive GRPO) and leads to steady, incremental performance gains. The final RL-tuned model demonstrates substantially improved results on advanced reasoning benchmarks (particularly math and coding tasks) while maintaining knowledge and language proficiency, successfully conducting its internal chain-of-thought entirely in Korean.

cs.CL

Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs

Multilingual large language models (LLMs) often exhibit language confusion, a tendency to generate responses in a dominant language irrespective of the prompt's language. To address this, we propose Smoothie-Qwen, a lightweight, post-hoc method that mitigates language bias without retraining. This technique selectively adjusts token-level output probabilities to effectively suppress undesired language generation. Applied to the Qwen model, our method reduces unintended Chinese output by over 95% while preserving task accuracy on multilingual benchmarks. This work provides a practical and efficient solution for enhancing the language controllability of LLMs, making them more reliable for global applications.

cs.CL

DNA 1.0 Technical Report

In this report, we present DNA 1.0 8B Instruct, a state-of-the-art bilingual language model optimized for Korean and English language tasks. By applying continual pre-training (CPT) with high-quality Korean datasets to Llama 3.1 8B and subsequent supervised fine-tuning (SFT), we create an instruction-following model with enhanced Korean language capabilities. This model is then merged with Llama 3.1 8B Instruct via spherical linear interpolation (SLERP) and undergoes further optimization through direct preference optimization (DPO) and knowledge distillation (KD). DNA 1.0 8B Instruct achieves state-of-the-art results on Korean-specific tasks, including KMMLU (53.26%), KoBEST (83.40%), and BELEBELE (57.99%), while maintaining strong English capabilities on MMLU (66.64%), MMLU-Pro (43.05%) and GSM8K (80.52%). As an open model, DNA 1.0 8B Instruct represents a significant advancement in bilingual language modeling. As an open model, DNA 1.0 8B Instruct is freely available through https://huggingface.co/dnotitia/Llama-DNA-1.0-8B-Instruct . For commercial licensing inquiries or feedback, please contact us at https://www.dnotitia.com/contact/post-form

cs.CL

The James Webb Space Telescope Mission: Optical Telescope Element Design, Development, and Performance

The James Webb Space Telescope (JWST) is a large, infrared space telescope that has recently started its science program which will enable breakthroughs in astrophysics and planetary science. Notably, JWST will provide the very first observations of the earliest luminous objects in the Universe and start a new era of exoplanet atmospheric characterization. This transformative science is enabled by a 6.6 m telescope that is passively cooled with a 5-layer sunshield. The primary mirror is comprised of 18 controllable, low areal density hexagonal segments, that were aligned and phased relative to each other in orbit using innovative image-based wavefront sensing and control algorithms. This revolutionary telescope took more than two decades to develop with a widely distributed team across engineering disciplines. We present an overview of the telescope requirements, architecture, development, superb on-orbit performance, and lessons learned. JWST successfully demonstrates a segmented aperture space telescope and establishes a path to building even larger space telescopes.

astro-ph.IM

MMT & Magellan Infrared Spectrograph

The MMT and Magellan infrared spectrograph (MMIRS) is a cryogenic multiple slit spectrograph operating in the wavelength range 0.9-2.4 micron. MMIRS' refractive optics offer a 6.9 by 6.9 arcmin field of view for imaging with a spatial resolution of 0.2 arcsec per pixel on a HAWAII-2 array. For spectroscopy, MMIRS can be used with long slits up to 6.9 arcmin long, or with custom slit masks having slitlets distributed over a 4 by 6.9 arcmin area. A range of dispersers offer spectral resolutions of 800 to 3000. MMIRS is designed to be used at the f/5 foci of the MMT or Magellan Clay 6.5m telescopes. MMIRS was commissioned in 2009 at the MMT and has been in routine operation at the Magellan Clay Telescope since 2010. MMIRS is being used for a wide range of scientific investigations from exoplanet atmospheres to Ly-alpha emitters.

astro-ph.IM

Fluorescence lifetime modification in Eu:Lu2O3 nanoparticles in the presence of silver nanoparticles

Europium-doped lutetium-oxide (Eu:Lu2O3) nanoparticles were synthesized using a combustion technique and a co-precipitation technique, and their properties were compared. Surface-modification utilizing small silane molecules and long chain polymers were explored to de-agglomerate and disperse the particles. Structural, morphological and optical properties were characterized with x-ray diffraction, scanning and transmission electron microscopy, and laser spectroscopy respectively to evaluate these materials. The luminescent behaviors were compared between the pristine and modified Eu:Lu2O3 nanoparticles to study the influence of surface ligands on emission properties. Subsequently, the Eu:Lu2O3 nanoparticles were placed on top of a thin film consisting of silver nanoparticles and combined with silver nanoparticles and dispersed in a polymer matrix. The presence of the silver nanoparticles led to a reduction of the fluorescence lifetime of 12-14%.

cond-mat.mtrl-sci