SearcharxivSearch

arXiv subjects

Jongho Park

Publications and source records attributed to Jongho Park.

At least 19 recordsLinked to original sources

Optimal block preconditioners for a mass-conserving mixed stress formulation of Stokes flow

We present optimal block diagonal and triangular preconditioners for a mass-conserving mixed stress formulation of Stokes flow. The algebraic formulation leads to a double saddle point system with unknowns corresponding to discrete stress, velocity, vorticity, and pressure. MINRES equipped with a block diagonal preconditioner for an augmented Lagrangian formulation of this system is analyzed and shown to be optimal, in the sense that the convergence rate is independent of key parameters such as mesh size and kinematic viscosity. GMRES equipped with a block triangular preconditioner is also analyzed using a field-of-values approach. Finally, we present numerical results for both two- and three-dimensional model problems to validate the parameter robustness of the proposed preconditioners.

math.NA

Looped Diffusion Language Models

Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDMs remains underexplored. In this paper, we show that selectively looping the early-middle transformer layers significantly improves both training efficiency and model performance in MDMs. We call this approach LoopMDM(Looped Masked Diffusion Model), which brings two key benefits: looping layers at training-time yields a depth-scaling effect without adding parameters, while varying the number of loops at inference-time enables flexible compute scaling. Despite the simplicity, the results are striking: across multiple pre-training corpora, LoopMDM matches the performance of same-size MDMs with up to 3.3 fewer training FLOPs, while its final performance outperforms them on various reasoning benchmarks, including up to 8.5 points on GSM8K. It even surpasses deeper non-looped MDMs trained with comparable per-step compute, indicating that selective looping is more effective than naive depth scaling. Furthermore, LoopMDM can scale inference-time compute by increasing the number of loops. Adaptively adjusting the number of loops throughout the sampling process further yields additional gains in compute efficiency while maintaining performance. Lastly, with attention analysis, we provide evidence that looping is effective in MDMs by promoting interactions among masked positions. Our code and weights will be publicly released.

cs.LG

Verification of the Polarimetric Capability of the East Asia VLBI Network

The East Asia VLBI Network (EAVN) has recently enabled dual-polarization observations at $22$ and $43\,\mathrm{GHz}$. We present the first systematic verification of its polarimetric performance using EAVN observations of M87, 3C 279, 3C 273, and OJ 287, calibrated with the GPCAL pipeline and evaluated against near-contemporaneous VLBA images at comparable frequencies. Most stations show stable polarimetric leakages with amplitudes of $5$-$10\%$ over monthly timescales. While several VERA stations exhibit D-term phase variations between epochs, we attribute these to field-rotator (FR) offsets and demonstrate that phase stability is restored after applying the analytically derived FR corrections. The resulting linear-polarization morphologies and EVPAs broadly agree with the VLBA results within uncertainties; fractional polarization measured by the EAVN tends to be slightly higher near polarization peaks. Although exact one-to-one comparisons are limited by moderate frequency and epoch differences, the combined evidence indicates robust EAVN polarimetric calibration and imaging capabilities at $22$ and $43\,\mathrm{GHz}$. These results support the scientific capability of EAVN polarimetry and lay the groundwork for expanded, higher-fidelity polarimetric studies in East Asia.

astro-ph.IM

A Unified Origin of Faraday Rotation toward 3C 84: The Circumnuclear Ambient Medium within the Parsec-Scale Bondi Radius of the Host Galaxy NGC 1275

We present multi-frequency polarimetric observations of 3C 84 obtained with the Korean VLBI Network at 43-141 GHz, the Very Long Baseline Array at 43 GHz, and the High Sensitivity Array at 8 GHz from 2015 to 2024. We find that the Faraday rotation measure (RM) decreases systematically with distance from the black hole over 1-8 pc, following a single power-law trend of RM proportional to r^{-2.7+/-0.2}. Notably, RM measurements from earlier studies across the same distance range follow the same relation. This consistency across epochs, frequencies, and independent datasets indicates a common and stable external Faraday screen. These results naturally identify the circumnuclear ambient medium within the parsec-scale Bondi radius of the host galaxy NGC 1275 as the origin of the Faraday rotation, thereby resolving a long-standing question about its physical origin. From the RM profile, we derive radial distributions of the electron density and magnetic-field strength in the circumnuclear ambient medium that are consistent with independent constraints. The derived density lies below that of the free-free absorption disk and, when extrapolated inward, remains below the density of the broad-line region. The magnetic-field strength gradually increases from 0.1-1.5 microgauss at the Bondi radius to milligauss-to-gauss levels toward the black hole, providing the first spatially resolved constraint on the magnetic-field strength at parsec-scale distances in an elliptical galaxy. Together, these results present a spatially resolved and physically consistent picture of the circumnuclear environment in NGC 1275.

astro-ph.GA

Transverse Oscillations and Wave Propagation in the Magnetically Dominated M87 Jet

We present an in-depth analysis of transverse oscillations in the M87 jet, as identified in our previous study (Ro et al. 2023a), which reported oscillatory patterns with a characteristic period of $\sim$1 year in the edge-brightened jet structure extending up to 12\,mas from the core. This work is based on high-cadence KaVA 22\,GHz observations conducted from December 2013 to June 2016. By analyzing the transverse velocity profiles and the spatial evolution of the oscillations, we find that the oscillations propagate downstream along the jet, with a wavelength of $\sim9-10$\,mas. A single-mode sinusoidal wave model applied to the ridge lines successfully reproduces the observed transverse oscillations and yields superluminal wave speeds of $\sim2.7-2.9\,c$, consistent with the bulk jet velocity in this region. These findings suggest that the transverse oscillations may be interpreted either as transverse MHD waves -- possibly excited by jet precession, nutation, or quasi-periodic magnetic flux eruptions near the central engine -- or as manifestations of jet instabilities, such as current-driven instabilities (CDIs). Further investigation is required to distinguish between these scenarios and to clarify the dominant physical mechanism.

astro-ph.HE

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritizes compliance, making models vulnerable to harmful prompts. To mitigate this safety degradation, recent approaches rely on external teacher distillation, yet this introduces a distributional discrepancy that degrades native reasoning. We formalize safety realignment as a KL projection onto the safe simplex and prove that the student's own safety-filtered distribution is the unique KL-optimal target, while any external teacher incurs an irreducible excess KL penalty. Guided by this analysis, we propose ThinkSafe, a self-generated alignment framework that restores safety without external teachers. Our key insight is that while compliance suppresses safety mechanisms, models often retain latent knowledge to identify harm. ThinkSafe unlocks this via lightweight refusal steering, which preserves the KL-optimal target while increasing the acceptance rate. Experiments on DeepSeek-R1-Distill and Qwen3 show ThinkSafe significantly improves safety while preserving reasoning proficiency, and achieves superior safety and comparable reasoning to GRPO with roughly an order of magnitude less compute. Code, models, and datasets are available at https://github.com/seanie12/ThinkSafe and https://huggingface.co/Seanie-lee/collections.

cs.AI

Locating the missing large-scale emission in the jet of M87* with short EHT baselines

In Very-Long Baseline Interferometric arrays, nearly co-located stations probe the largest scales and typically cannot resolve the observed source. In the absence of large-scale structure, closure phases constructed with these stations are zero and, since they are independent of station-based errors, they can be used to probe data issues. Here, we show with an expansion about co-located stations, how these trivial closure phases become non-zero with brightness distribution on smaller scales than their short baseline would suggest. When applied to sources that are made up of a bright compact and large-scale diffuse component, the trivial closure phases directly measure the centroid relative to the compact source and higher-order image moments. We present a technique to measure these image moments with minimal model assumptions and validate it on synthetic Event Horizon Telescope (EHT) data. We then apply this technique to 2017 and 2018 EHT observations of M87* and find a weak preference for extended emission in the direction of the large-scale jet. We also apply it to 2021 EHT data and measure the source centroid about 1 mas northwest of the compact ring, consistent with the jet observed at lower frequencies.

astro-ph.HE

A high-order augmented Lagrangian method with arbitrarily fast convergence

We propose a high-order version of the augmented Lagrangian method for solving convex optimization problems with linear constraints, which achieves arbitrarily fast -- and even superlinear -- convergence rates. First, we analyze the convergence rates of the high-order proximal point method under certain uniform convexity assumptions on the energy functional. We then introduce the high-order augmented Lagrangian method and analyze its convergence by leveraging the convergence results of the high-order proximal point method. Finally, we present applications of the high-order augmented Lagrangian method to various problems arising in the sciences, including data fitting, flow in porous media, and scientific machine learning.

math.OC

A polynomial dimension-dependence analysis of Bramble--Pasciak--Xu preconditioners

We investigate the dimension dependence of Bramble--Pasciak--Xu (BPX) preconditioners for high-dimensional partial differential equations and establish that the condition numbers of BPX-preconditioned systems grow only polynomially with the spatial dimension. Our analysis requires a careful derivation of the dimension dependence of several fundamental tools in the theory of finite element methods, including elliptic regularity, the Bramble--Hilbert lemma, trace inequalities, and inverse inequalities. We further analyze an averaged Scott--Zhang-type quasi-interpolation operator, and show that its associated constants scale polynomially with the dimension. Building on these ingredients, we prove a multilevel norm equivalence theorem and derive a BPX preconditioner with explicit polynomial bounds on its dimensional dependence. The analysis is motivated in part by recent tensor and quantum finite element methods, where dimension-explicit conditioning estimates for BPX preconditioners play an important role.

math.NA

Probing jet base emission of M87* with the 2021 Event Horizon Telescope observations

We investigate the presence and spatial characteristics of the jet base emission in M87* at 230 GHz, enabled by the enhanced uv coverage in the 2021 Event Horizon Telescope (EHT) observations. The addition of the 12-m Kitt Peak Telescope and NOEMA provides two key intermediate-length baselines to SMT and the IRAM 30-m, giving sensitivity to emission structures at scales of $\sim250~\mu$as and $\sim2500~\mu$as (0.02 pc and 0.2 pc). Without these baselines, earlier EHT observations lacked the capability to constrain emission on large scales, where a "missing flux" of order $\sim1$ Jy is expected. To probe these scales, we analyzed closure phases, robust against station-based gain errors, and modeled the jet base emission using a simple Gaussian offset from the compact ring emission at separations $>100~\mu$as. Our analysis reveals a Gaussian feature centered at ($\Delta$RA $\approx320~\mu$as, $\Delta$Dec $\approx60~\mu$as), a projected separation of $\approx5500$ AU, with a flux density of only $\sim60$ mJy, implying that most of the missing flux in previous studies must arise from larger scales. Brighter emission at these scales is ruled out, and the data do not favor more complex models. This component aligns with the inferred direction of the large-scale jet and is consistent with emission from the jet base. While our findings indicate detectable jet base emission at 230 GHz, coverage from only two intermediate baselines limits reconstruction of its morphology. We therefore treat the recovered Gaussian as an upper limit on the jet base flux density. Future EHT observations with expanded intermediate-baseline coverage will be essential to constrain the structure and nature of this component.

astro-ph.HE

Bayesian polarization calibration and imaging in very long baseline interferometry

Extracting polarimetric information from very long baseline interferometry (VLBI) data is demanding but vital for understanding the synchrotron radiation process and the magnetic fields of celestial objects, such as active galactic nuclei (AGNs). However, conventional CLEAN-based calibration and imaging methods provide suboptimal resolution without uncertainty estimation of calibration solutions, while requiring manual steering from an experienced user. We present a Bayesian polarization calibration and imaging method using Bayesian imaging software resolve for VLBI data sets, that explores the posterior distribution of antenna-based gains, polarization leakages, and polarimetric images jointly from pre-calibrated data. We demonstrate our calibration and imaging method with observations of the quasar 3C273 with the VLBA at 15 GHz and the blazar OJ287 with the GMVA+ALMA at 86 GHz. Compared to the CLEAN method, our approach provides physically realistic images that satisfy positivity of flux and polarization constraints and can reconstruct complex source structures composed of various spatial scales. Our method systematically accounts for calibration uncertainties in the final images and provides uncertainties of Stokes images and calibration solutions. The automated Bayesian approach for calibration and imaging will be able to obtain high-fidelity polarimetric images using high-quality data from next-generation radio arrays. The pipeline developed for this work is publicly available.

astro-ph.IM

Helical Magnetic Field in the Acceleration--Collimation Zone of the M87 Jet

Relativistic jets from supermassive black holes are expected to be magnetically launched and guided, with magnetic energy systematically converted to bulk kinetic energy throughout an extended acceleration-collimation zone (ACZ). A key prediction of magnetohydrodynamic (MHD) models is a transition from poloidally dominated fields near the engine to toroidally dominated fields downstream, yet direct tests within the ACZ are hampered by weak polarization and strong Faraday rotation. We report quasi-simultaneous, high-sensitivity, multifrequency very long baseline interferometric polarimetry of M87 spanning 1.4-24.4GHz. We present high-fidelity, Faraday rotation-corrected maps of intrinsic linear polarization that continuously resolve the ACZ in the de-projected distance range of ~9e3 to ~3.6e5 gravitational radii from the black hole. The maps reveal pronounced north-south asymmetries in fractional linear polarization and electric vector position angle (EVPA), peaking in the inner ACZ at a projected distance of ~20mas along the jet and remaining prominent out to ~100mas. These signatures are best reproduced by models with a large-scale, ordered helical field that retains a substantial poloidal component-contrary to the rapid toroidal dominance expected under steady, ideal MHD. This tension implies ongoing magnetic dissipation that limits toroidal buildup over the ACZ. The handedness of the helix provides an independent constraint on the black hole's spin direction, supporting a spin vector oriented away from the observer, consistent with the orientation inferred from horizon-scale imaging. Farther downstream, the asymmetries diminish, and the EVPA and fractional polarization distributions become more symmetric; we tentatively interpret this as evolution toward a more poloidally dominated configuration, while noting current sensitivity and dynamic-range limits.

astro-ph.HE

Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement

Test-time scaling through reward-guided generation remains largely unexplored for discrete diffusion models despite its potential as a promising alternative. In this work, we introduce Iterative Reward-Guided Refinement (IterRef), a novel test-time scaling method tailored to discrete diffusion that leverages reward-guided noising-denoising transitions to progressively refine misaligned intermediate states. We formalize this process within a Multiple-Try Metropolis (MTM) framework, proving convergence to the reward-aligned distribution. Unlike prior methods that assume the current state is already aligned with the reward distribution and only guide the subsequent transition, our approach explicitly refines each state in situ, progressively steering it toward the optimal intermediate distribution. Across both text and image domains, we evaluate IterRef on diverse discrete diffusion models and observe consistent improvements in reward-guided generation quality. In particular, IterRef achieves striking gains under low compute budgets, far surpassing prior state-of-the-art baselines.

cs.LG

Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models

Masked Diffusion Models (MDMs) as language models generate by iteratively unmasking tokens, yet their performance crucially depends on the inference time order of unmasking. Prevailing heuristics, such as confidence based sampling, are myopic: they optimize locally, fail to leverage extra test-time compute, and let early decoding mistakes cascade. We propose Lookahead Unmasking (LookUM), which addresses these concerns by reformulating sampling as path selection over all possible unmasking orders without the need for an external reward model. Our framework couples (i) a path generator that proposes paths by sampling from pools of unmasking sets with (ii) a verifier that computes the uncertainty of the proposed paths and performs importance sampling to subsequently select the final paths. Empirically, erroneous unmasking measurably inflates sequence level uncertainty, and our method exploits this to avoid error-prone trajectories. We validate our framework across six benchmarks, such as mathematics, planning, and coding, and demonstrate consistent performance improvements. LookUM requires only two to three paths to achieve peak performance, demonstrating remarkably efficient path selection. The consistent improvements on both LLaDA and post-trained LLaDA 1.5 are particularly striking: base LLaDA with LookUM rivals the performance of RL-tuned LLaDA 1.5, while LookUM further enhances LLaDA 1.5 itself showing that uncertainty based verification provides orthogonal benefits to reinforcement learning and underscoring the versatility of our framework. Code will be publicly released.

cs.LG

Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models

While 4-bit quantization has emerged as a memory-optimal choice for non-reasoning models and zero-shot tasks across scales, we show that this universal prescription fails for reasoning models, where the KV cache rather than model size can dominate memory. Through systematic experiments across 1,700 inference scenarios on AIME25 and GPQA-Diamond, we find a scale-dependent trade-off: models with an effective size below 8-bit 4B parameters achieve better accuracy by allocating memory to more weights rather than longer generation, while larger models achieve better accuracy by allocating memory to longer generations. This scale threshold also determines when parallel scaling becomes memory-efficient and whether KV cache eviction outperforms KV quantization. Our findings show that memory optimization for LLMs cannot be scale-agnostic, while providing principled guidelines: for small reasoning models, prioritize model capacity over test-time compute, while for larger ones, maximize test-time compute. Our results suggest that optimizing reasoning models for deployment requires fundamentally different strategies from those established for non-reasoning models.

cs.LG

Monitoring of 3C 286 with ALMA, IRAM, and SMA from 2006 to 2025: Stability, Synchrotron Ages, and Frequency-Dependent Polarization Attributed to Core-Shift

We present the results of multi-frequency monitoring of the radio quasar 3C 286, conducted using three instruments: ALMA at 91.5, 103.5, 233.0, and 343.4 GHz, the IRAM 30-m Telescope at 86 and 229 GHz, and SMA at 225 GHz. The IRAM measurements from 2006 to 2024 show that the total flux of 3C 286 is stable within measurement uncertainties, indicating long-term stability up to 229 GHz, when applying a fixed Kelvin-to-Jansky conversion factor throughout its dataset. ALMA data from 2018 to 2024 exhibit a decrease in flux, which up to 4% could be attributed to an apparent increase in the absolute brightness of Uranus, the primary flux calibrator for ALMA with the ESA4 model. Taken together, these results suggest that the intrinsic total flux of 3C 286 has remained stable up to 229 GHz over the monitoring period. The polarization properties of 3C 286 are stable across all observing frequencies. The electric vector position angle (EVPA) gradually rotates as a function of wavelength squared, which is well described by a single power-law over the full frequency range. We therefore propose using the theoretical EVPA values from this model curve for absolute EVPA calibration between 5 and 343.4 GHz. The Faraday rotation measure increases as a function of frequency up to (3.2+/-1.5)x10^4 rad m^-2, following RM proportional to nu^alpha with alpha = 2.05+/-0.06. This trend is consistent with the core-shift effect expected in a conical jet.

astro-ph.GA

Unified analysis of saddle point problems via auxiliary space theory

We present sharp estimates for the extremal eigenvalues of the Schur complements arising in saddle point problems. These estimates are derived using the auxiliary space theory, in which a given iterative method is interpreted as an equivalent but more elementary iterative method on an auxiliary space, enabling us to obtain sharp convergence estimates. The proposed framework improves or refines several existing results, which can be recovered as corollaries of our results. To demonstrate the versatility of the framework, we present various applications from scientific computing: the augmented Lagrangian method, mixed finite element methods, and nonoverlapping domain decomposition methods. In all these applications, the condition numbers of the corresponding Schur complements can be estimated in a straightforward manner using the proposed framework.

math.NA

Auxiliary space theory for the analysis of iterative methods for semidefinite linear systems

We present an auxiliary space theory that provides a unified framework for analyzing various iterative methods for solving linear systems that may be semidefinite. By interpreting a given iterative method for the original system as an equivalent, yet more elementary, iterative method for an auxiliary system defined on a larger space, we derive sharp convergence estimates using elementary linear algebra. In particular, we establish identities for the error propagation operator and the condition number associated with iterative methods, which generalize and refine existing results. The proposed auxiliary space theory is applicable to the analysis of numerous advanced numerical methods in scientific computing. To illustrate its utility, we present three examples -- subspace correction methods, Hiptmair--Xu preconditioners, and auxiliary grid methods -- and demonstrate how the proposed theory yields refined analyses for these cases.

math.NA