SearcharxivSearch

arXiv subjects

Sangyeol Lee

Publications and source records attributed to Sangyeol Lee.

6 recordsLinked to original sources

A.X K2 Technical Report

We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percentage points on some benchmarks, reflecting large gains in token efficiency. To support long contexts efficiently, we introduce Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training. SGA is trained natively at 128K through a \emph{sparse} indexer warmup that optimizes the indexer against its own sparse top-$k$ selection rather than the dense attention distribution, making adaptation markedly cheaper: each query reads only 2,048 positions, yet long-context quality is unchanged and A.X K2 scores 94.6 on RULER out to 256K. The outlier suppression of GN in turn keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A simple yet effective Think-Fusion recipe further lets users switch between thinking and non-thinking modes within a single unified model. Extensive evaluations show that A.X K2 performs competitively against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.

cs.AI

A.X K1 Technical Report

We introduce A.X K1, a 519B-parameter Mixture-of-Experts (MoE) language model trained from scratch. Our design leverages scaling laws to optimize training configurations and vocabulary size under fixed computational budgets. A.X K1 is pre-trained on a corpus of approximately 10T tokens, curated by a multi-stage data processing pipeline. Designed to bridge the gap between reasoning capability and inference efficiency, A.X K1 supports explicitly controllable reasoning to facilitate scalable deployment across diverse real-world scenarios. We propose a simple yet effective Think-Fusion training recipe, enabling user-controlled switching between thinking and non-thinking modes within a single unified model. Extensive evaluations demonstrate that A.X K1 achieves performance competitive with leading open-source models, while establishing a distinctive advantage in Korean-language benchmarks.

cs.CL

Fourier-type monitoring procedures for strict stationarity

We consider model-free monitoring procedures for strict stationarity of a given time series. The new criteria are formulated as L2-type statistics incorporating the empirical characteristic function. Asymptotic as well as Monte Carlo results are presented. The new methods are also employed in order to test for possible stationarity breaks in time-series data from the financial sector.

math.ST

On the Identifiability Conditions in Some Nonlinear Time Series Models

In this study, we consider the identifiability problem for nonlinear time series models. Special attention is paid to smooth transition GARCH, nonlinear Poisson autoregressive, and multiple regime smooth transition autoregressive models. Some sufficient conditions are obtained to establish the identifiability of these models.

math.ST

Quantile Regression for Location-Scale Time Series Models with Conditional Heteroscedasticity

This paper considers quantile regression for a wide class of time series models including ARMA models with asymmetric GARCH (AGARCH) errors. The classical mean-variance models are reinterpreted as conditional location-scale models so that the quantile regression method can be naturally geared into the considered models. The consistency and asymptotic normality of the quantile regression estimator is established in location-scale time series models under mild conditions. In the application of this result to ARMA-AGARCH models, more primitive conditions are deduced to obtain the asymptotic properties. For illustration, a simulation study and a real data analysis are provided.

stat.ME

Test for tail index change in stationary time series with Pareto-type marginal distribution

The tail index, indicating the degree of fatness of the tail distribution, is an important component of extreme value theory since it dominates the asymptotic distribution of extreme values such as the sample maximum. In this paper, we consider the problem of testing for a change in the tail index of time series data. As a test, we employ the cusum test and investigate its null limiting distribution. Further, we derive the null limiting distribution of the cusum test based on the residuals from autoregressive models. Simulation results are provided for illustration.

math.ST