SearcharxivSearch

arXiv subjects

Taehyeong Kim

Publications and source records attributed to Taehyeong Kim.

At least 19 recordsLinked to original sources

FRAME: Factored Retrieval via Attribute Readouts for Object-Centric Scene Memory

Language-guided robots need persistent scene memories to follow instructions, revisit objects, and resolve references to objects encountered over time. While much of language-guided scene-memory retrieval has emphasized spatial or relational references, many everyday object references specify objects by multiple persistent attributes, such as category, material, size, or surface appearance. We formalize this problem as attribute-compositional retrieval, where a fixed object-centric scene memory is queried with natural language to retrieve the object satisfying the requested attributes. To investigate this capability directly, we introduce a controlled evaluation protocol with fixed scene memories and attribute-defined targets, separating retrieval from perception and annotation ambiguities. We then propose FRAME, which turns language into query-relevant attribute weights, uses learned readouts to estimate per-attribute evidence from object embeddings, and ranks objects by aggregating this evidence according to the query. Across held-out scenes and object assets, FRAME outperforms representative scene-memory retrieval baselines while reducing post-decomposition object scoring to lightweight matrix-vector computation. These results position attribute-compositional retrieval as a complementary scene-memory capability for language-guided robots, showing that persistent object attributes can be exposed as composable evidence for accurate and efficient multi-attribute retrieval.

cs.CV

Weighted singular vectors in common-base self-similar sets

We prove a lower bound for the Hausdorff dimension of weighted totally irrational singular vectors in affine-spanning common-base integral self-similar sets satisfying the open set condition. For the middle-third Cantor square, the bound improves the previously known explicit lower bound.

math.NT

Structured matrix factorization length

Every (resp. a generic) complex $n \times n$ matrix can be expressed as a product of $2n+5$ (resp. $\lfloor n/2 \rfloor +1$) Toeplitz matrices. Motivated by this result, it is natural to ask the following question: what is the minimum number of Toeplitz matrices required to factor a given matrix? We generalize this question from Toeplitz structure to more general structures. In this paper, we introduce the notion of structured matrix factorization length when the set of matrices with a given structure is an affine variety $X \subseteq \mathbb{C}^{n \times n}$. Then we introduce the $r$-th $X$-factorization variety, defined as the Zariski closure of the set of products of $r$ matrices in $X$, and use it to define the border structured matrix factorization length. In particular, we study the cases in which $X$ is the affine variety of Toeplitz, Hankel, bidiagonal, tridiagonal, skew-symmetric or companion matrices. We calculate the dimension of the $X$-factorization varieties for all these cases, and discuss how numerical algebraic geometry can be used to obtain computational evidence for the degrees of $X$-factorization varieties with an example. In addition, we propose methods for deriving lower and upper bounds for (border) structured matrix factorization length. For lower bounds, we develop a method based on displacement rank, which can also be used to obtain some defining equations of the $r$-th $X$-factorization variety; for upper bounds, we suggest an approach using alternating minimization.

math.AG

Hamiltonian and Symplectic Tensors in the T-product Algebra

We study Hamiltonian and symplectic tensor structures in the T-product algebra. We define T-Hamiltonian and T-symplectic tensors and characterize them through their Fourier-domain slices. For T-Hamiltonian tensors we establish the standard block form and the spectral symmetry of T-eigenvalues, while for T-symplectic tensors we derive the inverse and exponential-map properties. Our main result is a constructive T-Williamson normal form for tensors whose Fourier-domain slices are real symmetric positive-definite matrices. We also show that, under the Hermitian symplectic convention adopted here, this decomposition does not extend directly to arbitrary Hermitian positive-definite Fourier-domain slices, and we derive a real-valued recovery criterion under Fourier conjugate symmetry. Numerical experiments verify the construction, exhibit runtime trends consistent with the slice-wise complexity $O(pn^3)$, and illustrate the framework on a Fourier-domain encoding of covariance-matrix families arising in continuous-variable quantum dynamics.

math.NA

SocraticKG: Knowledge Graph Construction via QA-Driven Fact Extraction

Constructing Knowledge Graphs (KGs) from unstructured text provides a structured framework for knowledge representation and reasoning, yet current LLM-based approaches struggle with a fundamental trade-off: factual coverage often leads to relational fragmentation, while premature consolidation causes information loss. To address this, we propose SocraticKG, an automated KG construction method that introduces question-answer pairs as a structured intermediate representation to systematically unfold document-level semantics prior to triple extraction. By employing 5W1H-guided QA expansion, SocraticKG captures contextual dependencies and implicit relational links typically lost in direct KG extraction pipelines, providing explicit grounding in the source document that helps mitigate implicit reasoning errors. Evaluation on the MINE benchmark and HotpotQA downstream task demonstrates that our approach effectively addresses the coverage-connectivity trade-off, achieving superior factual retention and structural cohesion while supporting complex multi-hop reasoning.

cs.CL

Improved identification of breakpoints in piecewise regression and its applications

Identifying breakpoints in piecewise regression is critical in enhancing the reliability and interpretability of data fitting. In this paper, we propose novel algorithms based on the greedy algorithm to accurately and efficiently identify breakpoints in piecewise polynomial regression. The algorithm updates the breakpoints to minimize the error by exploring the neighborhood of each breakpoint. It has a fast convergence rate and stability to find optimal breakpoints. Moreover, it can determine the optimal number of breakpoints. The computational results for real and synthetic data show that its accuracy is better than any existing methods. The real-world datasets demonstrate that breakpoints through the proposed algorithm provide valuable data information.

stat.ML

Singular systems of linear forms over global function fields

In this paper, we consider singular systems of linear forms over global function fields of class number one and give an upper bound for the Hausdorff dimension of the set of singular systems of linear forms by constructing an appropriate Margulis height function on the space of lattices over global function fields.

math.DS

Low T-Phase Rank Approximation of Third Order Tensors

We study low T-phase-rank approximation of sectorial third-order tensors $\mathscr{A}\in\mathbb{C}^{n\times n\times p}$ under the tensor T-product. We introduce canonical T-phases and T-phase rank, and formulate the approximation task as minimizing a symmetric gauge of the canonical phase vector under a T-phase-rank constraint. Our main tool is a tensor phase-majorization inequality for the geometric mean, obtained by lifting the matrix inequality through the block-circulant representation. In the positive-imaginary regime, this yields an exact optimal-value formula and an explicit optimal half-phase truncation family. We further establish tensor counterparts of classical matrix phase inequalities and derive a tensor small phase theorem for MIMO linear time-invariant systems.

math.NA

Tensor CUR Decomposition under the Linear-Map-Based Tensor-Tensor Multiplication

The factorization of three-dimensional data continues to gain attention due to its relevance in representing and compressing large-scale datasets. The linear-map-based tensor-tensor multiplication is a matrix-mimetic operation that extends the notion of matrix multiplication to higher order tensors, and which is a generalization of the T-product. Under this framework, we introduce the tensor CUR decomposition, show its performance in video foreground-background separation for different linear maps and compare it to a robust matrix CUR decomposition, another tensor approximation and the slice-based singular value decomposition (SS-SVD). We also provide a theoretical analysis of our tensor CUR decomposition, extending classical matrix results to establish exactness conditions and perturbation bounds.

math.NA

Mi:dm 2.0 Korea-centric Bilingual Language Models

We introduce Mi:dm 2.0, a bilingual large language model (LLM) specifically engineered to advance Korea-centric AI. This model goes beyond Korean text processing by integrating the values, reasoning patterns, and commonsense knowledge inherent to Korean society, enabling nuanced understanding of cultural contexts, emotional subtleties, and real-world scenarios to generate reliable and culturally appropriate responses. To address limitations of existing LLMs, often caused by insufficient or low-quality Korean data and lack of cultural alignment, Mi:dm 2.0 emphasizes robust data quality through a comprehensive pipeline that includes proprietary data cleansing, high-quality synthetic data generation, strategic data mixing with curriculum learning, and a custom Korean-optimized tokenizer to improve efficiency and coverage. To realize this vision, we offer two complementary configurations: Mi:dm 2.0 Base (11.5B parameters), built with a depth-up scaling strategy for general-purpose use, and Mi:dm 2.0 Mini (2.3B parameters), optimized for resource-constrained environments and specialized tasks. Mi:dm 2.0 achieves state-of-the-art performance on Korean-specific benchmarks, with top-tier zero-shot results on KMMLU and strong internal evaluation results across language, humanities, and social science tasks. The Mi:dm 2.0 lineup is released under the MIT license to support extensive research and commercial use. By offering accessible and high-performance Korea-centric LLMs, KT aims to accelerate AI adoption across Korean industries, public services, and education, strengthen the Korean AI developer community, and lay the groundwork for the broader vision of K-intelligence. Our models are available at https://huggingface.co/K-intelligence. For technical inquiries, please contact midm-llm@kt.com.

cs.CL

Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation

With the rapid advancement of mathematical reasoning capabilities in Large Language Models (LLMs), AI systems are increasingly being adopted in educational settings to support students' comprehension of problem-solving processes. However, a critical component remains underexplored in current LLM-generated explanations: multimodal explanation. In real-world instructional contexts, human tutors routinely employ visual aids, such as diagrams, markings, and highlights, to enhance conceptual clarity. To bridge this gap, we introduce the multimodal solution explanation task, designed to evaluate whether models can identify visual keypoints, such as auxiliary lines, points, angles, and generate explanations that incorporate these key elements essential for understanding. To evaluate model performance on this task, we propose ME2, a multimodal benchmark consisting of 1,000 math problems annotated with visual keypoints and corresponding explanatory text that references those elements. Our empirical results show that current models struggle to identify visual keypoints. In the task of generating keypoint-based explanations, open-source models also face notable difficulties. This highlights a significant gap in current LLMs' ability to perform mathematical visual grounding, engage in visually grounded reasoning, and provide explanations in educational contexts. We expect that the multimodal solution explanation task and the ME2 dataset will catalyze further research on LLMs in education and promote their use as effective, explanation-oriented AI tutors.

cs.CL

High entropy measures on the space of lattices with escape of mass

For any diagonal element $a$ with two eigenvalues, we construct a sequence of $a$-invariant probability measures on the space of unimodular lattices with high entropy but converging to the zero measure. This extends the result of Kadyrov [Ergodic Theory Dynam. Systems, 32(1) (2012)].

math.DS

Infinitely badly approximable affine forms

A pair $(A,\mathbf{b})$ of a real $m\times n$ matrix $A$ and $\mathbf{b}\in\mathbb{R}^m$ is said to be $\textit{infinitely badly approximable}$ if \[ \liminf_{\mathbf{q}\in\mathbb{Z}^n, \|\mathbf{q}\|\to\infty} \|\mathbf{q}\|^{\frac{n}{m}}\|A\mathbf{q}-\mathbf{b}\|_{\mathbb{Z}} =\infty, \] where $\|\cdot\|_\mathbb{Z}$ denotes the distance from the nearest integer vector. In this article, we introduce a novel concept of singularity for $(A,\mathbf{b})$ and characterize the infinitely badly approximable property by this singular property. As an application, we compute the Hausdorff dimension of the infinitely badly approximable set. We also discuss dynamical interpretations on the space of grids in $\mathbb{R}^{m+n}$.

math.NT

Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks

Stable and efficient training of ReLU networks with large depth is highly sensitive to weight initialization. Improper initialization can cause permanent neuron inactivation dying ReLU and exacerbate gradient instability as network depth increases. Methods such as He, Xavier, and orthogonal initialization preserve variance or promote approximate isometry. However, they do not necessarily regulate the pre-activation mean or control activation sparsity, and their effectiveness often diminishes in very deep architectures. This work introduces an orthogonal initialization specifically optimized for ReLU by solving an optimization problem on the Stiefel manifold, thereby preserving scale and calibrating the pre-activation statistics from the outset. A family of closed-form solutions and an efficient sampling scheme are derived. Theoretical analysis at initialization shows that prevention of the dying ReLU problem, slower decay of activation variance, and mitigation of gradient vanishing, which together stabilize signal and gradient flow in deep architectures. Empirically, across MNIST, Fashion-MNIST, multiple tabular datasets, few-shot settings, and ReLU-family activations, our method outperforms previous initializations and enables stable training in deep networks.

cs.LG

Knowledge Synthesis of Photosynthesis Research Using a Large Language Model

The development of biological data analysis tools and large language models (LLMs) has opened up new possibilities for utilizing AI in plant science research, with the potential to contribute significantly to knowledge integration and research gap identification. Nonetheless, current LLMs struggle to handle complex biological data and theoretical models in photosynthesis research and often fail to provide accurate scientific contexts. Therefore, this study proposed a photosynthesis research assistant (PRAG) based on OpenAI's GPT-4o with retrieval-augmented generation (RAG) techniques and prompt optimization. Vector databases and an automated feedback loop were used in the prompt optimization process to enhance the accuracy and relevance of the responses to photosynthesis-related queries. PRAG showed an average improvement of 8.7% across five metrics related to scientific writing, with a 25.4% increase in source transparency. Additionally, its scientific depth and domain coverage were comparable to those of photosynthesis research papers. A knowledge graph was used to structure PRAG's responses with papers within and outside the database, which allowed PRAG to match key entities with 63% and 39.5% of the database and test papers, respectively. PRAG can be applied for photosynthesis research and broader plant science domains, paving the way for more in-depth data analysis and predictive capabilities.

cs.CL

Hausdorff dimension of singular vectors in function fields

We compute the Hausdorff dimension of the set of singular vectors in function fields and bound the Hausdorff dimension of the set of $\varepsilon$-Dirichlet improvable vectors in this setting. This is a function field analogue of the results of Cheung and Chevallier [Duke Math. J. 165 (2016), 2273--2329].

math.NT

On the rate of convergence of continued fraction statistics of random rationals

We show that the statistics of the continued fraction expansion of a randomly chosen rational in the unit interval, with a fixed large denominator $q$, approaches the Gauss-Kuzmin statistics with polynomial rate in $q$. This improves on previous results giving the convergence without rate. As an application of this effective rate of convergence, we show that the statistics of a randomly chosen rational in the unit interval, with a fixed large denominator $q$ and prime numerator, also approaches the Gauss-Kuzmin statistics. Our results are obtained as applications of improved non-escape of mass and equidistribution statements for the geodesic flow on the space $SL_2(\mathbb{R})/SL_2(\mathbb{Z})$.

math.DS

Click-Gaussian: Interactive Segmentation to Any 3D Gaussians

Interactive segmentation of 3D Gaussians opens a great opportunity for real-time manipulation of 3D scenes thanks to the real-time rendering capability of 3D Gaussian Splatting. However, the current methods suffer from time-consuming post-processing to deal with noisy segmentation output. Also, they struggle to provide detailed segmentation, which is important for fine-grained manipulation of 3D scenes. In this study, we propose Click-Gaussian, which learns distinguishable feature fields of two-level granularity, facilitating segmentation without time-consuming post-processing. We delve into challenges stemming from inconsistently learned feature fields resulting from 2D segmentation obtained independently from a 3D scene. 3D segmentation accuracy deteriorates when 2D segmentation results across the views, primary cues for 3D segmentation, are in conflict. To overcome these issues, we propose Global Feature-guided Learning (GFL). GFL constructs the clusters of global feature candidates from noisy 2D segments across the views, which smooths out noises when training the features of 3D Gaussians. Our method runs in 10 ms per click, 15 to 130 times as fast as the previous methods, while also significantly improving segmentation accuracy. Our project page is available at https://seokhunchoi.github.io/Click-Gaussian

cs.CV