SearcharxivSearch

arXiv subjects

Eunchan Kim

Publications and source records attributed to Eunchan Kim.

3 recordsLinked to original sources

SCOPE-FE: Structured Control of Operator and Pairwise Exploration for Feature Engineering via Quality-Aware Candidate-Space Reduction

Automatic feature engineering can improve predictive performance on tabular data by generating diverse feature transformations. However, the candidate space induced by combinations of input features and operators grows rapidly with dimensionality, resulting in substantial computational cost. We propose SCOPE-FE, a framework that controls the search space before candidate generation. SCOPE-FE combines FeatureClustering, a structural pair gate based on mixed-type feature association, with OperatorProbing, a dataset-specific utility control over operators. Unlike conventional expand-and-reduce approaches that generate a large candidate set and prune it afterward, SCOPE-FE focuses computation on a smaller, data-dependent candidate pool. Across ten OpenFE benchmark datasets, SCOPE-FE achieves a median candidate-space reduction of 82.9% and lowers component-summed feature-engineering time-including separately measured FeatureClustering overhead-on all ten datasets, yielding a geometric-mean speedup of 2.66x and a maximum speedup of 5.48x. Despite this reduction, SCOPE-FE is within the stated practical-equivalence margin of OpenFE on 8 of 10 datasets. An exhaustive candidate audit shows enrichment above uniform-random expectation on 8 of 10 datasets, with a median enrichment of 1.35x. Against Random-Pair, SCOPE-FE has higher enrichment on 6 of 10 datasets, with 5 of 10 significant; against Random-Operator and Random-Joint, it has higher enrichment on 8 of 10 datasets, with 8 of 10 significant for each. These results demonstrate that pre-generation search-space control can substantially reduce feature-engineering time while retaining a utility-enriched candidate pool under the OpenFE-compatible evaluation protocol.

stat.ML

The Association Between SOC and Land Prices Considering Spatial Heterogeneity Based on Finite Mixture Modeling

An understanding of how Social Overhead Capital (SOC) is associated with the land value of the local community is important for effective urban planning. However, even within a district, there are multiple sections used for different purposes; the term for this is spatial heterogeneity. The spatial heterogeneity issue has to be considered when attempting to comprehend land prices. If there is spatial heterogeneity within a district, land prices can be managed by adopting the spatial clustering method. In this study, spatial attributes including SOC, socio-demographic features, and spatial information in a specific district are analyzed with Finite Mixture Modeling (FMM) in order to find (a) the optimal number of clusters and (b) the association among SOCs, socio-demographic features, and land prices. FMM is a tool used to find clusters and the attributes' coefficients simultaneously. Using the FMM method, the results show that four clusters exist in one district and the four clusters have different associations among SOCs, demographic features, and land prices. Policymakers and managerial administration need to look for information to make policy about land prices. The current study finds the consideration of closeness to SOC to be a significant factor on land prices and suggests the potential policy direction related to SOC.

stat.AP

ALBERT with Knowledge Graph Encoder Utilizing Semantic Similarity for Commonsense Question Answering

Recently, pre-trained language representation models such as bidirectional encoder representations from transformers (BERT) have been performing well in commonsense question answering (CSQA). However, there is a problem that the models do not directly use explicit information of knowledge sources existing outside. To augment this, additional methods such as knowledge-aware graph network (KagNet) and multi-hop graph relation network (MHGRN) have been proposed. In this study, we propose to use the latest pre-trained language model a lite bidirectional encoder representations from transformers (ALBERT) with knowledge graph information extraction technique. We also propose to applying the novel method, schema graph expansion to recent language models. Then, we analyze the effect of applying knowledge graph-based knowledge extraction techniques to recent pre-trained language models and confirm that schema graph expansion is effective in some extent. Furthermore, we show that our proposed model can achieve better performance than existing KagNet and MHGRN models in CommonsenseQA dataset.

cs.CL