SearcharxivSearch

arXiv subjects

Hidetoshi Matsui

Publications and source records attributed to Hidetoshi Matsui.

13 recordsLinked to original sources

Clustering-based aggregate value regression

In various practical situations, forecasting of aggregate values rather than individual ones is often our main focus. For instance, electricity companies are interested in forecasting the total electricity demand in a specific region to ensure reliable grid operation and resource allocation. However, to our knowledge, statistical learning specifically for forecasting aggregate values has not yet been well-established. In particular, the relationship between forecast error and the number of clusters has not been well studied, as clustering is usually treated as unsupervised learning. This study introduces a novel forecasting method specifically focused on the aggregate values in the linear regression model. We call it the Aggregate Value Regression (AVR), and it is constructed by combining all regression models into a single model. With the AVR, we must estimate a huge number of parameters when the number of regression models to be combined is large, resulting in overparameterization. To address the overparameterization issue, we introduce a hierarchical clustering technique, referred to as AVR-C (C stands for clustering). In this approach, several clusters of regression models are constructed, and the AVR is performed within each cluster. The AVR-C introduces a novel bias-variance trade-off theory under the assumption of a misspecified model. In this framework, the number of clusters characterizes model complexity. Monte Carlo simulation is conducted to investigate the behavior of training and test errors of our proposed clustering technique. The bias-variance trade-off theory is also demonstrated through the analysis of electricity demand forecasting.

stat.ME

Density Ratio-based Causal Discovery from Bivariate Continuous-Discrete Data

We address the problem of inferring the causal direction between a continuous variable $X$ and a discrete variable $Y$ from observational data. For the model $X \to Y$, we adopt the threshold model used in prior work. For the model $Y \to X$, we consider two cases: (1) the conditional distributions of $X$ given different values of $Y$ form a location-shift family, and (2) they are mixtures of generalized normal distributions with independently parameterized components. We establish identifiability of the causal direction through three theoretical results. First, we prove that under $X \to Y$, the density ratio of $X$ conditioned on different values of $Y$ is monotonic. Second, we establish that under $Y \to X$ with non-location-shift conditionals, monotonicity of the density ratio holds only on a set of Lebesgue measure zero in the parameter space. Third, we show that under $X \to Y$, the conditional distributions forming a location-shift family requires a precise coordination between the causal mechanism and input distribution, which is non-generic under the principle of independent mechanisms. Together, these results imply that monotonicity of the density ratio characterizes the direction $X \to Y$, whereas non-monotonicity or location-shift conditionals characterizes $Y \to X$. Based on this, we propose Density Ratio-based Causal Discovery (DRCD), a method that determines causal direction by testing for location-shift conditionals and monotonicity of the estimated density ratio. Experiments on synthetic and real-world datasets demonstrate that DRCD outperforms existing methods.

cs.LG

Reconciling Functional Data Regression with Excess Bases

As the development of measuring instruments and computers has accelerated the collection of massive amounts of data, functional data analysis (FDA) has experienced a surge of attention. The FDA methodology treats longitudinal data as a set of functions on which inference, including regression, is performed. Functionalizing data typically involves fitting the data with basis functions. In general, the number of basis functions smaller than the sample size is selected. This paper casts doubt on this convention. Recent statistical theory has revealed the so-called double-descent phenomenon in which excess parameters overcome overfitting and lead to precise interpolation. Applying this idea to choosing the number of bases to be used for functional data, we show that choosing an excess number of bases can lead to more accurate predictions. Specifically, we explored this phenomenon in a functional regression context and examined its validity through numerical experiments. In addition, we introduce two real-world datasets to demonstrate that the double-descent phenomenon goes beyond theoretical and numerical experiments, confirming its importance in practical applications.

stat.ME

Sparse estimation in ordinary kriging for functional data

We introduce a sparse estimation in the ordinary kriging for functional data. The functional kriging predicts a feature given as a function at a location where the data are not observed by a linear combination of data observed at other locations. To estimate the weights of the linear combination, we apply the lasso-type regularization in minimizing the expected squared error. We derive an algorithm to derive the estimator using the augmented Lagrange method. Tuning parameters included in the estimation procedure are selected by cross-validation. Since the proposed method can shrink some of the weights of the linear combination toward zeros exactly, we can investigate which locations are necessary or unnecessary to predict the feature. Simulation and real data analysis show that the proposed method appropriately provides reasonable results.

stat.ME

Variable screening using factor analysis for high-dimensional data with multicollinearity

Screening methods are useful tools for variable selection in regression analysis when the number of predictors is much larger than the sample size. Factor analysis is used to eliminate multicollinearity among predictors, which improves the variable selection performance. We propose a new method, called Truncated Preconditioned Profiled Independence Screening (TPPIS), that better selects the number of factors to eliminate multicollinearity. The proposed method improves the variable selection performance by truncating unnecessary parts from the information obtained by factor analysis. We confirmed the superior performance of the proposed method in variable selection through analysis using simulation data and real datasets.

stat.ME

Truncated estimation for varying-coefficient functional linear model

Varying-coefficient functional linear models consider the relationship between a response and a predictor, where the response depends not only the predictor but also an exogenous variable. It then accounts for the relation of the predictors and the response varying with the exogenous variable. We consider the method for truncated estimation for varying-coefficient models. The proposed method can clarify to what time point the functional predictor relates to the response at any value of the exogenous variable by investigating the coefficient function. To obtain the truncated model, we apply the penalized likelihood method with the sparsity inducing penalty. Simulation studies are conducted to investigate the effectiveness of the proposed method. We also report the application of the proposed method to the analysis of crop yield data to investigate when an environmental factor relates to the crop yield at any season.

stat.ME

Classification from Positive and Biased Negative Data with Skewed Labeled Posterior Probability

The binary classification problem has a situation where only biased data are observed in one of the classes. In this paper, we propose a new method to approach the positive and biased negative (PbN) classification problem, which is a weakly supervised learning method to learn a binary classifier from positive data and negative data with biased observations. We incorporate a method to correct the negative impact due to skewed confidence, which represents the posterior probability that the observed data are positive. This reduces the distortion of the posterior probability that the data are labeled, which is necessary for the empirical risk minimization of the PbN classification problem. We verified the effectiveness of the proposed method by numerical experiments and real data analysis.

stat.ME

Sparse varying-coefficient functional linear model

We consider the problem of variable selection in varying-coefficient functional linear models, where multiple predictors are functions and a response is a scalar and depends on an exogenous variable. The varying-coefficient functional linear model is estimated by the penalized maximum likelihood method with the sparsity-inducing penalty. Tuning parameters that controls the degree of the penalization are determined by a model selection criterion. The proposed method can reveal which combination of functional predictors relates to the response, and furthermore how each predictor relates to the response by investigating coefficient surfaces. Simulation studies are provided to investigate the effectiveness of the proposed method. We also apply it to the analysis of crop yield data to investigate which combination of environmental factors relates to the amount of a crop yield.

stat.ME

Varying-coefficient functional additive models

We extend the varying coefficient functional linear model to the nonlinear model and propose a varying coefficient functional additive model. The proposed method can represent the relationship between functional predictors and a scalar response where the response depends on an exogenous variable.It captures the nonlinear structure between variables and also provides interpretable relationship of them. The model is estimated through basis expansions and penalized likelihood method, and then the tuning parameters included at the estimation procedure are selected by a model selection criterion. Simulation studies are provided to show the effectiveness of the proposed method. We also apply it to the analysis of crop yield data and then investigate how and when the environmental factor relates to the amount of the crop yield.

stat.ME

Quadratic regression for functional response models

We consider the problem of constructing a regression model with a functional predictor and a functional response. We extend the functional linear model to the quadratic model, where the quadratic term also takes the interaction between the argument of the functional data into consideration. We assume that the predictor and the coefficient functions are expressed by basis expansions, and then parameters included in the model are estimated by the penalized likelihood method assuming that the error function follows a Gaussian process. Monte Carlo simulations are conducted to illustrate the efficacy of the proposed method. Finally, we apply the proposed method to the analysis of meteorological data and explore the results.

stat.ME

Selection of variables and decision boundaries for functional data via bi-level selection

Sparsity-inducing penalties are useful tools for variable selection and they are also effective for regression settings where the data are functions. We consider the problem of selecting not only variables but also decision boundaries in logistic regression models for functional data, using the sparse regularization. The functional logistic regression model is estimated by the framework of the penalized likelihood method with the sparse group lasso-type penalty, and then tuning parameters are selected using the model selection criterion. The effectiveness of the proposed method is investigated through real data analysis.

stat.ME

Model selection criteria for nonlinear mixed effects modeling

We consider constructing model selection criteria for evaluating nonlinear mixed effects models via basis expansions. Mean functions and random functions in the mixed effects model are expressed by basis expansions, then they are estimated by the maximum likelihood method. In order to select numbers of basis we derive a Bayesian model selection criterion for evaluating nonlinear mixed effects models estimated by the maximum likelihood method. Simulation results shows the effectiveness of the mixed effects modeling.

stat.ME

Varying-coefficient modeling via regularized basis functions

We address the problem of constructing varying-coefficient models based on basis expansions along with the technique of regularization. A crucial point in our modeling procedure is the selection of smoothing parameters in the regularization method. In order to choose the parameters objectively, we derive model selection criteria from the viewpoints of information-theoretic and Bayesian approach. We demonstrate the effectiveness of proposed modeling strategy through Monte Carlo simulations and analyzing a real data set.

stat.ME