SearcharxivSearch

arXiv subjects

Yuya Sasaki

Publications and source records attributed to Yuya Sasaki.

At least 19 recordsLinked to original sources

Inference for High-Dimensional Network Data

We develop a novel method of inference for network-dependent high-dimensional random vectors. Dependence is characterized via a functional dependence measure based on graph distance, allowing the approximation theory to capture the interaction between the decay of dependence and the growth of network neighborhoods. We establish Gaussian approximation results for the maximum norm under finite-moment and sub-Weibull conditions, providing explicit conditions under which the dimension may increase with the network size. We also propose a high-dimensional network HAC covariance estimator and establish its convergence properties, yielding a feasible procedure for simultaneous inference. Simulation studies demonstrate favorable finite-sample performance of the proposed method. We apply the procedure to study how spillover effects vary with an index of network homophily by constructing confidence bands for the conditional spillover-effect function. The application reveals heterogeneity and local significance that would be obscured by conventional low-dimensional inference.

econ.EM

Declining Modularity of Intellectual Bases During the Emergence of Research Areas

Understanding how research areas emerge can help identify nascent areas early and inform research strategy, yet how the intellectual base of a field restructures as an area takes shape remains unclear. We hypothesize that the emergence of a research area is accompanied by the integration of largely separate knowledge communities, observable as a decline in the modularity of its co-citation network, which represents its intellectual base. We propose a framework that tracks this modularity over time, evaluates the statistical robustness of its changes, and identifies the papers highly associated with the decline. We applied it to three areas with different modes of growth: higher-order network science, superstring theory, and graph representation learning. In all three, modularity declined in correspondence with each area's emergence or transformation, and in superstring theory, the decline aligns with an independently documented transition. Further analysis of higher-order network science shows that its decline reflects a cross-disciplinary integration. In graph representation learning, the gradual decline is followed by a rise, which we interpret as a re-differentiation after the emergence period. Our results suggest that a decline in the modularity of a co-citation network can serve as a structural signature that retrospectively characterizes this integrative mode of emergence.

physics.soc-ph

Bounds for Standard Errors in Combined Data

We propose methods for constructing lower bounds on the standard errors of parameters estimated from moment conditions obtained across different samples. Sharp explicit bounds are derived by exploiting geometric inequalities when no information about correlations across samples is available. Furthermore, we develop computationally tractable sharp bounds for more general settings with no or partial correlation information, which can be obtained by solving a simple semidefinite program. Finally, we illustrate the practical usefulness of our method through three empirical cases: two macroeconomics examples involving menu cost and Heterogeneous Agent New-Keynesian models; and a two sample instrumental variable microeconomic study.

econ.EM

Choosing A Headline Estimand from Matching, DID, and Hybrid Designs: A Minimax-Regret Approach

Researchers using panel data to estimate causal effects routinely choose among three approaches to using past outcomes: difference-in-differences (DID), conditioning on lagged outcomes (matching, M), and a hybrid that does both (DIDM). The corresponding identifying assumptions are non-nested, leaving little guidance on which to report. We give conditions under which the corresponding estimands are ordered, with DIDM bracketed between matching and DID. This makes DIDM the minimax-regret choice among the three under a broad class of loss functions. We recommend reporting DIDM as the headline estimate, with matching and DID as bounds. We illustrate in applications.

econ.EM

Graph Neural Network leveraging Higher-order Class Label Connectivity for Heterophilous Graphs

Node classification in graph neural networks (GNNs) has been widely applied in various fields of graph analysis. GNNs achieve high-accuracy node classification in homophilous graphs, where nodes with the same class label tend to be connected. However, their performance remains limited in heterophilous graphs, where nodes with different class labels are more likely to be connected. In particular, current GNNs derived from graph convolutional networks cannot capture higher-order class label connectivity, which is frequently observed in real-world heterophilous graphs. To address this issue, we propose a novel classifier, Label Context Classifier (LCC), designed to capture higher-order class label connectivity in directed graphs. LCC estimates the class label of a target node by leveraging label context embeddings that are generated through four distinct types of walks. In addition, our approach allows the integration of LCC and any GNN by adaptively learning their importance. Experimental results demonstrate that GNNs integrated with LCC outperform SOTA methods and the label context embeddings improve the node classification performance in heterophilous directed graphs.

cs.LG

Workload acceleration by optimizing materialized view selection using local search

The growing size of database workloads has made view selection a key performance challenge. Materializing frequent sub-queries in workloads improves query efficiency, but it incurs significant view maintenance costs due to updates. Although existing methods such as BIGSUBS address this trade-off between the benefit of using materialized views and the overhead of view maintenance, they have two drawbacks: insufficient maintenance cost modeling and ineffective view selection due to probabilistic techniques. We propose a novel view selection method that incorporates incremental view maintenance cost directly into the optimization objective of an integer linear program and applies local search to efficiently explore the solution space. In order to apply local search to the view selection problem, we develop neighboring solutions using sub-query containment, and select initial solutions based on sub-query frequency, utility, or utility per storage unit. Experiments using Redbench, a benchmark simulating real-world query workloads on Amazon Redshift, show that our approach outperforms BIGSUBS in both optimization utility and the quality of selected views.

cs.DB

Root-$n$ Asymptotically Normal Maximum Score Estimation

The maximum score method (Manski, 1975, 1985) is a powerful approach for binary choice models, yet it is known to face both practical and theoretical challenges. In particular, the estimator converges at a slower-than-root-$n$ rate to a nonstandard limiting distribution. We investigate conditions under which strictly concave surrogate score functions can be employed to achieve identification through a smooth criterion function. This criterion enables root-$n$ convergence to a normal limiting distribution. While the conditions to guarantee these desired properties are nontrivial, we characterize them in terms of primitive conditions. Extensive simulation studies support, the root-$n$ convergence rate, the asymptotic normality, and the validity of the standard inference methods.

econ.EM

Systemic Gendered Citation Imbalance in Computer Science: Evidence from Conferences and Journals

Gender imbalance persists across science, technology, engineering, and mathematics (STEM) fields, including computer science, where it appears in researcher demographics, productivity, recognition, hiring, and career progression. Given computer science's rapid expansion and global influence, addressing this imbalance is essential for broadening participation and fueling innovation. Although journal-oriented disciplines exhibit consistent gender imbalances in citation practices, it remains unclear whether similar patterns arise in the conference-centric culture of computer science. Here, we systematically investigate gender imbalance in citations of conference and journal papers in computer science. We find that papers for which a woman is listed as either first or last author receive fewer citations than expected, partly because of homophilic citation tendencies (i.e., authors tend to cite papers that share specific attributes). This imbalance is especially pronounced for conference papers--particularly those published at top-tier venues--relative to journals. Moreover, we find that the prominence of the first or last author and the structure of their local co-authorship networks are potential drivers of these imbalances. By exploring how conference-centric publishing practices can amplify systemic imbalances in computer science, our study offers insights that may inform efforts to foster more equitable representation in academia.

cs.DL

Heterogeneous Effects of Endogenous Treatments with Interference and Spillovers in a Large Network

This paper studies the identification and estimation of heterogeneous effects of an endogenous treatment under interference and spillovers in a large single-network setting. We model endogenous treatment selection as an equilibrium outcome that explicitly accounts for spillovers and derive conditions guaranteeing the existence and uniqueness of this equilibrium. We then identify heterogeneous marginal exposure effects (MEEs), which may vary with both the treatment status of neighboring nodes and unobserved heterogeneity. We develop estimation strategies and establish their large-sample properties. Equipped with these tools, we analyze the heterogeneous effects of import competition on U.S. local labor markets in the presence of interference and spillovers. We find negative MEEs, consistent with the existing literature. However, these effects are amplified by spillovers in the presence of treated neighbors and among localities that tend to select into lower levels of import competition. These additional empirical findings are novel and would not be credibly obtainable without the econometric framework proposed in this paper.

econ.EM

Learning Multi-Order Block Structure in Higher-Order Networks

Higher-order networks, naturally described as hypergraphs, are essential for modeling real-world systems involving interactions among three or more entities. Stochastic block models offer a principled framework for characterizing mesoscale organization, yet their extension to hypergraphs involves a trade-off between expressive power and computational complexity. A recent simplification, a single-order model, mitigates this complexity by assuming a single affinity pattern governs interactions of all orders. This universal assumption, however, may overlook order-dependent structural details. Here, we propose a framework that relaxes this assumption by introducing a multi-order block structure, in which different affinity patterns govern distinct subsets of interaction orders. Our framework is based on a multi-order stochastic block model and searches for the optimal partition of the set of interaction orders that maximizes out-of-sample hyperlink prediction performance. Analyzing a diverse range of real-world networks, we find that multi-order block structures are prevalent. Accounting for them not only yields better predictive performance over the single-order model but also uncovers sharper, more interpretable mesoscale organization. Our findings reveal that order-dependent mechanisms are a key feature of the mesoscale organization of real-world higher-order networks.

cs.SI

Nonparametric Uniform Inference in Binary Classification and Policy Values

We develop methods for nonparametric uniform inference in cost-sensitive binary classification, a framework that encompasses maximum score estimation, predicting utility maximizing actions, and policy learning. These problems are well known for slow convergence rates and non-standard limiting behavior, even under point identified parametric frameworks. In nonparametric settings, they may further suffer from failures of identification. To address these challenges, we introduce a strictly convex surrogate loss that point-identifies a representative nonparametric policy function. We then estimate this representative policy function to conduct inference on both the optimal classification policy and the optimal policy value. This approach enables Gaussian inference, substantially simplifying empirical implementation relative to working directly with the original classification problem. In particular, we establish root-$n$ asymptotic normality for the optimal policy value and derive a Gaussian approximation for the optimal classification policy at the standard nonparametric rate. Extensive simulation studies corroborate the theoretical findings. We apply our method to the National JTPA Study to conduct inference on the optimal treatment assignment policy and its associated welfare.

econ.EM

Benchmarking Fairness-aware Graph Neural Networks in Knowledge Graphs

Graph neural networks (GNNs) are powerful tools for learning from graph-structured data but often produce biased predictions with respect to sensitive attributes. Fairness-aware GNNs have been actively studied for mitigating biased predictions. However, no prior studies have evaluated fairness-aware GNNs on knowledge graphs, which are one of the most important graphs in many applications, such as recommender systems. Therefore, we introduce a benchmarking study on knowledge graphs. We generate new graphs from three knowledge graphs, YAGO, DBpedia, and Wikidata, that are significantly larger than the existing graph datasets used in fairness studies. We benchmark inprocessing and preprocessing methods in different GNN backbones and early stopping conditions. We find several key insights: (i) knowledge graphs show different trends from existing datasets; clearer trade-offs between prediction accuracy and fairness metrics than other graphs in fairness-aware GNNs, (ii) the performance is largely affected by not only fairness-aware GNN methods but also GNN backbones and early stopping conditions, and (iii) preprocessing methods often improve fairness metrics, while inprocessing methods improve prediction accuracy.

cs.LG

Robust Econometrics for Growth-at-Risk

The Growth-at-Risk (GaR) framework has garnered attention in recent econometric literature, yet current approaches implicitly assume a constant Pareto exponent. We introduce novel and robust econometrics to estimate the tails of GaR based on a rigorous theoretical framework and establish validity and effectiveness. Simulations demonstrate consistent outperformance relative to existing alternatives in terms of predictive accuracy. We perform a long-term GaR analysis that provides accurate and insightful predictions, effectively capturing financial anomalies better than current methods.

econ.EM

Algorithmic Fairness: Not a Purely Technical but Socio-Technical Property

The rapid trend of deploying artificial intelligence (AI) and machine learning (ML) systems in socially consequential domains has raised growing concerns about their trustworthiness, including potential discriminatory behaviours. Research in algorithmic fairness has generated a proliferation of mathematical definitions and metrics, yet persistent misconceptions and limitations -- both within and beyond the fairness community -- limit their effectiveness, such as an unreached consensus on its understanding, prevailing measures primarily tailored to binary group settings, and superficial handling for intersectional contexts. Here we critically remark on these misconceptions and argue that fairness cannot be reduced to purely technical constraints on models; we also examine the limitations of existing fairness measures through conceptual analysis and empirical illustrations, showing their limited applicability in the face of complex real-world scenarios, challenging prevailing views on the incompatibility between accuracy and fairness as well as that among fairness measures themselves, and outlining three worth-considering principles in the design of fairness measures. We believe these findings will help bridge the gap between technical formalisation and social realities and meet the challenges of real-world AI/ML deployment.

cs.LG

GMM and M Estimation under Network Dependence

This paper presents GMM and M estimators and their asymptotic properties for network-dependent data. To this end, I build on Kojevnikov, Marmer, and Song (KMS, 2021) and develop a novel uniform law of large numbers (ULLN), which is essential to ensure desired asymptotic behaviors of nonlinear estimators (e.g., Newey and McFadden, 1994, Section 2). Using this ULLN, I establish the consistency and asymptotic normality of both GMM and M estimators. For practical convenience, complete estimation and inference procedures are also provided.

econ.EM

Matching $\leq$ Hybrid $\leq$ Difference in Differences

Since LaLonde's (1986) seminal paper, there has been ongoing interest in estimating treatment effects using pre- and post-intervention data. Scholars have traditionally used experimental benchmarks to evaluate the accuracy of alternative econometric methods, including Matching, Difference-in-Differences (DID), and their hybrid forms (e.g., Heckman et al., 1998b; Dehejia and Wahba, 2002; Smith and Todd, 2005). We revisit these methodologies in the evaluation of job training and educational programs using four datasets (LaLonde, 1986; Heckman et al., 1998a; Smith and Todd, 2005; Chetty et al., 2014a; Athey et al., 2020), and show that the inequality relationship, Matching $\leq$ Hybrid $\leq$ DID, appears as a consistent norm, rather than a mere coincidence. We provide a formal theoretical justification for this puzzling phenomenon under plausible conditions such as negative selection, by generalizing the classical bracketing (Angrist and Pischke, 2009, Section 5). Consequently, when treatments are expected to be non-negative, DID tends to provide optimistic estimates, while Matching offers more conservative ones.

econ.EM

Mining Path Association Rules in Large Property Graphs (with Appendix)

How can we mine frequent path regularities from a graph with edge labels and vertex attributes? The task of association rule mining successfully discovers regular patterns in item sets and substructures. Still, to our best knowledge, this concept has not yet been extended to path patterns in large property graphs. In this paper, we introduce the problem of path association rule mining (PARM). Applied to any \emph{reachability path} between two vertices within a large graph, PARM discovers regular ways in which path patterns, identified by vertex attributes and edge labels, co-occur with each other. We develop an efficient and scalable algorithm PIONEER that exploits an anti-monotonicity property to effectively prune the search space. Further, we devise approximation techniques and employ parallelization to achieve scalable path association rule mining. Our experimental study using real-world graph data verifies the significance of path association rules and the efficiency of our solutions.

cs.DB