SearcharxivSearch

arXiv subjects

Junu Lee

Publications and source records attributed to Junu Lee.

7 recordsLinked to original sources

The ALMA survey to Resolve exoKuiper belt Substructures (ARKS) II. The radial structure of debris discs

The ALMA survey to Resolve exoKuiper belt Substructures (ARKS) was recently completed to cover the lack of high-resolution observations of debris discs and to investigate the prevalence of substructures such as radial gaps and rings in a sample of 24 discs. This study characterises the radial structure of debris discs in the ARKS programme. To identify and quantify the disc substructures, we modelled all discs with a range of non-parametric and parametric approaches. We find that of the 24 discs in the sample, 5 host multiple rings, 7 are single rings that display halos or additional low-amplitude rings, and 12 are single rings with at most tentative evidence of additional substructures. The fractional ring widths that we measured are significantly narrower than previously derived values, and they follow a distribution similar to the fractional widths of individual rings resolved in protoplanetary discs. However, there exists a population of rings in debris discs that are significantly wider than those in protoplanetary discs. We also find that discs with steep inner edges consistent with planet sculpting tend to be found at smaller (<100 au) radii, while more radially extended discs tend to have shallower edges more consistent with collisional evolution. An overwhelming majority of discs have radial profiles well-described by either a double power law or double-Gaussian parametrisation. While our findings suggest that it may be possible for some debris discs to inherit their structures directly from protoplanetary discs, there exists a sizeable population of broad debris discs that cannot be explained in this way. Assuming that the distribution of millimetre dust reflects the distribution of planetesimals, mechanisms that cause rings in protoplanetary discs to migrate or debris discs to broaden soon after formation may be at play, possibly mediated by planetary migration or scattering.

astro-ph.EP

The ALMA survey to Resolve exoKuiper belt Substructures (ARKS) III: The vertical structure of debris disks

Debris disks -- collisionally sustained belts of dust and sometimes gas around main sequence stars -- are remnants of planet formation processes and are found in systems ${\gtrsim}10$ Myr old. Millimeter-wavelength observations are particularly important, as the grains probed by these observations are not strongly affected by radiation pressure and stellar winds, allowing them to probe the dynamics of large bodies producing dust. The ALMA survey to Resolve exoKuiper belt Substructures (ARKS) is analyzing high-resolution observations of 24 debris disks to enable the characterization of debris disk substructures across a large sample for the first time. For the most highly inclined disks, it is possible to recover the vertical structure of the disk. We aim to model and analyze the most highly inclined systems in the ARKS sample in order to uniformly extract the vertical dust distributions for a sample of well-resolved debris disks. We employed both parametric and nonparametric methods to constrain the vertical dust distributions for the most highly inclined ARKS targets. We find a broad range of aspect ratios, revealing a wide diversity in vertical structure, with a range of best-fit parametric values of $0.0026 \leq h_{\rm HWHM} \leq 0.193$ and a median best-fit value of $h_{\rm HWHM}=0.021$. The results obtained by nonparametric modeling are generally consistent with the parametric modeling results. We find that five of the 13 disks are consistent with having total disk masses less than that of Neptune (17 $M_{\oplus}$), assuming stirring by internal processes (self-stirring and collisional and frictional damping). Furthermore, most systems show a significant preference for a Lorentzian vertical profile rather than a Gaussian.

astro-ph.EP

Power of masking methods for adaptive testing in a multivariate normal means problem

Many large-scale testing procedures learn signal structure from the data to boost power. Direct data reuse can inflate Type-I error ("double dipping"), so a common remedy is masking: withholding some information during learning and using it for testing. Sample splitting masks by withholding observations for testing, while null augmentation (e.g., knockoffs or full-conformal outlier detection) masks by appending null samples or variables and withholding their identities until testing. In many settings, little is known about how the power of masking methods compares across mechanisms, across tuning choices, or against more data-efficient non-masking alternatives. We study these questions in a stylized two-groups multivariate normal means model with an unknown signal direction learned from the data. Within this testbed, we develop a transparent, unified set of asymptotic power expressions for three parallel methods differing in masking choices: a sample splitting method, a full-conformal-style null augmentation method, and an oracle in-sample benchmark. Our main findings are: (1) the augmentation method is more powerful than the splitting method with matched tuning; (2) the power-optimal number of null samples for the augmentation method is a vanishing fraction of the number of tests, in which case its power approaches that of the in-sample benchmark; and (3) for a tractable approximation to the augmentation method, the optimal number of null samples scales as the square root of the number of tests, with empirical evidence suggesting a similar scaling for the method itself. These results characterize masking-induced power trade-offs in a tractable model and suggest qualitative lessons for other settings.

math.ST

Full-conformal novelty detection

This paper presents a powerful methodology for flexible full-data nonparametric novelty detection that offers distribution-free false discovery rate (FDR) control guarantees. Building on the full conformal inference framework and the concept of e-values, we introduce full conformal e-values to quantify evidence for novelty relative to a given reference dataset. These e-values are then utilized by carefully crafted multiple testing procedures to identify a set of novel units out-of-sample with provable finite-sample FDR control. We showcase several instantiations of e-values, including those which employ a data-driven model selection strategy to amplify power. Furthermore, our framework is extended to address distribution shift, accommodating scenarios where novelty detection must be performed on data drawn from a shifted distribution relative to the reference dataset. In all settings, our method can perform powerfully -- outperforming existing novelty detection methods -- even with limited amounts of reference data; this is illustrated by empirical evaluations on synthetic data and an application to a malicious LLM prompts dataset.

stat.ME

A general condition for bias attenuation by a nondifferentially mismeasured confounder

In real-world studies, the collected confounders may suffer from measurement error. Although mismeasurement of confounders is typically unintentional -- originating from sources such as human oversight or imprecise machinery -- deliberate mismeasurement also occurs and is becoming increasingly more common. For example, in the 2020 U.S. Census, noise was added to measurements to assuage privacy concerns. Sensitive variables such as income or age are oftentimes partially censored and are only known up to a range of values. In such settings, obtaining valid estimates of the causal effect of a binary treatment can be impossible, as mismeasurement of confounders constitutes a violation of the no unmeasured confounding assumption. A natural question is whether the common practice of simply adjusting for the mismeasured confounder is justifiable. In this article, we answer this question in the affirmative and demonstrate that in many realistic scenarios not covered by previous literature, adjusting for the mismeasured confounders reduces bias compared to not adjusting.

stat.ME

Boosting e-BH via conditional calibration

The e-BH procedure is an e-value-based multiple testing procedure that provably controls the false discovery rate (FDR) under any dependence structure between the e-values. Despite this appealing theoretical FDR control guarantee, the e-BH procedure often suffers from low power in practice. In this paper, we propose a general framework that boosts the power of e-BH without sacrificing its FDR control under arbitrary dependence. This is achieved by the technique of conditional calibration, where we take as input the e-values and calibrate them to be a set of "boosted e-values" that are guaranteed to be no less -- and are often more -- powerful than the original ones. Our general framework is explicitly instantiated in three classes of multiple testing problems: (1) testing under parametric models, (2) conditional independence testing under the model-X setting, and (3) model-free conformalized selection. Extensive numerical experiments show that our proposed method significantly improves the power of e-BH while continuing to control the FDR. We also demonstrate the effectiveness of our method through an application to an observational study dataset for identifying individuals whose counterfactuals satisfy certain properties.

stat.ME

Navigating Data Heterogeneity in Federated Learning A Semi-Supervised Federated Object Detection

Federated Learning (FL) has emerged as a potent framework for training models across distributed data sources while maintaining data privacy. Nevertheless, it faces challenges with limited high-quality labels and non-IID client data, particularly in applications like autonomous driving. To address these hurdles, we navigate the uncharted waters of Semi-Supervised Federated Object Detection (SSFOD). We present a pioneering SSFOD framework, designed for scenarios where labeled data reside only at the server while clients possess unlabeled data. Notably, our method represents the inaugural implementation of SSFOD for clients with 0% labeled non-IID data, a stark contrast to previous studies that maintain some subset of labels at each client. We propose FedSTO, a two-stage strategy encompassing Selective Training followed by Orthogonally enhanced full-parameter training, to effectively address data shift (e.g. weather conditions) between server and clients. Our contributions include selectively refining the backbone of the detector to avert overfitting, orthogonality regularization to boost representation divergence, and local EMA-driven pseudo label assignment to yield high-quality pseudo labels. Extensive validation on prominent autonomous driving datasets (BDD100K, Cityscapes, and SODA10M) attests to the efficacy of our approach, demonstrating state-of-the-art results. Remarkably, FedSTO, using just 20-30% of labels, performs nearly as well as fully-supervised centralized training methods.

cs.CV