SearcharxivSearch

arXiv subjects

Junjie Dong

Publications and source records attributed to Junjie Dong.

12 recordsLinked to original sources

Interpretable Clustering: A Survey

In recent years, much of the research on clustering algorithms has primarily focused on enhancing their accuracy and efficiency, frequently at the expense of interpretability. However, as these methods are increasingly being applied in high-stakes domains such as healthcare, finance, and autonomous systems, the need of transparent and interpretable clustering outcomes has become a critical concern. This is not only necessary for gaining user trust but also for satisfying the growing ethical and regulatory demands in these fields. Ensuring that decisions derived from clustering algorithms can be clearly understood and justified is now a fundamental requirement. To address this need, this paper provides a comprehensive and structured review of the current state of explainable clustering algorithms, identifying key criteria to distinguish between various methods. These insights can effectively assist researchers in making informed decisions about the most suitable explainable clustering methods for specific application contexts, while also promoting the development and adoption of clustering algorithms that are both efficient and transparent. For convenient access and reference, an open repository organizes representative and emerging interpretable clustering methods under the taxonomy proposed in this survey, available at https://hulianyu.xyz/Interpretable-Clustering-Repository

cs.LG

Metal Saturation and the Redistribution of Hydrogen in Earth's Mantle

Iron disproportionation reactions in mantle silicates can produce metallic iron that drives Earth's deep mantle toward metal saturation under reduced conditions. Subducting slabs transport hydrated silicates to these depths, where interactions with metallic iron can reduce structurally bound hydrogen in silicates to reduced hydrogen-bearing phases, such as molecular hydrogen or iron hydrides, leaving mantle rocks in effect dry. Using the thermodynamic code HeFESTo with its latest self-consistent treatment of iron-bearing mantle phases, we investigate the stability and distribution of metallic iron in Earth's pyrolitic mantle across a broad range of oxidation states, represented by whole-rock Fe3+/$Σ$Fe ratio from 1% to 10%. We find that metallic iron is present through much of the lower mantle across this range and, under very reduced compositions of whole-rock Fe3+/$Σ$Fe = 1-3%, extends into the upper mantle. Where subducted water meets metal-saturated regions, hydrous melts may form and migrate upward, rehydrating the overlying mantle or pooling near the transition zone. Metal saturation can thus redistribute hydrogen internally, creating a sharp contrast between a wet shallow mantle and a dry deep mantle. This redox-driven redistribution can decrease mantle silicate water storage capacity by 64-96% today, to only 0.1-0.8 modern ocean masses, and may explain the viscosity contrast near the upper-lower mantle boundary. Although quantitative estimates of metal abundance and distribution depend on thermodynamic assumptions and remain uncertain above 50 GPa, our results reveal the role of redox reactions between disproportionated iron and subducted water in governing the speciation and redistribution of hydrogen in Earth's mantle.

physics.geo-ph

Towards understanding evolution of science through language model series

We introduce AnnualBERT, a series of language models designed specifically to capture the temporal evolution of scientific text. Deviating from the prevailing paradigms of subword tokenizations and "one model to rule them all", AnnualBERT adopts whole words as tokens and is composed of a base RoBERTa model pretrained from scratch on the full-text of 1.7 million arXiv papers published until 2008 and a collection of progressively trained models on arXiv papers at an annual basis. We demonstrate the effectiveness of AnnualBERT models by showing that they not only have comparable performances in standard tasks but also achieve state-of-the-art performances on domain-specific NLP tasks as well as link prediction tasks in the arXiv citation network. We then utilize probing tasks to quantify the models' behavior in terms of representation learning and forgetting as time progresses. Our approach enables the pretrained models to not only improve performances on scientific text processing tasks but also to provide insights into the development of scientific discourse over time. The series of the models is available at https://huggingface.co/jd445/AnnualBERTs.

cs.CL

Structure and Melting of Fe, MgO, SiO2, and MgSiO3 in Planets: Database, Inversion, and Phase Diagram

We present globally inverted pressure-temperature (P-T) phase diagrams up to 5,000 GPa for four fundamental planetary materials, Fe, MgO, SiO2, and MgSiO3, derived from logistic regression and supervised learning, together with an experimental phase equilibria database. These new P-T phase diagrams provide a solution to long-standing disputes about their melting curves. Their implications extend to the melting and freezing of rocky materials in the interior of giant planets and super-Earth exoplanets, contributing to the refinement of their internal structure models.

astro-ph.EP

Nonlinearity of the post-spinel transition and its expression in slabs and plumes worldwide

At the interface of Earth's upper and lower mantle, the post-spinel transition boundary controls the dynamics and morphologies of downwelling slabs and upwelling plumes, and its Clapeyron slope is hence one of the most important constraints on mantle convection. In this study, we reported a new in situ experimental dataset on phase stability in Mg$_{2}$SiO$_{4}$ at mantle transition zone pressures from laser-heated diamond anvil cell experiments, along with a compilation of corrected in situ experimental datasets from the literature. We presented a machine learning framework for high-pressure phase diagram determination and focused on its application to constrain the location and Clapeyron slope of the post-spinel transition: ringwoodite $\leftrightarrow$ bridgmanite + periclase. We found that the post-spinel boundary is nonlinear and its Clapeyron slope varies locally from $-2.3_{-1.4}^{+0.6}$ MPa/K at 1900 K, to $-1.0_{-1.7}^{+1.3}$ MPa/K at 1700 K, and to $0.0_{-2.0}^{+1.7}$ MPa/K at 1500 K. We applied the temperature-dependent post-spinel Clapeyron slope to estimate its lateral variation across the "660-km" seismic discontinuity in subducting slabs and hotspot-associated plumes worldwide, as well as the ambient mantle. We found that, in the present-day mantle, the average post spinel Clapeyron slope in the plumes is three times more negative than that in slabs, and we then discussed the effects of a nonlinear post-spinel transition on the dynamics of Earth's mantle.

physics.geo-ph

Clusterability test for categorical data

The objective of clusterability evaluation is to check whether a clustering structure exists within the data set. As a crucial yet often-overlooked issue in cluster analysis, it is essential to conduct such a test before applying any clustering algorithm. If a data set is unclusterable, any subsequent clustering analysis would not yield valid results. Despite its importance, the majority of existing studies focus on numerical data, leaving the clusterability evaluation issue for categorical data as an open problem. Here we present TestCat, a testing-based approach to assess the clusterability of categorical data in terms of an analytical $p$-value. The key idea underlying TestCat is that clusterable categorical data possess many strongly associated attribute pairs and hence the sum of chi-squared statistics of all attribute pairs is employed as the test statistic for $p$-value calculation. We apply our method to a set of benchmark categorical data sets, showing that TestCat outperforms those solutions based on existing clusterability evaluation methods for numeric data. To the best of our knowledge, our work provides the first way to effectively recognize the clusterability of categorical data in a statistically sound manner.

cs.LG

Uranus Study Report: KISS

Determining the internal structure of Uranus is a key objective for planetary science. Knowledge of Uranus's bulk composition and the distribution of elements is crucial to understanding its origin and evolutionary path. In addition, Uranus represents a poorly understood class of intermediate-mass planets (intermediate in size between the relatively well studied terrestrial and gas giant planets), which appear to be very common in the Galaxy. As a result, a better characterization of Uranus will also help us to better understand exoplanets in this mass and size regime. Recognizing the importance of Uranus, a Keck Institute for Space Studies (KISS) workshop was held in September 2023 to investigate how we can improve our knowledge of Uranus's internal structure in the context of a future Uranus mission that includes an orbiter and a probe. The scientific goals and objectives of the recently released Planetary Science and Astrobiology Decadal Survey were taken as our starting point. We reviewed our current knowledge of Uranus's interior and identified measurement and other mission requirements for a future Uranus spacecraft, providing more detail than was possible in the Decadal Survey's mission study and including new insights into the measurements to be made. We also identified important knowledge gaps to be closed with Earth-based efforts in the near term that will help guide the design of the mission and interpret the data returned.

astro-ph.IM

Conjunction Subspaces Test for Conformal and Selective Classification

In this paper, we present a new classifier, which integrates significance testing results over different random subspaces to yield consensus p-values for quantifying the uncertainty of classification decision. The null hypothesis is that the test sample has no association with the target class on a randomly chosen subspace, and hence the classification problem can be formulated as a problem of testing for the conjunction of hypotheses. The proposed classifier can be easily deployed for the purpose of conformal prediction and selective classification with reject and refine options by simply thresholding the consensus p-values. The theoretical analysis on the generalization error bound of the proposed classifier is provided and empirical studies on real data sets are conducted as well to demonstrate its effectiveness.

cs.LG

Hamming Encoder: Mining Discriminative k-mers for Discrete Sequence Classification

Sequence classification has numerous applications in various fields. Despite extensive studies in the last decades, many challenges still exist, particularly in pattern-based methods. Existing pattern-based methods measure the discriminative power of each feature individually during the mining process, leading to the result of missing some combinations of features with discriminative power. Furthermore, it is difficult to ensure the overall discriminative performance after converting sequences into feature vectors. To address these challenges, we propose a novel approach called Hamming Encoder, which utilizes a binarized 1D-convolutional neural network (1DCNN) architecture to mine discriminative k-mer sets. In particular, we adopt a Hamming distance-based similarity measure to ensure consistency in the feature mining and classification procedure. Our method involves training an interpretable CNN encoder for sequential data and performing a gradient-based search for discriminative k-mer combinations. Experiments show that the Hamming Encoder method proposed in this paper outperforms existing state-of-the-art methods in terms of classification accuracy.

cs.LG

Interpretable Sequence Clustering

Categorical sequence clustering plays a crucial role in various fields, but the lack of interpretability in cluster assignments poses significant challenges. Sequences inherently lack explicit features, and existing sequence clustering algorithms heavily rely on complex representations, making it difficult to explain their results. To address this issue, we propose a method called Interpretable Sequence Clustering Tree (ISCT), which combines sequential patterns with a concise and interpretable tree structure. ISCT leverages k-1 patterns to generate k leaf nodes, corresponding to k clusters, which provides an intuitive explanation on how each cluster is formed. More precisely, ISCT first projects sequences into random subspaces and then utilizes the k-means algorithm to obtain high-quality initial cluster assignments. Subsequently, it constructs a pattern-based decision tree using a boosting-based construction strategy in which sequences are re-projected and re-clustered at each node before mining the top-1 discriminative splitting pattern. Experimental results on 14 real-world data sets demonstrate that our proposed method provides an interpretable tree structure while delivering fast and accurate cluster assignments.

cs.LG

A Data Science Approach to Study the Water Storage Capacity in Rocky Planet Mantles: Earth, Mars, and Exoplanets

Nominally anhydrous minerals (NAMs) are the primary carriers of water in rocky planet mantles. Therefore, studying water solubilities of major NAMs in the mantle can help us estimate the water storage capacities of rocky planet mantles and indirectly constrain the actual water contents of their interiors. By using data science methods such as statistics and statistical learning algorithms, in this paper, current modeling studies on the mantle water storage capacities of Earth, Mars, and exoplanets have been introduced and summarized. Firstly, the thermodynamic model for mantle water storage capacity has been reviewed. Then, based on the two case studies on Earth and Mars, how to translate atomic-scale experimental data of water solubility and their measurement errors into planetary-scale models of mantle water storage capacity has been explored by using robust regression, Monte Carlo methods, and bootstrap aggregation algorithms. Thirdly, how the large sample data from the exoplanet observational campaigns can help us understand the statistical properties of the mantle water storage capacities of rocky exoplanets has been introduced. Finally, the application limitations of data science methods in mineral physics research have been discussed, and how to better combine statistics and statistical algorithms with mineral physics data research has been prospected.

physics.geo-ph

Water storage capacity of the Martian mantle through time

Water has been stored in the Martian mantle since its formation, primarily in nominally anhydrous minerals. The short-lived early hydrosphere and intermittently flowing water on the Martian surface may have been supplied and replenished by magmatic degassing of water from the mantle. Estimating the water storage capacity of the solid Martian mantle places important constraints on its water inventory and helps elucidate the sources, sinks, and temporal variations of water on Mars. In this study, we applied a bootstrap aggregation method to investigate the effects of iron on water storage capacities in olivine, wadsleyite, and ringwoodite, based on high-pressure experimental data compiled from the literature, and we provide a quantitative estimate of the upper bound of the bulk water storage capacity in the FeO-rich solid Martian mantle. Along a series of areotherms at different mantle potential temperatures ($T_{p}$), we estimated a water storage capacity equal to $9.0_{-2.2} ^{+2.8}$ km Global Equivalent Layer (GEL) for the present-day Martian mantle at $T_{p}$ = 1600 K and $4.9_{-1.5}^{+1.7}$ km GEL for the initial Martian mantle at $T_{p}$ = 1900 K. The water storage capacity of the Martian mantle increases with secular cooling through time, but due to the lack of an efficient water recycling mechanism on Mars, its actual mantle water content may be significantly lower than its water storage capacity today.

physics.geo-ph