SearcharxivSearch

arXiv subjects

Xinyi Fang

Publications and source records attributed to Xinyi Fang.

At least 19 recordsLinked to original sources

Uniform non-homogeneous bundles on quadrics

Let $X$ be an $n$-dimensional generalized Grassmannian not isomorphic to $\mathbb{P}^n$. We prove that $k(X)\le n-1$, where $k(X)$ denotes the maximal integer such that every uniform bundle on $X$ of rank at most $k(X)$ is homogeneous. In particular, for smooth quadrics $\mathbb{Q}^n$, we have $k(\mathbb{Q}^n)=n-1$ for odd $n$, and $n-2\le k(\mathbb{Q}^n)\le n-1$ for even $n$. We classify uniform rank $n$ bundles on $\mathbb{Q}^{n}$ for $n=3$, $5$. Furthermore, we characterize projective spaces among generalized Grassmannians in terms of uniform bundles.

math.AG

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LLM, a multi granularity framework that unifies semantic projection hashing, attention weighted MinHash, contrastive boundary learning, and selective LLM based adjudication. The method combines character, token, and document level signals through gated fusion, then applies a cascaded filtering pipeline for efficient candidate reduction. Semantic projection hashing learns compact binary codes in distilled LLM embedding space, while attention weighted Min- Hash suppresses boilerplate and emphasizes informative content. Adaptive decision boundaries and uncertainty estimation further improve robustness across template pollution, short text perturbation, containment, and viral fragments. Experiments show that SemHash LLM achieves strong duplicate detection quality with less than one percent neural verification cost.

cs.AI

Synergizing chemical and AI communities for advancing laboratories of the future

The development of automated experimental facilities and the digitization of experimental data have introduced numerous opportunities to radically advance chemical laboratories. As many laboratory tasks involve predicting and understanding previously unknown chemical relationships, machine learning (ML) approaches trained on experimental data can substantially accelerate the conventional design-build-test-learn process. This outlook article aims to help chemists understand and begin to adopt ML predictive models for a variety of laboratory tasks, including experimental design, synthesis optimization, and materials characterization. Furthermore, this article introduces how artificial intelligence (AI) agents based on large language models can help researchers acquire background knowledge in chemical or data science and accelerate various aspects of the discovery process. We present three case studies in distinct areas to illustrate how ML models and AI agents can be leveraged to reduce time-consuming experiments and manual data analysis. Finally, we highlight existing challenges that require continued synergistic effort from both experimental and computational communities to address.

stat.AP

Fast phase prediction of charged polymer blends by white-box machine learning surrogates

Compatibilized polymer blends are a complex, yet versatile and widespread category of material. When the components of a binary blend are immiscible, they are typically driven towards a macrophase-separated state, but with the introduction of electrostatic interactions, they can be either homogenized or shifted to microphase separation. However, both experimental and simulation approaches face significant challenges in efficiently exploring the vast design space of charge-compatibilized polymer blends, encompassing chemical interactions, architectural properties, and composition. In this work, we introduce a white-box machine learning approach integrated with polymer field theory to predict the phase behavior of these systems, which is significantly more accurate than conventional black-box machine learning approaches. The random phase approximation (RPA) calculation is used as a testbed to determine polymer phases. Instead of directly predicting the polymer phase output of RPA calculations from a large input space by a machine learning model, we build a parallel partial Gaussian process model to predict the most computationally intensive component of the RPA calculation that only involves polymer architecture parameters as inputs. This approach substantially reduces the computational cost of the RPA calculation across a vast input space with nearly 100% accuracy for out-of-sample prediction, enabling rapid screening of polymer blend charge-compatibilization designs. More broadly, the white-box machine learning strategy offers a promising approach for dramatic acceleration of polymer field-theoretic methods for mapping out polymer phase behavior.

cond-mat.soft

Neural Operators for Forward and Inverse Potential-Density Mappings in Classical Density Functional Theory

Neural operators are capable of capturing nonlinear mappings between infinite-dimensional functional spaces, offering a data-driven approach to modeling complex functional relationships in classical density functional theory (cDFT). In this work, we evaluate the performance of several neural operator architectures in learning the functional relationships between the one-body density profile $\rho(x)$, the one-body direct correlation function $c_1(x)$, and the external potential $V_{ext}(x)$ of inhomogeneous one-dimensional (1D) hard-rod fluids, using training data generated from analytical solutions of the underlying statistical-mechanical model. We compared their performance in terms of the Mean Squared Error (MSE) loss in establishing the functional relationships as well as in predicting the excess free energy across two test sets: (1) a group test set generated via random cross-validation (CV) to assess interpolation capability, and (2) a newly constructed dataset for leave-one-group CV to evaluate extrapolation performance. Our results show that FNO achieves the most accurate predictions of the excess free energy, with the squared ReLU activation function outperforming other activation choices. Among the DeepONet variants, the Residual Multiscale Convolutional Neural Network (RMSCNN) combined with a trainable Gaussian derivative kernel (GK-RMSCNN-DeepONet) demonstrates the best performance. Additionally, we applied the trained models to solve for the density profiles at various external potentials and compared the results with those obtained from the direct mapping $V_{ext} \mapsto \rho$ with neural operators, as well as with Gaussian Process Regression (GPR) combined with Active Learning by Error Control (ALEC), which has shown strong performance in previous studies.

physics.chem-ph

Homogeneous Ulrich bundles on isotropic flag varieties

In this paper, we consider the existence problem of Ulrich bundles on a rational homogeneous space $G/P$ of type $B$, $C$ or $D$. We show that if the Picard number of $G/P$ is greater than or equal to $2$, then there are no irreducible homogeneous Ulrich bundles on $G/P$ with respect to the minimal ample class.

math.AG

Uniform bundles on quadrics

We show that there exist only constant morphisms from $\mathbb{Q}^{2n+1}(n\geq 1)$ to $\mathbb{G}(l,2n+1)$ if $l$ is even $(0<l<2n)$ and $(l,2n+1)$ is not $ (2,5)$. As an application, we prove on $\mathbb{Q}^{2m+1}$ and $\mathbb{Q}^{2m+2}(m\geq 3)$, any uniform bundle of rank at most $2m$ splits, which improves the upper bound of splitting for uniform bundles obtained by Kachi and Sato. We classify all unsplit uniform bundles of minimal rank on $B_n/P_k$ $(k=\frac{2n}{3},k\ge6)$ and $D_n/P_k$ $(k=\frac{2n-2}{3},k\ge 6)$. We partially answer a conjecture of Ellia, which predicts that some uniform bundles of special splitting types on $\mathbb{P}^n$ necessarily split and we find some restrictions on the splitting types of unsplit uniform bundles of minimal rank.

math.AG

The inverse Kalman filter

We introduce the inverse Kalman filter, which enables exact matrix-vector multiplication between a covariance matrix from a dynamic linear model and any real-valued vector with linear computational cost. We integrate the inverse Kalman filter with the conjugate gradient algorithm, which substantially accelerates the computation of matrix inversion for a general form of covariance matrix, where other approximation approaches may not be directly applicable. We demonstrate the scalability and efficiency of the proposed approach through applications in nonparametric estimation of particle interaction functions, using both simulations and cell trajectories from microscopy data.

stat.ME

Analysis of the Two-Step Heterogeneous Transfer Learning for Laryngeal Blood Vessel Classification: Issue and Improvement

Accurate classification of laryngeal vascular as benign or malignant is crucial for early detection of laryngeal cancer. However, organizations with limited access to laryngeal vascular images face challenges due to the lack of large and homogeneous public datasets for effective learning. Distinguished from the most familiar works, which directly transfer the ImageNet pre-trained models to the target domain for fine-tuning, this work pioneers exploring two-step heterogeneous transfer learning (THTL) for laryngeal lesion classification with nine deep-learning models, utilizing the diabetic retinopathy color fundus images, semantically non-identical yet vascular images, as the intermediate domain. Attention visualization technique, Layer Class Activate Map (LayerCAM), reveals a novel finding that yet the intermediate and the target domain both reflect vascular structure to a certain extent, the prevalent radial vascular pattern in the intermediate domain prevents learning the features of twisted and tangled vessels that distinguish the malignant class in the target domain, summarizes a vital rule for laryngeal lesion classification using THTL. To address this, we introduce an enhanced fine-tuning strategy in THTL called Step-Wise Fine-Tuning (SWFT) and apply it to the ResNet models. SWFT progressively refines model performance by accumulating fine-tuning layers from back to front, guided by the visualization results of LayerCAM. Comparison with the original THTL approach shows significant improvements. For ResNet18, the accuracy and malignant recall increases by 26.1% and 79.8%, respectively, while for ResNet50, these indicators improve by 20.4% and 62.2%, respectively.

cs.CV

Category-wise Fine-Tuning: Resisting Incorrect Pseudo-Labels in Multi-Label Image Classification with Partial Labels

Large-scale image datasets are often partially labeled, where only a few categories' labels are known for each image. Assigning pseudo-labels to unknown labels to gain additional training signals has become prevalent for training deep classification models. However, some pseudo-labels are inevitably incorrect, leading to a notable decline in the model classification performance. In this paper, we propose a novel method called Category-wise Fine-Tuning (CFT), aiming to reduce model inaccuracies caused by the wrong pseudo-labels. In particular, CFT employs known labels without pseudo-labels to fine-tune the logistic regressions of trained models individually to calibrate each category's model predictions. Genetic Algorithm, seldom used for training deep models, is also utilized in CFT to maximize the classification performance directly. CFT is applied to well-trained models, unlike most existing methods that train models from scratch. Hence, CFT is general and compatible with models trained with different methods and schemes, as demonstrated through extensive experiments. CFT requires only a few seconds for each category for calibration with consumer-grade GPUs. We achieve state-of-the-art results on three benchmarking datasets, including the CheXpert chest X-ray competition dataset (ensemble mAUC 93.33%, single model 91.82%), partially labeled MS-COCO (average mAP 83.69%), and Open Image V3 (mAP 85.31%), outperforming the previous bests by 0.28%, 2.21%, 2.50%, and 0.91%, respectively. The single model on CheXpert has been officially evaluated by the competition server, endorsing the correctness of the result. The outstanding results and generalizability indicate that CFT could be substantial and prevalent for classification model development. Code is available at: https://github.com/maxium0526/category-wise-fine-tuning.

cs.CV

Reliable emulation of complex functionals by active learning with error control

A statistical emulator can be used as a surrogate of complex physics-based calculations to drastically reduce the computational cost. Its successful implementation hinges on an accurate representation of the nonlinear response surface with a high-dimensional input space. Conventional "space-filling" designs, including random sampling and Latin hypercube sampling, become inefficient as the dimensionality of the input variables increases, and the predictive accuracy of the emulator can degrade substantially for a test input distant from the training input set. To address this fundamental challenge, we develop a reliable emulator for predicting complex functionals by active learning with error control (ALEC). The algorithm is applicable to infinite-dimensional mapping with high-fidelity predictions and a controlled predictive error. The computational efficiency has been demonstrated by emulating the classical density functional theory (cDFT) calculations, a statistical-mechanical method widely used in modeling the equilibrium properties of complex molecular systems. We show that ALEC is much more accurate than conventional emulators based on the Gaussian processes with "space-filling" designs and alternative active learning methods. Besides, it is computationally more efficient than direct cDFT calculations. ALEC can be a reliable building block for emulating expensive functionals owing to its minimal computational cost, controllable predictive error, and fully automatic features.

physics.chem-ph

Uniform bundles on generalised Grassmannians

Let $E$ be a uniform bundle on an arbitrary generalised Grassmannian $X$ defined over $\mathbb{C}$. We show that if the rank of $E$ is at most $e.d.(\mathrm{VMRT})$, then $E$ necessarily splits. For some generalised Grassmannians, we prove that the upper bounds $e.d.(\mathrm{VMRT})$ are optimal and classify all unsplit uniform bundles of minimal ranks. Under some special assumptions, we show that morphisms to some generalised flag varieties must be constant, which partially answered a conjecture of Kumar.

math.AG

Homogeneous ACM bundles on rational homogeneous spaces

In this paper, we characterize homogeneous arithmetically Cohen-Macaulay (ACM) bundles and Ulrich bundles on rational homogeneous spaces. %with respect to general polarizations. From this result, we see that there are only finitely many irreducible homogeneous ACM bundles (up to twist) and Ulrich bundles on these varieties. Moreover, we give numerical criteria for some special irreducible homogeneous bundles to be ACM bundles.

math.AG

Homogeneous ACM bundles on exceptional Grassmannians

In this paper, we characterize homogeneous arithmetically Cohen-Macaulay (ACM) bundles over exceptional Grassmannians in terms of their associated data. We show that there are only finitely many irreducible homogeneous ACM bundles by twisting line bundles over exceptional Grassmannians. As a consequence, we prove that some exceptional Grassmannians are of wild representation type.

math.AG

Morphisms from $\mathbb{P}^m$ to flag varieties

In this paper, we consider the morphisms from projective spaces to flag varieties. We show that the morphisms can only be constant under some special conditions. As a consequence, we prove that the splitting types of unsplit uniform $r$-bundles on $\mathbb{P}^m$ can not be $(a_1,\dots,a_1,a_2,\dots,a_{r-k+1})$ for $1\le k\le m-2$.

math.AG

Data-driven model construction for anisotropic dynamics of active matter

The dynamics of cellular pattern formation is crucial for understanding embryonic development and tissue morphogenesis. Recent studies have shown that human dermal fibroblasts cultured on liquid crystal elastomers can exhibit an increase in orientational alignment over time, accompanied by cell proliferation, under the influence of the weak guidance of a molecularly aligned substrate. However, a comprehensive understanding of how this order arises remains largely unknown. This knowledge gap may be attributed, in part, to a scarcity of mechanistic models that can capture the temporal progression of the complex nonequilibrium dynamics during the cellular alignment process. The orientational alignment occurs primarily when cells reach a high density near confluence. Therefore, for accurate modeling, it is crucial to take into account both the cell-cell interaction term and the influence from the substrate, acting as a one-body external potential term. To fill in this gap, we develop a hybrid procedure that utilizes statistical learning approaches to extend the state-of-the-art physics models for quantifying both effects. We develop a more efficient way to perform feature selection that avoids testing all feature combinations through simulation. The maximum likelihood estimator of the model was derived and implemented in computationally scalable algorithms for model calibration and simulation. By including these features, such as the non-Gaussian, anisotropic fluctuations, and limiting alignment interaction only to neighboring cells with the same velocity direction, this model quantitatively reproduce the key system-level parameters--the temporal progression of the velocity orientational order parameters and the variability of velocity vectors, whereas models missing any of the features fail to capture these temporally dependent parameters.

physics.bio-ph

Molecular-scale substrate anisotropy and crowding drive long-range nematic order of cell monolayers

The ability of cells to reorganize in response to external stimuli is important in areas ranging from morphogenesis to tissue engineering. Elongated cells can co-align due to steric effects, forming states with local order. We show that molecular-scale substrate anisotropy can direct cell organization, resulting in the emergence of nematic order on tissue scales. To quantitatively examine the disorder-order transition, we developed a high-throughput imaging platform to analyze velocity and orientational correlations for several thousand cells over days. The establishment of global, seemingly long-ranged order is facilitated by enhanced cell division along the substrate's nematic axis, and associated extensile stresses that restructure the cells' actomyosin networks. Our work, which connects to a class of systems known as active dry nematics, provides a new understanding of the dynamics of cellular remodeling and organization in weakly interacting cell collectives. This enables data-driven discovery of cell-cell interactions and points to strategies for tissue engineering.

physics.bio-ph

Homogeneous ACM bundles on isotropic Grassmannians

In this paper, we characterize homogeneous arithmetically Cohen-Macaulay (ACM) bundles over isotropic Grassmannians of types $B$, $C$ and $D$ in term of step matrices. We show that there are only finitely many irreducible homogeneous ACM bundles by twisting line bundles over these isotropic Grassmannians. So we classify all homogeneous ACM bundles over isotropic Grassmannians combining the results on usual Grassmannians by Costa and Mir{ó}-Roig. Moreover, if the irreducible initialized homogeneous ACM bundles correspond to some special highest weights, then they can be characterized by succinct forms.

math.AG