SearcharxivSearch

arXiv subjects

Tiankai Chen

Publications and source records attributed to Tiankai Chen.

3 recordsLinked to original sources

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.

cs.LG

Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation

Out-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detection methods have shown progress for 2D image data, extending these to 3D environments involves unique obstacles. This paper introduces a training-free framework that leverages Vision-Language Models (VLMs) for effective OOD detection in 3D point clouds. By constructing a graph based on class prototypes and testing data, we exploit the data manifold structure to enhancing the effectiveness of VLMs for 3D OOD detection. We propose a novel Graph Score Propagation (GSP) method that incorporates prompt clustering and self-training negative prompting to improve OOD scoring with VLM. Our method is also adaptable to few-shot scenarios, providing options for practical applications. We demonstrate that GSP consistently outperforms state-of-the-art methods across synthetic and real-world datasets 3D point cloud OOD detection.

cs.CV

Deep Learning Accelerated Gold Nanocluster Synthesis

The understanding of inorganic reactions, especially those far from the equilibrium state, is relatively limited due to their inherent complexity. Poor understandings on the underlying synthetic chemistry have constrained the design of efficient synthesis routes towards desired final products, especially those inorganic materials at atomic precision. In this work, using the synthesis of atomically precise gold nanoclusters as a demonstration platform, we have successfully developed a deep learning framework for guiding material synthesis and accelerating the whole workflow. With only 54 examples, the proposed Graph Convolutional Neural Networks (GCNN) plus Siamese Neural Networks (SNN) classification model with the basic descriptors have been trained. The capability of predicting the target synthesis results has been demonstrated with a successful experimental validation. In addition, understandings in the synthesis process can be acquired from a decision tree trained by a large amount of generated data from the well-trained classification model. This study not only provides a data-driven method accelerating gold nanocluster synthesis, but also sheds light on understanding complex inorganic materials synthesis with low data amount.

physics.comp-ph