SearcharxivSearch

arXiv subjects

Antong Zhang

Publications and source records attributed to Antong Zhang.

7 recordsLinked to original sources

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method. Data-DPO observes the local training feedback of the target model on different samples through one-step probing, transforms activation differences among samples into pairwise data preferences, and trains a lightweight reward model to learn target-model-aware data preferences. In the final selection stage, Data-DPO further combines target model preference, external quality scores, and marginal diversity to construct a more stable and effective training subset. Experimental results on Vision-Flan and LLaVA-CoT show that Data-DPO consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance.

cs.LG

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS. MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder. Experiments on Vision Flan and LLaVA-CoT show that MASS consistently outperforms strong data selection baselines across multiple budgets, and in several settings matches or surpasses full data training with only a small subset of data.

cs.LG

A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) has emerged as a paradigm for enhancing large language models (LLMs) with external knowledge, yet existing graph-based methods face a fundamental limitation: entity-centric and chunk-centric approaches operate on representations anchored to original text without true knowledge fusion. While entity-centric methods connect logically related content and chunk-centric methods preserve context, both retrieve information separately through similarity search, missing emergent understanding from their synthesis. In this paper, we propose HyGRAG, a hierarchical graph RAG framework that transcends source documents by addressing three core challenges: constructing summaries that genuinely integrate contextual and relational information, leveraging these synthesized representations to access emergent knowledge during retrieval, and efficiently updating hierarchical structures for dynamic corpora. Specifically, we design hierarchical index structures over hybrid graphs with both chunk and entity nodes, then iteratively cluster them and generate LLM-based summaries. Then, we design context and relation-aware retrieval that searches across all abstraction levels while expanding through community membership. Moreover, we enable dynamic knowledge update through attachment-based algorithms with only local re-summarization. Experimental results show that HyGRAG improves the average accuracy of multi-hop reasoning tasks by 9.7%, while maintaining reasonable efficiency.

cs.AI

BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation

We introduce BEACON--Best-Effort Adaptation for Cross-Domain Co-Training--a theory-driven framework for training generative robot policies with abundant source demonstrations and limited target demonstrations. BEACON casts cross-domain co-training as a discrepancy-aware importance-reweighting problem, jointly learning a diffusion-based visuomotor policy and per-sample source weights that minimize an objective informed by target-domain generalization guarantees. To make best-effort adaptation practical for high-dimensional sequence policies, we develop scalable instance-level discrepancy estimators, stochastic alternating updates for policy and weights, and a multi-source extension that balances heterogeneous source domains. Across sim-to-sim, sim-to-real, and multi-source manipulation settings, BEACON improves robustness and data efficiency over target-only, fixed-ratio co-training, and feature-alignment baselines. Importantly, even without an explicit alignment objective, BEACON achieves feature alignment as an implicit result of discrepancy-aware cross-domain co-training.

cs.RO

SpecMol: A Spectroscopy-Grounded Foundation Model for Multi-Task Molecular Learning

Large language models have emerged as transformative tools in molecular science, demonstrating remarkable potential in molecular property prediction and de novo molecular design. However, their application to spectroscopy remains notably limited, despite its foundational role in experimental molecular characterization and structural validation. Progress in spectroscopy-grounded reasoning has been hindered by the lack of standardized spectral representations and comprehensive evaluation protocols, making cross-study comparisons difficult. To bridge this gap, we present a unified framework for spectroscopy-grounded molecular modeling and evaluation. At its core, the SpecMol foundation model integrates spectral interpretation, molecular representation learning, and three-dimensional structure generation within a single interface. Complementing this, we establish SpecMol-Bench as a systematic evaluation protocol encompassing cross-modal tasks: spectra-to-structure elucidation, structure-to-spectra simulation, and SMILES-to-3D conformation generation. Under this unified framework, SpecMol achieves accurate spectra-driven structure elucidation and reproduces experimental nuclear magnetic resonance characteristics with high fidelity. The model also generates chemically valid three-dimensional conformations directly from SMILES strings and consistently outperforms existing general-purpose molecular language models across standardized evaluation metrics. Code is available at https://github.com/Eurekashen/SpecMol

cs.LG

2060: Civilization, Energy, and Progression of Mankind on the Kardashev Scale

Energy has been propelling the development of human civilization for millennia, and technologies acquiring energy beyond human and animal power have been continuously advanced and transformed. In 1964, the Kardashev Scale was proposed to quantify the relationship between energy consumption and the development of civilizations. Human civilization presently stands at Type 0.7276 on this scale. Projecting the future energy consumption, estimating the change of its constituting structure, and evaluating the influence of possible technological revolutions are critical in the context of civilization development. In this study, we use two machine learning models, random forest (RF) and autoregressive integrated moving average (ARIMA), to simulate and predict energy consumption on a global scale. We further project the position of human civilization on the Kardashev Scale in 2060. The result shows that the global energy consumption is expected to reach 928-940 EJ in 2060, with a total growth of over 50% in the coming 40 years, and our civilization is expected to achieve Type 0.7474 on the Kardashev Scale, still far away from a Type 1 civilization. Additionally, we discuss the potential energy segmentation change before 2060 and present the influence of the advent of nuclear fusion in this context.

cs.CY

Avoiding the Great Filter: Predicting the Timeline for Humanity to Reach Kardashev Type I Civilization

The level of technological development of any civilization can be gaged in large part by the amount of energy they produce for their use, but also encompasses that civilization's stewardship of their home world. Following the Kardashev definition, a Type I civilization is able to store and use all the energy available on its planet. In this study, we develop a model based on Carl Sagan's K formula and use this model to analyze the consumption and energy supply of the three most important energy sources: fossil fuels (e.g., coal, oil, natural gas, crude, NGL and feedstocks), nuclear energy and renewable energy. We also consider environmental limitations suggested by United Nations Framework Convention on Climate Change, the International Energy Agency, and those specific to our calculations to predict when humanity will reach the level of a Kardashev scale Type I civilization. Our findings suggest that the best estimate for this day will not come until year 2371.

physics.soc-ph