SearcharxivSearch

arXiv subjects

Yuzhe Zhou

Publications and source records attributed to Yuzhe Zhou.

7 recordsLinked to original sources

Rethinking Representativeness and Diversity in Dynamic Data Selection

Dynamic data selection accelerates training by sampling a changing subset of the dataset while preserving accuracy. We rethink two core notions underlying sample evaluation: representativeness and diversity. Instead of local geometric centrality, we define representativeness as coverage of dataset-level common or high-frequency feature factors. Instead of within-subset dispersion, we define diversity at the process level, requiring the selection trajectory to gradually include complementary rare factors over training. Based on this view, we propose a dynamic selection framework with three components. First, we score representativeness in a plug-in feature space to prioritize samples covering frequent factors. We instantiate this with a sparse autoencoder trained on the target dataset, using sparse unit activations to summarize both individual samples and dataset-wide factor statistics. Second, we realize process-level diversity by combining rare-factor sampling with a Usage-Frequency Penalty that promotes sample rotation, provably discourages monopoly, and reduces gradient bias. Third, we couple the two-dimensional scoring with a smooth scheduler that transitions selection from core-pattern consolidation to rare-factor exploration, without extra gradients, influence estimates, or second-order computations on the training model. Extensive experiments on five benchmarks across vision and text tasks demonstrate improved accuracy-efficiency trade-offs across models. Our method matches or exceeds full-data accuracy with over 2x training acceleration. Code will be released.

cs.AI

MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis

Continual instruction tuning(CIT) during the post-training phase is crucial for adapting multimodal large language models (MLLMs) to evolving real-world demands. However, the progress is hampered by the lack of benchmarks with rigorous, protocol-consistent evaluation. To bridge this gap, we introduce MLLM-CTBench, a comprehensive benchmark for CIT of MLLMs, covering seven challenging tasks across six diverse domains. MLLM-CTBench makes three key contributions. First, we establish a multidimensional evaluation framework that jointly assesses final-answer accuracy and process-level reasoning quality, where Chain-of-Thought (CoT) traces serve as an observable signal to diagnose catastrophic forgetting beyond answer-only evaluation. Second, we conduct a large-scale evaluation of continual learning methods by systematically assessing eight representative algorithms from four major families under a unified protocol across task orders, providing actionable insights for algorithm design. Third, we expand the scope from Supervised Fine-Tuning (SFT) to Reinforcement Fine-Tuning (RFT) in CIT. By investigating GRPO, an on-policy RL algorithm that stabilizes updates through explicit KL-divergence control to a prior policy, we aim to analyze how this mechanism affects cross-task knowledge retention. Our experiments yield several findings:(1) Process-level reasoning quality is often more resilient to catastrophic forgetting than final-answer accuracy, and forgetting is primarily driven by degradation in domain knowledge. (2) Model capability is critical factor influencing continual learning outcomes, with stronger baseline models exhibiting greater resistance to catastrophic forgetting. (3) On-policy RFT (GRPO), with its inherent KL control, achieves more stable cross-task retention than SFT. While removing KL control can amplify forgetting despite potential gains on new ones.

cs.CL

CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding Residues

Accurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of the protein or the global 3D geometric information. Consequently, these approaches may result in incomplete or inaccurate protein analysis. To address the above issue, in this paper, we present CrossBind, a novel collaborative cross-modal approach for identifying binding residues by exploiting both protein geometric structure and its sequence prior knowledge extracted from a large-scale protein language model. Specifically, our multi-modal approach leverages a contrastive learning technique and atom-wise attention to capture the positional relationships between atoms and residues, thereby incorporating fine-grained local geometric knowledge, for better binding residue prediction. Extensive experimental results demonstrate that our approach outperforms the next best state-of-the-art methods, GraphSite and GraphBind, on DNA and RNA datasets by 10.8/17.3% in terms of the harmonic mean of precision and recall (F1-Score) and 11.9/24.8% in Matthews correlation coefficient (MCC), respectively. We release the code at https://github.com/BEAM-Labs/CrossBind.

q-bio.BM

Future Communication Model for High-speed Railway Based on Unmanned Aerial Vehicles

High-speed railway is playing an important role in mass transportation, due to its lower energy consumption, less environmental pollution, larger capacity and higher safety features. The development of high-speed railway makes people's life more and more convenient. Meanwhile, providing high quality of service broadband communications for fast-moving users still remains unsolved, despite the fact that new solutions of incremental improvements are keeping up with this unprecedented communication requirement growth. This article proposes a communication system infrastructure based on airborne relay for high-speed trains in the further Cyber-Physical Systems. Comparisons and feasibility analysis are provided as well as discussions of key wireless technologies and obstacles in this system.

cs.NI

Quality of Service Improvement for High-Speed Railway Communications

With the fast development of high-speed railways, a call for fulfilling the notion of communication at "anytime, anywhere" for high-speed train passengers in the Train Operating Control System is on the way. In order to make a realization of that, new railway wireless communication networks are needed. The most promising one is the Long Term Evolution for Railway which will provide broadband access, fast handover, and reliable communication for high mobility users. However, with the increase of speed, the system is subjected to high bit error rate, Doppler frequency shift and handover failure just like other system does. This paper is trying to solve these problems by employing MIMO technique. Specifically, the goal is to provide higher data rate, higher reliability, less delay, and other relative quality of services for passengers. MIMO performance analysis, resource allocation, and access control for handover and various services in a two-hop model are proposed in this paper. Analytical results and simulation results show that the proposed model and schemes perform well in improving the system performances.

cs.IT

Provide High-QoS of the High-Speed Railway Mobile Communications in Cyber-Physical Systems

Technical advances in networks, embedded computing, and wireless communications are leading to the next generation of complex intelligent systems called Cyber-Physical Systems (CPS). CPS promises to transform the way we interact with the physical world. Efficient and reliable operation of the CTCS-3 (Chinese Train Control System Level 3) is great protection of the national economy and public safety. The CTCS-3 is based on GSM-R (GSM for Railway) to achieve a continuous and two-way transmission of information between ground and the train. To ensure the growing needs of safety, fastness and service diversity of China's railway, the pursuit of high-QoS has been the key of the relative study. This paper examines the main characteristics of GSM-R and the requirements of all-layers' QoS indicators for GSM-R. Several main technologies of improving QoS indicators of all-layers are summarized. As a solution, a comprehensive scheme is proposed to improve the delay and packet loss indicators. An example is also presented that illustrates the real-time features of the proposed solution. Based on the CPS characteristics are highly correlated with the QoS indicators, conclusions are made that GSM-R can provide a reliable and real-time way for message.

cs.NI

Evaluation of High-speed Train Communication Handover Models Based on DEA

Broadband communications for high speed train is becoming a main trend in high mobility communications. The main bottleneck of this communication network is handover, since the handover occurs so frequently and the delays are so long that broadband real-time communication cannot apply. Various handover models have been developed and studied recently. However, no comprehensive evaluation method for these models is employed. To this end, we borrow Data Envelopment Analysis (DEA) method to evaluate six typical handover system models. Handover models that to be evaluated are introduced. A brief presentation of DEA and its characters is provided. A specific procedure of the evaluation is proposed. Then the results of the evaluation are obtained by running the DEA. Finally, we give our comments and conclusions to all the handover models. We hope our work will supply a gap in the system evaluation area.

cs.NI