Searcharxiv⌕ Search

arXiv subjects

Yucheng Qiao

Publications and source records attributed to Yucheng Qiao.

6 recordsLinked to original sources

Complementary Roles of Activation and Parametric Memory in Few-Shot Learning

At test time, large language models (LLMs) can encode historical information in activation memory (i.e., KV caches) and parametric memory (i.e., updated parameters). While activation memory is generally considered effective for factual recall and parametric memory for learning new tasks, their interplay remains unclear. In this work, we systematically investigate the role of memory in few-shot learning through controlled experiments. We find that activation memory is superior for recalling facts, whereas parametric memory does not consistently outperform activation memory in task learning. Moreover, our experiments show that the composite task, Conditional Arithmetic, requires the synergy of both memory types. Through neuron-level analysis, we find that the model activates distinct sets of neurons when accessing the same historical information through activation versus parametric memory. When both memory types are combined, the model recruits neurons from both sets, which is crucial for solving Conditional Arithmetic. These findings suggest that neither memory mechanism alone is sufficient for this composite task, highlighting the importance of their collaboration.

cs.CL↗

M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment.

cs.SD↗

SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling

Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive. MoE upcycling mitigates this cost by converting pretrained dense models into sparse MoE models. However, existing upcycling methods typically rely on large-scale continued training and often perform poorly under data-constrained supervised adaptation, due to either homogeneous experts or overly disruptive perturbations to pretrained parameters. In this setting, effective upcycling must leverage pretrained weight structure while introducing sufficient diversity among routed experts. To this end, we propose SVD-Partitioned Residual Initialization (SPRI), which distributes SVD-partitioned residuals derived from pretrained feed-forward network (FFN) weights across routed experts, introducing controlled expert diversity grounded in pretrained spectral structure. We further introduce a two-stage training strategy to improve adaptation stability. We evaluate SPRI on multilingual speech-to-text translation, where limited supervised data challenges MoE upcycling and multiple target languages provide natural routing heterogeneity. On CoVoST2 across 15 En-to-XX directions, SPRI improves average BLEU and COMET over fully fine-tuned dense models by 2.58 and 3.32 points, respectively, and outperforms the prior best MoE upcycling baseline by 3.39 BLEU and 4.34 COMET points.

cs.LG↗

A Remote Quantum Error-correcting Code Preparation Protocol on Cluster State

The blind quantum computation (BQC) protocol allows for privacy-preserving remote quantum computations. In this paper, we introduce a remote quantum error correction code preparation protocol for BQC using a cluster state and analyze its blindness in the measurement-based quantum computation model. Our protocol requires fewer quantum resources than previous methods, as it only needs weak coherent pulses, eliminating the need for quantum memory and limited quantum computing. The results of our theoretical analysis and simulations show that our protocol requires fewer quantum resources compared to non-coding methods with the same qubit error rate.

quant-ph↗

Secure bound analysis of quantum key distribution with non-uniform random seed of privacy amplification

Precise quantum key distribution (QKD) secure bound analysis is essential for practical QKD systems. The effect of uniformity of random number seed for privacy amplification is not considered in existing secure bound analysis. In this paper, we propose and prove the quantum leftover hash lemma with non-uniform random number seeds based on the min-entropy, and we give a precise QKD secure bound analysis with non-uniform random number seeds on this basis. We take the two-decoy BB84 protocol as an example to simulate the effect of random number seed uniformity on the secure bound of a QKD system. The experimental results indicate that when the average min-entropy of the random number generator is below 0.95, the secure bound of a QKD system will be seriously affected.

quant-ph↗

Light Source Monitoring in Quantum Key Distribution with Single Photon Detector at Room Temperature

Photon number resolving monitoring is a practical light source monitoring scheme in QKD systems, which reduces the impacts from untrusted sources effectively. This scheme requires a single photon detector, normally working at low temperature to suppress its dark count rate. In this paper, we use a room-temperature detector and show that the dark count rate is irrelevant to the monitoring performance in our scheme, which can sufficiently relax requirements on the detector's working conditions as well as integration complexity, and this would be highly demanded for practical systems. Furthermore, influences of parameter drifts at room temperature are analyzed, and the monitoring scheme is testified in a real QKD system.

quant-ph↗