SearcharxivSearch

arXiv subjects

Weijun Yao

Publications and source records attributed to Weijun Yao.

5 recordsLinked to original sources

Skill Weaving: Efficient LLM Improvement via Modular Skillpacks

Large language models increasingly require specialization across diverse domains, yet existing approaches struggle to balance multi-domain capacities with strict memory and inference constraints. In this work, we introduce SkillWeave, a modular improvement framework that enables LLMs to specialize under fixed memory budgets. SkillWeave partitions full capabilities of a general-purpose model into skillpacks -- lightweight, domain-specific delta modules -- that reorganize and refine the model's internal knowledge. For efficient deployment, SkillWeave integrates SkillZip to compress skillpacks into compact and inference-ready format, enabling strong multi-domain performance with low-latency execution. On multi-task and agentic benchmarks, a 9B SkillWeave model outperforms several baselines and even surpasses a 32B monolithic LLM, while achieving up to 4x speedup.

cs.AI

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated user histories, neglecting the essential role of inter-user differences. We propose C-BPO, a framework that personalizes LLMs via preference-calibrated binary signals. By treating target user data as positive feedback and other users' data as an auxiliary set of implicit negative signals, C-BPO captures distinct inter-user differences. To mitigate the preference overlap issue, where shared task knowledge is erroneously penalized, we derive an objective grounded in Positive-Unlabeled (PU) learning theory. This approach purifies negative signals by subtracting ``positive bias'', ensuring alignment with unique idiosyncrasies without compromising general helpfulness. Empirical experiments across various personalization tasks and backbone LLMs show C-BPO consistently outperforms baselines, demonstrating the efficacy of preference-calibrated binary signals in modeling inter-user differences.

cs.CL

Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination

Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet their reliability is persistently undermined by hallucinations-generating text that contradicts visual input. Recent studies often attribute these errors to inadequate visual attention. In this work, we analyze the attention mechanisms via the logit lens, uncovering a distinct anomaly we term Vocabulary Hijacking. We discover that specific visual tokens, defined as Inert Tokens, disproportionately attract attention. Crucially, when their intermediate hidden states are projected into the vocabulary space, they consistently decode to a fixed set of unrelated words (termed Hijacking Anchors) across layers, revealing a rigid semantic collapse. Leveraging this semantic rigidity, we propose Hijacking Anchor-Based Identification (HABI), a robust strategy to accurately localize these Inert Tokens. To quantify the impact of this phenomenon, we introduce the Non-Hijacked Visual Attention Ratio (NHAR), a novel metric designed to identify attention heads that remain resilient to hijacking and are critical for factual accuracy. Building on these insights, we propose Hijacking-Aware Visual Attention Enhancement (HAVAE), a training-free intervention that selectively strengthens the focus of these identified heads on salient visual content. Extensive experiments across multiple benchmarks demonstrate that HAVAE significantly mitigates hallucinations with no additional computational overhead, while preserving the model's general capabilities. Our code is publicly available at https://github.com/lab-klc/HAVAE.

cs.MM

Electric Charging Effects on Insulating Surfaces in Cryogenic Liquids

This paper presents a new technique to study the adsorption and desorption of ions and electrons on insulating surfaces in the presence of strong electric fields in cryoliquids. The experimental design consists of a compact cryostat coupled with a sensitive electro-optical Kerr device to monitor the stability of the electric fields. The behavior of nitrogen and helium ions on a poly(methyl methacrylate) (PMMA) surface was compared to a PMMA surface coated with a mixture of deuterated polystyrene and deuterated polybutadiene. Ion accumulation and removal on these surfaces were unambiguously observed. Within the precision of the data, both surfaces behave similarly for the physisorbed ions. The setup was also used to measure the (quasi-)static dielectric constant of PMMA at T = 70 K. The impact of the ion adsorption on the search for a neutron permanent electric dipole moment in a cryogenic environment, like the nEDM@SNS experiment, is discussed.

physics.app-ph

Quantum computing with single electron bubbles in helium

An electron inside liquid helium forms a bubble of 17 Åin radius. In an external magnetic field, the two-level system of a spin 1/2 electron is ideal for the implementation of a qubit for quantum computing. The electron spin is well isolated from other thermal reservoirs so that the qubit should have very long coherence time. By confining a chain of single electron bubbles in a linear RF quadrupole trap, a multi-bit quantum register can be implemented. All spins in the register can be initialized to the ground state either by establishing thermal equilibrium at a temperature around 0.1 K and at a magnetic field of 1 T or by sorting the bubbles to be loaded into the trap with magnetic separation. Schemes are designed to address individual spins and to do two-qubit CNOT operations between the neighboring spins. The final readout can be carried out through a measurement similar to the Stern-Gerlach experiment.

cond-mat.other