SearcharxivSearch

arXiv subjects

Chen Qiu

Publications and source records attributed to Chen Qiu.

At least 19 recordsLinked to original sources

SPD: Single Pass Decoding for Generative Reranking

Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce SPD (Single Forward Pass), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. SPD reads an $N \times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with auto-regressive LLM ranking distillation reaches 28 ms end-to-end inference, a speed-up of 64x while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.

cs.LG

MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. On multi-turn instruction following, static guidance pushes a 4B student's constraint satisfaction 2.5 points below its no-guidance baseline; a training-free refresh every 16 tokens changes only the memory content and restores a 2.0-point gain over that baseline. We propose MentorPulse to keep guidance fresh at practical cost: it compresses mentor states into a capped slot memory, incrementally processes newly generated tokens, and updates the memory that the student reads through gated cross-attention without resetting the student's KV cache. Windowed Refresh Training exposes the bridge to prefix-conditioned memory. Across thirteen datasets, MentorPulse closes 52.2% of the mentor-student gap on macro average, outperforming C2C, T2T, and equal-budget LoRA, with the largest gains on long outputs. It performs best on all eleven mentor-student pairs from three model families, with margins that narrow as the capability gap grows, and a lightweight read-pattern check predicts the gain before deployment. Measured costs identify refresh intervals that dominate text guidance on long outputs.

cs.CL

Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions

Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns. Federated learning (FL) offers a decentralized training paradigm that enables clients to collaboratively train a learning model without sharing raw data, making it a promising solution for privacy-preserving LLM training and reasoning. This paper presents a comprehensive survey of federated prompt learning (FPL) to review recent advances in integrating the federated learning paradigm and large language models, answering the following research questions: RQ1: The fundamental motivations, characteristics, and enabling technologies of FPL, and how it differs from conventional FL and full-model federated fine-tuning; RQ2: The trade-offs FPL approaches exhibit in performance, communication efficiency, computational overhead, scalability, personalization, and heterogeneity handling; RQ3: The remaining security, privacy, robustness, and system challenges, along with key future research directions. To this end, we systematically examine existing FPL methods across the full model lifecycle: pre-training, fine-tuning, and practical applications, while discussing security, privacy, and robustness issues and summarizing existing defense mechanisms. Finally, we highlight open challenges and future directions, aiming to help readers understand how the insights drive research in FPL.

cs.LG

Local Asymptotics for Treatment Choice with Partial Identification

We provide a new asymptotic framework to derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its \emph{least-favorable} configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a normal location shift model with a suitable limiting identified set. We apply our results to treatment choice problems with contaminated outcomes, to robust welfare analyses with partially identified consumer surplus, and to the problem of aggregating experimental estimates for policy adoption.

econ.EM

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models

KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and benchmark with three contributions. First, a 25-method taxonomy groups methods into five mechanism families and links them to eight verified implementations and their valid diagnostic measurements. Second, for every supported method setting, we evaluate all sources in each fixed split against a per-source FullCache control before selecting FullCache-correct/compressed-wrong (C-to-W) rows separately for each method-setting, so no compressor defines another's test set. Third, a common record format links paired outputs and run metadata to cache, likelihood, attention, and decoding measurements with explicit applicability states. On Qwen3-8B, four evidence-aware workloads yield 59 800 supported compressed runs over 2600 sources and 12 520 C-to-W rows. Under fixed diagnostic rules, 63.2% have low or partial measured/projected coverage. Only 19 rows (0.2%) combine high measured/projected coverage with strong likelihood drift; another 2,126 (17.0%) preserve structural position addressability, for which representation fidelity remains unknown, while showing the same drift. Against C-to-C success controls, all ten diagnostics separate failed from successful compression (stratified AUROC 0.684-0.871). Among 96 reproducible low-EAR failures, a controlled 4x evidence-attention boost repairs 29.2%, versus 6.3% under a count-matched sham intervention and 3.3% degradation on matched C-to-C controls. Code and data are available at https://github.com/ChosenQC/KVDiagnosis.

cs.AI

Entwined lattice of atoms and anionic electrons in layered electride LaCl

Controlling the lattice geometry that governs electronic structure is a central theme in condensed-matter physics, yet in crystalline solids this geometry is usually fixed by the atomic framework. Electrides offer an alternative route to electronic structure design in which their excess electrons can organize into anionic electron lattice (AEL) and provide a lattice-like degree of freedom. Recent work has highlighted the standalone limit, where the AEL in YCl yields bands well described by the dice-lattice model. Here, using angle-resolved photoemission spectroscopy (ARPES), we show that LaCl, although isostructural to YCl, realizes a qualitatively different regime where the AEL is entwined with the La cation framework, producing a fully reconstructed electronic structure. Combining the ARPES result with tight-binding model analysis, we demonstrate that this radical divergence stems from the activation of direct hopping channels between the AEL and the La atomic lattice. This coupling reshapes the effective lattice geometry, reconstructs the electronic states, and modifies the associated Chern band topology, transforming the bipartite dice-lattice network in YCl into a tripartite structure in LaCl. Our findings demonstrate that the coupling between the AEL and the atomic lattice can actively shape the effective lattice geometry that governs the electronic structure. This coupling can act as a powerful tuning knob for electronic structure design that is inaccessible in conventional materials.

cond-mat.str-el

Bi-S network origin of cation-disorder stability and dispersive band edges in AgBiS2

Cation-disordered AgBiS2 is a promising lead-free optoelectronic material, but both its ordered structure and the microscopic origin of its favorable electronic properties remain debated. Theory has proposed a mixed-coordination tendency with tetrahedral AgS4 and octahedral BiS6 units, whereas experiments mainly report octahedrally coordinated ordered and cation-disordered phases, together with local cation off-centering. Here, we combine a machine-learning interatomic potential with a deep-learning Hamiltonian to resolve the coupled structural and electronic evolution of AgBiS2 at large length scales. We identify the three-dimensional Bi-S network as the central structural motif governing both disorder stability and band-edge electronic states. At weak disorder, Ag/Bi exchange competes with the off-centering tendency of the Ag sublattice, producing strongly distorted local environments and convoluted diffraction signatures that hinder the identification of the ordered phase. With increasing disorder, BiS6-like units connect into a continuous Bi-S network, which stabilizes the rocksalt-like disordered phase. Despite strong cation disorder, AgBiS2 retains clear semiconductor-like band dispersion and develops a direct band gap. The connected Bi:p-S:p states supported by the Bi-S network preserve a dispersive conduction-band edge and a small electron effective mass. In contrast, mobile Ag disrupts the long-range periodicity of Ag-S bonding, leading to strongly localized valence states. These results clarify the structural controversy in ordered AgBiS2 and establish a unified physical picture of disorder stability and optoelectronic response in nonisovalent semiconductor alloys.

cond-mat.mtrl-sci

ART: Attention Run-time Termination for Efficient Large Language Model Decoding

Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depend jointly on keys and values, most existing KV management methods rely on key-only pruning, since incorporating values incurs prohibitive overhead. In this paper, we propose Attention Run-time Termination (ART), a lightweight run-time mechanism that tracks accumulated attention outputs during kernel execution and terminates subsequent KV block accesses once further contributions become negligible. Rather than replacing KV selection, ART dynamically terminates redundant KV traversal on top of existing dense or sparse attention policies. We introduce a stability-based criterion that monitors both magnitude and directional changes of intermediate attention outputs and provideds a theoretical characterization of the resulting truncation error. Experiments on the LongBench and RULER Needle-in-a-Haystack tasks show that ART increases the generation throughput of existing KV-cache methods by up to 20%, without compromising the result quality.

cs.CL

Learning Versatile Humanoid Manipulation with Touch Dreaming

Humanoid robots promise general-purpose assistance, yet real-world humanoid loco-manipulation remains challenging because it requires whole-body stability, end-effector dexterity, and contact-aware interaction under frequent contact changes. In this work, we study dexterous, contact-rich humanoid loco-manipulation. We first develop an RL-based lower-body controller that serves as the stability backbone for whole-body execution during complex manipulation. Building on this controller, we develop a VR-based whole-body humanoid data collection system that integrates dexterous hands and tactile sensing for contact-rich manipulation. We then propose Humanoid Transformer with Touch Dreaming (HTD), a multimodal encoder-decoder Transformer that models touch as a core modality alongside multi-view vision and proprioception. HTD is trained in a single stage with behavioral cloning augmented by touch dreaming: in addition to predicting action chunks, the policy predicts future hand-joint forces and future tactile latents, with tactile-latent targets provided by an exponential moving average target encoder without requiring a separate tactile pretraining stage. This encourages the policy to learn contact-aware representations for dexterous manipulation. Across five real-world contact-rich tasks, HTD achieves a 90.9% relative improvement in average success rate over the stronger baseline for each task. Ablation results further show that latent-space tactile prediction is more effective than raw tactile prediction, yielding a 30% relative gain in success rate. These results demonstrate that our touch-dreaming-enhanced learning system enables versatile, high-dexterity humanoid manipulation in the real world. More information and open-source materials are available at humanoid-touch-dream.github.io.

cs.RO

Generative Channel Knowledge Base With Environmental Information for Joint Source-Channel Coding in Semantic Communications

Semantic knowledge bases are regarded as a promising technology for upcoming 6G communications. However, existing studies mainly focus on source-side semantic modeling while overlooking the structural impact of propagation environments on semantic transmission performance. To address this issue, we propose a generative channel knowledge base (CKB) with environmental information to facilitate joint source-channel coding (JSCC) in semantic communications (SemCom) systems. First, to enable the construction of the CKB, an environment-aware dataset is established by collecting spatial position information, global image features, fine-grained semantic features, and the corresponding channel matrices. A region-of-interest (ROI)-based filtering algorithm is further designed to remove semantic components that are irrelevant to signal propagation. Second, a Transformer-based generative framework is developed to learn the mapping between multidimensional environmental information and channel matrices. A self-attention mechanism is introduced to adaptively fuse heterogeneous features, enabling the construction of a structured CKB. Third, a CKB-driven JSCC SemCom architecture is proposed, where the generated channel knowledge is injected into both of the encoder and decoder to jointly exploit source semantics and channel-environment priors in an end-to-end manner. Experimental results demonstrate that the proposed multidimensional feature fusion method achieves a channel matrix estimation error at the $10^{-3}$ level. Moreover, the CKB-driven JSCC SemCom framework integrated into SemCom systems significantly outperforms existing benchmark schemes in terms of transmission performance.

cs.IT

Delta6: A Low-Cost, 6-DOF Force-Sensing Flexible End-Effector

This paper presents Delta6, a low-cost, six-degree-of-freedom (6-DOF) force/torque end-effector that combines antagonistic springs with magnetic encoders to deliver accurate wrench sensing while remaining as simple to assemble as flat-pack furniture. A fully 3D-printed prototype, assembled entirely from off-the-shelf parts, withstands peak forces above +/-14.4 N and torques of +/-0.33 N.m per axis; these limits can be further extended by leveraging the proposed parametric analytical model. Without calibration, Delta6 attains a 99th-percentile error of 7% full scale (FS). With lightweight sequence models, the error is reduced to 3.8% FS by the best-performing network. Benchmarks on multiple computing platforms confirm that the device's bandwidth is adjustable, enabling balanced trade-offs among update rate, accuracy, and cost, while durability, thermal drift, and zero-calibration tests confirm its robustness. With Delta6 mounted on a robot arm governed by a force-impedance controller, the system successfully performs two contact-rich tasks: buffing curved surfaces and tight assemblies. Experiments validate the design, showing that Delta6 is a robust, low-cost alternative to existing 6-DOF force sensing solutions. Open-source site: https://wings-robotics.github.io/delta6 .

cs.RO

Statistical Decisions and Partial Identification: With Application to Boundary Discontinuity Design

We are delighted to respond to the excellent surveys by Cattaneo et al. (2026) and Hirano (2026). Our discussion will attempt two things: first, we show how statistical decision theory can be applied to situations with partial identification; second, we connect the surveys' themes by applying these insights to an imagined policy experiment in one of Cattaneo et al.'s (2025) applications. To do so, we lay out a stylized scenario of statistical decision making under partial identification and, drawing on our own and others' earlier work, provide a complete solution for that scenario. We then apply these results to a hypothetical reduction (modelled on actual policies) in eligibility for educational subsidies. We will see that something of interest can be said, but also that bringing the theory to the application involves some leaps of faith and leaves some questions open. This leads to the final section, where we discuss what we see as the main open challenges in statistical decision theory under partial identification.

econ.EM

Approximate Least-Favorable Distributions and Nearly Optimal Tests via Stochastic Mirror Descent

We consider a class of hypothesis testing problems where the null hypothesis postulates $M$ distributions for the observed data, and there is only one possible distribution under the alternative. We show that one can use a stochastic mirror descent routine for convex optimization to provably obtain - after finitely many iterations - both an approximate least-favorable distribution and a nearly optimal test, in a sense we make precise. Our theoretical results yield concrete recommendations about the algorithm's implementation, including its initial condition, its step size, and the number of iterations. Importantly, our suggested algorithm can be viewed as a slight variation of the algorithm suggested by Elliott, M\"uller, and Watson (2015), whose theoretical performance guarantees are unknown.

econ.EM

Effect of Concentration Fluctuations on Material Properties of Disordered Alloys

Alloying compound AX with another compound BX is widely used to tune material properties. For disordered alloys, due to the lack of periodicity, it has been challenging to calculate and study their material properties. Special quasi-random structure (SQS) method has been developed and widely used to treat this issue by matching averaged atomic correlation functions to those of ideal random alloys, enabling accurate predictions of macroscopic material properties such as total energy and volume. However, in AxB1-x alloys, statistically allowed local concentration fluctuations can give rise to defect-like minority configurations, such as bulk-like AX or BX regions in the extreme, which could strongly affect calculation of some of the material properties such as semiconductor bandgap, if it is not defined properly, leading to significant discrepancies between theory and experiment. In this work, taking the bandgap as an example, we demonstrate that the calculated alloy bandgap can be significantly underestimated in standard SQS calculations when the SQS cell size is increased to improve the structural model and the bandgap is defined conventionally as the energy difference between the lowest unoccupied state and the highest occupied state, because the rare event motifs can lead to wavefunction localization and become the dominant factor in determining the "bandgap", contrary to experiment. To be consistent with experiment, we show that the bandgap of the alloy should be extracted from the majority configurations using a density-of-states fitting (DOSF) method. This DOSF approach resolves the long-standing issue of calculating electronic structure of disordered semiconductor alloys. Similar approaches should also be developed to treat material properties that depends on localized alloy wavefunctions.

cond-mat.mtrl-sci

Visual Self-Refinement for Autoregressive Models

Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequential dependencies of next-token prediction, leading to suboptimal results. This work proposes a plug-and-play refinement module to enhance the complex spatial correspondence modeling within the generated visual sequence. This module operates as a post-pretraining step to jointly refine all generated tokens of autoregressive model, enhancing vision-language modeling under a shared sequential prediction framework. By leveraging global context and relationship across the tokens, our method mitigates the error accumulation issue within the sequential generation. Experiments demonstrate that the proposed method improves the generation quality, enhancing the model's ability to produce semantically consistent results.

cs.CV

Epsilon-Minimax Solutions of Statistical Decision Problems

A decision rule is epsilon-minimax if it is minimax up to an additive factor epsilon. We present an algorithm for provably obtaining epsilon-minimax solutions for a class of statistical decision problems. In particular, we are interested in problems where the statistician chooses randomly among I decision rules. The minimax solution of these problems admits a convex programming representation over the (I-1)-simplex. Our suggested algorithm is a well-known mirror subgradient descent routine, designed to approximately solve the convex optimization problem that defines the minimax decision rule. This iterative routine is known in the computer science literature as the hedge algorithm and is used in algorithmic game theory as a practical tool to find approximate solutions of two-person zero-sum games. We apply the suggested algorithm to different minimax problems in the econometrics literature. An empirical application to the problem of optimally selecting sites to maximize the external validity of an experimental policy evaluation illustrates the usefulness of the suggested procedure.

econ.EM

Experimental realization of dice-lattice flat band at the Fermi level in layered electride YCl

Flat electronic bands, where interactions among electrons overwhelm their kinetic energies, hold the promise for exotic correlation physics. The dice lattice has long been theorized as a host of flat bands with intriguing band topology. However, to date, no material has ever been found to host the characteristic flat bands of a dice lattice. Here, using angle-resolved photoemission spectroscopy (ARPES), we discover a dice-lattice flat band at $E_F$ in the van der Waals (vdW) electride [YCl]$^{2+}$: 2e-. In this system, excess valence electrons from Y deconfine from the cation framework to form an interstitial anionic electron lattice that constitutes the dice lattice. Our ARPES measurements unambiguously identify two sets of dice-lattice bands in YCl, including a nearly dispersionless band at the Fermi level. The flat bands and other dispersive bands observed in ARPES find excellent agreement with first-principles calculations, and theoretical analysis reveals that the near-$E_F$ electronic structure is well captured by a simple dice-lattice model. Our findings thus end the long quest of a real dice flat band material and establish vdW electride YCl as a prototype of dice metals. Our results further demonstrate the anionic electron lattice as a novel scheme for realizing lattice geometries and electronic structures rare to find in conventional crystalline systems.

cond-mat.str-el

VQ-VAE Based Digital Semantic Communication with Importance-Aware OFDM Transmission

Semantic communication (SemCom) significantly reduces redundant data and improves transmission efficiency by extracting the latent features of information. However, most of the conventional deep learning-based SemCom systems focus on analog transmission and lack in compatibility with practical digital communications. This paper proposes a vector quantized-variational autoencoder (VQ-VAE) based digital SemCom system that directly transmits the semantic features and incorporates the importance-aware orthogonal frequency division multiplexing (OFDM) transmission to enhance the SemCom performance, where the VQ-VAE generates a discrete codebook shared between the transmitter and receiver. At transmitter, the latent semantic features are firstly extracted by VQ-VAE, and then the shared codebook is adopted to match these features, which are subsequently transformed into a discrete version to adapt the digital transmission. To protect the semantic information, an importance-aware OFDM transmission strategy is proposed to allocate the key features near the OFDM reference signals, where the feature importance is derived from the gradient-based method. At the receiver, the features are rematched with the shared codebook to further correct errors. Finally, experimental results demonstrate that our proposed scheme outperforms the conventional DeepSC and achieves better reconstruction performance under low SNR region.

eess.SP