SearcharxivSearch

arXiv subjects

Hualei Zhang

Publications and source records attributed to Hualei Zhang.

7 recordsLinked to original sources

OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models

Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.

cs.CL

Mitigating Bias in Low-SNR Financial Reinforcement Learning via Quantum Representations

The financial market is a typical low signal-to-noise ratio (SNR) setting, which often destabilizes off-policy maximum-entropy methods like Soft Actor-Critic (SAC). Specifically, noisy state representations may produce unreliable Q-value estimates, and bootstrapping amplifies these errors, forming a failure mode we call the "Financial Entropy Trap". In this paper, we propose FPQC-SAC, an efficient and plug-and-play SAC variant that places a compact and bounded Parameterized Quantum Circuit (PQC) before the actor and critic networks to constrain feature propagation at the representation level, rather than filtering raw inputs or regularizing Q-values after bootstrapping. Notably, FPQC-SAC reduces the impact of extreme market fluctuations on Bellman target estimation, while trainable quantum entanglement preserves flexible cross-asset interactions. Empirical evaluations on real-world portfolio management tasks demonstrate that FPQC-SAC substantially enhances out-of-sample stability and cumulative returns by achieving a 66.89% relative gain in cumulative return over standard unconstrained SAC and outperforms the best continuous-control deep reinforcement learning baseline by approximately 27%. Open-source code is available at https://github.com/ZeyuLIU-UST/FPQC-SAC-main.

cs.LG

VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model

Visual content generation has advanced from single-image to multi-image workflows, yet existing agents remain largely plan-driven and lack systematic reflection mechanisms to correct mid-trajectory visual errors. To address this limitation, we propose VisionCreator-R1, a native visual generation agent with explicit reflection, together with a Reflection-Plan Co-Optimization (RPCO) training methodology. Through extensive experiments and trajectory-level analysis, we uncover reflection-plan optimization asymmetry in reinforcement learning (RL): planning can be reliably optimized via plan rewards, while reflection learning is hindered by noisy credit assignment. Guided by this insight, our RPCO first trains on the self-constructed VCR-SFT dataset with reflection-strong single-image trajectories and planning-strong multi-image trajectories, then co-optimization on VCR-RL dataset via RL. This yields our unified VisionCreator-R1 agent, which consistently outperforms Gemini2.5Pro on existing benchmarks and our VCR-bench covering single-image and multi-image tasks.

cs.CV

LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning

Large Language Models (LLMs) exhibit enhanced capabilities by Chain-of-Thought reasoning. However, the extended reasoning sequences introduce significant GPU memory overhead due to increased key-value (KV) cache. Existing KV cache compression methods mitigate memory bottlenecks but struggle in long reasoning tasks. In this paper, we analyze attention patterns in reasoning tasks and reveal a Token Importance Recurrence phenomenon: a large proportion of tokens regain high attention after multiple decoding steps, which is failed to capture by existing works and may lead to unpredictable eviction on such periodically critical tokens. To address this, we propose LazyEviction, an observation window-based lagged eviction framework retaining latent recurring tokens by prioritized eviction based on tokens' recurrence patterns. Extensive experiments demonstrate that LazyEviction reduces KV cache by 50%~70% while maintaining comparable accuracy, outperforming existing KV cache compression baselines. Our implementation code can be found at https://github.com/Halo-949/LazyEviction.

cs.LG

Theory of transformation-mediated twinning

High-density and nanosized deformation twins in face-centered cubic (fcc)materials can effectively improve the combination of strength and ductility. However, the microscopic dislocation mechanisms enabling a high twinnability remain elusive. Twinning usually occurs via continuous nucleation and gliding of twinning partial dislocations on consecutive close-packed atomic planes. Here we unveil a completely different twinning mechanism being active in metastable fcc materials. The transformation-mediated twinning (TMT) is featured by a preceding displacive transformation from the fcc phase to the hexagonal close-packed (hcp) one, followed by a second-step transformation from the hcp phase to the fcc twin. The nucleation of the intermediate hcp phase is driven by the thermodynamic instability and the negative stacking fault energy of the metastable fcc phase. The intermediate hcp structure is characterized by the easy slips of Shockley partial dislocations on the basal planes, which leads to both fcc and fcc twin platelets during deformation, creating more twin boundaries and further enhancing the prosperity of twins. The disclosed fundamental understanding of the complex dislocation mechanism of deformation twinning in metastable alloys paves the road to design novel materials with outstanding mechanical properties.

cond-mat.mtrl-sci

Can experiment determine the stacking fault energy of metastable alloys?

Stacking fault energy (SFE) plays an important role in deformation mechanisms and mechanical properties of face-centered cubic (fcc) metals and alloys. In metastable fcc alloys, the SFEs determined from density functional theory (DFT) calculations and experimental methods often have opposite signs. Here, we show that the negative SFE by DFT reflects the thermodynamic instability of the fcc phase relative to the hexagonal close-packed one; while the experimentally determined SFEs are restricted to be positive by the models behind the indirect measurements. We argue that the common models underlying the experimental measurements of SFE fail in metastable alloys. In various concentrated solid solutions, we demonstrate that the SFEs obtained by DFT calculations correlate well with the primary deformation mechanisms observed experimentally, showing a better resolution than the experimentally measured SFEs. Furthermore, we believe that the negative SFE is important for understanding the abnormal behaviors of partial dislocations in metastable alloys under deformation. The present work advances the fundamental understanding of SFE and its relation to plastic deformations, and sheds light on future alloy design by physical metallurgy.

cond-mat.mtrl-sci

Tensile strain-induced softening of iron at high temperature

In weakly ferromagnetic materials, already small changes in the atomic configuration triggered by temperature or chemistry can alter the magnetic interactions responsible for the non-random atomic-spin orientation. Different magnetic states, in turn, can give rise to substantially different macroscopic properties. A classical example is iron, which exhibits a great variety of properties as one gradually removes the magnetic long-range order by raising the temperature towards and beyond its Curie point of $T_{\text{C}}^{0}=1043$\,K. Using first-principles theory, here we demonstrate that uniaxial tensile strain can also destabilize the magnetic order in iron and eventually lead to a ferromagnetic to paramagnetic transition at temperatures far below $T_{\text{C}}^{0}$. In consequence, the intrinsic strength of the ideal single-crystal body-centered cubic iron dramatically weakens above a critical temperature of $\sim 500$\,K. The discovered strain-induced magneto-mechanical softening provides a plausible atomic-level mechanism behind the observed drop of the measured strength of Fe whiskers around $300-500$\,K. Alloying additions which have the capability to partially restore the magnetic order in the strained Fe lattice, push the critical temperature for the strength-softening scenario towards the magnetic transition temperature of the undeformed lattice. This can result in a surprisingly large alloying-driven strengthening effect at high temperature as illustrated here in the case of Fe-Co alloy.

cond-mat.mtrl-sci