SearcharxivSearch

arXiv subjects

Jianze Wang

Publications and source records attributed to Jianze Wang.

7 recordsLinked to original sources

TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.

cs.CL

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.

cs.CL

A Novel FeFET Differential Bit-Cell With Hybrid Volatile and Non-Volatile Memory Modes

Non-volatile SRAM (nvSRAM) designs have been investigated to address the high leakage power of CMOS-based SRAM and the large write latency of emerging non-volatile memory (eNVM) technologies. However, prior nvSRAM designs that combine SRAM with eNVM devices typically require backup and restore (B\&R) operations and incur significant cell-area overhead. Here, we propose a differential memory bit-cell consisting of a pair of cross-coupled ferroelectric field-effect transistors (FeFETs) and a pair of access transistors, resulting in a four-transistor (4T) structure, which is smaller than conventional 6T SRAM and many prior nvSRAM designs. The proposed bit-cell can be configured to operate in either volatile or non-volatile mode by adjusting the write conditions. In the non-volatile mode, the proposed nvSRAM achieves a store power of 0.13~$μ$W with a 2~ns store time, and no explicit B\&R operation is required. The proposed bit-cell can also be viewed as a cross-coupled gain cell, enabling further applications.

cs.ET

Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation

Structural modeling is a fundamental component of computational engineering science, in which even minor physical inconsistencies or specification violations may invalidate downstream simulations. The potential of large language models (LLMs) for automatic generation of modeling code has been demonstrated. However, non-executable or physically inconsistent outputs remain prevalent under stringent engineering constraints. A framework for physics-consistent automatic building modeling is therefore proposed, integrating domain knowledge construction, constraint-oriented model alignment, and verification-driven evaluation. CivilInstruct is introduced as a domain-specific dataset that formalizes structural engineering knowledge and constraint reasoning to enable simulation-ready model generation. A two-stage fine-tuning strategy is further employed to enforce constraint satisfaction and application programming interface compliance, substantially reducing hallucinated and non-conforming outputs. MBEval is presented as a verification-driven benchmark that evaluates executability and structural dynamics consistency through closed-loop validation. Experimental results show consistent improvements over baselines across rigorous verification metrics. Our code is available at https://github.com/Jovanqing/AutoBM.

cs.SE

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single-teacher capability ceiling: when the teacher errs, the student inherits the error. OPD also remains largely unexplored in agentic tasks, where per-step errors compound across long trajectories and destabilize training. We propose MAD-OPD (Multi-Agent Debate-driven On-Policy Distillation), which breaks this ceiling by recasting the distillation teacher as a deliberative collective of teachers that debate over the student's on-policy state; the debate produces an emergent collective intelligence that supplies token-level supervision, with each teacher's contribution weighted by its post-debate confidence. To extend OPD to agentic tasks, we also introduce On-Policy Agentic Distillation (OPAD), which adds step-level sampling to stabilize training under multi-step error compounding. We additionally derive a task-adaptive divergence principle, selecting JSD (Jensen-Shannon divergence) for agentic stability and reverse KL (Kullback-Leibler) divergence for code generation, and verify it both theoretically and empirically. Across six teacher-student configurations (Qwen3 and Qwen3.5; 1.7B-14B students, 8B-32B teachers) and five agentic and code benchmarks, MAD-OPD ranks first across all six configurations; on the 14B+8B$\to$4B setting it lifts the agentic average by $+2.4\%$ and the code average by $+3.7\%$ over the stronger single-teacher OPD.

cs.CL

Bilayer-skyrmion-based design of neuron and synapse for spiking neural network

Magnetic skyrmion technology is promising for the next-generation spintronics-based memory and neuromorphic computing due to their small size, non-volatility and low depinning current density. However, the Magnus force originating from the skyrmion Hall effect causes the skyrmion to move along a curved trajectory, which may lead to the annihilation of the skyrmion in a nanotrack during current-induced skyrmion motion. Consequently, circuits utilizing skyrmionic motion need to be designed to limit the impact of the skyrmion Hall effect. In this work, we propose a design of an artificial neuron, and a synapse using the bilayer device consisting of two antiferromagnetically exchange coupled ferromagnetic layers, which achieves robustness against the skyrmion Hall effect by nullifying the Magnus force. Using micromagnetic simulations, we show that the bilayer device can work as an artificial neuron and also as a synapse by modifying its uniaxial anisotropy. We also demonstrate that our proposed skyrmionic synapse has an intrinsic property of perfectly linear and symmetric weight update, which is highly desirable for the synapse operation. A spiking neural network implemented using our proposed synapse and neuron was simulated and showed to achieve 96.23\% accuracy in performing classification on the MNIST handwritten digit dataset.

cond-mat.mes-hall

Design of Spintronics-based Neuronal and Synaptic Devices for Spiking Neural Network Circuits

Topologically stable magnetic skyrmion has a much lower depinning current density that may be useful for memory as well as neuromorphic computing. However, skyrmion-based devices suffer from the Magnus force originating from the skyrmion Hall effect, which may result in unwanted skyrmion annihilation if the magnitude of the driving current gets too large. A design of an artificial neuron and a synapse using a synthetic antiferromagnetically coupled bilayer device, which nullifies the Magnus force, is demonstrated in this work. The leak term in the artificial leaky integrate-and-fire neuron is achieved by engineering the uniaxial anisotropy profile of the neuronal device. The synaptic device has a similar structure as the neuronal device but has a constant uniaxial anisotropy. The synaptic device also has a linear and symmetric weight update, which is a highly desirable trait of an artificial synapse. Neuronal and synaptic devices based on magnetic domain-wall (DW) motion are also studied and compared to skyrmionic devices. Our simulation results show the energy required to perform such operation in DW or skyrmion-based devices is on the order of a few fJ.

cond-mat.mes-hall