SearcharxivSearch

arXiv subjects

Haojin Wang

Publications and source records attributed to Haojin Wang.

7 recordsLinked to original sources

MentorCollab: Large-to-Small Inference-Time Mentorship for Concise Reasoning in Language Models

Large reasoning models (LRMs) have demonstrated impressive reasoning capabilities, but their solutions are often verbose and computationally expensive, and taxing for users to read. In contrast, small language models (SLMs) produce concise outputs with lower inference costs, yet they frequently struggle on challenging multi-step reasoning tasks. Existing inference-time collaboration methods attempt to bridge this gap through imitation, encouraging SLMs to follow the reasoning process of LRMs. However, the student often inherits the mentor's overthinking, producing long and reflective reasoning chains while still falling short in accuracy. We propose MentorCollab, a collaboration method based on mentorship: the SLM remains the primary generator and consults the LRM only when additional reasoning support is needed. At sparsely sampled token positions, we probe for divergence between the two models and use a lightweight verifier to decide whether the SLM should follow a short lookahead segment from its mentor or continue on its own. Across 15 SLM-LRM pairs and 3 domains, our method achieves an average gain of 3.0%, with improvements of up to 8.0% in 12 settings. The resulting traces remain shorter than the mentor's, using only a small fraction of its tokens. These results demonstrate that selective, verified mentorship can boost reasoning accuracy while preserving concise generation.

cs.CL

Magnetoelectric effect in multiferroic metals via a direct spin-charge interaction

Much is known about the magnetoelectric effect of multiferroic insulators, yet little is understood about multiferroic metals. In this work, we propose a stacking engineering strategy based on monolayer magnets to construct multiferroic metals, with validation through first-principles calculations on six experimentally synthesized materials. Such multiferroic metals exhibit predominantly linear magnetoelectric response, originating from direct spin-charge interactions as a result of external field-modulated Fermi energy. This fundamentally differs from spin-charge-lattice or spin-orbit coupling mechanisms in multiferroic insulators, offering application advantages in the field of high-speed response. We derive a universal formula for understanding the magnetoelectric coupling in these multiferroic metals. Our work provides insights for exploring magnetoelectric coupling mechanisms and designing functional materials with strong magnetoelectric coupling.

cond-mat.mtrl-sci

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement learning (RL) stage. However, SFT is often optimized in isolation to maximize SFT performance alone. We show that, after identical RL training, models initialized from stronger SFT checkpoints can significantly underperform those initialized from weaker ones. We attribute this to a mismatch typical in current SFT-RL pipelines: the distribution that generates the offline SFT data can differ substantially from the policy optimized during online RL, which learns from its own rollouts. We propose PEAR (Policy Evaluation-inspired Algorithm for Offline Learning Loss Re-weighting), an SFT-stage method that corrects this mismatch and better prepares the model for RL. PEAR uses importance sampling to reweight the SFT loss, with three variants operating at the token, block, and sequence levels. It can be used to augment standard SFT objectives and incurs little additional training overhead once probabilities for the offline data are collected. We conduct controlled experiments on verifiable reasoning games and mathematical reasoning tasks on Qwen 2.5 and 3 and DeepSeek-distilled models. PEAR consistently improves post-RL performance over canonical SFT, with pass at 8 gains up to a 14.6 percent on AIME2025. Our results suggest that PEAR is an effective step toward more holistic LLM post-training by designing and evaluating SFT with downstream RL in mind rather than in isolation.

cs.LG

MoCo: A One-Stop Shop for Model Collaboration Research

Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multiple LMs collaborate, compose, and complement each other. Existing research on this topic has mostly been disparate and disconnected, from different research communities, and lacks rigorous comparison. To consolidate existing research and establish model collaboration as a school of thought, we present MoCo: a one-stop Python library of executing, benchmarking, and comparing model collaboration algorithms at scale. MoCo features 26 model collaboration methods, spanning diverse levels of cross-model information exchange such as routing, text, logit, and model parameters. MoCo integrates 25 evaluation datasets spanning reasoning, QA, code, safety, and more, while users could flexibly bring their own data. Extensive experiments with MoCo demonstrate that most collaboration strategies outperform models without collaboration in 61.0% of (model, data) settings on average, with the most effective methods outperforming by up to 25.8%. We further analyze the scaling of model collaboration strategies, the training/inference efficiency of diverse methods, highlight that the collaborative system solves problems where single LMs struggle, and discuss future work in model collaboration, all made possible by MoCo. We envision MoCo as a valuable toolkit to facilitate and turbocharge the quest for an open, modular, decentralized, and collaborative AI future.

cs.CL

MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning

Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical VQA through educationally-validated diagnostic sequences. To construct this dataset, we develop a scalable pipeline that leverages narrative transcripts from medical education videos to align visual frames with textual concepts, automatically producing 2,851 high-quality multi-image VQA pairs with explicit, transcript-grounded reasoning chains. Our evaluation of 11 advanced MLLMs (including reasoning models) exposes severe deficiencies in multi-image synthesis, where accuracies mostly fall below 50% and exhibit instability across varying image counts. Error analysis demonstrates that models often treat images as isolated instances, failing to track pathological progression or cross-reference anatomical shifts. MedFrameQA provides a rigorous standard for evaluating the next generation of MLLMs in handling complex, temporally grounded medical narratives.

cs.CV

Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce

Autoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt. In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit than others. Specifically, for any target next-token distribution over the vocabulary, we attempt to find a prompt that induces the LM to output a distribution as close as possible to the target, using either soft or hard gradient-based prompt tuning. We find that (1) in general, distributions with very low or very high entropy are easier to approximate than those with moderate entropy; (2) among distributions with the same entropy, those containing ''outlier tokens'' are easier to approximate; (3) target distributions generated by LMs -- even LMs with different tokenizers -- are easier to approximate than randomly chosen targets. These results offer insights into the expressiveness of LMs and the challenges of using them as probability distribution proposers.

cs.CL

Are there type-III multiferroics?

Multiferroics are known to be classified into two types. However, type-I lacks sufficient magnetoelectric coupling and type-II lacks sufficient electric polarization, making both practically difficult. In this work, we explore the possibility of type-III multiferroics, where the origins of ferroelectricity and magnetism are highly intertwined but not causally related, with a combination of strong magnetoelectric coupling and large polarization. Our first-principles calculations predict that monolayer TiCdO$_{4}$ is such a type-III ferroelectric-ferromagnetic multiferroics with both electronic and magnetic orders originating from competing electron populations on oxygen atoms. It shows an electric polarization of 50 $μ$C/m$^{2}$ while the maximum linear and quadratic magnetoelectric response are as high as 35000 ps/m and 1.59 $\times$ 10$^{-14}$ s/A, respectively. Our study opens up new perspectives for the discovery and design of much-anticipated multiferroics that can be used for cross-modulation.

cond-mat.mtrl-sci