SearcharxivSearch

arXiv subjects

Zhen Zhu

Publications and source records attributed to Zhen Zhu.

At least 19 recordsLinked to original sources

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: https://github.com/avaxiao/ReToken

cs.CV

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface. While diffusion-based systems dominate UVG due to strong quality and controllability, their iterative sampling incurs substantial inference latency, limiting practical deployment. To address these limitations, we propose UniGen-AR, a framework that pairs a general-purpose multi-modal language model (MLLM) with an efficient next-scale visual auto-regressive (VAR) decoder. This design retains the flexibility of MLLM-based conditioning while leveraging the sampling efficiency and latent unification properties of VAR models. In our framework, the MLLM encodes free-form instructions and control signals into a unified sequence, which guides the VAR decoder to generate image-valued outputs for over 15 tasks spanning four families. Empirically, UniGen-AR achieves up to $19 \times$ lower inference latency than diffusion-based baselines while maintaining or improving output quality. Our ablations further reveal that VQ-VAE tokenizer design, particularly codebook size and hierarchy, is a critical factor for VAR scalability in UVG. These results establish visual auto-regressive modeling as a compelling and efficient backbone for unified visual generation. Our project page is at https://zpbao.github.io/projects/unigenar.

cs.CV

Supercurrent effect in a charge density wave intertwined superconductor

The energy-momentum (E-k) dispersion of quasiparticles constitutes a fundamental concept in condensed matter systems. The ability to modify the E-k dispersion, exemplified by supercurrent-induced Doppler shifts of Bogoliubov quasiparticle spectra in superconductors, enables manipulation of various emergent quantum properties. However, investigations into the supercurrent effect on superconductors intertwined with charge orders remain scarce. Here, we report that the Meissner current, generated by the diamagnetic response to an applied in-plane magnetic field, can tailor Bogoliubov quasiparticle excitations at the precursor charge density wave (CDW) vectors. Our scanning tunneling spectroscopic imaging reveals a field-driven symmetry breaking of CDW modulations, specifically a C3v-to-Cs transition, in superconducting NbSe2. Model calculations suggest that the observed anisotropy originates from a selective Doppler-shift-induced E-k dispersion reconstruction. Furthermore, altering the field direction enables on-demand tuning of anisotropic CDW modulations and visualization of their momentum-space distribution. These results highlight a novel mechanism for controlling emergent electronic phases through momentum-space engineering.

cond-mat.supr-con

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint, which commits every retrieval direction to the same stability-plasticity trade-off. We propose AlphaWiSE, a post-hoc weight-space interpolation method that composes two frozen source checkpoints. For each aligned parameter tensor identified by its checkpoint key, AlphaWiSE fits one scalar interpolation coefficient shared by all tensor entries. The coefficients are fitted on a smaller exemplar memory and used to materialize one interpolated checkpoint. The deployed model has the same architecture and parameter count as either source checkpoint, which does not require additional inference time. Extensive experiments on audio-image-text retrieval show consistent improvements over strong continual-learning baselines across multiple retrieval directions and evaluation metrics.

cs.CV

Why Fine-Tuning Encourages Hallucinations and How to Fix It

Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acquired during pre-training. Since these errors arise as a by-product of knowledge degradation, we explore whether established continual learning tools can mitigate them. We propose a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t.~pre-existing knowledge by regularizing output-distribution drift. We also show that when new knowledge acquisition is unnecessary, suppressing factual plasticity by freezing parameter groups preserves task performance while reducing hallucinations. Lastly, we investigate the mechanism, contrasting capacity limitations, behavior cloning, and localized interference. Our experiments show that a main driver is interference among overlapping semantic representations, which self-distillation mitigates and an associative-memory model explains: forgetting grows with the overlap between new and stored facts.

cs.CL

Topological surface states revealed by the Zeeman effect in superconducting UTe2

Intrinsic topological superconductors with protected boundary modes obeying non-Abelian statistics constitute a vanishingly small class of quantum materials. A defining spectroscopic signature of such phases is the presence of in-gap topological surface states (TSS). However, despite extensive theoretical proposals, their unambiguous experimental identification has remained elusive. Here we use vector magnetic-field scanning tunnelling microscopy to obtain direct spectroscopic evidence of TSS in the spin-triplet superconductor UTe2. Atomic-scale spectroscopy reveals striking site-dependent superconductivity: Te sites host a large in-gap density of states that nearly fills the superconducting gap, whereas neighboring atomic sites remain gapped. Upon application of a magnetic field, the in-gap states on the Te sites are selectively suppressed, yielding a spatially homogeneous superconducting state with a markedly deeper gap relative to zero field. This site-selective gap evolution is in quantitative agreement with theoretical predictions for TSS in UTe2 that possess dominant Te-orbital character. Spectral-function calculations incorporating the Zeeman coupling reproduce the observed magnetic-field response. Our results provide a spectroscopic fingerprint of the long-sought TSS in superconductors and establish UTe2 as a compelling system for exploring intrinsic topological superconductivity.

cond-mat.supr-con

Evidence of intertwined pair density and charge density wave orders in UTe2

The strongly correlated spin-triplet superconductor UTe2 hosts an unusual landscape of magnetic-field-sensitive charge density wave (CDW) phases, positioning it as a compelling system for studying intertwined electronic orders. A central challenge is determining whether the observed charge modulations arise from a triplet pair density wave (PDW) order and, if so, how the anisotropic magnetic field response of triplet superconductivity is manifested in the CDW response. Here, using a scanning tunneling microscope equipped with a vector magnetic field, we systematically investigate the evolution and interrelation of distinct CDW orders. Complementing the previously identified incommensurate CDW peaks (qi=1,2,3), we resolve an additional set of nondispersive modulations (pi=1,2,3 and h1,2) with distinct temperature and magnetic field dependencies. The pi CDW peaks vanish near Tc, while the qi peaks survive well above Tc but are progressively suppressed by magnetic field in an anisotropic manner. The critical fields of the qi peaks mirror the directional hierarchy of Hc2, which suggests a PDW is present above the bulk Tc. This is consistent with a Landau free-energy picture where PDWs with wavevectors pi form above the bulk Tc, leading to composite CDW orders with wavevector qi. Below Tc, the coupling of PDWs and uniform superconductivity leads to the pi CDWs. Together, these findings establish UTe2 as a rare platform where both the parent PDW and descendant orders are directly resolved, enabling access to both the fundamental and emergent manifestations of PDW physics.

cond-mat.supr-con

Decoupling Intrinsic Molecular Efficacy from Platform Effects: An Interpretable Machine Learning Framework for Unbiased Perovskite Passivator Discovery

Rational design of interface passivators for perovskite solar cells is hindered by the entanglement of intrinsic molecular efficacy with extrinsic platform-dependent performance - a confounding factor that obscures true chemical advances. Here, we present a generalizable, interpretable machine learning framework that decouples these effects via an asymptotic saturation model, enabling unbiased discovery of molecules with genuine intrinsic gains. Trained on a curated dataset of 240 experimental entries, our model identifies hydrogen bond acceptor strength and electrostatic potential difference as key descriptors. Guided by these insights, we screened >121 million PubChem compounds using a hierarchical strategy integrating diversity clustering and uncertainty quantification. Five dual-functional candidates (e.g., TDZ-S, TZC-F) are identified, exhibiting superior predicted efficacy (surpassing experimental benchmarks) and high confidence. First-principles calculations confirm strong chemisorption (Eads<-1.7 eV), net electron donation, and optimized interfacial energetics. Crucially, our closed-loop "data-interpretation-screening-verification" pipeline establishes a transferable paradigm for rational materials design, extendable to other optoelectronic interfaces beyond perovskites.

cond-mat.mtrl-sci

TacMamba: A Tactile History Compression Adapter Bridging Fast Reflexes and Slow VLA Reasoning

In visually ambiguous manipulation such as detecting button click tactile feedback is often the sole source of ground truth. However, fusing tactile data poses a significant challenge due to a spatiotemporal mismatch: tactile perception requires high-frequency processing with long-horizon memory (System 1), whereas visual policies operate at low control frequencies (System 2). Existing architectures struggle to bridge this gap: Transformers are computationally prohibitive for high-frequency loops (>100Hz), while LSTMs suffer from forgetting over extended interaction histories. In this paper, we introduce TacMamba, a hierarchical architecture that aligns high-bandwidth tactile reflexes with low-frequency visual planning. Our approach comprises three core contributions: (1) a custom high-frequency tactile interface designed for flexible integration; (2) a Mamba-based Tactile History Compressor that encodes continuous force history into a compact state with O(1) inference latency (0.45 ms), enabling plug-and-play fusion with VLA models without joint pre-training and (3) a Tactile-Guided Dual-Stage Training strategy that leverages temporal discrimination for self-supervised representation learning and phase-uniform sampling to mitigate data sparsity. Experiments on discrete counting and implicit state switching demonstrate that TacMamba achieves 100% success rates, significantly outperforming the visual-only pi_0.5 baseline, while strictly satisfying hard real-time constraints.

cs.RO

CTransformer: Deep-transformer-based 3D cell membrane tracking with subcellular-resolved molecular quantification

Deep learning segmentation and fluorescence imaging techniques allow the cellular morphology of living embryos to be constructed spatiotemporally. These development processes involve numerous molecules distributed at the subcellular scale, such as cell adhesion (E-cadherin), which accumulate at cell-cell interfaces to regulate intercellular connection. However, quantifying molecular distributions within specific subcellular regions across the entire embryo, where cell movement and molecular redistribution occur rapidly, is challenging due to the need for simultaneous cell morphology reconstruction and lineage tracing due to photobleaching and phototoxicity. We report a transformer-based pipeline, CTransformer, that establishes a 4D cellular morphology map before the 550-cell (late) stage. CTransformer constructed 4D cellular morphology atlases, reaching 80% accuracy at the 550-cell stage. Through this advanced architecture, we use only one channel to reconstruct cell morphology and achieve cell tracing. With each cell's morphology as a reference, the distribution of specific molecules throughout the cell body and at cell interfaces can be quantitatively measured in another fluorescence channel. We apply this methodology to track E-cadherin during embryonic development of the worm Caenorhabditis elegans, from fertilization to gastrulation. Our results reveal that E-cadherin is tightly regulated across individual embryos, both within single cells and at cell-cell interfaces, displaying an anterior-posterior gradient and cell- and lineage-specific patterns. Furthermore, its spatiotemporal heterogeneity influences cell mechanics and embryonic morphogenesis, helping explain how C. elegans achieves stereotypical developmental patterns at cellular resolution.

physics.bio-ph

Identifying the Catalytic Descriptor of Single-Atom Catalysts in Nitrate Reduction Reaction: An Interpretable Machine-Learning Method

Elucidating the catalytic descriptor that accurately characterizes the structure-activity relationships of typical catalysts for various important heterogeneous catalytic reactions is pivotal for designing high-efficient catalytic systems. Here, an interpretable machine learning technique was employed to identify the key determinants governing the nitrate reduction reaction ($\rm NO_3RR$) performance across 286 single-atom catalysts (SACs) with the active sites anchored on double-vacancy $\rm BC_3$ monolayers. Through Shapley Additive Explanations (SHAP) analysis with reliable predictive accuracy, we quantitatively demonstrated that, favorable $\rm NO_3RR$ activity stems from a delicate balance among three critical factors: low $\rm N_V$, moderate $\rm D_N$, and specific doping patterns. Building upon these insights, we established a descriptor ($\psi$) that integrates the intrinsic catalytic properties and the intermediate O-N-H angle ($\theta$), effectively capturing the underlying structure-activity relationship. Guided by this, we further identified 16 promising catalysts with predicted low limiting potential ($U_{\rm L}$). Importantly, these catalysts are composed of cost-effective non-precious metal elements and are predicted to surpass most reported catalysts, with the best-performing Ti-V-1N1 is predicted to have an ultra-low $U_{\rm L}$ of $-0.10$ V.

physics.chem-ph

CryoDyna: Multiscale end-to-end modeling of cryo-EM macromolecule dynamics with physics-aware neural network

Single-particle cryo-EM has transformed structural biology but still faces challenges in resolving conformational heterogeneity at atomic resolution. Existing cryo-EM heterogeneity analysis methods either lack atomic details or tend to subject to overfitting due to image noise and limited information in single views. To obtain atomic detailed multiple conformations and make full use of particle images of different orientations, we present here CryoDyna, a deep learning framework to infer macromolecular dynamics directly from 2D projections by integrating cross-view attention and multi-scale deformation modeling. Combining coarse-grained MARTINI representation with atomic backmapping, CryoDyna achieves near-atomic interpretation of protein conformational landscapes. Validated on multiple simulated and experimental datasets, CryoDyna demonstrates improved modeling accuracy and robustly recovers multi-scale complex structure changes hidden in the cryo-EM particle stacks. As examples, we generated protein-RNA coordinated motions, resolved dynamics in the unseen region of RAG signal end complex, mapped translocating ribosome states in a one-shot manner, and revealed step-wise closure of a membrane-anchored protein multimer. This work bridges the gap between cryo-EM heterogeneity analysis and atomic-scale structural dynamics, offering a promising tool for exploration of complex biological mechanisms.

q-bio.BM

How to Teach Large Multimodal Models New Skills

How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills while monitoring general ability on eight held-out benchmarks across three model families. Surprisingly, we find that performance lost on held-out tasks after fine-tuning on one skill can partly recover when the model is subsequently tuned on a different skill. We trace this behavior to a measurable shift in the output token distribution, manifested through a simple counting-bias probe that shows the shift co-varies with forgetting. Guided by this insight, we identify two simple, robust tuning recipes that learn strongly while limiting drift: (i) updating only the self-attention projection layers (SA Proj., $\Delta$ learning +24.9 / $\Delta$ held-out forgetting -0.6), and (ii) updating only the MLP Gate&Up while freezing the Down projection (+30.5 / -2.1). Both substantially outperform full-LLM tuning (+31.8 / -23.3) in the learning-forgetting trade-off. We also compare against common forgetting mitigation methods: Learning without Forgetting (LwF), LoRA, Mixture-of-Experts, and weight-space interpolation (WiSE-FT), and find that our selective tuning recipes match or exceed their learning-stability balance while remaining simpler, requiring no replay, auxiliary parameters, or per-stage tuning. These results hold across LLaVA-OneVision, LLaVA-NeXT, and Qwen2.5-VL, confirming that the key to teaching LMMs new skills without forgetting lies in controlling output distribution shift by choosing which components to tune. Code will be made available.

cs.AI

Economic Bidding Strategy of Electric Vehicles in Real-Time Electricity Markets based on Marginal Opportunity Value

The participation of electric vehicle (EV) aggregators in real-time electricity markets offers promising revenue opportunities through price-responsive energy arbitrage. A central challenge in economic bidding lies in quantifying the marginal opportunity value of EVs' charging and discharging decisions. This value is implicitly defined and dynamically shaped by uncertainties in electricity prices and availability of EV resources. In this paper, we propose an efficient bidding strategy that enables EV aggregators to generate market-compliant bids based on the underlying marginal value of energy. The approach first formulates the EV aggregator's power scheduling problem as a Markov decision process, linking the opportunity value of energy to the value function. Building on this formulation, we derive the probability distributions of marginal opportunity values across EVs' different energy states under stochastic electricity prices. These are then used to construct closed-form expressions for marginal charging values and discharging costs under both risk-neutral and risk-averse preferences. The resulting expressions support a fully analytical bid construction procedure that transforms marginal valuations into stepwise price-quantity bids without redundant computation. Case studies using real-world EV charging data and market prices demonstrate the effectiveness and adaptability of the proposed strategy.

math.OC

CrowdAgent: Multi-Agent Managed Multi-Source Annotation System

High-quality annotated data is a cornerstone of modern Natural Language Processing (NLP). While recent methods begin to leverage diverse annotation sources-including Large Language Models (LLMs), Small Language Models (SLMs), and human experts-they often focus narrowly on the labeling step itself. A critical gap remains in the holistic process control required to manage these sources dynamically, addressing complex scheduling and quality-cost trade-offs in a unified manner. Inspired by real-world crowdsourcing companies, we introduce CrowdAgent, a multi-agent system that provides end-to-end process control by integrating task assignment, data annotation, and quality/cost management. It implements a novel methodology that rationally assigns tasks, enabling LLMs, SLMs, and human experts to advance synergistically in a collaborative annotation workflow. We demonstrate the effectiveness of CrowdAgent through extensive experiments on six diverse multimodal classification tasks. The source code and video demo are available at https://github.com/QMMMS/CrowdAgent.

cs.AI

Quantum-size effect induced Andreev bound states in ultrathin metallic islands proximitized by a superconductor

While Andreev bound states (ABSs) have been realized in engineered superconducting junctions, their direct observation in normal metal/superconductor heterostructures-enabled by quantum confinement-remains experimentally elusive. Here, we report the detection of ABSs in ultrathin metallic islands (Bi, Ag, and SnTe) grown on the s-wave superconductor NbN. Using high-resolution scanning tunneling microscopy and spectroscopy, we clearly reveal in-gap ABSs with energies symmetric about the Fermi level. While the energies of these states show no position dependence, their wave functions exhibit spatial oscillations, demonstrating a quantum size effect. Both the energy levels and spatial distribution of the ABSs can be reproduced by our effective model in which a metallic island is coupled to the superconducting substrate via the proximity effect. We demonstrate that the coupling strength plays a critical role in determining the ABS energies. Our work introduces a novel physical platform for implementing ABSs, which hold promise for significant device applications.

cond-mat.supr-con

InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow

We propose a fast text-guided image editing method called InstantEdit based on the RectifiedFlow framework, which is structured as a few-step editing process that preserves critical content while following closely to textual instructions. Our approach leverages the straight sampling trajectories of RectifiedFlow by introducing a specialized inversion strategy called PerRFI. To maintain consistent while editable results for RectifiedFlow model, we further propose a novel regeneration method, Inversion Latent Injection, which effectively reuses latent information obtained during inversion to facilitate more coherent and detailed regeneration. Additionally, we propose a Disentangled Prompt Guidance technique to balance editability with detail preservation, and integrate a Canny-conditioned ControlNet to incorporate structural cues and suppress artifacts. Evaluation on the PIE image editing dataset demonstrates that InstantEdit is not only fast but also achieves better qualitative and quantitative results compared to state-of-the-art few-step editing methods.

cs.CV

Training-free Geometric Image Editing on Diffusion Models

We tackle the task of geometric image editing, where an object within an image is repositioned, reoriented, or reshaped while preserving overall scene coherence. Previous diffusion-based editing methods often attempt to handle all relevant subtasks in a single step, proving difficult when transformations become large or structurally complex. We address this by proposing a decoupled pipeline that separates object transformation, source region inpainting, and target region refinement. Both inpainting and refinement are implemented using a training-free diffusion approach, FreeFine. In experiments on our new GeoBench benchmark, which contains both 2D and 3D editing scenarios, FreeFine outperforms state-of-the-art alternatives in image fidelity, and edit precision, especially under demanding transformations. Code and benchmark are available at: https://github.com/CIawevy/FreeFine

cs.CV