SearcharxivSearch

arXiv subjects

Xianfeng Wu

Publications and source records attributed to Xianfeng Wu.

18 recordsLinked to original sources

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.

cs.CR

Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation

Discrete diffusion language models (dLLMs) recover masked tokens in parallel, offering significant speedups over autoregressive (AR) generation. However, such promising frameworks face a fundamental architectural design dilemma: \ding{182} Adopting bidirectional attention achieves strong generation quality by allowing each position to access the full context, but is inherently incompatible with KV caching, limiting inference throughput in batch-serving scenarios; \ding{183} Conversely, causal attention enables efficient cached inference but loses all right-side context, substantially degrading generation quality. This paper introduces Bifocal dLLMs, a new paradigm that resolves this dilemma through \emph{asymmetric bidirectional context}. Analogous to bifocal lenses, we instantiate the paradigm as \textbf{R2LM} (Right-to-Left Mamba), which combines two complementary mechanisms: $a$) standard causal attention providing precise left-context with full KV cache compatibility, while $b$) a lightweight reverse Mamba SSM sidecar supplying compressed right-side context without breaking cacheability. Comprehensive experiments on continued pretraining of Qwen3-1.7B with 60B tokens demonstrate that R2LM achieves $2.4\times$ to $12.9\times$ higher throughput than bidirectional dLLMs and $1.9\times$ to $2.9\times$ speedup over AR baselines in batch serving through parallel decoding with KV caching, while exceeding the causal baseline on most benchmarks and surpassing the bidirectional dLLM on average.

cs.IR

ScalingAR: Scaling Confidence for Autoregressive Image Generation

Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation remains largely underexplored. Existing test-time scaling (TTS) methods for visual autoregressive models (VAR) rely on frequent partial decoding and external reward models, which are inefficient and often ineffective for NTP-based image generation due to the inherent instability of intermediate decoding results. To address these limitations, we propose ScalingAR, a novel test-time scaling framework tailored for NTP-based AR image generation. ScalingAR introduces token entropy as a confidence signal and operates at two complementary levels: (i) Profile Level, integrates intrinsic uncertainty and conditional utilization into a unified confidence state, and (ii) Policy Level, leverages this state for adaptive trajectory pruning and dynamic guidance scheduling. Without requiring early decoding or auxiliary rewards, ScalingAR achieves significant improvements across diverse benchmarks. Experiments show that ScalingAR (I) improves base models by $12.5\%$ on GenEval and $15.2\%$ on TIIF-Bench, (II) reduces visual token consumption by $62.0\%$ while outperforming baselines, and (III) enhances robustness, mitigating performance degradation by $26.0\%$ in challenging scenarios. These results establish ScalingAR as a robust and efficient test-time scaling solution for autoregressive image generation.

cs.CV

Bipolar-doped superconducting infinite-layer cuprates

Distilling the intrinsic physics of the superconducting CuO2 plane from the complexities of charge-reservoir layers is a defining challenge in high-temperature superconductivity. While superconducting electron-doped infinite-layer cuprates have been synthesized, controllable and uniform hole doping has long remained elusive despite exploratory attempts, limiting spectroscopic insights. Here, we realize bipolar doping across infinite-layer (Sr,Eu)CuO2 and (Ca,Li)CuO2+δ single-crystalline thin films, mapping the electronic phase diagram. Both electron- and hole-doped films show pronounced electrical resistance anisotropy, indicating the quasi-two-dimensional nature of the CuO2 planes. Angle-resolved photoemission spectroscopy across electron- and hole-doped regimes reveals persistent antiferromagnetic band folding coexisting with superconductivity. Remarkably, at a hole doping ~0.07 determined by Luttinger volume, the antiferromagnetic folding emerges from Fermi arcs within the film's single Fermi surface, with the onset superconducting transition temperature exceeding 60 K. These findings redefine the interplay between magnetic order and superconductivity and establish a definitive platform to investigate the intrinsic mechanism of high-temperature superconducting cuprates.

cond-mat.supr-con

Field re-entrant superconductivity in Eu-doped infinite-layer nickelates

Intertwined superconducting and magnetic orders may give rise to exotic quantum phases, including field-induced and re-entrant superconductivity. However, such magnetism-enhanced superconductivity has remained elusive in superconductors with higher transition temperatures. While infinite-layer nickelates represent a new class of unconventional superconductors, the impact of rare-earth magnetism on superconducting properties remains largely unexplored. Here, we show that Eu-doped infinite-layer nickelate Sm$_{0.95-x}$Ca$_{0.05}$Eu$_x$NiO$_2$ exhibits a magnetic-field-induced re-entrant superconducting phase in the Eu-rich over-doped regime. Zero-resistance transport and high-field diamagnetic screening confirm the superconducting nature of this phase, which emerges after the initial suppression of low-field superconductivity and remains robust across a broad range of temperatures, fields and field orientations. In the same doping range, we observe nonlinear Hall transport and hysteretic magnetoresistance, indicating the unconventional nature of the re-entrant behaviour. While partially consistent with a compensation mechanism between the Eu-derived exchange field and the applied field, our data reveal pronounced deviations from this model at the highest-doping levels. Our findings establish infinite-layer nickelates as a fertile platform for exploring magnetically driven high-field superconductivity in strongly correlated oxides.

cond-mat.supr-con

$3d_{z^2}$ orbital delocalization and magnetic collapse in superconducting (La,Pr)$_3$Ni$_2$O$_{7-δ}$ films

The recent discovery of Ruddlesden--Popper (RP) nickelate thin-film superconductors has opened a new frontier in unconventional superconductivity. Its realization requires both compressive epitaxial strain and highly oxidative growth conditions, yet the microscopic pathway from the parent phase to the superconducting phase remains elusive. Here, X-ray absorption spectra and resonant inelastic X-ray scattering are employed to track this evolution by independently tuning strain and oxygen content in (La,Pr)$_3$Ni$_2$O$_{7-δ}$ thin films. We uncover a remarkable two-step narrative. First, signatures of delocalization emerge in the same way upon two independent tunings: Spectral weight transfers from a ''Upper Hubbard''-like peak to the hole-like peak associated with O $2p_z$ state, and in parallel, the initially localized Ni $3d_{z^2}$ orbital becomes more itinerant followed by the broadening and weakening of $dd$ orbital excitations. Second, as itinerancy increases, long-range spin-density-wave (SDW) order is suppressed in both intensity and correlation length, indicating direct competition with superconductivity. Yet, short-range magnons persist: they become damped but their bandwidth stays unchanged. Our results paint a coherent picture that both strain and oxygenation drive the RP bilayer nickelates towards the superconducting instability, where the O $2p_z$ and Ni $3d_{z^2}$ orbitals become delocalized. Concomitantly, the long-range magnetic order loses coherence and gets suppressed. These findings establish an orbital-selective route to RP nickelate superconductivity, in which the delocalization of the $2p_z$ and $3d_{z^2}$ orbitals and the robust short-range magnons upon the melting of SDW order are prerequisites, providing strong constraints for theory and the roadmap for designing nickelate superconductors.

cond-mat.supr-con

Electronic structures across superconductor-insulator transition in Ruddlesden-Popper bilayer nickelate films

High-transition-temperature ($T_{C}$) superconductivity is recently discovered in Ruddlesden-Popper (RP) nickelate films with extraordinarily strong oxidation. While investigating phase diagrams is essential for uncovering the superconducting mechanism, the oxygen-tuned superconductor-insulator transition (SIT) in RP nickelates differs fundamentally from that in cuprates or iron-based systems. Here, we unveil the evolution of electronic structure in RP bilayer nickelate thin films across the SIT, combining angle-resolved photoemission spectroscopy (ARPES) and X-ray absorption spectroscopy (XAS) for both occupied and unoccupied states. In the superconducting state, a coherent quasiparticle band near Fermi level ($E_{F}$) coexists with an incoherent waterfall feature at high energy, paralleling that in cuprates. Approaching the insulating state with oxygen deficiency, the spectral weight of the occupied coherent quasiparticle band is gradually suppressed, accompanied by pronounced density of states redistribution and orbital reconfiguration in unoccupied states. These results reveal the electronic origin of the SIT in the phase diagram, which transcends carrier doping effects and oxygen vacancy states. Our findings point to a decisive role of oxygen in shaping the essential electronic landscape of RP bilayer nickelates, offering crucial insights into the superconducting mechanism.

cond-mat.supr-con

Learning Latent Proxies for Controllable Single-Image Relighting

Single-image relighting is highly under-constrained: small illumination changes can produce large, nonlinear variations in shading, shadows, and specularities, while geometry and materials remain unobserved. Existing diffusion-based approaches either rely on intrinsic or G-buffer pipelines that require dense and fragile supervision, or operate purely in latent space without physical grounding, making fine-grained control of direction, intensity, and color unreliable. We observe that a full intrinsic decomposition is unnecessary and redundant for accurate relighting. Instead, sparse but physically meaningful cues, indicating where illumination should change and how materials should respond, are sufficient to guide a diffusion model. Based on this insight, we introduce LightCtrl that integrates physical priors at two levels: a few-shot latent proxy encoder that extracts compact material-geometry cues from limited PBR supervision, and a lighting-aware mask that identifies sensitive illumination regions and steers the denoiser toward shading relevant pixels. To compensate for scarce PBR data, we refine the proxy branch using a DPO-based objective that enforces physical consistency in the predicted cues. We also present ScaLight, a large-scale object-level dataset with systematically varied illumination and complete camera-light metadata, enabling physically consistent and controllable training. Across object and scene level benchmarks, our method achieves photometrically faithful relighting with accurate continuous control, surpassing prior diffusion and intrinsic-based baselines, including gains of up to +2.4 dB PSNR and 35% lower RMSE under controlled lighting shifts.

cs.CV

Bosonic phases across the superconductor-insulator transition in infinite-layer samarium nickelate

Superconductivity arises from the global phase coherence of Cooper pairs. Modulation of phase coherence leads to quantum phase transitions, serving as an important tool for studying unconventional superconductivity. Here, we demonstrate bosonic phases across the superconductor-insulator transition in infinite-layer nickelate superconducting films by the control of spatially periodic network patterns. Magnetoresistance oscillations with a periodicity of h/2e provide direct evidence of 2e Cooper pairing in nickelates. The phase transition is predominantly driven by enhanced superconducting fluctuations, and Cooper pairs are involved in charge transport across the transition. Notably, we observe two types of anomalous metallic phases, emerging respectively at finite magnetic fields and down to zero magnetic field. They can be characterized by bosonic excitations, suggesting the dynamic roles of vortices in the ground states. Our work establishes nickelates as a key platform for investigating the rich landscape of bosonic phases controlled via the phase coherence of Cooper pairs.

cond-mat.supr-con

Superconductor-insulator transitions in infinite-layer nickelates controlled via ${operando}$ monitored reduction

Nickelates represent an emerging class of superconductors that demand innovative approaches for structural and electronic phase modulations. Continuous control over superconductor-insulator transition (SIT) in nickelates remains particularly challenging, hindering both fundamental understanding and potential applications. Here, we demonstrate SIT in infinite-layer nickelate superconductors utilizing multiple techniques, including an ${operando}$ monitored reduction (OMR) method. OMR enables ultrawide-range continuous modulation of the Ni 3${d}$ orbital electron occupancy from ~3${d}^7$ to ~3${d}^9$. The 3${d}$ occupancy is calibrated through systematic synchrotron X-ray absorption (XAS), combined with scanning transmission electron microscopy (STEM) annular bright field (ABF) analysis of oxygen atoms. SIT is further modulated via ionic liquid gating and magnetic field. Strikingly different from cuprates, our Nernst effect measurements show that pairing initiates at the onset of the resistive drop. The subsequent emergence of the Meissner effect at zero resistance marks the establishment of global phase coherence. Angle-dependent magnetotransport within the transition temperature regime indicates a mixture of two-dimensional (2D) and three-dimensional (3D) superconducting characters, suggesting the observed SIT deviates from the canonical 2D model. Our results provide a unique perspective on the interplay of structural and electronic phase transitions in the infinite-layer nickelates across the oxygen content-magnetic field-temperature parameter space.

cond-mat.supr-con

4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex scenarios. However, existing approaches to 4D semantic field construction primarily rely on scene-specific Gaussian splatting, which requires per-scene optimization, exhibits limited generalization, and is difficult to scale to real-world applications. To address these limitations, we propose 4DLangVGGT, the first Transformer-based feed-forward unified framework for 4D language grounding, that jointly integrates geometric perception and language alignment within a single architecture. 4DLangVGGT has two key components: the 4D Visual Geometry Transformer, StreamVGGT, which captures spatio-temporal geometric representations of dynamic scenes; and the Semantic Bridging Decoder (SBD), which projects geometry-aware features into a language-aligned semantic space, thereby enhancing semantic interpretability while preserving structural fidelity. Unlike prior methods that depend on costly per-scene optimization, 4DLangVGGT can be jointly trained across multiple dynamic scenes and directly applied during inference, achieving both deployment efficiency and strong generalization. This design significantly improves the practicality of large-scale deployment and establishes a new paradigm for open-vocabulary 4D scene understanding. Experiments on HyperNeRF and Neu3D datasets demonstrate that our approach not only generalizes effectively but also achieves state-of-the-art performance, achieving up to 2% gains under per-scene training and 1% improvements under multi-scene training. Our code released in https://github.com/hustvl/4DLangVGGT

cs.CV

Enhanced Superconductivity and Mixed-dimensional Behaviour in Infinite-layer Samarium Nickelate Thin Films

Rare-earth infinite-layer nickelates represent an emerging class of unconventional superconductors, with materials synthesis largely limited to early lanthanide compounds. Here, we report the synthesis and characterization of phase-pure superconducting samarium-based infinite-layer nickelate thin films, including the first demonstration of Sm$_{1-x}$Sr$_x$NiO$_2$, along with co-doped variants incorporating europium and calcium. These films, grown on LSAT (001) substrates, exhibit coherent lattice structures up to $\sim$ 9 nm thickness with minimal stacking faults. The co-doped compounds achieve a record-small $c$-axis parameter of 3.26 Å and display remarkable superconducting transition temperatures up to 32.5 K. These results establish a clear correlation between decreasing $c$-axis parameter and increasing critical temperature across different rare-earth systems. In addition, angle-dependent magnetoresistance investigations reveal the existence of a hybrid mixture of 2D and 3D superconductivity in this novel system with enhanced coupling between the rare-earth 5d and Ni 3d orbitals, confirmed by resonant inelastic X-ray scattering experiments. As the concentration of Eu increases, the system exhibits a clear tendency towards 3D superconductivity. Furthermore, we observe distinctive negative magnetoresistance in the europium-containing samples. These findings advocate clear materials design principles for higher transition temperatures and exotic physics in infinite-layer nickelate superconductors through structural engineering of the rare-earth site.

cond-mat.supr-con

Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

Recent advances in diffusion models have demonstrated remarkable capabilities in video generation. However, the computational intensity remains a significant challenge for practical applications. While feature caching has been proposed to reduce the computational burden of diffusion models, existing methods typically overlook the heterogeneous significance of individual blocks, resulting in suboptimal reuse and degraded output quality. To this end, we address this gap by introducing ProfilingDiT, a novel adaptive caching strategy that explicitly disentangles foreground and background-focused blocks. Through a systematic analysis of attention distributions in diffusion models, we reveal a key observation: 1) Most layers exhibit a consistent preference for either foreground or background regions. 2) Predicted noise shows low inter-step similarity initially, which stabilizes as denoising progresses. This finding inspires us to formulate a selective caching strategy that preserves full computation for dynamic foreground elements while efficiently caching static background features. Our approach substantially reduces computational overhead while preserving visual fidelity. Extensive experiments demonstrate that our framework achieves significant acceleration (e.g., 2.01 times speedup for Wan2.1) while maintaining visual fidelity across comprehensive quality metrics, establishing a viable method for efficient video generation.

cs.CV

Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

Multimodal Large Language Models (MLLMs) have displayed remarkable performance in multi-modal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a benchmark with 12 categories and design evaluation metrics that assess the degree of error in MLLM responses even when the visual content is seemingly understood. Based on this benchmark, we test 15 leading MLLMs and analyze the distribution of attention maps and logits of some MLLMs. Our investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that 'directly' relate to the visual content, leading to a bias in MLLMs' responses to other indirect questions, and 2) MLLMs' attention to visual tokens is notably lower than to system and question tokens. We further observe that attention scores between questions and visual tokens as well as the model's confidence in the answers are lower in response to misleading questions than to straightforward ones. To address the first challenge, we introduce a paired positive and negative data construction pipeline to diversify the dataset. For the second challenge, we propose to enhance the model's focus on visual content during decoding by refining the text and visual prompt. For the text prompt, we propose a content guided refinement strategy that performs preliminary visual content analysis to generate structured information before answering the question. Additionally, we employ a visual attention refinement strategy that highlights question-relevant visual tokens to increase the model's attention to visual content that aligns with the question. Extensive experiments demonstrate that these challenges can be significantly mitigated with our proposed dataset and techniques.

cs.CV

Temporal Regularization Makes Your Video Generator Stronger

Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal augmentation in video generation for the first time, and introduce FluxFlow for initial investigation, a strategy designed to enhance temporal quality. Operating at the data level, FluxFlow applies controlled temporal perturbations without requiring architectural modifications. Extensive experiments on UCF-101 and VBench benchmarks demonstrate that FluxFlow significantly improves temporal coherence and diversity across various video generation models, including U-Net, DiT, and AR-based architectures, while preserving spatial fidelity. These findings highlight the potential of temporal augmentation as a simple yet effective approach to advancing video generation quality.

cs.CV

LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization

Recent advances in text-to-image generation have primarily relied on extensive datasets and parameter-heavy architectures. These requirements severely limit accessibility for researchers and practitioners who lack substantial computational resources. In this paper, we introduce \model, an efficient training paradigm for image generation models that uses knowledge distillation (KD) and Direct Preference Optimization (DPO). Drawing inspiration from the success of data KD techniques widely adopted in Multi-Modal Large Language Models (MLLMs), LightGen distills knowledge from state-of-the-art (SOTA) text-to-image models into a compact Masked Autoregressive (MAR) architecture with only $0.7B$ parameters. Using a compact synthetic dataset of just $2M$ high-quality images generated from varied captions, we demonstrate that data diversity significantly outweighs data volume in determining model performance. This strategy dramatically reduces computational demands and reduces pre-training time from potentially thousands of GPU-days to merely 88 GPU-days. Furthermore, to address the inherent shortcomings of synthetic data, particularly poor high-frequency details and spatial inaccuracies, we integrate the DPO technique that refines image fidelity and positional accuracy. Comprehensive experiments confirm that LightGen achieves image generation quality comparable to SOTA models while significantly reducing computational resources and expanding accessibility for resource-constrained environments. Code is available at https://github.com/XianfengWu01/LightGen

cs.CV

FSC: Few-point Shape Completion

While previous studies have demonstrated successful 3D object shape completion with a sufficient number of points, they often fail in scenarios when a few points, e.g. tens of points, are observed. Surprisingly, via entropy analysis, we find that even a few points, e.g. 64 points, could retain substantial information to help recover the 3D shape of the object. To address the challenge of shape completion with very sparse point clouds, we then propose Few-point Shape Completion (FSC) model, which contains a novel dual-branch feature extractor for handling extremely sparse inputs, coupled with an extensive branch for maximal point utilization with a saliency branch for dynamic importance assignment. This model is further bolstered by a two-stage revision network that refines both the extracted features and the decoder output, enhancing the detail and authenticity of the completed point cloud. Our experiments demonstrate the feasibility of recovering 3D shapes from a few points. The proposed Few-point Shape Completion (FSC) model outperforms previous methods on both few-point inputs and many-point inputs, and shows good generalizability to different object categories.

cs.CV

Completing point cloud from few points by Wasserstein GAN and Transformers

In many vision and robotics applications, it is common that the captured objects are represented by very few points. Most of the existing completion methods are designed for partial point clouds with many points, and they perform poorly or even fail completely in the case of few points. However, due to the lack of detail information, completing objects from few points faces a huge challenge. Inspired by the successful applications of GAN and Transformers in the image-based vision task, we introduce GAN and Transformer techniques to address the above problem. Firstly, the end-to-end encoder-decoder network with Transformers and the Wasserstein GAN with Transformer are pre-trained, and then the overall network is fine-tuned. Experimental results on the ShapeNet dataset show that our method can not only improve the completion performance for many input points, but also keep stable for few input points. Our source code is available at https://github.com/WxfQjh/Stability-point-recovery.git.

cs.CV