SearcharxivSearch

arXiv subjects

Fei Peng

Publications and source records attributed to Fei Peng.

At least 19 recordsLinked to original sources

Nearly Isotropic Vortex Solid in $\mathbf{(La,Pr)_{3}Ni_{2}O_{7}}$ Thin Films

The discovery of superconductivity in bulk bilayer nickelates has established a new platform for exploring high-$T_c$ superconductivity beyond the cuprates. The role of the Ni $3d_{z^2}$-derived $\gamma$ band in the superconductivity of bilayer nickelates remains unresolved. By performing simultaneous resistance and diamagnetism measurements on (La,Pr)$_3$Ni$_2$O$_7$ thin films, we map the vortex melting phase diagram for both in-plane and out-of-plane magnetic fields. For $H\parallel c$, the geometric confinement effect gives rise to pancake vortices. Remarkably, the anisotropy parameter of the vortex melting field $\gamma_{H_m} \equiv H_m^{ab}/H_m^c$ decreases monotonically with decreasing temperature and approaches unity at low temperatures. Within the anisotropic Ginzburg--Landau scaling, $H_m^{ab}/H_m^c = \sqrt{\rho_s^{ab}/\rho_s^c}$ tracks the superfluid-density anisotropy. Such a vortex solid implies a nearly isotropic superfluid density, which is irreconcilable with the strictly two-dimensional $3d_{x^2-y^2}$-derived bands, but naturally explained by a substantial interlayer superfluid contribution from the $3d_{z^2}$-derived $\gamma$ band. Our results provide thermodynamic evidence for a substantial contribution of the $\gamma$ band to superconductivity in bilayer nickelate thin films.

cond-mat.supr-con

Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces. This produces synergistic cost amplification, mechanistically distinct from loop-dependent failures, exhibiting negligible loop incidence in our experiments. Evaluating five open-source VLM families on MS COCO and ImageNet under an 8/255 infinity-norm budget, JPPO achieves over 4.6x latency and 5.3x energy amplification on Qwen2.5-VL-7B, and over 36.6x latency with 32.7x energy amplification on BLIP-2. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities. These findings reveal structural blind spots in current VLM serving defenses, motivating cost-aware robustness evaluation as a first-class security requirement for multimodal deployments.

cs.AI

Remarks on diagonal dimension for algebraic stacks

This note is concerned with the Rouquier dimension of the bounded derived category of coherent complexes on a Noetherian algebraic stack. Specifically, we study the diagonal dimension of a morphism, which can be used to produce upper bounds on Rouquier dimension. First, we obtain an explicit upper bound for smooth morphisms with a regular target. Second, we identify classical generators of a fibre product, recovering a result of Elagin--Lunts--Schn\"{u}rer. Finally, we show that the diagonal dimension of a variety in arbitrary characteristic with mild singularities is at most twice its Krull dimension.

math.AG

$3d_{z^2}$ orbital delocalization and magnetic collapse in superconducting (La,Pr)$_3$Ni$_2$O$_{7-\delta}$ films

The recent discovery of Ruddlesden--Popper (RP) nickelate thin-film superconductors has opened a new frontier in unconventional superconductivity. Its realization requires both compressive epitaxial strain and highly oxidative growth conditions, yet the microscopic pathway from the parent phase to the superconducting phase remains elusive. Here, X-ray absorption spectra and resonant inelastic X-ray scattering are employed to track this evolution by independently tuning strain and oxygen content in (La,Pr)$_3$Ni$_2$O$_{7-\delta}$ thin films. We uncover a remarkable two-step narrative. First, signatures of delocalization emerge in the same way upon two independent tunings: Spectral weight transfers from a ''Upper Hubbard''-like peak to the hole-like peak associated with O $2p_z$ state, and in parallel, the initially localized Ni $3d_{z^2}$ orbital becomes more itinerant followed by the broadening and weakening of $dd$ orbital excitations. Second, as itinerancy increases, long-range spin-density-wave (SDW) order is suppressed in both intensity and correlation length, indicating direct competition with superconductivity. Yet, short-range magnons persist: they become damped but their bandwidth stays unchanged. Our results paint a coherent picture that both strain and oxygenation drive the RP bilayer nickelates towards the superconducting instability, where the O $2p_z$ and Ni $3d_{z^2}$ orbitals become delocalized. Concomitantly, the long-range magnetic order loses coherence and gets suppressed. These findings establish an orbital-selective route to RP nickelate superconductivity, in which the delocalization of the $2p_z$ and $3d_{z^2}$ orbitals and the robust short-range magnons upon the melting of SDW order are prerequisites, providing strong constraints for theory and the roadmap for designing nickelate superconductors.

cond-mat.supr-con

A H.265/HEVC Fine-Grained ROI Video Encryption Algorithm Based on Coding Unit and Prompt Segmentation

ROI (Region of Interest) video selective encryption based on H.265/HEVC is a technology that protects the sensitive regions of videos by perturbing the syntax elements associated with target areas. However, existing methods typically adopt Tile (with a relatively large size) as the minimum encryption unit, which suffers from problems such as inaccurate encryption regions and low encryption precision. This low-precision encryption makes them difficult to apply in sensitive fields such as medicine, military, and remote sensing. In order to address the aforementioned problem, this paper proposes a fine-grained ROI video selective encryption algorithm based on Coding Units (CUs) and prompt segmentation. First, to achieve a more precise ROI acquisition, we present a novel ROI mapping approach based on prompt segmentation. This approach enables precise mapping of ROIs to small $8\times8$ CU levels, significantly enhancing the precision of encrypted regions. Second, we propose a selective encryption scheme based on multiple syntax elements, which distorts syntax elements within high-precision ROI to effectively safeguard ROI security. Finally, we design a diffusion isolation based on Pulse Code Modulation (PCM) mode and MV restriction, applying PCM mode and MV restriction strategy to the affected CU to address encryption diffusion during prediction. The above three strategies break the inherent mechanism of using Tiles in existing ROI encryption and push the fine-grained level of ROI video encryption to the minimum $8\times8$ CU precision. The experimental results demonstrate that the proposed algorithm can accurately segment ROI regions, effectively perturb pixels within these regions, and eliminate the diffusion artifacts introduced by encryption. The method exhibits great potential for application in medical imaging, military surveillance, and remote areas.

eess.IV

A Video Steganography for H.265/HEVC Based on Multiple CU Size and Block Structure Distortion

Video steganography based on block structure, which embeds secret information by modifying Coding Unit (CU) block structure of I-frames, is currently a research hotspot. However, the existing algorithms still suffer from the limitation of poor anti-steganalysis, which results from significantly disrupting the original CU block structure after embedding secret information. To overcome this limitation, this paper proposes a video steganography algorithm based on multiple CU size and block structure distortion. Our algorithm introduces three key innovations: 1) a CU Block Structure Stability Metric (CBSSM) based on CU block structure restoration phenomenon to reveal the reasons for the insufficient anti-steganalysis performance of current algorithms. 2) a novel mapping rule based on multiple CU size to reduce block structure change and enhance embedding capacity. 3) a three-level distortion function based on block structure to better guide the secret information embedding. This triple strategy ensures that the secret information embedding minimizes disruption to the original CU block structure while concealing it primarily in areas where block structure changes occur after recompression, ultimately enhancing the algorithm's anti-steganalysis. Comprehensive experimental results highlight the crucial role of the proposed CBSSM in evaluating anti-steganalysis performance even at a low embedding rate. Meanwhile, compared to State-of-the-Art video steganography algorithms based on block structure, our proposed steganography algorithm exhibits greater anti-steganalysis, as well as further improving visual quality, bitrate increase ratio and embedding capacity.

cs.MM

H.265/HEVC Video Steganalysis Based on CU Block Structure Gradients and IPM Mapping

Existing H.265/HEVC video steganalysis research mainly focuses on detecting the steganography based on motion vectors, intra prediction modes, and transform coefficients. However, there is currently no effective steganalysis method capable of detecting steganography based on Coding Unit (CU) block structure. To address this issue, we propose, for the first time, a H.265/HEVC video steganalysis algorithm based on CU block structure gradients and intra prediction mode mapping. The proposed method first constructs a new gradient map to explicitly describe changes in CU block structure, and combines it with a block level mapping representation of IPM. It can jointly model the structural perturbations introduced by steganography based on CU block structure. Then, we design a novel steganalysis network called GradIPMFormer, whose core innovation is an integrated architecture that combines convolutional local embedding with Transformer-based token modeling to jointly capture local CU boundary perturbations and long-range cross-CU structural dependencies, thereby effectively enhancing the capability to perceive CU block structure embedding. Experimental results show that under different quantization parameters and resolution settings, the proposed method consistently achieves superior detection performance across multiple steganography methods based on CU block structure. This study provides a new CU block structure steganalysis paradigm for H.265/HEVC and has significant research value for covert communication security detection.

eess.IV

Robust, High-Contrast, Recyclable Zinc-Based Dynamic Windows via Synergistic Electrolyte and Interfacial Engineering

Zinc-based electrochromic devices offer a sustainable route for dynamic optical management but are plagued by poor cycling stability due to irreversible zinc plating/stripping and side reactions. Herein, we report a robust, high-contrast, and recyclable zinc-based dynamic window enabled by a synergistic electrolyte and interfacial engineering strategy. Molecular dynamics simulations and electrochemical analyses reveal a dual-ion cooperative mechanism that governs the reversibility: anions with the strongest binding affinity guide uniform Zn deposition by stabilizing the inner solvation shell, while formate anions co-enriched at the interface facilitate smooth stripping via protonation during the oxidation process. This orchestrated interplay effectively eliminates "dead Zn" accumulation and dendrite growth. Consequently, the device demonstrates a record-high lifespan of 15,000 cycles with negligible degradation and maintains a large optical modulation of >50%, along with multiple optical states (transparent, gray, black, and mirror). Furthermore, it achieves a large reflectance modulation (>50%) stable for over 2,000 cycles. This work establishes the recyclable zinc-based dynamic window as a scalable, high-performance alternative to conventional electrochromic systems, advancing the feasibility of solution-processed energy-saving windows in sustainable buildings.

cond-mat.mtrl-sci

Frobenius generation for algebraic stacks

We introduce a notion of $F$-finiteness for algebraic stacks in positive characteristic. Our main result shows that sufficiently many Frobenius pushforwards generate the bounded derived categories of coherent sheaves on Noetherian concentrated $F$-finite algebraic stacks with quasi-finite and separated diagonal. This generalizes, and independently recovers, a result of Ballard--Iyengar--Lank--Mukhopadhyay--Pollitz.

math.AG

Superconductivity onset above 60 K in ambient-pressure nickelate films

Ambient-pressure superconductivity in nickelates has been capped at an onset transition temperature ($T_{c}^{onset}$) of ~50 K, a value that remains lower than the cuprate (~133 K) and iron-based (~55 K) counterparts, despite the promise shown under high pressure. Here, we report ambient-pressure superconductivity onset at ~63 K in epitaxial (La,Pr)3Ni2O7 thin films grown under compressive strain on SrLaAlO4 substrates. This $T_{c}$ leap is enabled by pushing our gigantic-oxidative atomic-layer-by-layer epitaxy (GAE) method into an extreme non-equilibrium growth regime. It simultaneously enhances kinetics via higher temperatures and achieves full oxygenation in situ without post-annealing. Synchrotron X-ray diffraction and scanning transmission electron microscopy confirm that this approach yields films of large-scale crystalline purity, overcoming the inherent metastability of the strained superconducting phase. Transport measurements reveal a zero-resistance temperature ($T_{c}^{zero}$) reaching ~37 K, while mutual inductance measurements demonstrate a robust diamagnetic transition starting at ~23 K. These films exhibit a systematic evolution in their normal-state resistivity-temperature curve: the power-law exponent $\alpha$ evolves from Fermi-liquid-like ($\alpha$ ~2) at lower $T_{c}^{onset}$ to strange-metal-like ($\alpha$ ~1) in higher $T_{c}^{onset}$ samples, directly linking the enhanced superconductivity to non-Fermi liquid behavior. Mapping the vortex melting phase diagram by the mutual inductance technique further reveals 2D melting limit suppressed to near zero, which demonstrates significantly stronger interlayer coupling than that of cuprates. These results identify the nickelates as an ambient-pressure strange-metal high-temperature superconductors with strong interlayer coupling.

cond-mat.supr-con

A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption

ROI selective encryption, as an efficient privacy protection technique, encrypts only the key regions in the video, thereby ensuring security while minimizing the impact on coding efficiency. However, existing ROI-based video encryption methods suffer from insufficient flexibility and lack of a unified evaluation system. To address these issues, we propose a visual perception-based tunable framework and evaluation benchmark for H.265/HEVC ROI encryption. Our scheme introduces three key contributions: 1) A ROI region recognition module based on visual perception network is proposed to accurately identify the ROI region in videos. 2) A three-level tunable encryption strategy is implemented while balancing security and real-time performance. 3) A unified ROI encryption evaluation benchmark is developed to provide a standardized quantitative platform for subsequent research. This triple strategy provides new solution and significant unified performance evaluation methods for ROI selective encryption field. Experimental results indicate that the proposed benchmark can comprehensively measure the performance of the ROI selective encryption. Compared to existing ROI encryption algorithms, our proposed enhanced and advanced level encryption exhibit superior performance in multiple performance metrics. In general, the proposed framework effectively meets the privacy protection requirements in H.265/HEVC and provides a reliable solution for secure and efficient processing of sensitive video content.

eess.IV

MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion

Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a significant challenge. In this work, we address the task of layout-controllable multi-subject synthesis (LMS), which requires both faithful reconstruction of reference subjects and their accurate placement in specified regions within a unified image. While recent advancements have separately improved layout control and subject synthesis, existing approaches struggle to simultaneously satisfy the dual requirements of spatial precision and identity preservation in this composite task. To bridge this gap, we propose MUSE, a unified synthesis framework that employs concatenated cross-attention (CCA) to seamlessly integrate layout specifications with textual guidance through explicit semantic space expansion. The proposed CCA mechanism enables bidirectional modality alignment between spatial constraints and textual descriptions without interference. Furthermore, we design a progressive two-stage training strategy that decomposes the LMS task into learnable sub-objectives for effective optimization. Extensive experiments demonstrate that MUSE achieves zero-shot end-to-end generation with superior spatial accuracy and identity consistency compared to existing solutions, advancing the frontier of controllable image synthesis. Our code and model are available at https://github.com/pf0607/MUSE.

cs.CV

Compact approximation and descent for algebraic stacks

This work focuses on approximation and generation for the derived category of complexes with quasi-coherent cohomology on algebraic stacks. Our methods establish that approximation by compact objects descends along covers that are quasi-finite and flat. This generalizes a result of Lipman--Neeman for schemes and extends a related result known for algebraic spaces. We also study the behavior of generation under the derived pushforward and pullback of a morphism between algebraic stacks.

math.AG

Categorical characterizations of regularity for algebraic stacks

For a Noetherian scheme $X$ of finite Krull dimension, Neeman recently established two characterizations of the regularity of $X$ using strong generators and bounded $t$-structures on $\operatorname{Perf}(X)$. In this note, we obtain variants of Neeman's results for large classes of Noetherian algebraic stacks. An important intermediate step is the fact that $X$ is regular if and only if $\operatorname{Perf}(X)=D_{\operatorname{coh}}^b(X)$, which we establish for Noetherian algebraic stacks. Our approach also yields a criterion for the existence of classical generators for the bounded derived categories of coherent sheaves on algebraic stacks, generalizing previous results for commutative rings and schemes.

math.AG

Orlov's theorem over a quasiexcellent ring

Following the approach of Kawamata and Canonaco-Stellari, we establish Orlov's representability theorem for smooth tame Deligne-Mumford stacks with projective coarse moduli spaces over a quasiexcellent ring of finite Krull dimension. This generalizes a previous result of Canonaco-Stellari for smooth projective varieties over a field.

math.AG

Group-Level Data Selection for Efficient Pretraining

In this paper, we introduce Group-MATES, an efficient group-level data selection approach to optimize the speed-quality frontier of language model pretraining. Specifically, Group-MATES parameterizes costly group-level selection with a relational data influence model. To train this model, we sample training trajectories of the language model and collect oracle data influences alongside. The relational data influence model approximates the oracle data influence by weighting individual influence with relationships among training data. To enable efficient selection with our relational data influence model, we partition the dataset into small clusters using relationship weights and select data within each cluster independently. Experiments on DCLM 400M-4x, 1B-1x, and 3B-1x show that Group-MATES achieves 3.5%-9.4% relative performance gains over random selection across 22 downstream tasks, nearly doubling the improvements achieved by state-of-the-art individual data selection baselines. Furthermore, Group-MATES reduces the number of tokens required to reach a certain downstream performance by up to 1.75x, substantially elevating the speed-quality frontier. Further analyses highlight the critical role of relationship weights in the relational data influence model and the effectiveness of our cluster-based inference. Our code is open-sourced at https://github.com/facebookresearch/Group-MATES.

cs.CL

A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation

Conventional 2D human pose estimation methods typically require extensive labeled annotations, which are both labor-intensive and expensive. In contrast, semi-supervised 2D human pose estimation can alleviate the above problems by leveraging a large amount of unlabeled data along with a small portion of labeled data. Existing semi-supervised 2D human pose estimation methods update the network through backpropagation, ignoring crucial historical information from the previous training process. Therefore, we propose a novel semi-supervised 2D human pose estimation method by utilizing a newly designed Teacher-Reviewer-Student framework. Specifically, we first mimic the phenomenon that human beings constantly review previous knowledge for consolidation to design our framework, in which the teacher predicts results to guide the student's learning and the reviewer stores important historical parameters to provide additional supervision signals. Secondly, we introduce a Multi-level Feature Learning strategy, which utilizes the outputs from different stages of the backbone to estimate the heatmap to guide network training, enriching the supervisory information while effectively capturing keypoint relationships. Finally, we design a data augmentation strategy, i.e., Keypoint-Mix, to perturb pose information by mixing different keypoints, thus enhancing the network's ability to discern keypoints. Extensive experiments on publicly available datasets, demonstrate our method achieves significant improvements compared to the existing methods.

cs.CV

An improved bound on Seymour's second neighborhood conjecture

Seymour's celebrated second neighborhood conjecture, now more than thirty years old, states that in every oriented digraph, there is a vertex $u$ such that the size of its second out-neighborhood $N^{++}(u)$ is at least as large as that of its first out-neighborhood $N^+(u)$. In this paper, we prove the existence of $u$ for which $|N^{++}(u)| \ge 0.715538 |N^+(u)|$. This result provides the first improvement to the best known constant factor in over two decades.

math.CO