SearcharxivSearch

arXiv subjects

Jianbo Zhang

Publications and source records attributed to Jianbo Zhang.

At least 19 recordsLinked to original sources

Radio-Frequency Method for Detecting Superconductivity Under High Pressure

We introduce a contactless technique for probing superconductivity and magnetic ordering transitions in micron-sized samples under extreme pressure. Utilizing a multistage Lenz lens system, directly sputtered onto diamond anvils, we realize a radio-frequency (RF, 50 kHz - 200 MHz) transformer with a sample of 50-100 $μ$m in diameter, as its core. This configuration enables efficient transfer and focusing of an electromagnetic field within the diamond anvil cell's chamber. Consequently, the transmitted RF signal exhibits high sensitivity to variations in the sample's surface conductivity and magnetic permeability. We validate this method by determining the critical temperatures ($T_{\text{c}}$) of known superconductors, including NbTi, MgB$_2$, Hg-1223, Bi-2212, YBCO, and REBCO in various magnetic fields, as well as the magnetic ordering temperatures of Gd and Tb. Notably, we apply this technique to the LaH$_{10-x}$, CeH$_{9-10}$, and (La,Ce)H$_{10-12}$ superhydrides at a pressure of about 1-1.5 Mbar. The observed superconducting transitions in Ce and La superhydrides at 90-110 K and 215-242 K, respectively, correlate with the $T_{\text{c}}$'s determined via traditional electrical-resistance measurements. Moreover, we show how multiple repetitions of the RF experiment with the La-Ce superhydride make it possible to detect the increase in $T_{\text{c}}$ over time up to $\approx$ 260-270 K. This finding indicates the possibility of reaching a critical $T_{\text{c}}$ around 0$^\circ$C in the La-based superhydrides.

cond-mat.supr-con

Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach

Recent fine-tuning techniques for diffusion models enable them to reproduce specific image sets, such as particular faces or artistic styles, but also introduce copyright and security risks. Dataset watermarking has been proposed to ensure traceability by embedding imperceptible watermarks into training images, which remain detectable in outputs even after fine-tuning. However, current methods lack a unified evaluation framework. To address this, this paper establishes a general threat model and introduces a comprehensive evaluation framework encompassing Universality, Transmissibility, and Robustness. Experiments show that existing methods perform well in universality and transmissibility, and exhibit some robustness against common image processing operations, yet still fall short under real-world threat scenarios. To reveal these vulnerabilities, the paper further proposes a practical watermark removal method that fully eliminates dataset watermarks without affecting fine-tuning, highlighting a key challenge for future research.

cs.CV

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.

cs.AI

A$^3$: Towards Advertising Aesthetic Assessment

Advertising images significantly impact commercial conversion rates and brand equity, yet current evaluation methods rely on subjective judgments, lacking scalability, standardized criteria, and interpretability. To address these challenges, we present A^3 (Advertising Aesthetic Assessment), a comprehensive framework encompassing four components: a paradigm (A^3-Law), a dataset (A^3-Dataset), a multimodal large language model (A^3-Align), and a benchmark (A^3-Bench). Central to A^3 is a theory-driven paradigm, A^3-Law, comprising three hierarchical stages: (1) Perceptual Attention, evaluating perceptual image signals for their ability to attract attention; (2) Formal Interest, assessing formal composition of image color and spatial layout in evoking interest; and (3) Desire Impact, measuring desire evocation from images and their persuasive impact. Building on A^3-Law, we construct A^3-Dataset with 120K instruction-response pairs from 30K advertising images, each richly annotated with multi-dimensional labels and Chain-of-Thought (CoT) rationales. We further develop A^3-Align, trained under A^3-Law with CoT-guided learning on A^3-Dataset. Extensive experiments on A^3-Bench demonstrate that A^3-Align achieves superior alignment with A^3-Law compared to existing models, and this alignment generalizes well to quality advertisement selection and prescriptive advertisement critique, indicating its potential for broader deployment. Dataset, code, and models can be found at: https://github.com/euleryuan/A3-Align.

cs.CV

Pressure Induced 18 K Superconductivity and Two Superconducting Phases in CuIr2S4

We report pressure-induced superconductivity in the spinel CuIr$_{2}$S$_{4}$ with a transition temperature ($T_{\text{c}}$) reaching \textbf{18.2 K}, establishing a new record for this class of materials and surpassing the decades-old limit of 13.7 K. Our electrical transport and synchrotron X-ray diffraction studies up to 224 GPa reveal the emergence of \textbf{two distinct superconducting phases} from a charge-ordered insulating state. The first phase (SC-I) appears around 18 GPa, and forms a dome-shaped superconducting region in which the resistivity exhibits a pronounced, field- and current-sensitive drop without reaching strict zero above our base temperature. Above 111.8 GPa, a second, lower-$T_{\text{c}}$ phase (SC-II) emerges and coexists with SC-I over a broad pressure range, and SC-II ultimately develops a true zero-resistance state above 122.2 GPa. These superconducting phases are intimately linked to a cascade of structural transitions that systematically distort the frustrated pyrochlore lattice of Ir atoms. Our results expand the potential for superconductivity in spinels and demonstrate a pathway to high-$T_{\text{c}}$ pairing directly from a correlated insulating state driven by lattice tuning.

cond-mat.supr-con

Embodied Image Compression

Image Compression for Machines (ICM) has emerged as a pivotal research direction in the field of visual data compression. However, with the rapid evolution of machine intelligence, the target of compression has shifted from task-specific virtual models to Embodied agents operating in real-world environments. To address the communication constraints of Embodied AI in multi-agent systems and ensure real-time task execution, this paper introduces, for the first time, the scientific problem of Embodied Image Compression. We establish a standardized benchmark, EmbodiedComp, to facilitate systematic evaluation under ultra-low bitrate conditions in a closed-loop setting. Through extensive empirical studies in both simulated and real-world settings, we demonstrate that existing Vision-Language-Action models (VLAs) fail to reliably perform even simple manipulation tasks when compressed below the Embodied bitrate threshold. We anticipate that EmbodiedComp will catalyze the development of domain-specific compression tailored for Embodied agents , thereby accelerating the Embodied AI deployment in the Real-world.

cs.CV

Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation

With the rapid progress in diffusion models, image synthesis has advanced to the stage of zero-shot image-to-image generation, where high-fidelity replication of facial identities or artistic styles can be achieved using just one portrait or artwork, without modifying any model weights. Although these techniques significantly enhance creative possibilities, they also pose substantial risks related to intellectual property violations, including unauthorized identity cloning and stylistic imitation. To counter such threats, this work presents Adapter Shield, the first universal and authentication-integrated solution aimed at defending personal images from misuse in zero-shot generation scenarios. We first investigate how current zero-shot methods employ image encoders to extract embeddings from input images, which are subsequently fed into the UNet of diffusion models through cross-attention layers. Inspired by this mechanism, we construct a reversible encryption system that maps original embeddings into distinct encrypted representations according to different secret keys. The authorized users can restore the authentic embeddings via a decryption module and the correct key, enabling normal usage for authorized generation tasks. For protection purposes, we design a multi-target adversarial perturbation method that actively shifts the original embeddings toward designated encrypted patterns. Consequently, protected images are embedded with a defensive layer that ensures unauthorized users can only produce distorted or encrypted outputs. Extensive evaluations demonstrate that our method surpasses existing state-of-the-art defenses in blocking unauthorized zero-shot image synthesis, while supporting flexible and secure access control for verified users.

cs.CV

Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling

Blind image quality assessment (BIQA) plays a crucial role in evaluating and optimizing visual experience. Most existing BIQA approaches fuse shallow and deep features extracted from backbone networks, while overlooking the unequal contributions to quality prediction. Moreover, while various vision encoder backbones are widely adopted in BIQA, the effective quality decoding architectures remain underexplored. To address these limitations, this paper investigates the contributions of shallow and deep features to BIQA, and proposes a effective quality feature decoding framework via GCN-enhanced \underline{l}ayer\underline{i}nteraction and MoE-based \underline{f}eature d\underline{e}coupling, termed \textbf{(Life-IQA)}. Specifically, the GCN-enhanced layer interaction module utilizes the GCN-enhanced deepest-layer features as query and the penultimate-layer features as key, value, then performs cross-attention to achieve feature interaction. Moreover, a MoE-based feature decoupling module is proposed to decouple fused representations though different experts specialized for specific distortion types or quality dimensions. Extensive experiments demonstrate that Life-IQA shows more favorable balance between accuracy and cost than a vanilla Transformer decoder and achieves state-of-the-art performance on multiple BIQA benchmarks.The code is available at: \href{https://github.com/TANGLONG2/Life-IQA/tree/main}{\texttt{Life-IQA}}.

cs.CV

Data Assessment for Embodied Intelligence

In embodied intelligence, datasets play a pivotal role, serving as both a knowledge repository and a conduit for information transfer. The two most critical attributes of a dataset are the amount of information it provides and how easily this information can be learned by models. However, the multimodal nature of embodied data makes evaluating these properties particularly challenging. Prior work has largely focused on diversity, typically counting tasks and scenes or evaluating isolated modalities, which fails to provide a comprehensive picture of dataset diversity. On the other hand, the learnability of datasets has received little attention and is usually assessed post-hoc through model training, an expensive, time-consuming process that also lacks interpretability, offering little guidance on how to improve a dataset. In this work, we address both challenges by introducing two principled, data-driven tools. First, we construct a unified multimodal representation for each data sample and, based on it, propose diversity entropy, a continuous measure that characterizes the amount of information contained in a dataset. Second, we introduce the first interpretable, data-driven algorithm to efficiently quantify dataset learnability without training, enabling researchers to assess a dataset's learnability immediately upon its release. We validate our algorithm on both simulated and real-world embodied datasets, demonstrating that it yields faithful, actionable insights that enable researchers to jointly improve diversity and learnability. We hope this work provides a foundation for designing higher-quality datasets that advance the development of embodied intelligence.

cs.RO

Image Quality Assessment for Embodied AI

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to predict human preferences for distorted images; however, there is no IQA method to assess the usability of an image in embodied tasks, namely, the perceptual quality for robots. To provide accurate and reliable quality indicators for future embodied scenarios, we first propose the topic: IQA for Embodied AI. Specifically, we (1) based on the Mertonian system and meta-cognitive theory, constructed a perception-cognition-decision-execution pipeline and defined a comprehensive subjective score collection process; (2) established the Embodied-IQA database, containing over 36k reference/distorted image pairs, with more than 5m fine-grained annotations provided by Vision Language Models/Vision Language Action-models/Real-world robots; (3) trained and validated the performance of mainstream IQA methods on Embodied-IQA, demonstrating the need to develop more accurate quality indicators for Embodied AI. We sincerely hope that through evaluation, we can promote the application of Embodied AI under complex distortions in the Real-world. Project page: https://github.com/lcysyzxdxc/EmbodiedIQA

cs.CV

Embodied Image Quality Assessment for Robotic Intelligence

Image Quality Assessment (IQA) of User-Generated Content (UGC) is a critical technique for human Quality of Experience (QoE). However, does the the image quality of Robot-Generated Content (RGC) demonstrate traits consistent with the Moravec paradox, potentially conflicting with human perceptual norms? Human subjective scoring is more based on the attractiveness of the image. Embodied agent are required to interact and perceive in the environment, and finally perform specific tasks. Visual images as inputs directly influence downstream tasks. In this paper, we explore the perception mechanism of embodied robots for image quality. We propose the first Embodied Preference Database (EPD), which contains 12,500 distorted image annotations. We establish assessment metrics based on the downstream tasks of robot. In addition, there is a gap between UGC and RGC. To address this, we propose a novel Multi-scale Attention Embodied Image Quality Assessment called MA-EIQA. For the proposed EPD dataset, this is the first no-reference IQA model designed for embodied robot. Finally, the performance of mainstream IQA algorithms on EPD dataset is verified. The experiments demonstrate that quality assessment of embodied images is different from that of humans. We sincerely hope that the EPD can contribute to the development of embodied AI by focusing on image quality assessment. The benchmark is available at https://github.com/Jianbo-maker/EPD_benchmark.

cs.CV

Static and Plugged: Make Embodied Evaluation Simple

Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.

cs.CV

Superconducting susceptibility signal captured in a record wide pressure range

In recent years, the resistance signature of the high temperature superconductivity above 250 K in highly compressed hydrides (more than 100 GPa) has garnered significant attention within the condensed matter physics community. This has sparked renewed optimism for achieving superconductivity under room-temperature conditions. However, the superconducting diamagnetism, another crucial property for confirming the superconductivity, has yet to be conclusively observed. The primary challenge arises from the weak diamagnetic signals detected from the small samples compressed in diamond anvil cells. Therefore, the reported results of superconducting diamagnetism in hydrides have sparked intense debate, highlighting the urgent need for workable methodology to assess the validity of the experimental results. Here, we are the first to report the ultrahigh pressure measurements of the superconducting diamagnetism on Nb0.44Ti0.56, a commercial superconducting alloy, in a record-wide pressure range from 5 GPa to 160 GPa. We present detailed results on factors such as sample size, the diamagnetic signal intensity, the signal-to-noise ratio and the superconducting transition temperature across various pressures and different pressure transmitting media. These comprehensive results clearly demonstrate that this alloy is an ideal reference sample for evaluating superconductivity in compressed hydrides,validating the credibility of the experimental systems and superconducting diamagnetic results, as well as determining the nature of the superconductivity of the investigated sample. In addition, these results also provide a valuable benchmark for studying the pressure-induced superconductivity in other material families.

cond-mat.supr-con

R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

The outstanding performance of Large Multimodal Models (LMMs) has made them widely applied in vision-related tasks. However, various corruptions in the real world mean that images will not be as ideal as in simulations, presenting significant challenges for the practical application of LMMs. To address this issue, we introduce R-Bench, a benchmark focused on the **Real-world Robustness of LMMs**. Specifically, we: (a) model the complete link from user capture to LMMs reception, comprising 33 corruption dimensions, including 7 steps according to the corruption sequence, and 7 groups based on low-level attributes; (b) collect reference/distorted image dataset before/after corruption, including 2,970 question-answer pairs with human labeling; (c) propose comprehensive evaluation for absolute/relative robustness and benchmark 20 mainstream LMMs. Results show that while LMMs can correctly handle the original reference images, their performance is not stable when faced with distorted images, and there is a significant gap in robustness compared to the human visual system. We hope that R-Bench will inspire improving the robustness of LMMs, **extending them from experimental simulations to the real-world application**. Check https://q-future.github.io/R-Bench for details.

cs.CV

Experimental evidence of crystal-field, Zeeman splitting, and spin-phonon excitations in the quantum supersolid Na2BaCo(PO4)2

Drawing inspiration from the recent breakthroughs in the \ce{Na_{2}BaCo(PO_{4})_{2}} quantum magnet, renowned for its spin supersolidity phase and its potential for revolutionary cooling applications, our study delves into the intricate interplay among lattice, spin, and orbital degrees of freedom within this intriguing compound. Using meticulous temperature, field, and pressure-dependent Raman scattering techniques, we present compelling experimental evidence revealing pronounced crystal-electric field (CEF) excitations, alongside the interplay of CEF-phonon interactions. Notably, our experiments elucidate all electronic transitions from $j_{1 / 2}$ to $j_{3 / 2}$ and from $j_{1 / 2}$ to $j_{5 / 2}$, with energy level patterns closely aligned with theoretical predictions based on point-charge models. Furthermore, the application of a magnetic field and pressure reveals Zeeman splittings characterized by Landé-g factors as well as the CEF-phonon resonances. The anomalous shift in coupled peak at low temperatures originates from the hybridization of CEF and phonon excitations due to their close energy proximity. These findings constitute a significant step towards unraveling the fundamental properties of this exotic quantum material for future research in fundamental physics or engineering application.

cond-mat.mtrl-sci

Study of topological quantities of lattice QCD with a modified Wasserstein generative adversarial network

We propose a modified Wasserstein generative adversarial network (M-WGAN) to study the distribution of the topological charge in lattice QCD based on Monte Carlo simulations. We construct new generator and discriminator in M-WGAN to support the generation of high-quality distribution. Our results show that the M-WGAN scheme of machine learning should be helpful for us to calculate efficiently the 1D distribution of topological charge compared with the method by the MC simulation alone.

hep-lat

A study of topological quantities of lattice QCD by a modified DCGAN frame

A modified deep convolutional generative adversarial network (M-DCGAN) frame is proposed to study the N-dimensional (ND) topological quantities in lattice QCD based on the Monte Carlo (MC) simulations. We construct a new scaling structure including fully connected layers to support the generation of high-quality high-dimensional images for the M-DCGAN. Our results show that the M-DCGAN scheme of the Machine learning should be helpful for us to calculate efficiently the 1D distribution of topological charge and the 4D topological charge density compared with the case by the MC simulation alone.

hep-lat

Unveiling a Novel Metal-to-Metal Transition in LuH2: Critically Challenging Superconductivity Claims in Lutetium Hydrides

Following the recent report by Dasenbrock-Gammon et al. (2023) of near-ambient superconductivity in nitrogen-doped lutetium trihydride (LuH3-δNε), significant debate has emerged surrounding the composition and interpretation of the observed sharp resistance drop. Here, we meticulously revisit these claims through comprehensive characterization and investigations. We definitively identify the reported material as lutetium dihydride (LuH2), resolving the ambiguity surrounding its composition. Under similar conditions (270-295 K and 1-2 GPa), we replicate the reported sharp decrease in electrical resistance with a 30% success rate, aligning with Dasenbrock-Gammon et al.'s observations. However, our extensive investigations reveal this phenomenon to be a novel, pressure-induced metal-to-metal transition intrinsic to LuH2, distinct from superconductivity. Intriguingly, nitrogen doping exerts minimal impact on this transition. Our work not only elucidates the fundamental properties of LuH2 and LuH3 but also critically challenges the notion of superconductivity in these lutetium hydride systems. These findings pave the way for future research on lutetium hydride systems while emphasizing the crucial importance of rigorous verification in claims of ambient temperature superconductivity.

cond-mat.supr-con