SearcharxivSearch

arXiv subjects

Zhi Zeng

Publications and source records attributed to Zhi Zeng.

At least 19 recordsLinked to original sources

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

cs.AI

From Patches to Evidence Balls: Class-Conditioned Evidence Retrieval for Few-Shot Whole Slide Image Classification

Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Existing MIL and vision-language methods aggregate a large pool of patch features into a single global slide representation. Under few-shot supervision, limited slide-level labels make it difficult to learn a reliable aggregation mechanism that organizes sparse local cues into compact and coherent diagnostic evidence. Moreover, a shared slide representation compresses evidence supporting a candidate class and its alternatives into the same feature, limiting class-specific reasoning and interpretability. To address these issues, we propose EviBall, a class-conditioned evidence retrieval framework for few-shot WSI classification. EviBall organizes local patches into Evidence Balls through semantic-spatial assignment and center refinement, yielding compact and spatially coherent evidence units under weak supervision. It then uses task-specific class queries, including language-guided queries for morphology-oriented tasks and molecular-guided queries for molecular endpoint prediction, to retrieve supporting evidence balls and produce class-conditioned evidence representations for direct class-wise prediction. By introducing structured evidence units and task-relevant semantic guidance, EviBall reduces the reliance on learning an unconstrained global aggregation mechanism from scarce slide-level labels. It therefore reformulates few-shot WSI classification as structured evidence retrieval and competition among candidate classes. Extensive experiments across four morphology-oriented and molecular endpoint WSI tasks demonstrate that EviBall consistently outperforms conventional and vision-language MIL baselines under diverse few-shot settings, while providing spatially localized and class-specific evidence for each prediction.

cs.CV

Hydrostatic Pressure-Induced Evolution of the Superconducting Transition Temperature of Bi-2212: Insights from First-Principles Calculations

High-pressure experiments on Bi$_2$Sr$_2$CaCu$_2$O$_{8+x}$ (Bi-2212) have reported apparently conflicting evolutions of the superconducting transition temperature $T_c$, ranging from weak enhancement to strong suppression and even a proposed second superconducting dome. To clarify the origin of these discrepancies, we combine first-principles density functional theory calculations with a pressure-dependent low-energy bilayer model solved by the slave-boson mean-field method together with a Berezinskii-Kosterlitz-Thouless estimate of phase coherence. Our results show that hydrostatic pressure induces a pronounced self-doping effect in Bi-2212: holes are transferred from the Bi-O charge-reservoir layers to the CuO$_2$ superconducting planes, leading to a systematic increase in the effective CuO$_2$-plane hole concentration $\delta_x$. At the same time, pressure enhances the pairing scale through the renormalization of the hopping and superexchange parameters. As a consequence, the pressure evolution of $T_c$ is governed by the competition between pressure-enhanced pairing and pressure-driven motion along the common $T_c$-$\delta_x$ dome, making $T_c(P)$ highly sensitive to the initial doping state. Even samples with very similar ambient-pressure $T_c$ but slightly different initial doping can therefore display qualitatively different pressure responses. This provides a unified interpretation of a large part of the disparate high-pressure behavior reported for Bi-2212 and suggests that slightly underdoped samples are more favorable than ambient-pressure optimal samples for achieving improved superconducting performance under pressure.

cond-mat.supr-con

M3D-Stereo: A Multiple-Medium and Multiple-Degradation Dataset for Stereo Image Restoration

Image restoration under adverse conditions, such as underwater, haze or fog, and low-light environments, remains a highly challenging problem due to complex physical degradations and severe information loss. Existing datasets are predominantly limited to a single degradation type or heavily rely on synthetic data without stereo consistency, inherently restricting their applicability in real-world scenarios. To address this, we introduce M3D-Stereo, a stereo dataset with 7904 high-resolution image pairs for image restoration research acquired in multiple media with multiple controlled degradation levels. It encompasses four degradation scenarios: underwater scatter, haze/fog, underwater low-light, and haze low-light. Each scenario forms a subset, and is divided into six levels of progressive degradation, allowing fine-grained evaluations of restoration methods with increasing severity of degradation. Collected via a laboratory setup, the dataset provides aligned stereo image pairs along with their pixel-wise consistent clear ground truths. Two restoration tasks, single-level and mixed-level degradation, were performed to verify its validity. M3D-Stereo establishes a better controlled and more realistic benchmark to evaluate image restoration and stereo matching methods in complex degradation environments. It is made public under LGPLv3 license.

cs.CV

From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the Wild

The rise of micro-videos has reshaped how misinformation spreads, amplifying its speed, reach, and impact on public trust. Existing benchmarks typically focus on a single deception type, overlooking the diversity of real-world cases that involve multimodal manipulation, AI-generated content, cognitive bias, and out-of-context reuse. Meanwhile, most detection models lack fine-grained attribution, limiting interpretability and practical utility. To address these gaps, we introduce WildFakeBench, a large-scale benchmark of over 10,000 real-world micro-videos covering diverse misinformation types and sources, each annotated with expert-defined attribution labels. Building on this foundation, we develop FakeAgent, a Delphi-inspired multi-agent reasoning framework that integrates multimodal understanding with external evidence for attribution-grounded analysis. FakeAgent jointly analyzes content and retrieved evidence to identify manipulation, recognize cognitive and AI-generated patterns, and detect out-of-context misinformation. Extensive experiments show that FakeAgent consistently outperforms existing MLLMs across all misinformation types, while WildFakeBench provides a realistic and challenging testbed for advancing explainable micro-video misinformation detection. Data and code are available at: https://github.com/Aiyistan/FakeAgent.

cs.SI

Physics-Informed Spatial-Temporal Transformer for Terahertz Near-Field Beam Tracking

Terahertz (THz) ultra-massive multiple-input multiple-output (UM-MIMO) promises ultra-high throughput, while its highly directional beams demand rapid and accurate beam tracking driven by precise user-state estimation. Moreover, large array apertures at high frequencies induce near-field propagation effects, where far-field modeling becomes inaccurate and near-field parametric channel estimation is costly. Bypassing near-field codebook, PAST-TT is proposed to bridge near-field tracking with low-overhead far-field codebook probing by exploiting parallax, amplified by widely spaced subarrays. With comb-type frequency-division multiplexing pilots, each subarray yields frequency-affine phase signatures whose frequency and temporal increments encode propagation delay and its variation between frames. Building on these signatures, a Parallax-Aware Spatial Transformer (PAST) compresses them and outputs per-frame position estimates with token reliability to downweight bad frames, regularized by a physics-in-the-loop consistency loss. A causal Temporal Transformer (TT) then performs reliability-aware filtering and prediction over a sliding window to initialize the beam of the next frame. Acting on short token sequences, PAST-TT avoids a monolithic spatial-temporal network over raw pilots, which keeps the model lightweight with a critical path latency of 0.61 ms. Simulations show that at 15 dB signal-to-noise ratio, PAST achieves 7.81 mm distance RMSE and 0.0588{\deg} angle RMSE. Even with a bad-frame rate of 0.1, TT reduces the distance and angle prediction RMSE by 23.1% and 32.8% compared with the best competing tracker.

eess.SP

Simulation Study on the Discrimination of $0\nu\beta\beta$ Events from Single-Electron Events Using Orthogonal-Strip HPGe Detectors

Neutrinoless double beta decay ($0\nu\beta\beta$) offers a sensitive probe of neutrino mass and its Majorana nature. Orthogonal-strip high-purity germanium (HPGe) detectors with high spatial resolution provide a promising approach for distinguishing $0\nu\beta\beta$ events from single-electron backgrounds. In this work, a simulation framework was developed to evaluate the discrimination performance of these detectors. The framework combined Geant4 simulations with a hybrid numerical-analytical approach to model charge cloud dynamics. A dual-branch convolutional neural network (CNN) was implemented to extract topological features for event classification. The impact of detector geometry on discrimination performance was quantitatively assessed. For a fixed crystal thickness of 15 mm, the background rejection efficiency decreased from 79.5\% to 59.0\% as the strip pitch increased from 0.1 mm to 0.5 mm. For a strip pitch of 0.25 mm, a crystal thickness of 20 mm was found to be optimal, balancing full-energy peak (FEP) efficiency with discrimination capability. These results demonstrate that orthogonal-strip HPGe detectors can effectively suppress single-electron backgrounds, and provide quantitative guidance for detector design in $^{76}$Ge $0\nu\beta\beta$ decay searches.

physics.ins-det

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

Foundation models have recently achieved impressive success in computational pathology, demonstrating strong generalization across diverse histopathology tasks. However, existing models overlook the heterogeneous and non-uniform organization of pathological regions of interest (ROIs) because they rely on natural image backbones not tailored for tissue morphology. Consequently, they often fail to capture the coherent tissue architecture beyond isolated patches, limiting interpretability and clinical relevance. To address these challenges, we present Cross-modal Adaptive Region Encoder (CARE), a foundation model for pathology that automatically partitions WSIs into several morphologically relevant regions. Specifically, CARE employs a two-stage pretraining strategy: (1) a self-supervised unimodal pretraining stage that learns morphological representations from 34,277 whole-slide images (WSIs) without segmentation annotations, and (2) a cross-modal alignment stage that leverages RNA and protein profiles to refine the construction and representation of adaptive regions. This molecular guidance enables CARE to identify biologically relevant patterns and generate irregular yet coherent tissue regions, selecting the most representative area as ROI. CARE supports a broad range of pathology-related tasks, using either the ROI feature or the slide-level feature obtained by aggregating adaptive regions. Based on only one-tenth of the pretraining data typically used by mainstream foundation models, CARE achieves superior average performance across 33 downstream benchmarks, including morphological classification, molecular prediction, and survival analysis, and outperforms other foundation model baselines overall.

cs.CV

The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence

To reliably assist human decision-making, LLMs must maintain factual internal beliefs against misleading injections. While current models resist explicit misinformation, we uncover a fundamental vulnerability to sophisticated, hard-to-falsify evidence. To systematically probe this weakness, we introduce MisBelief, a framework that generates misleading evidence via collaborative, multi-round interactions among multi-role LLMs. This process mimics subtle, defeasible reasoning and progressive refinement to create logically persuasive yet factually deceptive claims. Using MisBelief, we generate 4,800 instances across three difficulty levels to evaluate 7 representative LLMs. Results indicate that while models are robust to direct misinformation, they are highly sensitive to this refined evidence: belief scores in falsehoods increase by an average of 93.0\%, fundamentally compromising downstream recommendations. To address this, we propose Deceptive Intent Shielding (DIS), a governance mechanism that provides an early warning signal by inferring the deceptive intent behind evidence. Empirical results demonstrate that DIS consistently mitigates belief shifts and promotes more cautious evidence evaluation.

cs.CL

MT-Mark: Rethinking Image Watermarking via Mutual-Teacher Collaboration with Adaptive Feature Modulation

Existing deep image watermarking methods follow a fixed embedding-distortion-extraction pipeline, where the embedder and extractor are weakly coupled through a final loss and optimized in isolation. This design lacks explicit collaboration, leaving no structured mechanism for the embedder to incorporate decoding-aware cues or for the extractor to guide embedding during training. To address this architectural limitation, we rethink deep image watermarking by reformulating embedding and extraction as explicitly collaborative components. To realize this reformulation, we introduce a Collaborative Interaction Mechanism (CIM) that establishes direct, bidirectional communication between the embedder and extractor, enabling a mutual-teacher training paradigm and coordinated optimization. Built upon this explicitly collaborative architecture, we further propose an Adaptive Feature Modulation Module (AFMM) to support effective interaction. AFMM enables content-aware feature regulation by decoupling modulation structure and strength, guiding watermark embedding toward stable image features while suppressing host interference during extraction. Under CIM, the AFMMs on both sides form a closed-loop collaboration that aligns embedding behavior with extraction objectives. This architecture-level redesign changes how robustness is learned in watermarking systems. Rather than relying on exhaustive distortion simulation, robustness emerges from coordinated representation learning between embedding and extraction. Experiments on real-world and AI-generated datasets demonstrate that the proposed method consistently outperforms state-of-the-art approaches in watermark extraction accuracy while maintaining high perceptual quality, showing strong robustness and generalization.

cs.CV

Ultra-Low background germanium spectrometers at the China Jinping Underground Laboratory

Four ultra-low background germanium spectrometers, called GeTHU, have been installed at the first phase of China Jinping Underground Laboratory (CJPL-I), and served for material screening of dark matter and neutrino experiments. Recently, a new multi-detector spectrometer with five germanium detectors has been developed and installed at the second phase of CJPL (CJPL-II) with a minimum detectable activity (MDA) of about 10 {\mu}Bq/kg. In addition, another fifteen GeTHU-like spectrometers have been installed at CJPL-II with an MDA of about 1 mBq/kg. This paper will introduce the ultra-low background germanium spectrometers including shielding design, background characteristics and application to material screening.

physics.ins-det

DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales

Generating textual rationales from large vision-language models (LVLMs) to support trainable multimodal misinformation detectors has emerged as a promising paradigm. However, its effectiveness is fundamentally limited by three core challenges: (i) insufficient diversity in generated rationales, (ii) factual inaccuracies due to hallucinations, and (iii) irrelevant or conflicting content that introduces noise. We introduce DiFaR, a detector-agnostic framework that produces diverse, factual, and relevant rationales to enhance misinformation detection. DiFaR employs five chain-of-thought prompts to elicit varied reasoning traces from LVLMs and incorporates a lightweight post-hoc filtering module to select rationale sentences based on sentence-level factuality and relevance scores. Extensive experiments on four popular benchmarks demonstrate that DiFaR outperforms four baseline categories by up to 5.9% and boosts existing detectors by as much as 8.7%. Both automatic metrics and human evaluations confirm that DiFaR significantly improves rationale quality across all three dimensions.

cs.CL

Three-dimensional position reconstruction of orthogonal-strip planar high-purity germanium detectors using maximum likelihood estimation

Orthogonal-strip planar high-purity germanium (HPGe) detectors can reconstruct three-dimensional (3D) positions of photon interactions through analysis of parameters extracted from multiple charge signals. The conventional method independently reconstructs positions in each dimension using amplitude-based parameters, leading to noise sensitivity and systematic biases. In this study, we propose a multi-parameter-joint reconstruction method based on maximum likelihood estimation (MLE) which establishes a mapping between pulse shape parameters and corresponding 3D positions. To mitigate the effects of electronic noise, we employ integral-based parameters. The reconstruction performance was evaluated using pulse shape simulations. For 100 keV photons under 1 keV root-mean-square (RMS) electronic noise, the maximum Z reconstruction bias was reduced from 0.4 mm to 0.02 mm in the central region and from 2 mm to 0.15 mm near the electrodes. The maximum reconstruction bias in the X/Y directions was reduced from 0.4 mm to 0.016 mm. Furthermore, the use of integral-based parameters mitigated the rapid degradation of resolution under high-noise conditions. The achieved position resolution ranged from 0.07 mm to 0.16 mm in the Z directions and from 0.07 mm to 0.44 mm in the X/Y direction. This method offers a promising approach to 3D position reconstruction with HPGe detectors for applications such as medical imaging and gamma-ray astronomy.

physics.ins-det

Pressure and doping control of magnetic order and metallization in Ruddlesden-Popper La2NiO4

The discovery of superconductivity in multilayer nickelates under pressure has intensified interest in understanding the magnetic and electronic properties of Ruddlesden-Popper nickelates. Using density functional theory with Hubbard corrections, we investigate the magnetic ground state, electronic structure evolution under pressure, and Sr-doping effects in La$_2$NiO$_4$. We find that at ambient pressure, tetragonal La$_2$NiO$_4$ exhibits G-type antiferromagnetic order with negligible interlayer magnetic coupling. Under hydrostatic pressure, the system undergoes a continuous insulator-metal transition at ~50 GPa while maintaining robust magnetic order up to 75 GPa, contrasting sharply with the rapid magnetic suppression in La$_3$Ni$_2$O$_7$. Sr doping induces a systematic evolution from G-type to A-type, to striped antiferromagnetic orders, and eventually to ferromagnetic order, accompanied by metallization. Furthermore, LaSrNiO$_4$ displays weak charge and orbital orders. These results reveal the unique pressure and doping effects of single-layer nickelates and provide insights into the magnetic mechanisms underlying nickelate superconductivity.

cond-mat.supr-con

Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection

Misinformation detection models often rely on superficial cues (i.e., \emph{shortcuts}) that correlate with misinformation in training data but fail to generalize to the diverse and evolving nature of real-world misinformation. This issue is exacerbated by large language models (LLMs), which can easily generate convincing misinformation through simple prompts. We introduce TruthOverTricks, a unified evaluation paradigm for measuring shortcut learning in misinformation detection. TruthOverTricks categorizes shortcut behaviors into intrinsic shortcut induction and extrinsic shortcut injection, and evaluates seven representative detectors across 14 popular benchmarks, along with two new factual misinformation datasets, NQ-Misinfo and Streaming-Misinfo. Empirical results reveal that existing detectors suffer severe performance degradation when exposed to both naturally occurring and adversarially crafted shortcuts. To address this, we propose SMF, an LLM-augmented data augmentation framework that mitigates shortcut reliance through paraphrasing, factual summarization, and sentiment normalization. SMF consistently enhances robustness across 16 benchmarks, encouraging models to rely on deeper semantic understanding rather than shortcut cues. To promote the development of misinformation detectors, we have published the resources publicly at https://github.com/whr000001/TruthOverTricks.

cs.CL

RealityAvatar: Towards Realistic Loose Clothing Modeling in Animatable 3D Gaussian Avatars

Modeling animatable human avatars from monocular or multi-view videos has been widely studied, with recent approaches leveraging neural radiance fields (NeRFs) or 3D Gaussian Splatting (3DGS) achieving impressive results in novel-view and novel-pose synthesis. However, existing methods often struggle to accurately capture the dynamics of loose clothing, as they primarily rely on global pose conditioning or static per-frame representations, leading to oversmoothing and temporal inconsistencies in non-rigid regions. To address this, We propose RealityAvatar, an efficient framework for high-fidelity digital human modeling, specifically targeting loosely dressed avatars. Our method leverages 3D Gaussian Splatting to capture complex clothing deformations and motion dynamics while ensuring geometric consistency. By incorporating a motion trend module and a latentbone encoder, we explicitly model pose-dependent deformations and temporal variations in clothing behavior. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach in capturing fine-grained clothing deformations and motion-driven shape variations. Our method significantly enhances structural fidelity and perceptual quality in dynamic human reconstruction, particularly in non-rigid regions, while achieving better consistency across temporal frames.

cs.CV

Isotropic superconductivity in pressurized trilayer nickelate La4Ni3O10

Evidence of superconductivity (SC) has recently been reported in pressurized La3Ni2O7 and La4Ni3O10, providing a new platform to explore high-temperature superconductivity. However, while zero resistance state has been observed, experimental characterization of the superconducting properties of pressurized nickelates is still limited and experimentally challenging. Here, we present the first full temperature dependence of the upper critical field Hc2 measurement in La4Ni3O10 single crystal, achieved by combining high magnetic field and high-pressure techniques. Remarkably, the Hc2 of La4Ni3O10 is nearly isotropic, with the anisotropic parameter monotonically increasing from 1.4 near Tc to 1 at lower temperatures. By analyzing the Hc2 using the two-band model, we uncover that the anisotropic diffusivity of the bands, primarily originating from d(z2 ) and d(x2-y2 ) orbitals, is well compensated, resulting in an unusually isotropic superconducting state. These findings provide critical experimental evidence that underscores the significant role of the d(z2 ) orbital in enabling superconductivity in pressurized Ruddlesden-Popper nickelates.

cond-mat.supr-con

Linear-optical four-dimensional Bell state measurement with two-photon interference

We theoretically investigate the distinguishability of a set of mutually orthogonal four-dimensional Bell states of photon system in path degree of freedom using only linear optics, resorting to the two-photon interference. With quantum interference effect, we find that the 16 four-dimensional Bell states can be classified into 7 groups, which can support the transmission of $log_2 7 = 2.81$ bits classical information with just sending one photon in the quantum superdense coding protocol. When an auxiliary two-dimensional polarization entanglement is introduced, the 16 four-dimensional Bell states then can be classified into 12 groups, which can promote the channel capacity to $log_2 12 = 3.58$ bits with encoding one photon. Our results are significant for photonic superdense coding, and can be useful for other quantum information technologies involved linear-optical high-dimensional Bell state measurement.

quant-ph