SearcharxivSearch

arXiv subjects

Xianjie Liu

Publications and source records attributed to Xianjie Liu.

18 recordsLinked to original sources

Pay More Attention To Text In High-Resolution MLLMs

Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.

cs.CV

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbf{training-free, single-visual-pass} evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V$^*$Bench, HRBench-4K, and HRBench-8K by \textit{+3.1}, \textit{+3.0}, and \textit{+2.7} points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit{+9.9}, \textit{+4.6}, and \textit{+5.5} points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V$^*$Bench inference time by \textbf{97.2\%} while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.

cs.CV

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often struggle with these videos because existing benchmarks focus primarily on general-purpose tasks and neglect the reasoning of commercial intent. In this work, we first propose a multi-modal information density assessment framework to quantify the complexity of this domain. Our evaluation reveals that e-commerce content exhibits substantially higher density across visual, audio, and textual modalities compared to mainstream datasets, establishing a more challenging frontier for video understanding. To address this gap, we introduce E-commerce Video Ads Benchmark, which is the first benchmark specifically designed for e-commerce short video understanding. We curated 3,961 high-quality videos from Taobao covering a wide range of product categories and used a multi-agent system to generate 19,785 open-ended Q&A pairs, which consist of five distinct tasks. Finally, we develop E-VAds-R1, an RL-based reasoning model featuring a multi-grained reward design called MG-GRPO. This strategy provides smooth guidance for early exploration while creating a non-linear incentive for expert-level precision. Experimental results demonstrate that E-VAds-R1 achieves a 109.2% performance gain in commercial intent reasoning with only a few hundred training samples. Data is available at https://github.com/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark.

cs.CV

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tasks. However, their performance on high-resolution images remains suboptimal. While existing approaches often attribute this limitation to perceptual constraints and argue that MLLMs struggle to recognize small objects, leading them to use "zoom in" strategies for better detail, our analysis reveals a different cause: the main issue is not object size, but rather caused by complex background interference. We systematically analyze this "zoom in" operation through a series of decoupling experiments and propose the Hierarchical Decoupling Framework (HiDe), a training-free framework that uses Token-wise Attention Decoupling (TAD) to decouple the question tokens and identify the key information tokens, then leverages their attention weights to achieve precise alignment with the target visual regions. Subsequently, it employs Layout-Preserving Decoupling (LPD) to decouple these regions from the background and reconstructs a compact representation that preserves essential spatial layouts while eliminating background interference. HiDe sets a new SOTA on V*Bench, HRBench4K, and HRBench8K, boosting Qwen2.5-VL 7B and InternVL3 8B to SOTA (92.1% and 91.6% on V*Bench), even surpassing RL methods. After optimization, HiDe uses 75% less memory than the previous training-free approach. Code is provided in https://tennine2077.github.io/HiDe.github.io/.

cs.CV

High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy

High-precision dichotomous image segmentation (DIS) is a task of extracting fine-grained objects from high-resolution images. Existing methods trade efficiency for accuracy: non-diffusion methods are fast but suffer from weak semantics and unstable spatial priors, causing false detections; diffusion-based methods offer high accuracy via strong generative priors but are computationally expensive. In depth maps, a complete object appears as a low variance region with a smooth interior and sharp boundaries, whereas the background exhibits a chaotic, high variance pattern due to disconnected surfaces at varying depths. We refer to this as the depth integrity-prior. Inspired by this, and noting that DIS currently lacks depth maps, we leverage pseudo-depth information from monocular depth estimation models to obtain essential semantic understanding, thereby rapidly revealing spatial differences across target objects and the background. To exploit this prior, we propose the Prior-guided Depth Fusion Network (PDFNet), which fuses RGB and pseudo-depth features for depth-aware structure perception. We further introduce a novel depth integrity-prior loss to enforce depth consistency in segmentation and a fine-grained enhancement module with adaptive patch selection to sharpen boundaries. Notably, PDFNet with DAM-v2 achieves SOTA (Fmax 0.915 on DIS-VD and 0.915 on DIS-TE) using less than half the params of diffusion-based methods. Our code is available at https://tennine2077.github.io/PDFNet.github.io/ .

cs.CV

Enzyme-free in situ polymerization of conductive polymers catalyzed by porous Au@Ag nanowires for stretchable neural electrodes

In situ polymerization of conductive polymers (CPs) represents a transformative approach in bioelectronics, by enabling the controlled growth of electrically active materials right at the tissue or device surface to create seamless biotic-abiotic interfaces. Traditional CP deposition techniques often use high anodic potentials, non-physiological electrolytes, or strong oxidants, making them harmful to adjacent tissues. A possible solution is enzymatic polymerization which operates under milder conditions, but it is limited by the stability and activity window of the enzyme catalysts, low throughput, and challenges in spatially confining polymer growth. To resolve these issues, here we developed one-dimensional porous Au-coated Ag nanowires with horseradish peroxidase (HRP)-like catalytic properties, thereby for the first time enabling mild in situ enzyme-free polymerization of conductive polymers near neutral pH. The enzyme-free polymerization is demonstrated both in aqueous dispersions at pH=6 and in situ onto porous Au coated Ag nanowire based stretchable electrodes. Following enzyme-free catalytic polymerization, the electrically conducting polymer coating on the electrode greatly improves the impedance and achieves an impedance of 2.6 kOhm at 1 kHz for 50x50 um large electrodes.

physics.chem-ph

Promoting Segment Anything Model towards Highly Accurate Dichotomous Image Segmentation

The Segment Anything Model (SAM) represents a significant breakthrough into foundation models for computer vision, providing a large-scale image segmentation model. However, despite SAM's zero-shot performance, its segmentation masks lack fine-grained details, particularly in accurately delineating object boundaries. Therefore, it is both interesting and valuable to explore whether SAM can be improved towards highly accurate object segmentation, which is known as the dichotomous image segmentation (DIS) task. To address this issue, we propose DIS-SAM, which advances SAM towards DIS with extremely accurate details. DIS-SAM is a framework specifically tailored for highly accurate segmentation, maintaining SAM's promptable design. DIS-SAM employs a two-stage approach, integrating SAM with a modified advanced network that was previously designed to handle the prompt-free DIS task. To better train DIS-SAM, we employ a ground truth enrichment strategy by modifying original mask annotations. Despite its simplicity, DIS-SAM significantly advances the SAM, HQ-SAM, and Pi-SAM ~by 8.5%, ~6.9%, and ~3.7% maximum F-measure. Our code at https://github.com/Tennine2077/DIS-SAM

cs.CV

Magnetic hysteresis control in thin film Fe/Si multilayers by incorporation of B4C

Magnetic hysteresis properties in Fe/Si multilayers have been studied as a function of the B4C content to control magnetization amplitude, coercivity, and hysteresis tilt, properties that are beneficial to tune for advancing applications in e.g. data storage, spintronics, and sensors. With an ion-assisted magnetron sputtering technique, 35 distinct thin film multilayer samples were prepared and their magnetic and structural properties were characterized by vibrating sample magnetometry, X-ray photoelectron spectroscopy, near edge X-ray absorption fine structure spectroscopy, and X-ray and neutron scattering methods. Key findings indicate that adding B4C lowers the coercivity and can decrease the saturation magnetization, demonstrating the tunability of magnetic responses based on composition. For samples with =30Å periodicity, 10-15% of B4C addition produces antiferromagnetically (AF) coupled multilayers, and such AF coupling strength increases with the B4C content. Our findings reveal that B atoms do not chemically bind within the Fe atoms but instead occupy interstitial positions, disrupting medium- to long-range crystallinity thereby inducing the amorphization. Thereon, the observed effects on magnetic properties are directly attributed to this amorphization process caused by the presence of B4C. The demonstrated ability to finely adjust magnetic properties by varying the B4C content offers a promising approach to overcome challenges in magnetic device performance and efficiency.

cond-mat.mtrl-sci

Uncertainty-informed Mutual Learning for Joint Medical Image Classification and Segmentation

Classification and segmentation are crucial in medical image analysis as they enable accurate diagnosis and disease monitoring. However, current methods often prioritize the mutual learning features and shared model parameters, while neglecting the reliability of features and performances. In this paper, we propose a novel Uncertainty-informed Mutual Learning (UML) framework for reliable and interpretable medical image analysis. Our UML introduces reliability to joint classification and segmentation tasks, leveraging mutual learning with uncertainty to improve performance. To achieve this, we first use evidential deep learning to provide image-level and pixel-wise confidences. Then, an Uncertainty Navigator Decoder is constructed for better using mutual features and generating segmentation results. Besides, an Uncertainty Instructor is proposed to screen reliable masks for classification. Overall, UML could produce confidence estimation in features and performance for each link (classification and segmentation). The experiments on the public datasets demonstrate that our UML outperforms existing methods in terms of both accuracy and robustness. Our UML has the potential to explore the development of more reliable and explainable medical image analysis models. We will release the codes for reproduction after acceptance.

cs.CV

HGT: A Hierarchical GCN-Based Transformer for Multimodal Periprosthetic Joint Infection Diagnosis Using CT Images and Text

Prosthetic Joint Infection (PJI) is a prevalent and severe complication characterized by high diagnostic challenges. Currently, a unified diagnostic standard incorporating both computed tomography (CT) images and numerical text data for PJI remains unestablished, owing to the substantial noise in CT images and the disparity in data volume between CT images and text data. This study introduces a diagnostic method, HGT, based on deep learning and multimodal techniques. It effectively merges features from CT scan images and patients' numerical text data via a Unidirectional Selective Attention (USA) mechanism and a graph convolutional network (GCN)-based feature fusion network. We evaluated the proposed method on a custom-built multimodal PJI dataset, assessing its performance through ablation experiments and interpretability evaluations. Our method achieved an accuracy (ACC) of 91.4\% and an area under the curve (AUC) of 95.9\%, outperforming recent multimodal approaches by 2.9\% in ACC and 2.2\% in AUC, with a parameter count of only 68M. Notably, the interpretability results highlighted our model's strong focus and localization capabilities at lesion sites. This proposed method could provide clinicians with additional diagnostic tools to enhance accuracy and efficiency in clinical practice.

cs.CV

A multimodal method based on cross-attention and convolution for postoperative infection diagnosis

Postoperative infection diagnosis is a common and serious complication that generally poses a high diagnostic challenge. This study focuses on PJI, a type of postoperative infection. X-ray examination is an imaging examination for suspected PJI patients that can evaluate joint prostheses and adjacent tissues, and detect the cause of pain. Laboratory examination data has high sensitivity and specificity and has significant potential in PJI diagnosis. In this study, we proposed a self-supervised masked autoencoder pre-training strategy and a multimodal fusion diagnostic network MED-NVC, which effectively implements the interaction between two modal features through the feature fusion network of CrossAttention. We tested our proposed method on our collected PJI dataset and evaluated its performance and feasibility through comparison and ablation experiments. The results showed that our method achieved an ACC of 94.71% and an AUC of 98.22%, which is better than the latest method and also reduces the number of parameters. Our proposed method has the potential to provide clinicians with a powerful tool for enhancing accuracy and efficiency.

cs.CV

Fully 3D-Printed Organic Electrochemical Transistors

Organic electrochemical transistors (OECTs) are currently being investigated for various applications, ranging from sensors to logics and neuromorphic hardware. The fabrication process must be compatible with flexible and scalable digital techniques to address this wide spectrum of applications. Here, we report a direct-write additive process to fabricate fully 3D printed OECTs. We developed 3D printable conducting, semiconducting, insulating, and electrolyte inks to achieve this. The 3D-printed OECTs, operating in the depletion mode, can be fabricated on thin and flexible substrates, yielding high mechanical and environmental stability. We also developed a 3D printable nanocellulose formulation for the OECT substrate, demonstrating one of the first examples of fully 3D printed electronic devices. Good dopamine biosensing capabilities (limit of detection down to 6 uM without metal gate electrodes) and long-term (~1 hour) synapses response underscore that the present OECT manufacturing strategy is suitable for diverse applications requiring rapid design change and digitally enabled direct-write techniques.

physics.app-ph

Accessing the conduction band dispersion in CH3NH3PbI3 single crystals

The conduction band structure in methylammonium lead iodide (CH3NH3PbI3) was studied both by angle-resolved two-photon photoemission spectroscopy (AR-2PPE) with low-photon intensity and angle-resolved low-energy inverse photoelectron spectroscopy (AR-LEIPS). Clear energy dispersion of the conduction band along the ΓM direction was observed by these independent methods under different temperatures, and the dispersion was found to be consistent with band calculations under the cubic phase. The effective mass of the electrons at the Γ point was estimated to be (0.20+-0.05)m0 at 90 K. The observed energy position was largely different between the AR-LEIPS and AR-2PPE, demonstrating the electron correlation effects on the band structures. The present results also indicate that the surface structure in CH3NH3PbI3 provides the cubic-dominated electronic property even at lower temperatures.

cond-mat.mtrl-sci

Surface geometry determined temperature-dependent band structure evolutions in organic halide perovskite single crystals

In this study, different electronic structure evolutions of perovskite single crystals are found via angle-resolved photoelectron spectroscopy (ARPES): (i) unchanged top valence band (VB) dispersions under different temperatures can be found in the CH3NH3PbI3, (ii) phase transitions induced the evolution of top VB dispersions, and even a top VB splitting with Rashba effects can be observed in the CH3NH3PbBr3. Combined with low-energy electron diffraction (LEED), metastable atom electron spectroscopy (MAES), and DFT calculation, we confirm different band structure evolutions observed in these two perovskite single crystals are originated from the cleaved top surface layers, where the different surface geometries with CH3NH3+-I in CH3NH3PbI3 and Pb-Br in CH3NH3PbBr3 are responsible for finding band dispersion change and appearing of the Rashba-type splitting. Such findings suggest that the top surface layer in organic halide perovskites should be carefully considered to create functional interfaces for developing perovskite devices.

cond-mat.mtrl-sci

Synergistically creating sulfur vacancies in semimetal-supported amorphous MoS2 for efficient hydrogen evolution

The presence of elemental vacancies in materials is inevitable according to statistical thermodynamics, which will decide the chemical and physical properties of the investigated system. However, the controlled manipulation of vacancies for specific applications is a challenge. Here we report a facile method for creating large concentrations of S vacancies in the inert basal plane of MoS2 supported on semimetal CoMoP2. With a small applied potential, S atoms can be removed in the form of H2S due to the optimized free energy of formation. The existence of vacancies favors electron injection from the electrode to the active site by decreasing the contact resistance. As a consequence, the activity is increased by 221 % with the vacancy-rich MoS2 as electrocatalyst for hydrogen evolution reaction (HER). A small overpotential of 75 mV is needed to deliver a current density of 10 mA cm-2, which is considered among the best values achieved for MoS2. It is envisaged that this work may provide a new strategy for utilizing the semimetal phase for structuring MoS2 into a multi-functional material.

cond-mat.mtrl-sci

11,11,12,12-tetracyanonaphtho-2,6-quinodimethane in Contact with Ferromagnetic Electrodes for Organic Spintronics

Spinterface engineering has shown quite important roles in organic spintronics as it can improve spin injection or extraction. In this study, 11,11,12,12-tetracyanonaptho-2,6-quinodimethane (TNAP) is introduced as an interfacial layer for a prototype interface of Fe/TNAP. We report an element-specific investigation of the electronic and magnetic structures of Fe/TNAP system by use of near edge X-Ray absorption fine structure (NEXAFS) and X-ray magnetic circular dichroism (XMCD). Strong hybridization between TNAP and Fe and induced magnetization of N atoms in TNAP molecule are observed. XMCD sum rule analysis demonstrates that the adsorption of TNAP reduces the spin moment of Fe by 12%. In addition, induced magnetization in N K-edge of TNAP has also been found with other commonly used ferromagnets in organic spintronics, such as La0.7Sr0.3MnO3 and permalloy, which makes TNAP a very promising molecule for spinterface engineering in organic spintronics.

cond-mat.mtrl-sci

Hybrid Interface States and Spin Polarization at Ferromagnetic Metal-Organic Heterojunctions: Interface Engineering for Efficient Spin Injection in Organic Spintronics

Ferromagnetic metal-organic semiconductor (FM-OSC) hybrid interfaces have shown to play an important role for spin injection in organic spintronics. Here, 11,11,12,12-tetracyanonaptho-2,6-quinodimethane (TNAP) is introduced as an interfacial layer in Co-OSCs heterojunction with an aim to tune the spin injection. The Co/TNAP interface is investigated by use of X-ray and ultraviolet photoelectron spectroscopy (XPS/UPS), near edge X-ray absorption fine structure (NEXAFS) and X-ray magnetic circular dichroism (XMCD). Hybrid interface states (HIS) are observed at Co/TNAP interface resulting from chemical interaction between Co and TNAP. The energy level alignment at Co/TNAP/OSCs interface is also obtained, and a reduction of the hole injection barrier is demonstrated. XMCD results confirm sizeable spin polarization at the Co/TNAP hybrid interface.

cond-mat.mtrl-sci

Orbital and spin magnetic moments of transforming 1D iron inside metallic and semiconducting carbon nanotubes

The orbital and spin magnetic properties of iron inside transforming metallic and semiconducting 1D carbon nanotube hybrids are studied by means of local x-ray magnetic circular dichroism (XMCD) and bulk superconducting quantum interference device (SQUID) measurements. Nanotube hybrids are initially ferrocene filled single-walled carbon nanotubes (SWCNT) of different metallicities. After a high temperature nanochemical reaction ferrocene molecules react with each other to form iron nano clusters. We show that the ferrocenes molecular orbitals interact differently with the SWCNT of different metallicities without significant XMCD response. This XMCD at various temperatures and magnetic fields reveals that the orbital and/or spin magnetic moments of the encapsulated iron are altered drastically as the transformation to 1D Fe nanoclusters takes place. The orbital and spin magnetic moments are both found to be larger in filled semiconducting nanotubes than in the metallic sample. This could mean that the magnetic polarizations of the encapsulated material is dependent on the metallicity of the tubes. From a comparison between the iron 3d magnetic moments and the bulk magnetism measured by SQUID, we conclude that the delocalized magnetisms dictate the magnetic properties of these 1D hybrid nanostructures.

cond-mat.mes-hall