SearcharxivSearch

arXiv subjects

Xue Wu

Publications and source records attributed to Xue Wu.

12 recordsLinked to original sources

Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors. However, these methods still face two key limitations. First, existing SD-based one-step and multi-step Real-ISR approaches adopt a unified processing paradigm for all input samples, ignoring the varying restoration difficulty across images. Second, the aggressive resolution reduction of the VAE in SD models (e.g., 8x downsampling) leads to irreversible loss of fine-scale details, which cannot be recovered by the subsequent diffusion process. To address these limitations, we propose a Difficulty-aware Dynamic Routing (DDR) strategy that overcomes the rigid, one-size-fits-all processing paradigm. Specifically, we first design a difficulty estimator to predict the restoration cost of each input image, enabling automatic assignment to a network of appropriate capacity. Then, we construct a set of Real-ISR networks with varying model capacities by modulating the spatial downsampling ratio of the VAE in the SD backbone, thereby preserving more high-frequency information for challenging cases while maintaining efficiency for simpler inputs. Extensive experiments have demonstrated the superior efficiency and effectiveness of the proposed model compared to recent state-of-the-art methods.

cs.CV

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning

Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability indices are coarse and provide limited insight into risk contexts. We propose SatBLIP, a satellite-specific vision-language framework for rural context understanding and feature identification that predicts county-level Social Vulnerability Index (SVI). SatBLIP addresses limitations of prior remote sensing pipelines-handcrafted features, manual virtual audits, and natural-image-trained VLMs-by coupling contrastive image-text alignment with bootstrapped captioning tailored to satellite semantics. We use GPT-4o to generate structured descriptions of satellite tiles (roof type/condition, house size, yard attributes, greenery, and road context), then fine-tune a satellite-adapted BLIP model to generate captions for unseen images. Captions are encoded with CLIP and fused with LLM-derived embeddings via attention for SVI estimation under spatial aggregation. Using SHAP, we identify salient attributes (e.g., roof form/condition, street width, vegetation, cars/open space) that consistently drive robust predictions, enabling interpretable mapping of rural risk environments.

cs.CV

Ir3Ge20: A 3n-Connected Cloverleaf-Shaped Supercluster

Group 14 Zintl clusters are promising molecular building blocks for nanoscale architecture. Endohedral variants, which encapsulate d/f-block metals within p-block semimetal cages, provide insights into intermetallic bonding and compound formation. In this study, experimental photoelectron spectroscopy and first-principles calculations were used to investigate the Ir-doped germanium cluster species. A new C2v symmetric building block, IrGe12, was identified, serving as the basis for designing the supercluster Ir3Ge20 with a cloverleaf-shaped, D3h symmetric architecture. This structure consists of three interconnected aromatic IrGe12 units linked by Ge-Ge sigma bonds, forming shielding cones. Ir3Ge20 follows the 5n rule with 100 valence electrons, featuring a core-Ir3 unit with a d10 closed-shell configuration sharing electrons with the Ge20 skeleton. The stability, chemical bonding, and aromaticity of Ir3Ge20 were confirmed, demonstrating a novel approach to precise atom manipulation in cluster-based materials and devices.

cond-mat.mtrl-sci

One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation

Diffusion-based models have been widely used in various visual generation tasks, showing promising results in image super-resolution (SR), while typically being limited by dozens or even hundreds of sampling steps. Although existing methods aim to accelerate the inference speed of multi-step diffusion-based SR methods through knowledge distillation, their generated images exhibit insufficient semantic alignment with real images, resulting in suboptimal perceptual quality reconstruction, specifically reflected in the CLIPIQA score. These methods still have many challenges in perceptual quality and semantic fidelity. Based on the challenges, we propose VPD-SR, a novel visual perception diffusion distillation framework specifically designed for SR, aiming to construct an effective and efficient one-step SR model. Specifically, VPD-SR consists of two components: Explicit Semantic-aware Supervision (ESS) and High-Frequency Perception (HFP) loss. Firstly, the ESS leverages the powerful visual perceptual understanding capabilities of the CLIP model to extract explicit semantic supervision, thereby enhancing semantic consistency. Then, Considering that high-frequency information contributes to the visual perception quality of images, in addition to the vanilla distillation loss, the HFP loss guides the student model to restore the missing high-frequency details in degraded images that are critical for enhancing perceptual quality. Lastly, we expand VPD-SR in adversarial training manner to further enhance the authenticity of the generated content. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed VPD-SR achieves superior performance compared to both previous state-of-the-art methods and the teacher model with just one-step sampling.

cs.CV

Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. However, they often struggle with complex reasoning tasks and are prone to hallucination. Recent research has shown promising results in leveraging knowledge graphs (KGs) to enhance LLM performance. KGs provide a structured representation of entities and their relationships, offering a rich source of information that can enhance the reasoning capabilities of LLMs. For this work, we have developed different techniques that tightly integrate KG structures and semantics into LLM representations. Our results show that we are able to significantly improve the performance of LLMs in complex reasoning scenarios, and ground the reasoning process with KGs. We are the first to represent KGs with programming language and fine-tune pretrained LLMs with KGs. This integration facilitates more accurate and interpretable reasoning processes, paving the way for more advanced reasoning capabilities of LLMs.

cs.CL

Resource Management for IRS-Assisted Full-Duplex Integrated Sensing, Communication and Computing Systems

In this paper, we investigate an intelligent reflecting surface (IRS) assisted full-duplex (FD) integrated sensing, communication and computing system. Specifically, an FD base station (BS) provides service for uplink and downlink transmission, and a local cache is connected to the BS through a backhaul link to store data. Meanwhile, active sensing elements are deployed on the IRS to receive target echo signals. On this basis, in order to evaluate the overall performance of the system under consideration, we propose a system utility maximization problem while ensuring the sensing quality, expressed as the difference between the sum of communication throughput, total computation bits (offloading bits and local computation bits) and the total backhaul cost for content delivery. This makes the problem difficult to solve due to the highly non-convex coupling of the optimization variables. To effectively solve this problem, we first design the most effective caching strategy. Then, we develop an algorithm based on weighted minimum mean square error, alternative direction method of multipliers, majorization-minimization framework, semi-definite relaxation techniques, and several complex transformations to jointly solve the optimization variables. Finally, simulation results are provided to verify the utility performance of the proposed algorithm and demonstrate the advantages of the proposed scheme compared with the baseline scheme.

cs.IT

B63: the most stable bilayer structure with dual aromaticity

The emergence of the first bilayer B48, which has been both theoretically predicted and experimentally observed, as well as the recent experimental synthesis of bilayer borophene on Ag and Cu, has generated tremendous curiosity in the bilayer structure of boron clusters. However, the connection between the bilayer cluster and the bilayer borophene remains unknown. By combining a genetic algorithm and density functional theory calculations, a global search for the low-energy structures of B63 clusters was conducted, revealing that the Cs bilayer structure with three interlayer B-B bonds was the most stable bilayer structure. This structure was further examined in terms of its structural stability, chemical bonding, and aromaticity. Interestingly, the interlayer bonds exhibited electronegativity and robust aromaticity. Furthermore, the double aromaticity stemmed from diatropic currents originating from virtual translational transitions at both the sigma and pi electrons. This new boron bilayer is anticipated to enrich the concept of double aromaticity and serve as a valuable precursor for bilayer borophene.

cond-mat.mtrl-sci

TransCC: Transformer Network for Coronary Artery CCTA Segmentation

The accurate segmentation of Coronary Computed Tomography Angiography (CCTA) images holds substantial clinical value for the early detection and treatment of Coronary Heart Disease (CHD). The Transformer, utilizing a self-attention mechanism, has demonstrated commendable performance in the realm of medical image processing. However, challenges persist in coronary segmentation tasks due to (1) the damage to target local structures caused by fixed-size image patch embedding, and (2) the critical role of both global and local features in medical image segmentation tasks.To address these challenges, we propose a deep learning framework, TransCC, that effectively amalgamates the Transformer and convolutional neural networks for CCTA segmentation. Firstly, we introduce a Feature Interaction Extraction (FIE) module designed to capture the characteristics of image patches, thereby circumventing the loss of semantic information inherent in the original method. Secondly, we devise a Multilayer Enhanced Perceptron (MEP) to augment attention to local information within spatial dimensions, serving as a complement to the self-attention mechanism. Experimental results indicate that TransCC outperforms existing methods in segmentation performance, boasting an average Dice coefficient of 0.730 and an average Intersection over Union (IoU) of 0.582. These results underscore the effectiveness of TransCC in CCTA image segmentation.

eess.IV

GroupRegNet: A Groupwise One-shot Deep Learning-based 4D Image Registration Method

Accurate deformable 4-dimensional (4D) (3-dimensional in space and time) medical images registration is essential in a variety of medical applications. Deep learning-based methods have recently gained popularity in this area for the significant lower inference time. However, they suffer from drawbacks of non-optimal accuracy and the requirement of a large amount of training data. A new method named GroupRegNet is proposed to address both limitations. The deformation fields to warp all images in the group into a common template is obtained through one-shot learning. The use of the implicit template reduces bias and accumulated error associated with the specified reference image. The one-shot learning strategy is similar to the conventional iterative optimization method but the motion model and parameters are replaced with a convolutional neural network (CNN) and the weights of the network. GroupRegNet also features a simpler network design and a more straightforward registration process, which eliminates the need to break up the input image into patches. The proposed method was quantitatively evaluated on two public respiratory-binned 4D-CT datasets. The results suggest that GroupRegNet outperforms the latest published deep learning-based methods and is comparable to the top conventional method pTVreg. To facilitate future research, the source code is available at https://github.com/vincentme/GroupRegNet.

eess.IV

Medium-sized Sin- (n=14-20) clusters: a combined study of photoelectron spectroscopy and DFT calculations

Size-selected anionic silicon clusters, Sin- (n=14-20), have been investigated by photoelectron spectroscopy and density functional theory (DFT) calculations. Low-energy structures of the clusters are globally searched for by using a genetic algorithm based on DFT calculations. The electronic density of states and VDEs have been simulated by using ten DFT functionals and compared to the experimental results. We systematically evaluated the DFT functionals for the calculation of the energetics of silicon clusters. CCSD(T) single-point energies based on MP2 optimized geometries for selected isomers of Sin- are also used as benchmark for the energy sequence. The HSE06 functional with aug-cc-pVDZ basis set is found to show the best performance. Our global minimum search corroborates that most of the lowest-energy structures of Sin- (n=14-20) clusters can be derived from assembling tricapped trigonal prisms (TTP) in various ways. For most sizes previous structures are confirmed, whereas for Si20- a new structure has been found.

physics.atm-clus

Record High Magnetic Anisotropy in Chemically Engineered Iridium Dimer

Exploring giant magnetic anisotropy in small magnetic nanostructures is of both fundamental interest and technological merit for information storage. To prevent spin flipping at room temperature due to thermal fluctuation, large magnetic anisotropy energy (MAE) over 50 meV in magnetic nanostructure is desired for practical applications. We chose one of the smallest magnetic nanostructures-Ir2 dimer, to investigate its magnetic properties and explore possible approach to engineer the magnetic anisotropy. Through systematic first-principles calculations, we found that the Ir2 dimer already possesses giant MAE of 77 meV. We proposed an effective way to enhance the MAE of the Ir2 dimer to 223~294 meV by simply attaching a halogen atom at one end of the Ir-Ir bond. The underlying mechanism for the record high MAE is attributed to the modification of the energy diagram of the Ir2 dimer by the additional halogen-Ir bonding, which alters the spin-orbit coupling Hamiltonian and hence the magnetic anisotropy. Our strategy can be generalized to design other magnetic molecules or clusters with giant magnetic anisotropy.

physics.atm-clus

Calculation of the minimum computational complexity based on information entropy

In order to find out the limiting speed of solving a specific problem using computer, this essay provides a method based on information entropy. The relationship between the minimum computational complexity and information entropy change is illustrated. A few examples are served as evidence of such connection. Meanwhile some basic rules of modeling problems are established. Finally, the nature of solving problems with computer programs is disclosed to support this theory and a redefinition of information entropy in this filed is proposed. This will develop a new field of science.

cs.CC