SearcharxivSearch

arXiv subjects

Yingbo Wang

Publications and source records attributed to Yingbo Wang.

18 recordsLinked to original sources

Skill Retrieval Augmentation for Agentic AI

As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their native parametric capabilities. In existing agent systems, the dominant strategy for incorporating skills is to explicitly enumerate available skills within the context window. However, this strategy fails to scale: as skill corpora expand, context budgets are consumed rapidly, and the agent becomes markedly less accurate in identifying the right skill. To this end, this paper formulates Skill Retrieval Augmentation (SRA), a new paradigm in which agents dynamically retrieve, incorporate, and apply relevant skills from large external skill corpora on demand. To make this problem measurable, we construct a large-scale skill corpus and introduce SRA-Bench, the first benchmark for decomposed evaluation of the full SRA pipeline, covering skill retrieval, skill incorporation, and end-task execution. SRA-Bench contains 5,400 capability-intensive test instances and 636 manually constructed gold skills, which are mixed with web-collected distractor skills to form a large-scale corpus of 26,262 skills. Extensive experiments show that retrieval-based skill augmentation can substantially improve agent performance, validating the promise of the paradigm. At the same time, we uncover a fundamental gap in skill incorporation: current LLM agents tend to load skills at similar rates, regardless of whether a gold skill is retrieved or whether the task actually requires external capabilities. This shows that the bottleneck in skill augmentation lies not only in retrieval but also in the base model's ability to determine which skill to load and when external loading is actually needed. These findings position SRA as a distinct research problem and establish a foundation for the scalable augmentation of capabilities in future agent systems.

cs.CL

SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

In medical image segmentation tasks, the domain gap caused by the difference in data collection between training and testing data seriously hinders the deployment of pre-trained models in clinical practice. Continual Test-Time Adaptation (CTTA) aims to enable pre-trained models to adapt to continuously changing unlabeled domains, providing an effective approach to solving this problem. However, existing CTTA methods often rely on unreliable supervisory signals, igniting a self-reinforcing cycle of error accumulation that culminates in catastrophic performance degradation. To overcome these challenges, we propose a CTTA via Semantic-Prompt-Enhanced Graph Clustering (SPEGC) for medical image segmentation. First, we design a semantic prompt feature enhancement mechanism that utilizes decoupled commonality and heterogeneity prompt pools to inject global contextual information into local features, alleviating their susceptibility to noise interference under domain shift. Second, based on these enhanced features, we design a differentiable graph clustering solver. This solver reframes global edge sparsification as an optimal transport problem, allowing it to distill a raw similarity matrix into a refined and high-order structural representation in an end-to-end manner. Finally, this robust structural representation is used to guide model adaptation, ensuring predictions are consistent at a cluster-level and dynamically adjusting decision boundaries. Extensive experiments demonstrate that SPEGC outperforms other state-of-the-art CTTA methods on two medical image segmentation benchmarks. The source code is available at https://github.com/Jwei-Z/SPEGC-for-MIS.

cs.CV

Boosting Meta-Learning for Few-Shot Text Classification via Label-guided Distance Scaling

Few-shot text classification aims to recognize unseen classes with limited labeled text samples. Existing approaches focus on boosting meta-learners by developing complex algorithms in the training stage. However, the labeled samples are randomly selected during the testing stage, so they may not provide effective supervision signals, leading to misclassification. To address this issue, we propose a \textbf{L}abel-guided \textbf{D}istance \textbf{S}caling (LDS) strategy. The core of our method is exploiting label semantics as supervision signals in both the training and testing stages. Specifically, in the training stage, we design a label-guided loss to inject label semantic information, pulling closer the sample representations and corresponding label representations. In the testing stage, we propose a Label-guided Scaler which scales sample representations with label semantics to provide additional supervision signals. Thus, even if labeled sample representations are far from class centers, our Label-guided Scaler pulls them closer to their class centers, thereby mitigating the misclassification. We combine two common meta-learners to verify the effectiveness of the method. Extensive experimental results demonstrate that our approach significantly outperforms state-of-the-art models. All datasets and codes are available at https://anonymous.4open.science/r/Label-guided-Text-Classification.

cs.LG

Signature of gate tunable superconducting network in twisted bilayer graphene

Twisted van der Waals materials provide a tunable platform for investigating two-dimensional superconductivity and quantum phases. Using spectra-imaging scanning tunneling microscopy, we study the superconducting states in twisted bilayer graphene and track their evolution from insulating phases. Gate-dependent spectroscopic measurements reveal two distinct regimes: under-doped (ν = -2.3) and optimally doped (ν = -2.6). In the under-doped regime, partial superconductivity arises, forming a network interspersed with non-gapped regions. At optimal doping, the entire unit cell demonstrates superconductivity, with gap size modulation showing an anti-correlation with the local density of states. This gate-dependent transition from an insulating phase to a modulated superconductor uncovers an unexpected spatial hierarchy in pairing behavior and offers direct microscopic insights to constrain theories of superconductivity in moiré systems.

cond-mat.supr-con

FILA: Fine-Grained Vision Language Models

Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common approach currently involves dynamically cropping the original high-resolution image into smaller sub-images, which are then fed into a vision encoder that was pre-trained on lower-resolution images. However, this cropping approach often truncates objects and connected areas in the original image, causing semantic breaks. To address this limitation, we introduce HyViLM, designed to process images of any resolution while retaining the overall context during encoding. Specifically, we: (i) Design a new visual encoder called Hybrid Encoder that not only encodes individual sub-images but also interacts with detailed global visual features, significantly improving the model's ability to encode high-resolution images. (ii) Propose an optimal feature fusion strategy for the dynamic cropping approach, effectively leveraging information from different layers of the vision encoder. Compared with the state-of-the-art MLLMs under the same setting, our HyViLM outperforms existing MLLMs in nine out of ten tasks. Specifically, HyViLM achieves a 9.6% improvement in performance on the TextVQA task and a 6.9% enhancement on the DocVQA task.

cs.CV

FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding

Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing limitations still hinder long-video comprehension. A common approach is video feature compression to reduce token input to large language models, yet many methods either fail to prioritize essential features, leading to redundant inter-frame information, or introduce computationally expensive modules.To address these issues, we propose FiLA(Fine-grained Vision Language Model)-Video, a novel framework that leverages a lightweight dynamic-weight multi-frame fusion strategy, which adaptively integrates multiple frames into a single representation while preserving key video information and reducing computational costs. To enhance frame selection for fusion, we introduce a keyframe selection strategy, effectively identifying informative frames from a larger pool for improved summarization. Additionally, we present a simple yet effective long-video training data generation strategy, boosting model performance without extensive manual annotation. Experimental results demonstrate that FiLA-Video achieves superior efficiency and accuracy in long-video comprehension compared to existing methods.

cs.CV

Spectral signature of periodic modulation and sliding of pseudogap state in moire system

The nature of the pseudogap state is widely believed as a key to understanding the pairing mechanism underlying unconventional superconductivity. Over the past two decades, significant efforts have been devoted to searching for spontaneous symmetry breaking or potential order parameters associated with these pseudogap states, aiming to better characterize their properties. Recently, pseudogap states have also been realized in moire systems with extensive gate tunability, yet their local electronic structure remains largely unexplored8. In this study, we report the observation of gate-tunable spontaneous symmetry breaking and sliding behavior of the pseudogap state in magic-angle twisted bilayer graphene (MAtBG) using spectroscopic imaging scanning tunneling microscopy. Our spectroscopy reveals a distinct pseudogap at 4.4 K within the doping range -3 < v < -2. Spectroscopic imaging highlights a gap size modulation at moire scale that is sensitive to the filling, indicative of a wave-like fluctuating pseudogap feature. Specifically, the positions of gap size minima (GSM) coincide with regions of the highest local density of states (LDOS) at the filling v = -2.63, but a unidirectional sliding behavior of GSM is observed for other fillings. In addition, the pseudogap size distribution at certain doping levels also causes a clear nematic order, or an anisotropic gap distribution. Our results have shed light on the complex nature of this pseudogap state, revealing critical insights into the phase diagram of correlated electron systems.

cond-mat.mes-hall

Polar Vortex Superstructure and Its Coupling with Correlated Electrons in Quasiperiodic Moire Crystal

Nanoscale polar structures are significant for understanding polarization processes in low-dimensional systems and hold potential for developing high-performance electronics. Here, we demonstrate a polar vortex superstructure arising from the reconstructed moiré patterns in twisted bilayer graphene aligned with hexagonal boron nitride. Scanning tunneling microscopy reveals spatially modulated charge polarization, while theoretical simulations indicate that the in-plane polarization field forms an array of polar vortices. Notably, this polar field is gate-tunable, exhibiting an unconventional gate-tunable polar sliding and screening process. Moreover, its interaction with electron correlations in twisted bilayer graphene leads to modulated correlated states. Our findings establish moiré pattern reconstruction as a powerful strategy for engineering nanoscale polar structures and emergent quantum phases in van der Waals materials.

cond-mat.mes-hall

Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning

Graphs are data structures used to represent irregular networks and are prevalent in numerous real-world applications. Previous methods directly model graph structures and achieve significant success. However, these methods encounter bottlenecks due to the inherent irregularity of graphs. An innovative solution is converting graphs into textual representations, thereby harnessing the powerful capabilities of Large Language Models (LLMs) to process and comprehend graphs. In this paper, we present a comprehensive review of methodologies for applying LLMs to graphs, termed LLM4graph. The core of LLM4graph lies in transforming graphs into texts for LLMs to understand and analyze. Thus, we propose a novel taxonomy of LLM4graph methods in the view of the transformation. Specifically, existing methods can be divided into two paradigms: Graph2text and Graph2token, which transform graphs into texts or tokens as the input of LLMs, respectively. We point out four challenges during the transformation to systematically present existing methods in a problem-oriented perspective. For practical concerns, we provide a guideline for researchers on selecting appropriate models and LLMs for different graphs and hardware constraints. We also identify five future research directions for LLM4graph.

cs.LG

Directly visualizing nematic superconductivity driven by the pair density wave in NbSe$_2$

Pair density wave (PDW) is a distinct superconducting state characterized by a periodic modulation of its order parameter in real space. Its intricate interplay with the charge density wave (CDW) state is a continuing topic of interest in condensed matter physics. While PDW states have been discovered in cuprates and other unconventional superconductors, the understanding of diverse PDWs and their interactions with different types of CDWs remains limited. Here, utilizing scanning tunneling microscopy, we unveil the subtle correlations between PDW ground states and two distinct CDW phases -- namely, anion-centered-CDW (AC-CDW) and hollow-centered-CDW (HC-CDW) -- in 2H-NbSe$_2$. In both CDW regions, we observe coexisting PDWs with a commensurate structure that aligns with the underlying CDW phase. The superconducting gap size, $Δ(r)$, related to the pairing order parameter is in phase with the charge density in both CDW regions. Meanwhile, the coherence peak height, $H(r)$, qualitatively reflecting the electron-pair density, exhibits a phase difference of approximately $2π/3$ relative to the CDW. The three-fold rotational symmetry is preserved in the HC-CDW region but is spontaneously broken in the AC-CDW region due to the PDW state, leading to the emergence of nematic superconductivity.

cond-mat.supr-con

Panoramic single-pixel imaging with megapixel resolution based on rotational subdivision

Single-pixel imaging (SPI) using a single-pixel detector is an unconventional imaging method, which has great application prospects in many fields to realize high-performance imaging. In especial, the recent proposed catadioptric panoramic ghost imaging (CPGI) extends the application potential of SPI to high-performance imaging at a wide field of view (FOV) with recent growing demands. However, the resolution of CPGI is limited by hardware parameters of the digital micromirror device (DMD), which may not meet ultrahigh-resolution panoramic imaging needs that require detailed information. Therefore, to overcome the resolution limitation of CPGI, we propose a panoramic SPI based on rotational subdivision (RSPSI). The key of the proposed RSPSI is to obtain the entire panoramic scene by the rotation-scanning with a rotating mirror tilted 45°, so that one single pattern that only covers one sub-Fov with a small FOV can complete a uninterrupted modulation on the entire panoramic FOV during a once-through pattern projection. Then, based on temporal resolution subdivision, images sequence of sub-Fovs subdivided from the entire panoramic FOV can be reconstructed with pixels-level or even subpixels-level horizontal shifting adjacently. Experimental results using a proof-of-concept setup show that the panoramic image can be obtained with 10428*543 of 5,662,404 pixels, which is more than 9.6 times higher than the resolution limit of the CPGI using the same DMD. To our best knowledge, the RSPSI is the first to achieve a megapixel resolution via SPI, which can provide potential applications in fields requiring the imaging with ultrahigh-resolution and wide FOV.

physics.optics

RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning

Recent advancements in Large Language Models (LLMs) and Large Multi-modal Models (LMMs) have shown potential in various medical applications, such as Intelligent Medical Diagnosis. Although impressive results have been achieved, we find that existing benchmarks do not reflect the complexity of real medical reports and specialized in-depth reasoning capabilities. In this work, we introduced RJUA-MedDQA, a comprehensive benchmark in the field of medical specialization, which poses several challenges: comprehensively interpreting imgage content across diverse challenging layouts, possessing numerical reasoning ability to identify abnormal indicators and demonstrating clinical reasoning ability to provide statements of disease diagnosis, status and advice based on medical contexts. We carefully design the data generation pipeline and proposed the Efficient Structural Restoration Annotation (ESRA) Method, aimed at restoring textual and tabular content in medical report images. This method substantially enhances annotation efficiency, doubling the productivity of each annotator, and yields a 26.8% improvement in accuracy. We conduct extensive evaluations, including few-shot assessments of 5 LMMs which are capable of solving Chinese medical QA tasks. To further investigate the limitations and potential of current LMMs, we conduct comparative experiments on a set of strong LLMs by using image-text generated by ESRA method. We report the performance of baselines and offer several observations: (1) The overall performance of existing LMMs is still limited; however LMMs more robust to low-quality and diverse-structured images compared to LLMs. (3) Reasoning across context and image content present significant challenges. We hope this benchmark helps the community make progress on these challenging tasks in multi-modal medical document understanding and facilitate its application in healthcare.

cs.CL

GACE: Learning Graph-Based Cross-Page Ads Embedding For Click-Through Rate Prediction

Predicting click-through rate (CTR) is the core task of many ads online recommendation systems, which helps improve user experience and increase platform revenue. In this type of recommendation system, we often encounter two main problems: the joint usage of multi-page historical advertising data and the cold start of new ads. In this paper, we proposed GACE, a graph-based cross-page ads embedding generation method. It can warm up and generate the representation embedding of cold-start and existing ads across various pages. Specifically, we carefully build linkages and a weighted undirected graph model considering semantic and page-type attributes to guide the direction of feature fusion and generation. We designed a variational auto-encoding task as pre-training module and generated embedding representations for new and old ads based on this task. The results evaluated in the public dataset AliEC from RecBole and the real-world industry dataset from Alipay show that our GACE method is significantly superior to the SOTA method. In the online A/B test, the click-through rate on three real-world pages from Alipay has increased by 3.6%, 2.13%, and 3.02%, respectively. Especially in the cold-start task, the CTR increased by 9.96%, 7.51%, and 8.97%, respectively.

cs.IR

CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation

The hybrid architecture of convolutional neural networks (CNNs) and Transformer are very popular for medical image segmentation. However, it suffers from two challenges. First, although a CNNs branch can capture the local image features using vanilla convolution, it cannot achieve adaptive feature learning. Second, although a Transformer branch can capture the global features, it ignores the channel and cross-dimensional self-attention, resulting in a low segmentation accuracy on complex-content images. To address these challenges, we propose a novel hybrid architecture of convolutional neural networks hand in hand with vision Transformers (CiT-Net) for medical image segmentation. Our network has two advantages. First, we design a dynamic deformable convolution and apply it to the CNNs branch, which overcomes the weak feature extraction ability due to fixed-size convolution kernels and the stiff design of sharing kernel parameters among different inputs. Second, we design a shifted-window adaptive complementary attention module and a compact convolutional projection. We apply them to the Transformer branch to learn the cross-dimensional long-term dependency for medical images. Experimental results show that our CiT-Net provides better medical image segmentation results than popular SOTA methods. Besides, our CiT-Net requires lower parameters and less computational costs and does not rely on pre-training. The code is publicly available at https://github.com/SR0920/CiT-Net.

eess.IV

Quantum Graph Learning: Frontiers and Outlook

Quantum theory has shown its superiority in enhancing machine learning. However, facilitating quantum theory to enhance graph learning is in its infancy. This survey investigates the current advances in quantum graph learning (QGL) from three perspectives, i.e., underlying theories, methods, and prospects. We first look at QGL and discuss the mutualism of quantum theory and graph learning, the specificity of graph-structured data, and the bottleneck of graph learning, respectively. A new taxonomy of QGL is presented, i.e., quantum computing on graphs, quantum graph representation, and quantum circuits for graph neural networks. Pitfall traps are then highlighted and explained. This survey aims to provide a brief but insightful introduction to this emerging field, along with a detailed discussion of frontiers and outlook yet to be investigated.

cs.LG

Imaging topological torus lattice from an electron crystal in twisted mono-bilayer graphene

A variety of exotic quantum phases of matter have been created by Van der Waals heterostructures. Moreover, these twisted heterostructures provide a feasible way of braiding correlation effect and nontrivial band topology together. Here, through a comprehensive spectrum study, we report the discovery of topological torus lattice in twisted mono-bilayer graphene. The strong Coulomb correlations give rise to an unusual charge localization behavior within the moiré supercell, leading to an electron crystal. The nontrivial band topology is encoded into the electron crystal, which would result in spatial modulated Chern numbers, and is evidenced by an emergent topological torus lattice state. Our result illustrates an efficient strategy for entwining and engineering topological physics with a strong electron correlation.

cond-mat.mes-hall

A Scalable AI Approach for Clinical Trial Cohort Optimization

FDA has been promoting enrollment practices that could enhance the diversity of clinical trial populations, through broadening eligibility criteria. However, how to broaden eligibility remains a significant challenge. We propose an AI approach to Cohort Optimization (AICO) through transformer-based natural language processing of the eligibility criteria and evaluation of the criteria using real-world data. The method can extract common eligibility criteria variables from a large set of relevant trials and measure the generalizability of trial designs to real-world patients. It overcomes the scalability limits of existing manual methods and enables rapid simulation of eligibility criteria design for a disease of interest. A case study on breast cancer trial design demonstrates the utility of the method in improving trial generalizability.

cs.CL

Using Ecological Propensity Score to Adjust for Missing Confounders in Small Area Studies

Small area ecological studies are commonly used in epidemiology to assess the impact of area level risk factors on health outcomes when data are only available in an aggregated form. However the resulting estimates are often biased due to unmeasured confounders, which typically are not available from the standard administrative registries used for these studies. Extra information on confounders can be provided through external datasets such as surveys or cohorts, where the data are available at the individual level rather than at the area level; however such data typically lack the geographical coverage of administrative registries. We develop a framework of analysis which combines ecological and individual level data from different sources to provide an adjusted estimate of area level risk factors which is less biased. Our method (i) summarises all available individual level confounders into an area level scalar variable, which we call ecological propensity score (EPS), (ii) implements a hierarchical structured approach to predict the values of EPS whenever they are missing, (iii) includes the estimated and predicted EPS into the ecological regression linking the risk factors to the health outcome. Through a simulation study we show that integrating individual level data into small area analyses via EPS is a promising method to reduce the bias intrinsic in ecological studies due to unmeasured confounders; we also apply the method to a real case study to evaluate the effect of air pollution on coronary heart disease hospital admissions in Greater London.

stat.AP