Searcharxiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 145 records · Page 8Linked to original sources

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Achieving better alignment between vision embeddings and Large Language Models (LLMs) is crucial for enhancing the abilities of Multimodal LLMs (MLLMs), particularly for recent models that rely on powerful pretrained vision encoders and LLMs. A common approach to connect the pretrained vision encoder and LLM is through a projector applied after the vision encoder. However, the projector is often trained to enable the LLM to generate captions, and hence the mechanism by which LLMs understand each vision token remains unclear. In this work, we first investigate the role of the projector in compressing vision embeddings and aligning them with word embeddings. We show that the projector significantly compresses visual information, removing redundant details while preserving essential elements necessary for the LLM to understand visual content. We then examine patch-level alignment -- the alignment between each vision patch and its corresponding semantic words -- and propose a *multi-semantic alignment hypothesis*. Our analysis indicates that the projector trained by caption loss improves patch-level alignment but only to a limited extent, resulting in weak and coarse alignment. To address this issue, we propose *patch-aligned training* to efficiently enhance patch-level alignment. Our experiments show that patch-aligned training (1) achieves stronger compression capability and improved patch-level alignment, enabling the MLLM to generate higher-quality captions, (2) improves the MLLM's performance by 16% on referring expression grounding tasks, 4% on question-answering tasks, and 3% on modern instruction-following benchmarks when using the same supervised fine-tuning (SFT) setting. The proposed method can be easily extended to other multimodal models.

cs.CV↗

STRCMP: Integrating Graph Structural Priors with Language Models for Combinatorial Optimization

Combinatorial optimization (CO) problems, central to operation research and theoretical computer science, present significant computational challenges due to their NP-hard nature. While large language models (LLMs) have emerged as promising tools for CO--either by directly generating solutions or synthesizing solver-specific codes--existing approaches often neglect critical structural priors inherent to CO problems, leading to suboptimality and iterative inefficiency. Inspired by human experts' success in leveraging CO structures for algorithm design, we propose STRCMP, a novel structure-aware LLM-based algorithm discovery framework that systematically integrates structure priors to enhance solution quality and solving efficiency. Our framework combines a graph neural network (GNN) for extracting structural embeddings from CO instances with an LLM conditioned on these embeddings to identify high-performing algorithms in the form of solver-specific codes. This composite architecture ensures syntactic correctness, preserves problem topology, and aligns with natural language objectives, while an evolutionary refinement process iteratively optimizes generated algorithm. Extensive evaluations across Mixed Integer Linear Programming and Boolean Satisfiability problems, using nine benchmark datasets, demonstrate that our proposed STRCMP outperforms five strong neural and LLM-based methods by a large margin, in terms of both solution optimality and computational efficiency. The code and learned model will be publicly available upon the acceptance of the paper.

cs.LG↗

Resolved ALMA [CII] 158 micron Observations at Cosmic Noon: ISM Structure and Dynamics of Starbursting QSO SDSSJ1000

We present spatially resolved Alma Band-9 observations of the [CII] 158 $μ$m fine structure line from an optically selected quasar, SDSS J100038.01+020822.4 (J1000), at z=1.8275. By utilizing [OI] 63 $μ$m line observations from Herschel/PACS and constructing a detailed dust SED using Herschel and Spitzer archival imaging data, we show that the [CII] line emission is well explained by a photodissociation region (PDR) model, in which the emission arises from the surfaces of molecular clouds exposed to far-UV radiation fields $\sim 5\cdot10^3$ times the local interstellar radiation field (G$_0$). We find a factor of 30 variation in spatially resolved [CII]/Far-IR continuum across the source which is explained by the reduced fraction of cooling via [CII] line emission at such high far-UV field strengths. By matching derived PDR parameters to the observed far-IR line and continuum intensities we derive cloud size-scales and find that typical cloud radii in J1000 are $\sim$3.5 pc perhaps indicating an ISM that is highly fractured due to intense star formation activity. We model the galaxy dynamically and find that the [CII] emission is contained within a compact, dynamically cold disk with v/$σ$=6.2, consistent with cosmological simulations. We also report the discovery of a companion galaxy to j1000 confirmed by the detection of [CII] and use recently obtained JWST/NirCAM imaging of the system to argue for J1000 being an interacting system. With total stellar mass $\sim 1.5 \times 10^{10}$ M$_\odot$ and main-component dynamical mass $\gtrsim 10^{11}$ M$_\odot$, the J1000 system is a progenitor to the most massive galaxies seen in the local Universe.

astro-ph.GA↗

Anti-classification results for conjugacy of diffeomorphisms on manifolds

We show that the topological conjugacy relation of diffeomorphisms on any manifold of dimension at least 2 is not classifiable by countable structures. This answers a question of Foreman and Gorodetski. We also prove that $E_0$ is reducible into the topological conjugacy relation of minimal diffeomorphisms on the 2-torus, which answers a question of Foreman.

math.DS↗

VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction

Objective: Ulcerative colitis (UC), characterized by chronic inflammation with alternating remission-relapse cycles, requires precise histological healing (HH) evaluation to improve clinical outcomes. To overcome the limitations of annotation-intensive deep learning methods and suboptimal multi-instance learning (MIL) in HH prediction, we propose VIGIL, the first vision-language guided MIL framework integrating white light endoscopy (WLE) and endocytoscopy (EC). Methods:VIGIL begins with a dual-branch MIL module KS-MIL based on top-K typical frames selection and similarity metric adaptive learning to learn relationships among frame features effectively. By integrating the diagnostic report text and specially designed multi-level alignment and supervision between image-text pairs, VIGIL establishes joint image-text guidance during training to capture richer disease-related semantic information. Furthermore, VIGIL employs a multi-modal masked relation fusion (MMRF) strategy to uncover the latent diagnostic correlations of two endoscopic image representations. Results:Comprehensive experiments on a real-world clinical dataset demonstrate VIGIL's superior performance, achieving 92.69\% accuracy and 94.79\% AUC, outperforming existing state-of-the-art methods. Conclusion: The proposed VIGIL framework successfully establishes an effective vision-language guided MIL paradigm for UC HH prediction, reducing annotation burdens while improving prediction reliability. Significance: The research outcomes provide new insights for non-invasive UC diagnosis and hold theoretical significance and clinical value for advancing intelligent healthcare development.

q-bio.QM↗

SAPIENT: Mastering Multi-turn Conversational Recommendation with Strategic Planning and Monte Carlo Tree Search

Conversational Recommender Systems (CRS) proactively engage users in interactive dialogues to elicit user preferences and provide personalized recommendations. Existing methods train Reinforcement Learning (RL)-based agent with greedy action selection or sampling strategy, and may suffer from suboptimal conversational planning. To address this, we present a novel Monte Carlo Tree Search (MCTS)-based CRS framework SAPIENT. SAPIENT consists of a conversational agent (S-agent) and a conversational planner (S-planner). S-planner builds a conversational search tree with MCTS based on the initial actions proposed by S-agent to find conversation plans. The best conversation plans from S-planner are used to guide the training of S-agent, creating a self-training loop where S-agent can iteratively improve its capability for conversational planning. Furthermore, we propose an efficient variant SAPIENT for trade-off between training efficiency and performance. Extensive experiments on four benchmark datasets validate the effectiveness of our approach, showing that SAPIENT outperforms the state-of-the-art baselines. Our code and data are accessible through https://github.com/ninglab/SAPIENT.

cs.CL↗

An Unsupervised Learning Method for Radio Interferometry Deconvolution

Given the incomplete sampling of spatial frequencies by radio interferometers, achieving precise restoration of astrophysical information remains challenging. To address this ill-posed problem, compressive sensing(CS) provides a robust framework for stable and unique recovery of sky brightness distributions in noisy environments, contingent upon satisfying specific conditions. We explore the applicability of CS theory and find that for radio interferometric telescopes, the conditions can be simplified to sparse representation. {Building on this insight, we develop a deep dictionary (realized through a convolutional neural network), which is designed to be multi-resolution and overcomplete, to achieve sparse representation and integrate it within the CS framework. The resulting method is a novel, fully interpretable unsupervised learning approach that combines} the mathematical rigor of CS with the expressive power of deep neural networks, effectively bridging the gap between deep learning and classical dictionary methods. {During the deconvolution process, the model image and the deep dictionary are updated alternatively.} This approach enables efficient and accurate recovery of extended sources with complex morphologies from noisy measurements. Comparative analyses with state-of-the-art algorithms demonstrate the outstanding performance of our method, i.e., achieving a dynamic range (DR) nearly 45 to 100 times higher than that of multiscale CLEAN (MS-CLEAN).

astro-ph.IM↗

Monolayer C$_{60}$ networks: A first-principles perspective

Monolayer fullerene (C$_{60}$) networks combine molecular-level rigidity with crystalline connectivity, offering a promising platform for numerous applications. In this Feature article, we review the physical and chemical properties of fullerene monolayers, focusing on first-principles studies. We first explore the structural stability of monolayer phases and investigate their thermal expansion behaviours. We then outline criteria for photocatalytic water splitting and introduce theoretical predictions which are supported by recent experimental verification. Finally, we show how interlayer stacking, molecular size, and dimensional tuning (from 2D monolayers into 3D crystals, 1D chains, or nanoribbons) offer versatile approaches to modulate their chemical functionality. Together, these insights establish fullerene networks as a novel class of carbon-based materials with tailored properties for catalysis, photovoltaics, and flexible electronics.

cond-mat.mtrl-sci↗

FEASTS Combined with Interferometry. III. The Low Column Density HI Around M51 and Possibility of Turbulent-mixing Gas Accretion

With a new joint-deconvolution pipeline, we combine the single-dish and interferometric atomic hydrogen (HI) data of M51 observed by the Five-hundred-meter Aperture Spherical radio Telescope (FAST) (FEASTS program) and the Very Large Array (VLA) (THINGS). The product data cube has a typical line width of $13\,\text{km}\,\text{s}^{-1}$ and a $2σ$ line-of-sight (LOS) sensitivity of HI column density $N_\text{HI}\sim3.2\times10^{18}\,\text{cm}^{-2}$ at a spatial resolution of ${\sim}18''$ (${\sim}0.7\,\text{kpc}$). Among the HI-detected LOSs extending to ${\sim}50\,\text{kpc}$, ${\sim}89\%$ consist of diffuse HI only, which is missed by previous VLA observations. The distribution of dense HI is reproduced by previous hydrodynamical simulations of this system, but the diffuse component is not, likely due to unresolved physics related to the interaction between the circumgalactic and interstellar media. With simple models, we find that these low-$N_\text{HI}$ structures could survive the background ultraviolet photoionization, but are susceptible to the thermal evaporation. We find a positive correlation between LOS velocity dispersion ($σ_v$) and $N_\text{HI}$ with a logarithmic index of ${\sim}0.5$. Based on existing turbulent mixing layer (TML) theories and simulations, we propose a scenario of hot gas cooling and accreting onto the disk through a TML, which could reproduce the observed power index of ${\sim}0.5$. We estimate the related cooling and accretion rates to be roughly one-third to two-thirds of the star-formation rate. A typical column density of diffuse HI (${\sim}10^{19}\,\text{cm}^{-2}$) can be accreted within $300\,\text{Myr}$, the interaction time scale previously estimated for the system. Such a gas accretion channel has been overlooked before, and may be important for gas-rich interacting systems and for high redshift galaxy evolution.

astro-ph.GA↗

Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detection

The increasing complexity of industrial anomaly detection (IAD) has positioned multimodal detection methods as a focal area of machine vision research. However, dedicated multimodal datasets specifically tailored for IAD remain limited. Pioneering datasets like MVTec 3D have laid essential groundwork in multimodal IAD by incorporating RGB+3D data, but still face challenges in bridging the gap with real industrial environments due to limitations in scale and resolution. To address these challenges, we introduce Real-IAD D3, a high-precision multimodal dataset that uniquely incorporates an additional pseudo3D modality generated through photometric stereo, alongside high-resolution RGB images and micrometer-level 3D point clouds. Real-IAD D3 features finer defects, diverse anomalies, and greater scale across 20 categories, providing a challenging benchmark for multimodal IAD Additionally, we introduce an effective approach that integrates RGB, point cloud, and pseudo-3D depth information to leverage the complementary strengths of each modality, enhancing detection performance. Our experiments highlight the importance of these modalities in boosting detection robustness and overall IAD performance. The dataset and code are publicly accessible for research purposes at https://realiad4ad.github.io/Real-IAD D3

cs.CV↗

Diffusion-empowered AutoPrompt MedSAM

MedSAM, a medical foundation model derived from the SAM architecture, has demonstrated notable success across diverse medical domains. However, its clinical application faces two major challenges: the dependency on labor-intensive manual prompt generation, which imposes a significant burden on clinicians, and the absence of semantic labeling in the generated segmentation masks for organs or lesions, limiting its practicality for non-expert users. To address these limitations, we propose AutoMedSAM, an end-to-end framework derived from SAM, designed to enhance usability and segmentation performance. AutoMedSAM retains MedSAM's image encoder and mask decoder structure while introducing a novel diffusion-based class prompt encoder. The diffusion-based encoder employs a dual-decoder structure to collaboratively generate prompt embeddings guided by sparse and dense prompt definitions. These embeddings enhance the model's ability to understand and process clinical imagery autonomously. With this encoder, AutoMedSAM leverages class prompts to embed semantic information into the model's predictions, transforming MedSAM's semi-automated pipeline into a fully automated workflow. Furthermore, AutoMedSAM employs an uncertainty-aware joint optimization strategy during training to effectively inherit MedSAM's pre-trained knowledge while improving generalization by integrating multiple loss functions. Experimental results across diverse datasets demonstrate that AutoMedSAM achieves superior performance while broadening its applicability to both clinical settings and non-expert users. Code is available at https://github.com/HP-ML/AutoPromptMedSAM.git.

eess.IV↗

Boosting Multi-View Stereo with Depth Foundation Model in the Absence of Real-World Labels

Learning-based Multi-View Stereo (MVS) methods have made remarkable progress in recent years. However, how to effectively train the network without using real-world labels remains a challenging problem. In this paper, driven by the recent advancements of vision foundation models, a novel method termed DFM-MVS, is proposed to leverage the depth foundation model to generate the effective depth prior, so as to boost MVS in the absence of real-world labels. Specifically, a depth prior-based pseudo-supervised training mechanism is developed to simulate realistic stereo correspondences using the generated depth prior, thereby constructing effective supervision for the MVS network. Besides, a depth prior-guided error correction strategy is presented to leverage the depth prior as guidance to mitigate the error propagation problem inherent in the widely-used coarse-to-fine network structure. Experimental results on DTU and Tanks & Temples datasets demonstrate that the proposed DFM-MVS significantly outperforms existing MVS methods without using real-world labels.

cs.CV↗

Electronic structure of fullerene nanoribbons

Using first-principles calculations, we examine the electronic structure of quasi-one-dimensional fullerene nanoribbons derived from two-dimensional fullerene networks. Depending on the edge geometry and width, these nanoribbons exhibit a rich variety of properties, including direct and indirect band gaps, positive and negative effective masses, as well as dispersive and flat bands. Our findings establish a comprehensive understanding of the electronic properties of fullerene nanoribbons, with potential implications for the design of future nanoscale devices.

cond-mat.mtrl-sci↗

Toeplitz subshifts of finite rank

In this paper we study some basic problems about Toeplitz subshifts of finite topological rank. We define the notion of a strong Toeplitz subshift of finite rank $K$ by combining the characterizations of Toeplitz-ness and of finite topological rank $K$ from the point of view of the Bratteli--Vershik representation or from the $\mathcal{S}$-adic point of view. The characterization problem asks if for every $K\geq 2$, every Toeplitz subshift of topological rank $K$ is a strong Toeplitz subshift of rank $K$. We give a negative answer to the characterization problem by constructing a Toeplitz subshift of topological rank $2$ which fails to be a strong Toeplitz subshift of rank $2$. However, we show that the set of all strong Toeplitz subshifts of finite rank is generic in the space of all infinite minimal subshifts. In the second part we consider several classification problems for Toeplitz subshifts of topological rank $2$ from the point of view of descriptive set theory. We completely determine the complexity of the conjugacy problem, the flip conjugacy problem, and the bi-factor problem by showing that, as equivalence relations, they are hyperfinite and not smooth. We also consider the inverse problem for all Toeplitz subshifts. We give a criterion for when a Toeplitz subshift is conjugate to its own inverse, and use it to show that the set of all such Toeplitz subshifts is a meager set in the space of all infinite minimal subshifts. Finally, we show that the automorphism group of any Toeplitz subshift of finite rank is isomorphic to $\mathbb{Z}\oplus C$ for some finite cyclic group $C$, and for every nontrivial finite cyclic group $C$, $\mathbb{Z}\oplus C$ can be realized as the isomorphism type of an automorphism group of a strong Toeplitz subshift of finite rank greater than $2$.

math.DS↗

Enhancing Adversarial Transferability via Component-Wise Transformation

Deep Neural Networks (DNNs) are highly vulnerable to adversarial examples, which pose significant challenges in security-sensitive applications. Among various adversarial attack strategies, input transformation-based attacks have demonstrated remarkable effectiveness in enhancing adversarial transferability. However, existing methods still perform poorly across different architectures, even though they have achieved promising results within the same architecture. This limitation arises because, while models of the same architecture may focus on different regions of the object, the variation is even more pronounced across different architectures. Unfortunately, current approaches fail to effectively guide models to attend to these diverse regions. To address this issue, this paper proposes a novel input transformation-based attack method, termed Component-Wise Transformation (CWT). CWT applies interpolation and selective rotation to individual image blocks, ensuring that each transformed image highlights different target regions, thereby improving the transferability of adversarial examples. Extensive experiments on the standard ImageNet dataset show that CWT consistently outperforms state-of-the-art methods in both attack success rates and stability across CNN- and Transformer-based models.

cs.CV↗

RWKV-7 "Goose" with Expressive Dynamic State Evolution

We present RWKV-7 "Goose", a new sequence modeling architecture with constant memory usage and constant inference time per token. Despite being trained on dramatically fewer tokens than other top models, our 2.9 billion parameter language model achieves a new 3B SoTA on multilingual tasks and matches the current 3B SoTA on English language downstream performance. RWKV-7 introduces a newly generalized formulation of the delta rule with vector-valued gating and in-context learning rates, as well as a relaxed value replacement rule. We show that RWKV-7 can perform state tracking and recognize all regular languages, while retaining parallelizability of training. This exceeds the capabilities of Transformers under standard complexity conjectures, which are limited to $\mathsf{TC}^0$. To demonstrate RWKV-7's language modeling capability, we also present an extended open source 3.1 trillion token multilingual corpus, and train four RWKV-7 models ranging from 0.19 billion to 2.9 billion parameters on this dataset. To foster openness, reproduction, and adoption, we release our models and dataset component listing at https://huggingface.co/RWKV, and our training and inference code at https://github.com/RWKV/RWKV-LM all under the Apache 2.0 License.

cs.CL↗

Unleashed from Constrained Optimization: Quantum Computing for Quantum Chemistry Employing Generator Coordinate Inspired Method

Hybrid quantum-classical approaches offer potential solutions to quantum chemistry problems, yet they often manifest as constrained optimization problems. Here, we explore the interconnection between constrained optimization and generalized eigenvalue problems through the Unitary Coupled Cluster (UCC) excitation generators. Inspired by the generator coordinate method, we employ these UCC excitation generators to construct non-orthogonal, overcomplete many-body bases, projecting the system Hamiltonian into an effective Hamiltonian, which bypasses issues such as barren plateaus that heuristic numerical minimizers often encountered in standard variational quantum eigensolver (VQE). Diverging from conventional quantum subspace expansion methods, we introduce an adaptive scheme that robustly constructs the many-body basis sets from a pool of the UCC excitation generators. This scheme supports the development of a hierarchical ADAPT quantum-classical strategy, enabling a balanced interplay between subspace expansion and ansatz optimization to address complex, strongly correlated quantum chemical systems cost-effectively, setting the stage for more advanced quantum simulations in chemistry.

quant-ph↗

GraphTEN: Graph Enhanced Texture Encoding Network

Texture recognition is a fundamental problem in computer vision and pattern recognition. Recent progress leverages feature aggregation into discriminative descriptions based on convolutional neural networks (CNNs). However, modeling non-local context relations through visual primitives remains challenging due to the variability and randomness of texture primitives in spatial distributions. In this paper, we propose a graph-enhanced texture encoding network (GraphTEN) designed to capture both local and global features of texture primitives. GraphTEN models global associations through fully connected graphs and captures cross-scale dependencies of texture primitives via bipartite graphs. Additionally, we introduce a patch encoding module that utilizes a codebook to achieve an orderless representation of texture by encoding multi-scale patch features into a unified feature space. The proposed GraphTEN achieves superior performance compared to state-of-the-art methods across five publicly available datasets.

cs.CV↗