SearcharxivSearch

arXiv subjects

Xinyun Zhang

Publications and source records attributed to Xinyun Zhang.

16 recordsLinked to original sources

Full Topological entropy of Xiong chaotic sets beyond the specification property

In this paper, we show that for a shift dynamical system $(X_{\mathcal{L}}, \sigma)$ with a weakened form of the specification property, there exists a Xiong chaotic set $C$ of $X_{\mathcal{L}}$ with full topological entropy everywhere (i.e. the intersection of $C$ and arbitrary non-empty open subset of $X_{\mathcal{L}}$ has full topological entropy). Moreover, $C$ satisfies that for any non-empty subset $A$ and any continuous map $F: A\rightarrow X_{\mathcal{L}},$ there exists an increasing sequence $\{p_{k}\}_{k=1}^{+\infty}$ of positive integers such that $\lim\limits_{k\rightarrow +\infty}\sigma^{p_{k}}(x)=F(x)$ holds for any $x\in A.$

math.DS

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessment (IEQA) methods remain primarily focused on single-image editing tasks and largely overlook the more challenging setting of MIE. This highlights the urgent need for a comprehensive and human-aligned benchmark for MIE. To this end, we introduce MIE-Bench, the first large-scale multiple image editing benchmark with fine-grained human preference annotations. Specifically, MIE-Bench includes 3,000 editing instances across 16 tasks, each involving more than two source images and an editing prompt, together with 36K edited images produced by 12 state-of-the-art editing models and over 108K mean opinion scores (MOSs) covering visual quality, instruction following, and attribute preservation. Based on MIE-Bench, we propose MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback for MIE. Extensive experiments show that MIEScore achieves state-of-the-art performance in aligning with human preferences and generalizes well across other IEQA datasets. Both the dataset and the model are available at https://github.com/IntMeGroup/MIEScore.

cs.CV

Exact approximation order of real numbers in Cantor series expansions

Let $Q = \{q_n\}_{n \ge 1}$ be a sequence of integers with $q_n \ge 2$ for all $n \in\mathbb{N}$. For any real number $x \in [0,1)$, it can be expanded into the following infinite series: $$x =\frac{\varepsilon_1(x)}{q_1}+ \frac{\varepsilon_2(x)}{q_1 q_2}+ \cdots+ \frac{\varepsilon_n(x)}{q_1 q_2 \cdots q_n}+ \cdots,$$ which is called the Cantor series expansion of $x$. We introduce the exact spproximation order in Cantor series expansions. It is analogous to the notion appearing in classical Diophantine approximation. More precisely, let $\omega_n(x)$ denote the $n$-th partial sum of the Cantor series expansion of $x$. For any monotonic function $\psi$, we study the metric theory of the set $E_c(\psi)$ of points that are exactly $\psi$-approximable by $\omega_n(x)$.

math.NT

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning

Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended modifications, and suboptimal aesthetics. Although several benchmarks and evaluation methods have been proposed, most existing approaches rely on scalar scores and lack interpretability. This limitation largely stems from the absence of high-quality interpretation datasets for TIE and effective reward models to train interpretable evaluators. To address these challenges, we introduce ReasonEdit-22K, the first dataset that combines 22K edited images with 113K Chain-of-Thought (CoT) samples, along with 1.3M human judgments assessing these interpretations in terms of logicality, accuracy, and usefulness. Building upon this dataset, we propose RE-Reward, a multimodal large language model (MLLM)-based reward model designed to provide human-aligned feedback for evaluating interpretable reasoning in image editing. Furthermore, we develop ReasonEdit, which is trained using reward signals derived from RE-Reward and the Group Relative Policy Optimization (GRPO) algorithm to learn an interpretable evaluation model. Extensive experiments demonstrate that ReasonEdit achieves superior alignment with human preferences and exhibits strong generalization across public benchmarks. In addition, it is capable of generating high-quality interpretable evaluation text, enabling more transparent and trustworthy assessment for image editing. The code is available at https://github.com/IntMeGroup/ReasonEdit.

cs.CV

EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing

Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks and methods have been proposed for evaluating edited images, scalable evaluation models are still lacking, which limits the development of human feedback reward models for image editing. To address the challenges, we first introduce \textbf{EditHF-1M}, a million-scale image editing dataset with over 29M human preference pairs and 148K human mean opinion ratings, both evaluated from three dimensions, \textit{i.e.}, visual quality, instruction alignment, and attribute preservation. Based on EditHF-1M, we propose \textbf{EditHF}, a multimodal large language model (MLLM) based evaluation model, to provide human-aligned feedback from image editing. Finally, we introduce \textbf{EditHF-Reward}, which utilizes EditHF as the reward signal to optimize the text-guided image editing models through reinforcement learning. Extensive experiments show that EditHF achieves superior alignment with human preferences and demonstrates strong generalization on other datasets. Furthermore, we fine-tune the Qwen-Image-Edit using EditHF-Reward, achieving significant performance improvements, which demonstrates the ability of EditHF to serve as a reward model to scale-up the image editing. Both the dataset and code will be released in our GitHub repository: https://github.com/IntMeGroup/EditHF.

cs.CV

Resonant Coupling Between Electromagnetic Waves and Protein Conformational Dynamics Revealed by Molecular Dynamics Simulations

The biological effects of electromagnetic fields on proteins remain controversial beyond well-established thermal mechanisms, particularly with respect to frequency-dependent responses. Here, we propose that electromagnetic waves can modulate protein conformation through resonant coupling with intrinsic protein dynamics. Molecular dynamics simulations were employed to characterize spontaneous conformational fluctuations in the absence of external fields, and a tiered screening strategy combined with fast Fourier transform analysis was used to identify dominant intrinsic frequencies associated with periodically fluctuating non-covalent atom or residue pairs. Oscillating external electric fields were subsequently applied at resonant and off-resonant frequencies to evaluate conformational responses across diverse protein systems. The results demonstrate that resonant excitation induces significantly enhanced backbone conformational deviations compared to off-resonant conditions, with the effect becoming more pronounced in structurally flexible and multichain proteins. These findings provide atomistic evidence for frequency-specific resonance between electromagnetic fields and protein conformational dynamics, offering mechanistic insight into frequency-dependent electromagnetic effects and a computational framework for electromagnetic wave-based modulation of protein function.

physics.chem-ph

Hausdorff dimension of the Cartesian product of exact approximation set in $\beta$-expansions

In this paper, we study the metrical theory of Cartesian products of exact approximation sets in $\beta$-expansions. More precisely, for an integer $d \ge 2$ and real numbers $\beta_i > 1$ $(1 \le i \le d)$, we consider the set of points $x_i \in [0,1)$ is approximable by its convergents in the $\beta_i$-expansion to order $\psi_i$, but not to any better order. For any non-increasing functions $\psi_i$, we determine the Hausdorff dimension of the Cartesian product of these sets.

math.NT

SWE-QA: Can Language Models Answer Repository-level Code Questions?

Understanding and reasoning about entire software repositories is an essential capability for intelligent software engineering tools. While existing benchmarks such as CoSQA and CodeQA have advanced the field, they predominantly focus on small, self-contained code snippets. These setups fail to capture the complexity of real-world repositories, where effective understanding and reasoning often require navigating multiple files, understanding software architecture, and grounding answers in long-range code dependencies. In this paper, we present SWE-QA, a repository-level code question answering (QA) benchmark designed to facilitate research on automated QA systems in realistic code environments. SWE-QA involves 576 high-quality question-answer pairs spanning diverse categories, including intention understanding, cross-file reasoning, and multi-hop dependency analysis. To construct SWE-QA, we first crawled 77,100 GitHub issues from 11 popular repositories. Based on an analysis of naturally occurring developer questions extracted from these issues, we developed a two-level taxonomy of repository-level questions and constructed a set of seed questions for each category. For each category, we manually curated and validated questions and collected their corresponding answers. As a prototype application, we further develop SWE-QA-Agent, an agentic framework in which LLM agents reason and act to find answers automatically. We evaluate six advanced LLMs on SWE-QA under various context augmentation strategies. Experimental results highlight the promise of LLMs, particularly our SWE-QA-Agent framework, in addressing repository-level QA, while also revealing open challenges and pointing to future research directions.

cs.CL

ChatEDA: A Large Language Model Powered Autonomous Agent for EDA

The integration of a complex set of Electronic Design Automation (EDA) tools to enhance interoperability is a critical concern for circuit designers. Recent advancements in large language models (LLMs) have showcased their exceptional capabilities in natural language processing and comprehension, offering a novel approach to interfacing with EDA tools. This research paper introduces ChatEDA, an autonomous agent for EDA empowered by an LLM, AutoMage, complemented by EDA tools serving as executors. ChatEDA streamlines the design flow from the Register-Transfer Level (RTL) to the Graphic Data System Version II (GDSII) by effectively managing task decomposition, script generation, and task execution. Through comprehensive experimental evaluations, ChatEDA has demonstrated its proficiency in handling diverse requirements, and our fine-tuned AutoMage model has exhibited superior performance compared to GPT-4 and other similar LLMs.

cs.AR

Hausdorff dimension of some exceptional sets in Lüroth expansions

In this paper, we study the metrical theory of the growth rate of digits in Lüroth expansions. More precisely, for $ x\in \left( 0,1 \right] $, let $ \left[ d_1\left( x \right) ,d_2\left( x \right) ,\cdots \right] $ denote the Lüroth expansion of $ x $, we completely determine the Hausdorff dimension of the following sets \begin{align*} E_{\mathrm{sup}}\left( ψ\right) =\Big\{ x\in \left( 0,1 \right] :\limsup\limits_{n\rightarrow \infty}\frac{\log d_n\left( x \right)}{ψ\left( n \right)}=1 \Big\} , \end{align*} \begin{align*} E\left( ψ\right) =\Big\{ x\in \left( 0,1 \right] :\lim_{n\rightarrow \infty}\frac{\log d_n\left( x \right)}{ψ\left( n \right)}=1 \Big\} \end{align*} and \begin{align*} E_{\mathrm{inf}}\left( ψ\right) =\Big\{ x\in \left( 0,1 \right] : \liminf_{n\rightarrow \infty}\frac{\log d_n\left( x \right)}{ψ\left( n \right)}=1 \Big\} , \end{align*} where $ ψ:\mathbb{N} \rightarrow \mathbb{R} ^+ $ is an arbitrary function satisfying $ ψ\left( n \right) \rightarrow \infty$ as $n\rightarrow \infty$.

math.NT

Towards Versatile and Efficient Visual Knowledge Integration into Pre-trained Language Models with Cross-Modal Adapters

Humans learn language via multi-modal knowledge. However, due to the text-only pre-training scheme, most existing pre-trained language models (PLMs) are hindered from the multi-modal information. To inject visual knowledge into PLMs, existing methods incorporate either the text or image encoder of vision-language models (VLMs) to encode the visual information and update all the original parameters of PLMs for knowledge fusion. In this paper, we propose a new plug-and-play module, X-adapter, to flexibly leverage the aligned visual and textual knowledge learned in pre-trained VLMs and efficiently inject them into PLMs. Specifically, we insert X-adapters into PLMs, and only the added parameters are updated during adaptation. To fully exploit the potential in VLMs, X-adapters consist of two sub-modules, V-expert and T-expert, to fuse VLMs' image and text representations, respectively. We can opt for activating different sub-modules depending on the downstream tasks. Experimental results show that our method can significantly improve the performance on object-color reasoning and natural language understanding (NLU) tasks compared with PLM baselines.

cs.CL

Existence of infinitely many solutions for a critical Hartree type equation with potential: local Pohožaev identities methods

This paper deals with the following equation $$-Δu =K(|x'|, x'')\Big(|x|^{-α}\ast (K(|x'|, x'')|u|^{2^{\ast}_α})\Big) |u|^{2^{\ast}_α-2}u\quad\mbox{in}\ \mathbb{R}^N,$$ where $N\geq5$, $α>5-\frac{6}{N-2}$, $2^{\ast}_α=\frac{2N-α}{N-2}$ is the so-called upper critical exponent in the Hardy-Littlewood-Sobolev inequality and $K(|x'|, x'')$, where $(x',x'')\in \mathbb{R}^2\times\mathbb{R}^{N-2}$, is bounded and nonnegative. Under proper assumptions on the potential function $K$, we obtain the existence of infinitely many solutions for the nonlocal critical equation by using a finite dimensional reduction argument and local Pohožaev identities. It is a remarkable fact that the order of the Riesz potential influences the existence/non-existence of solutions.

math.AP

p-Laplacian Adaptation for Generative Pre-trained Vision-Language Models

Vision-Language models (VLMs) pre-trained on large corpora have demonstrated notable success across a range of downstream tasks. In light of the rapidly increasing size of pre-trained VLMs, parameter-efficient transfer learning (PETL) has garnered attention as a viable alternative to full fine-tuning. One such approach is the adapter, which introduces a few trainable parameters into the pre-trained models while preserving the original parameters during adaptation. In this paper, we present a novel modeling framework that recasts adapter tuning after attention as a graph message passing process on attention graphs, where the projected query and value features and attention matrix constitute the node features and the graph adjacency matrix, respectively. Within this framework, tuning adapters in VLMs necessitates handling heterophilic graphs, owing to the disparity between the projected query and value space. To address this challenge, we propose a new adapter architecture, $p$-adapter, which employs $p$-Laplacian message passing in Graph Neural Networks (GNNs). Specifically, the attention weights are re-normalized based on the features, and the features are then aggregated using the calibrated attention matrix, enabling the dynamic exploitation of information with varying frequencies in the heterophilic attention graphs. We conduct extensive experiments on different pre-trained VLMs and multi-modal tasks, including visual question answering, visual entailment, and image captioning. The experimental results validate our method's significant superiority over other PETL methods.

cs.CV

Nondegeneracy of positive solutions for a biharmonic hartree equation and its applications

In this paper, we are interested in some problems related to the following biharmonic hartree equation \begin{equation*} Δ^{2} u=(|x|^{-α}\ast |u|^{p})u^{p-1},\sp \text{in}\quad\R^N. \end{equation*} where $p=\frac{2N-α}{N-4}$, $N\geq 9$ and $0<α<N$. First, by using the spherical harmonic decomposition and the Funk-Heck formula of the spherical harmonic functions, we prove the nondegeneracy of the positive solutions of the above biharmonic equation. As applications, we investigate the stability of a version of nonlocal Sobolev inequality \begin{equation} \int_{\R^N}|Δu|^{2} \geq S^{*}\left( \int_{\R^N}\Big(|x|^{-α}\ast |u|^{p}\Big)u^{p} dx\right)^{\frac{1}{p}}, \end{equation} and give a gradient form remainder. Moreover, by applying a finite dimension reduction and local Poho$\check{z}$aev identity, we can also construct multi-bubble solutions for the following equation with potential \begin{equation*} Δ^2 u+V(|x'|, x'')u =\Big(|x|^{-α}\ast |u|^{p}\Big)u^{p-1}\hspace{4.14mm}x\in \mathbb{R}^N. \end{equation*} where $N\geq9$, $(x',x'')\in \mathbb{R}^2\times\mathbb{R}^{N-2}$ and $V(|x'|, x'')$ is a bounded and nonnegative function. We will show what is the role fo the dimension and the order of the Riesz potential in proving the existence result. In fact, the existence result is restricted to the range $6-\frac{12}{N-4}\leqα<N$.

math.AP

Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts

Meetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content. To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and efficient meeting summarization. RbS first leverages a self-supervised paradigm to annotate essential contents by reconstructing the meeting transcripts. Secondly, we propose a relative positional bucketing (RPB) algorithm to equip (conventional) summarization models to generate the summary. Despite the additional reconstruction process, our proposed RPB significantly compressed the input, leading to faster processing and reduced memory consumption compared to traditional summarization methods. We validate the effectiveness and efficiency of our method through extensive evaluations and analysis. On two meeting summarization datasets, AMI and ICSI, our approach outperforms previous state-of-the-art approaches without relying on large-scale pre-training or expert-grade annotating tools.

cs.CL

Out-of-distribution Few-shot Learning For Edge Devices without Model Fine-tuning

Few-shot learning (FSL) via customization of a deep learning network with limited data has emerged as a promising technique to achieve personalized user experiences on edge devices. However, existing FSL methods primarily assume independent and identically distributed (IID) data and utilize either computational backpropagation updates for each task or a common model with task-specific prototypes. Unfortunately, the former solution is infeasible for edge devices that lack on-device backpropagation capabilities, while the latter often struggles with limited generalization ability, especially for out-of-distribution (OOD) data. This paper proposes a lightweight, plug-and-play FSL module called Task-aware Normalization (TANO) that enables efficient and task-aware adaptation of a deep neural network without backpropagation. TANO covers the properties of multiple user groups by coordinating the updates of several groups of the normalization statistics during meta-training and automatically identifies the appropriate normalization group for a downstream few-shot task. Consequently, TANO provides stable but task-specific estimations of the normalization statistics to close the distribution gaps and achieve efficient model adaptation. Results on both intra-domain and out-of-domain generalization experiments demonstrate that TANO outperforms recent methods in terms of accuracy, inference speed, and model size. Moreover, TANO achieves promising results on widely-used FSL benchmarks and data from real applications.

cs.CV