SearcharxivSearch

arXiv subjects

Zhongyang Li

Publications and source records attributed to Zhongyang Li.

At least 19 recordsLinked to original sources

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs

Multimodal large language models (MLLMs) typically employ resampling-based projectors to transform dense visual features into a compact token sequence for language modeling. Most existing resamplers adopt a single, fixed aggregation scope via global cross-attention, which can blur fine-grained local evidence and limit the ability to capture both local details and global context within a fixed token budget. In this work, we propose MS-Resampler, a multi-scope visual resampling framework for MLLMs. MS-Resampler instantiates multiple scope-specific resamplers by injecting explicit spatial scope priors into the resampling attention, enabling each branch to aggregate visual information at a particular granularity from local to global. The outputs of these scope-specific resamplers are then adaptively fused to produce the final visual representations for language modeling. Extensive experiments on ten public multimodal benchmarks show that MS-Resampler consistently improves visual understanding and multimodal reasoning over conventional single-scope resamplers, while introducing only minimal computational overhead.

cs.CV

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how humans interact with information. However, frontier systems remain proprietary, while existing open agents often generalize poorly across different task types, leaving unclear how to train a broadly capable deep research agent. We release QUEST, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of long-horizon search tasks, with strong capabilities in fact seeking, citation grounding, and report synthesis. To build QUEST, we propose an effective training recipe combining mid-training, supervised fine-tuning, and reinforcement learning. Central to this recipe is a curated data synthesis pipeline based on unified rubric trees, which applies to different task types and enables synthesizing training data with verifiable rewards without human annotation. In addition, QUEST incorporates a built-in context management mechanism that enables effective long-horizon reasoning and knowledge synthesis. Using only 8K synthesized tasks, QUEST approaches or even surpasses frontier closed-source agents across eight deep research benchmarks spanning diverse task types, and achieves the best overall performance among recent open-weight agents. We released everything: models, data, and training scripts.

cs.CL

QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression

Multimodal large language models suffer from severe computational and memory bottlenecks, as the number of visual tokens far exceeds that of textual tokens. While recent methods employ projector modules to align and compress visual tokens into text-aligned features, they typically depend on fixed heuristics that limit adaptability across diverse scenarios. In this paper, we first propose Query Guided Mixture-of-Projector (QMoP), a novel and flexible framework that adaptively compresses visual tokens via three collaborative branches: (1) a pooling-based branch for coarse-grained global semantics, (2) a resampler branch for extracting high-level semantic representations, and (3) a pruning-based branch for fine-grained token selection to preserve critical visual detail. To adaptively coordinate these branches, we introduce the Query Guided Router (QGR), which dynamically selects and weights the outputs from different branches based on both visual input and textual queries. A Mixture-of-Experts-style fusion mechanism is designed to aggregate the outputs, harnessing the strengths of each strategy while suppressing noise. To systematically evaluate the effects of Visual Token Compression, we also develop VTCBench, a dedicated benchmark for evaluating the information loss induced by visual token compression. Extensive experiments demonstrate that despite relying on fundamental compression modules, QMoP outperforms strong baselines and delivers significant savings in memory, computation, and inference time.

cs.CV

Scene-Aware Memory Discrimination: Deciding Which Personal Knowledge Stays

Intelligent devices have become deeply integrated into everyday life, generating vast amounts of user interactions that form valuable personal knowledge. Efficient organization of this knowledge in user memory is essential for enabling personalized applications. However, current research on memory writing, management, and reading using large language models (LLMs) faces challenges in filtering irrelevant information and in dealing with rising computational costs. Inspired by the concept of selective attention in the human brain, we introduce a memory discrimination task. To address large-scale interactions and diverse memory standards in this task, we propose a Scene-Aware Memory Discrimination method (SAMD), which comprises two key components: the Gating Unit Module (GUM) and the Cluster Prompting Module (CPM). GUM enhances processing efficiency by filtering out non-memorable interactions and focusing on the salient content most relevant to application demands. CPM establishes adaptive memory standards, guiding LLMs to discern what information should be remembered or discarded. It also analyzes the relationship between user intents and memory contexts to build effective clustering prompts. Comprehensive direct and indirect evaluations demonstrate the effectiveness and generalization of our approach. We independently assess the performance of memory discrimination, showing that SAMD successfully recalls the majority of memorable data and remains robust in dynamic scenarios. Furthermore, when integrated into personalized applications, SAMD significantly enhances both the efficiency and quality of memory construction, leading to better organization of personal knowledge.

cs.CL

Self-avoiding walks on cubic graphs and local transformations

Despite its elementary definition, the self-avoiding walk (SAW) poses notoriously hard enumerative problems: exact connective constants are known for only a handful of infinite graphs, notably the honeycomb lattice \cite{ds}. We establish a general substitution principle for SAWs on infinite connected quasi-transitive cubic graphs under port-transitive vertex replacements, where each degree-$3$ vertex is replaced by a fixed finite three-port gadget. Writing $g(x)$ for the associated two-port SAW series, we prove that for $G_1=\phi(G)$, \[ \mu(G)^{-1}=g\bigl(\mu(G_1)^{-1}\bigr), \] equivalently $\mu(G_1)^{-1}$ is the unique solution $x\in(0,1)$ of $g(x)=\mu(G)^{-1}$, thereby extending the Fisher-triangle relation of Grimmett--Li to arbitrary symmetric three-port gadgets. We also obtain the corresponding identity for bipartite graphs when one or both colour classes are transformed, and show that the critical exponents $\gamma$ and $\eta$ (and $\nu$ under a standard regularity hypothesis) are invariant. For explicit gadget families, including complete-graph gadgets $K_N$ and Fisher-type constructions, these identities turn base graphs with known $\mu$ into infinite families of new quasi-transitive graphs whose connective constants are determined exactly as the unique roots of explicit algebraic equations.

math.CO

Recursive Packing Bounds for Supercritical Disconnection in Bernoulli Site Percolation

For Bernoulli site percolation on an infinite, connected, locally finite graph $G=(V,E)$, we obtain quantitative upper bounds on the supercritical disconnection probability \[ \mathbb{P}_p(S\nleftrightarrow\infty) \] for arbitrary finite or infinite sets $S\subset V$ and all $p>p^{\mathrm{site}}_c(G)$. The key quantity is a recursive packing number $\mathbf{PK}_{p,\eps,c}(S)$. It is the maximal number of vertices that can be extracted from $S$ so that, after deleting witness balls around the previously chosen vertices, each selected vertex still connects to infinity with probability at least $c$, while its failure to connect to infinity is already detected, up to a factor $1+\eps$, by failure to reach the inner boundary of its witness ball. Thus $\mathbf{PK}_{p,\eps,c}(S)$ counts essentially independent local witnesses for the global event $\{S\nleftrightarrow\infty\}$. We prove the structural estimate \[ \mathbb{P}_p(S\nleftrightarrow\infty) \le \frac{\eps(1-c)}{c} +(1-c)^{\mathbf{PK}_{p,\eps,c}(S)}. \] Combining this bound with the local functional characterization of $p^{\mathrm{site}}_c(G)$ from \cite{ZL24} yields an explicit supercritical estimate valid on every infinite, connected, locally finite graph. We also illustrate the packing number on ray-homogeneous trees. In particular, sparse finite subsets of a distinguished ray have packing number equal to their cardinality, both for regular trees and for a non-regular decorated spine. This shows that the packing number is explicit on concrete graph families.

math.PR

Planar Site Percolation, End Structure, and the Benjamini-Schramm Conjecture

Let $G$ be an infinite, connected, locally finite planar graph and consider i.i.d.\ Bernoulli$(p)$ site percolation. Write $p_c^{\mathrm{site}}(G)$ and $p_u^{\mathrm{site}}(G)$ for the critical and uniqueness thresholds. Using a well--separated Freudenthal embedding $G\hookrightarrow\mathbb S^2$, we introduce a cycle--separation equivalence on ends and associated ``directional'' thresholds $p^{\mathrm{site}}_{c,F}(G)$. When the set of end--equivalence classes is countable, we show that $p_c^{\mathrm{site}}(G)=\inf_F p^{\mathrm{site}}_{c,F}(G)$ and that for every $p\in\bigl(\tfrac12,\,1-p_c^{\mathrm{site}}(G)\bigr)$ there are almost surely infinitely many infinite open clusters. Combined with the $0/\infty$ theorem of Glazman--Harel--Zelesko for $p\le \tfrac12$, this yields non--uniqueness throughout the full coexistence interval $\bigl(p_c^{\mathrm{site}}(G),\,1-p_c^{\mathrm{site}}(G)\bigr)$, and hence $p_u^{\mathrm{site}}(G)\ge 1-p_c^{\mathrm{site}}(G)$ in this setting. This resolves the extension problem posed by Glazman--Harel--Zelesko for the upper half of the coexistence regime under a natural countability hypothesis. In contrast, for graphs with uncountably many end--equivalence classes we give criteria guaranteeing infinitely many infinite clusters above criticality, and we construct an explicit locally finite planar graph of minimum degree at least $7$ for which $p_u^{\mathrm{site}}(G)<1-p_c^{\mathrm{site}}(G)$. Consequently, the Benjamini--Schramm conjecture (Conjecture 7 in \cite{bs96}) that planarity together with minimal vertex degree at least 7 forces infinitely many infinite clusters for all $p\in(p_c,1-p_c)$ does not hold in full generality. Our proofs combine a cutset characterization of $p_c^{\mathrm{site}}$ with a planar alternating--arm exploration organized by an end--adapted boundary decomposition.

math.PR

Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs

Sparse Mixture-of-Experts (MoE) have been widely adopted in recent large language models since it can efficiently scale up the model capability without increasing the inference cost. However, evaluations on broad downstream tasks reveal a consistent suboptimality of the routers in existing MoE LLMs, which results in a severe performance gap (e.g., 10-20% in accuracy) to the optimal routing. In this paper, we show that aligning the manifold of routing weights with that of task embedding can effectively reduce the gap and improve MoE LLMs' generalization performance. Our method, "Routing Manifold Alignment (RoMA)", introduces an additional manifold regularization term in the post-training objective and only requires lightweight finetuning of routers (with other parameters frozen). Specifically, the regularization encourages the routing weights of each sample to be close to those of its successful neighbors (whose routing weights lead to correct answers) in a task embedding space. Consequently, samples targeting similar tasks will share similar expert choices across layers. Building such bindings between tasks and experts over different samples is essential to achieve better generalization. Moreover, RoMA demonstrates the advantage of unifying the task understanding (by embedding models) with solution generation (by MoE LLMs). In experiments, we finetune routers in OLMoE, DeepSeekMoE, and Qwen3-MoE using RoMA. Evaluations on diverse benchmarks and extensive comparisons with baselines show the substantial improvement brought by RoMA.

cs.LG

AIM 2025 challenge on Inverse Tone Mapping Report: Methods and Results

This paper presents a comprehensive review of the AIM 2025 Challenge on Inverse Tone Mapping (ITM). The challenge aimed to push forward the development of effective ITM algorithms for HDR image reconstruction from single LDR inputs, focusing on perceptual fidelity and numerical consistency. A total of \textbf{67} participants submitted \textbf{319} valid results, from which the best five teams were selected for detailed analysis. This report consolidates their methodologies and performance, with the lowest PU21-PSNR among the top entries reaching 29.22 dB. The analysis highlights innovative strategies for enhancing HDR reconstruction quality and establishes strong benchmarks to guide future research in inverse tone mapping.

cs.CV

MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents

Modern task-oriented dialogue (TOD) systems increasingly rely on large language model (LLM) agents, leveraging Retrieval-Augmented Generation (RAG) and long-context capabilities for long-term memory utilization. However, these methods are primarily based on semantic similarity, overlooking task intent and reducing task coherence in multi-session dialogues. To address this challenge, we introduce MemGuide, a two-stage framework for intent-driven memory selection. (1) Intent-Aligned Retrieval matches the current dialogue context with stored intent descriptions in the memory bank, retrieving QA-formatted memory units that share the same goal. (2) Missing-Slot Guided Filtering employs a chain-of-thought slot reasoner to enumerate unfilled slots, then uses a fine-tuned LLaMA-8B filter to re-rank the retrieved units by marginal slot-completion gain. The resulting memory units inform a proactive strategy that minimizes conversational turns by directly addressing information gaps. Based on this framework, we introduce the MS-TOD, the first multi-session TOD benchmark comprising 132 diverse personas, 956 task goals, and annotated intent-aligned memory targets, supporting efficient multi-session task completion. Evaluations on MS-TOD show that MemGuide raises the task success rate by 11% (88% -> 99%) and reduces dialogue length by 2.84 turns in multi-session settings, while maintaining parity with single-session benchmarks.

cs.CL

Doubly Free-Boundary Macdonald Processes: Reflection Identities and Jack Asymptotics

We introduce a doubly free-boundary Macdonald process on rail-yard interlacings and develop a reflection calculus for its observables. Boundary Cauchy--Littlewood identities, combined with Negu\c t operators, yield exact multipoint contour formulas for arbitrary \(L/R\) words. Under the Jack scaling \[ q=t^\alpha,\qquad t=e^{-n\beta\epsilon}, \] and piecewise-periodic data, these formulas imply a Laplace-transform law of large numbers and a weak slope-measure limit shape at \(L\)-type columns. For arbitrary piecewise-periodic \(L/R\) backgrounds and finitely many \(L\)-type marked columns, under the stated contour, branch, and normal-convergence hypotheses, the centered height-Laplace observables converge jointly to a Gaussian vector. Its covariance exhibits a boundary--deformation separation: the Jack parameter and the microscopic rail-yard data enter through the one-point spectral factors and the normalization, whereas the two-point interaction is the logarithmic derivative of an annular prime function generated by the two boundary reflections. Thus the deformation changes the spectral map while preserving the annular image geometry of the Schur specialization. For \(\beta=1\), in the all-\(L\) sector and under explicit signed zero--pole and root-localization hypotheses, we characterize regular liquid and frozen points through the nonreal-root structure of the characteristic equation and show that nondegenerate regular interfaces lie on the real double-root locus. The half-space Macdonald-process formulas are recovered in the continuous degeneration \(v\downarrow0\), which forces the right boundary partition to be empty.

math.PR

MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents

Recently, large language model based (LLM-based) agents have been widely applied across various fields. As a critical part, their memory capabilities have captured significant interest from both industrial and academic communities. Despite the proposal of many advanced memory models in recent research, however, there remains a lack of unified implementations under a general framework. To address this issue, we develop a unified and modular library for developing advanced memory models of LLM-based agents, called MemEngine. Based on our framework, we implement abundant memory models from recent research works. Additionally, our library facilitates convenient and extensible memory development, and offers user-friendly and pluggable memory usage. For benefiting our community, we have made our project publicly available at https://github.com/nuster1128/MemEngine.

cs.AI

C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing

Mixture-of-Experts (MoE) Large Language Models (LLMs) suffer from severely sub-optimal expert pathways-our study reveals that naive expert selection learned from pretraining leaves a surprising 10-20% accuracy gap for improvement. Motivated by this observation, we develop a novel class of test-time optimization methods to re-weight or "re-mixing" the experts in different layers jointly for each test sample. Since the test sample's ground truth is unknown, we propose to optimize a surrogate objective defined by the sample's "successful neighbors" from a reference set of samples. We introduce three surrogates and algorithms based on mode-finding, kernel regression, and the average loss of similar reference samples/tasks. To reduce the cost of optimizing whole pathways, we apply our algorithms merely to the core experts' mixing weights in critical layers, which enjoy similar performance but save significant computation. This leads to "Critical-Layer, Core-Expert, Collaborative Pathway Optimization (C3PO)". We apply C3PO to two recent MoE LLMs and examine it on six widely-used benchmarks. It consistently improves the base model by 7-15% in accuracy and outperforms widely used test-time learning baselines, e.g., in-context learning and prompt/prefix tuning, by a large margin. Moreover, C3PO enables MoE LLMs with 1-3B active parameters to outperform LLMs of 7-9B parameters, hence improving MoE's advantages on efficiency. Our thorough ablation study further sheds novel insights on achieving test-time improvement on MoE.

cs.LG

EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue Systems

Large language models have improved dialogue systems, but often process conversational turns in isolation, overlooking the event structures that guide natural interactions. Hence we introduce EventWeave, a framework that explicitly models relationships between conversational events to generate more contextually appropriate dialogue responses. EventWeave constructs a dynamic event graph that distinguishes between core events (main goals) and supporting events (interconnected details), employing a multi-head attention mechanism to selectively determine which events are most relevant to the current turn. Unlike summarization or standard graph-based approaches, our method captures three distinct relationship types between events, allowing for more nuanced context modeling. Experiments on three dialogue datasets demonstrate that EventWeave produces more natural and contextually appropriate responses while requiring less computational overhead than models processing the entire dialogue history. Ablation studies confirm improvements stem from better event relationship modeling rather than increased information density. Our approach effectively balances comprehensive context understanding with generating concise responses, maintaining strong performance across various dialogue lengths through targeted optimization techniques.

cs.CL

A Survey on Transformer Context Extension: Approaches and Evaluation

Large language models (LLMs) based on Transformer have been widely applied in the filed of natural language processing (NLP), demonstrating strong performance, particularly in handling short text tasks. However, when it comes to long context scenarios, the performance of LLMs degrades due to some challenges. To alleviate this phenomenon, there is a number of work proposed recently. In this survey, we first list the challenges of applying pre-trained LLMs to process long contexts. Then systematically review the approaches related to long context and propose our taxonomy categorizing them into four main types: positional encoding, context compression, retrieval augmented, and attention pattern. In addition to the approaches, we focus on the evaluation of long context, organizing relevant data, tasks, and metrics based on existing long context benchmarks. Finally, we summarize unresolved issues in the long context domain and put forward our views on future developments.

cs.CL

R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts

In large multimodal models (LMMs), the perception of non-language modalities (e.g., visual representations) is usually not on par with the large language models (LLMs)' powerful reasoning capabilities, deterring LMMs' performance on challenging downstream tasks. This weakness has been recently mitigated by replacing the vision encoder with a mixture-of-experts (MoE), which provides rich, multi-granularity, and diverse representations required by diverse downstream tasks. The performance of multimodal MoE largely depends on its router, which reweights and mixes the representations of different experts for each input. However, we find that the end-to-end trained router does not always produce the optimal routing weights for every test sample. To bridge the gap, we propose a novel and efficient method "Re-Routing in Test-Time (R2-T2)" that locally optimizes the vector of routing weights in test-time by moving it toward those vectors of the correctly predicted samples in a neighborhood of the test sample. We propose three R2-T2 strategies with different optimization objectives and neighbor-search spaces. R2-T2 consistently and greatly improves state-of-the-art LMMs' performance on challenging benchmarks of diverse tasks, without training any base-model parameters.

cs.LG

A Survey of Personalized Large Language Models: Progress and Future Directions

Large Language Models (LLMs) excel in handling general knowledge tasks, yet they struggle with user-specific personalization, such as understanding individual emotions, writing styles, and preferences. Personalized Large Language Models (PLLMs) tackle these challenges by leveraging individual user data, such as user profiles, historical dialogues, content, and interactions, to deliver responses that are contextually relevant and tailored to each user's specific needs. This is a highly valuable research topic, as PLLMs can significantly enhance user satisfaction and have broad applications in conversational agents, recommendation systems, emotion recognition, medical assistants, and more. This survey reviews recent advancements in PLLMs from three technical perspectives: prompting for personalized context (input level), finetuning for personalized adapters (model level), and alignment for personalized preferences (objective level). To provide deeper insights, we also discuss current limitations and outline several promising directions for future research. Updated information about this survey can be found at the https://github.com/JiahongLiu21/Awesome-Personalized-Large-Language-Models.

cs.AI

VAGeo: View-specific Attention for Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) aims to locate an object of interest in a captured ground- or drone-view image within the satellite image. However, existing works treat ground-view and drone-view query images equivalently, overlooking their inherent viewpoint discrepancies and the spatial correlation between the query image and the satellite-view reference image. To this end, this paper proposes a novel View-specific Attention Geo-localization method (VAGeo) for accurate CVOGL. Specifically, VAGeo contains two key modules: view-specific positional encoding (VSPE) module and channel-spatial hybrid attention (CSHA) module. In object-level, according to the characteristics of different viewpoints of ground and drone query images, viewpoint-specific positional codings are designed to more accurately identify the click-point object of the query image in the VSPE module. In feature-level, a hybrid attention in the CSHA module is introduced by combining channel attention and spatial attention mechanisms simultaneously for learning discriminative features. Extensive experimental results demonstrate that the proposed VAGeo gains a significant performance improvement, i.e., improving acc@0.25/acc@0.5 on the CVOGL dataset from 45.43%/42.24% to 48.21%/45.22% for ground-view, and from 61.97%/57.66% to 66.19%/61.87% for drone-view.

cs.CV