SearcharxivSearch

arXiv subjects

Shijun Li

Publications and source records attributed to Shijun Li.

At least 19 recordsLinked to original sources

STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs significant storage overhead, raises privacy concerns, and fails to adequately capture the true data distribution with sparse exemplars. Inspired by human cognitive mechanisms, we propose a novel framework termed Semantic Text-Anchored Incremental Learning (STAIL) for sequential clinical tasks. To overcome the rehearsal bottleneck, STAIL introduces an asymmetric semantic consolidation buffer (SCB). By incorporating a minimal set of image anchors and extensive textual descriptions, the SCB enables dense semantic reconstruction of old tasks at a minimal storage cost. Furthermore, we design an LLM-derived Semantic Anchoring Mechanism (LSAM) that leverages the stable semantic space of frozen large language models as developmental priors. This mechanism explicitly anchors evolving visual features to textual representations, guiding and constraining plasticity and stability at both macroscopic and microscopic levels. Extensive experiments across three heterogeneous medical datasets, covering fundus, ultrasound, and X-ray imaging, demonstrate that STAIL acts as a highly effective plug-and-play module. It comprehensively enhances the performance of various existing baselines, achieving average gains of 2.24\% in AAA-AUC for sustained performance and 3.55\% in BWT-AUC for reduced forgetting. Code is available.

cs.CV

MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling

Multimodal irregular time series (MITS) consist of asynchronous and irregularly sampled observations from heterogeneous numerical and textual channels. In healthcare, for example, patients' electronic health records (EHR) include irregular lab measurements and clinical notes. The irregular timing and channel patterns of observations carry predictive signal alongside the numerical values and textual content. LLMs are natural candidates for processing such heterogeneous data, given their extensive pretrained knowledge spanning textual and numerical domains. We introduce MILM (Multimodal Irregular time series Language Model), which represents MITS as time-ordered triplets in Extensible Markup Language (XML) format and fine-tunes an LLM through a two-stage strategy for MITS classification. The first stage trains on value-redacted MITS to predict from sampling patterns alone, and the second stage trains on full MITS to jointly model sampling patterns and observed values. Our two-stage model (MILM-2S) and its single-stage counterpart (MILM-Direct) achieve the best and second-best average performance on multiple EHR datasets. Further value redaction evaluations confirm that sampling patterns carry predictive signal and that MILM-2S learns to exploit them. In the value pending evaluation we introduce, where some values are unavailable at prediction time, MILM-2S outperforms MILM-Direct by a larger margin compared to standard evaluation. For MILM-2S, preserving the time and channel of value-pending observations as additional sampling information further improves in-hospital mortality prediction.

cs.LG

RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation

Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and natural-language reasoning abilities. Despite recent progress, current LLM-based recommenders still face key challenges in constructing decision-relevant contexts from heterogeneous evidence. First, existing methods often rely on fixed context construction strategies: collaborative behavioral evidence and item-side metadata are typically incorporated through predefined prompts, static retrieval pipelines, or handcrafted injection mechanisms, making it difficult to determine what information is truly beneficial for each instance. Second, heterogeneous evidence introduces a severe context-efficiency bottleneck. Rich metadata and collaborative interaction records can quickly overwhelm the context window, while aggressive compression or heuristic filtering may discard fine-grained evidence critical for accurate recommendation. To address these challenges, we propose RRCM, a ranking-driven retrieval-and-reasoning framework over collaborative and metadata memories for LLM-based agentic recommendation. RRCM starts from a lightweight user-history context and learns whether to recommend directly, retrieve collaborative evidence, retrieve item metadata, or interleave both through reasoning. Both memories are represented in natural language and accessed through a unified retrieval interface, enabling flexible evidence acquisition without handcrafted CF injection or fixed retrieval rules. We optimize this memory-reading policy with an outcome-only ranking reward, instantiated using group relative policy optimization, so that retrieval decisions are directly driven by final top-k recommendation quality. Extensive experiments show that RRCM significantly outperforms traditional baselines and diverse LLM-based recommendation approaches.

cs.IR

Goal-Conditioned Supervised Learning for LLM Fine-Tuning

Large language models often require fine-tuning to better align their behavior with user intent at deployment. Existing approaches are commonly divided into online and offline paradigms. Online methods, such as RL-based alignment, can directly optimize outcome quality but typically rely on external reward models and iterative rollouts, making them costly and difficult to deploy in many cases. Offline methods are more efficient, but prevailing approaches such as supervised fine-tuning (SFT) and direct preference optimization (DPO) remain limited: SFT typically collapses graded feedback into binary supervision, while DPO depends on paired preference data that is often unavailable or expensive to construct. In this paper, we propose goal-conditioned supervised learning (GCSL) as an offline fine-tuning framework for LLMs. Our core idea is to treat feedback signals directly as an explicit goal and train the model, purely through supervised learning, to generate responses that achieve that goal. To better exploit graded feedback, we further introduce a novel goal formulation that defines learning as consistently pursuing outcomes above a target quality threshold, rather than imitating samples from a selected high-quality subset. This design mitigates the bounded-learning effect of SFT and classic GCSL by explicitly guiding the model to learn the directional progression of quality. We also propose natural-language goal representations to better leverage the semantic understanding and reasoning capabilities of LLMs. We evaluate our method on three tasks: non-toxic generation, code generation, and LLM for recommendation. Results show that our approach consistently outperforms standard offline fine-tuning baselines while retaining the efficiency, scalability, and simple data requirements of supervised learning.

cs.LG

Renormalized Solution for the Nonlinear Parabolic Problem with Lower Order Terms

In this paper, we consider the following problem: \[ \begin{cases} -\nabla\cdot A(x,u,\nabla u) + H(x,u,\nabla u) = f(x), & x \in Ω, u = 0, & x \in \partial Ω, \end{cases} \] in a bounded open set \( Ω\subset \mathbb{R}^N \). We have established certain gradient estimates and proved the existence of a renormalized solution for the equation.

math.AP

Renormalized Solution for the Nonlinear Parabolic Problem with Lower Order Terms

In this paper, we consider the following nonlinear parabolic equation with non-coercive terms in \(R^N\) space \[ \dfrac{\partial u}{\partial t} -\nabla \cdot (a(x,t,u,\nabla u)+ Φ(x,t,\nabla u))=f, \text{ in }Ω\times (0,T). \] Here \(Ω\) is a bounded open set of \(R^N\) with the boundary \(\partial Ω\) satisfying Lipschitz condition. The Carathéodory function \(Φ\) is restricted by $|Φ(x,t,s)|\le c(x,t)|s|^γ$ with parameters depending on $p$ and $N$. And the initial value $u(x,0)=u_0(x)$. For convenience, we define the domain $Q := Ω\times (0,T)$ and the boundary similarly. Then for $f\in L^1(Q)$ and $u_0\in L^1(Ω)$, we prove the existence and uniqueness of a renormalized solution via truncation methods, monotone operator theory, and a prior gradient estimates.

math.AP

Renormalized Solutions for a Class of Nonlinear Parabolic Equation with a Lower Order Term and Variable Exponents

We consider a class of nonlinear parabolic equations \[ \dfrac{\partial}{\partial t} b(u)-\nabla \cdot (A(x,t,u,\nabla u))+H(x,t,\nabla u)=f , \] where $H$ is a nonlinear lower order term satisfied the Carath$\acute{e}$odory condition and \[ \left\lvert H(x,t,\nabla u)\right\rvert\leqslant g(x,t)\left\lvert \nabla u\right\rvert^{δ(x)} \] with \[ δ(x)=\frac{p(x)(N+1)-N}{(N+2)(p(x)-1)}(p^--1) \quad \text{and} \quad p^-=\underset{x\in\barΩ}{min}\,p(x). \] By virtue of truncation metheod,the monotone operator theory and a gradient estimate we prove existence of renormalized solutions without coercivity condition on lower order term in the framework of variable exponents.

math.AP

Finite-time blow-up in a class of chemotaxis systems with spatially heterogeneous diffusion sensitivity

\indent In this paper, we study a class of parabolic-elliptic Keller-Segel systems with diffusion sensitivity dependent on spatial position, given by type \begin{equation} \left\{ \begin{array}{ll} u_{t} = \bigtriangledown\cdot(|x|^β \bigtriangledown u)-\bigtriangledown\cdot(u^α \bigtriangledown v), 0=\bigtriangleup v-μ+u, \qquad μ:=\frac{1}{|Ω|}\int_Ωudx,\end{array}\right. \end{equation} under homogeneous Neumann conditions in a ball $Ω=B_{R}(0)\subset \mathbb{R}^{n}$ with $α\ge 1$, $β>0$ and $n\ge 2$.\par \indent It is proved that any nonconstant nonnegative radial initial data $u_{0}\in C^θ(\overlineΩ)$, where $θ\in (0,1)$, there exists a radially symmetric classical solution of the system (0.1) in $(Ω\setminus \{ 0 \})\times (0,T)$ for some $T>0$; moreover, if the initial values $u_{0}\in C^{1+θ}(\overlineΩ)$ for some $θ\in (0,1)$ and satisfy a certain compatibility criterion and are radially decreasing, then this solution is bounded and unique in $(Ω\setminus \{ 0 \})\times (0,T^{*})$ with $T^{*}<T$.\par Finally, it is found that the initial mass corresponding to this parabolic-elliptic problem (0.1) is sufficiently concentrated to allow the solution to blow up in finite time.

math.AP

Approximation Analysis of a Parabolic-Parabolic Chemotaxis Model with Logarithmic Nonlinearity

We consider the Keller-Segel system with logical source \begin{align*} \begin{cases} u_t = \nabla \cdot (ϕ(u)\nabla u) - \nabla \cdot (ψ(u)\nabla v)+f(u), & x \in Ω, \; t > 0, v_t = Δv - v + u, & x \in Ω, \; t > 0, \end{cases} \end{align*} in a smooth bounded domain \(Ω\subset \mathbb{R}^n\) with \(n \geq 2\), the Neumann initial-boundary value problem admits a globally defined, uniformly bounded classic solution for all sufficiently regular non-negative initial data \(u_0\) and \(v_0\). In the first equation, assume that \(ϕ\) and \(ψ\) are dominated by a logarithmic function and a polynomial respectively. The logical source \(f\) representing the natural growth and decay of cells satisfies \(f \in W^{1,\infty}_{\mathrm{loc}}(Ω)\) and \(f(0) \geq 0\). Then we will see that the unique solution \(u \in C^{2,1}((\overlineΩ) \times [0,T] )\) and \(v \in W^{1,q}([0,T] ; C^{2,1}(\overlineΩ))\).

math.AP

LLM Reasoning for Cold-Start Item Recommendation

Large Language Models (LLMs) have shown significant potential for improving recommendation systems through their inherent reasoning capabilities and extensive knowledge base. Yet, existing studies predominantly address warm-start scenarios with abundant user-item interaction data, leaving the more challenging cold-start scenarios, where sparse interactions hinder traditional collaborative filtering methods, underexplored. To address this limitation, we propose novel reasoning strategies designed for cold-start item recommendations within the Netflix domain. Our method utilizes the advanced reasoning capabilities of LLMs to effectively infer user preferences, particularly for newly introduced or rarely interacted items. We systematically evaluate supervised fine-tuning, reinforcement learning-based fine-tuning, and hybrid approaches that combine both methods to optimize recommendation performance. Extensive experiments on real-world data demonstrate significant improvements in both methodological efficacy and practical performance in cold-start recommendation contexts. Remarkably, our reasoning-based fine-tuned models outperform Netflix's production ranking model by up to 8% in certain cases.

cs.IR

Vague Preference Policy Learning for Conversational Recommendation

Conversational recommendation systems (CRS) commonly assume users have clear preferences, leading to potential over-filtering of relevant alternatives. However, users often exhibit vague, non-binary preferences. We introduce the Vague Preference Multi-round Conversational Recommendation (VPMCR) scenario, employing a soft estimation mechanism to accommodate users' vague and dynamic preferences while mitigating over-filtering. In VPMCR, we propose Vague Preference Policy Learning (VPPL), consisting of Ambiguity-aware Soft Estimation (ASE) and Dynamism-aware Policy Learning (DPL). ASE captures preference vagueness by estimating scores for clicked and non-clicked options, using a choice-based approach and time-aware preference decay. DPL leverages ASE's preference distribution to guide the conversation and adapt to preference changes for recommendations or attribute queries. Extensive experiments demonstrate VPPL's effectiveness within VPMCR, outperforming existing methods and setting a new benchmark. Our work advances CRS by accommodating users' inherent ambiguity and relative decision-making processes, improving real-world applicability.

cs.IR

OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration

Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Recently, Multi-Agent Systems (MAS), wherein multiple agents collaborate and communicate with each other, have exhibited enhanced capabilities in complex tasks, such as high-quality code generation and arithmetic reasoning. However, the development of such systems often relies on handcrafted methods, and the literature on systematic design and optimization of LLM-based MAS remains limited. In this work, we introduce \textbf{OMAC}, a general framework designed for holistic optimization of LLM-based MAS. Specifically, we identify five key optimization dimensions for MAS, encompassing both agent functionality and collaboration structure. Building upon these dimensions, we first propose a general algorithm, utilizing two actors termed the Semantic Initializer and the Contrastive Comparator, to optimize any single dimension. Then, we present an algorithm for joint optimization across multiple dimensions. Extensive experiments demonstrate the superior performance of OMAC on diverse tasks against recent approaches.

cs.MA

Goal-Conditioned Supervised Learning for Multi-Objective Recommendation

Multi-objective learning endeavors to concurrently optimize multiple objectives using a single model, aiming to achieve high and balanced performance across diverse objectives. However, this often entails a more complex optimization problem, particularly when navigating potential conflicts between objectives, leading to solutions with higher memory requirements and computational complexity. This paper introduces a Multi-Objective Goal-Conditioned Supervised Learning (MOGCSL) framework for automatically learning to achieve multiple objectives from offline sequential data. MOGCSL extends the conventional GCSL method to multi-objective scenarios by redefining goals from one-dimensional scalars to multi-dimensional vectors. It benefits from naturally eliminating the need for complex architectures and optimization constraints. Moreover, MOGCSL effectively filters out uninformative or noisy instances that fail to achieve desirable long-term rewards across multiple objectives. We also introduces a novel goal-selection algorithm for MOGCSL to model and identify "high" achievable goals for inference. While MOGCSL is quite general, we focus on its application to the next action prediction problem in commercial-grade recommender systems. In this context, any viable solution needs to be reasonably scalable and also be robust to large amounts of noisy data that is characteristic of this application space. We show that MOGCSL performs admirably on both counts by extensive experiments on real-world recommendation datasets. Also, analysis and experiments are included to explain its strength in discounting the noisier portions of training data in recommender systems with multiple objectives.

cs.LG

LM4LV: A Frozen Large Language Model for Low-level Vision Tasks

The success of large language models (LLMs) has fostered a new research trend of multi-modality large language models (MLLMs), which changes the paradigm of various fields in computer vision. Though MLLMs have shown promising results in numerous high-level vision and vision-language tasks such as VQA and text-to-image, no works have demonstrated how low-level vision tasks can benefit from MLLMs. We find that most current MLLMs are blind to low-level features due to their design of vision modules, thus are inherently incapable for solving low-level vision tasks. In this work, we purpose $\textbf{LM4LV}$, a framework that enables a FROZEN LLM to solve a range of low-level vision tasks without any multi-modal data or prior. This showcases the LLM's strong potential in low-level vision and bridges the gap between MLLMs and low-level vision tasks. We hope this work can inspire new perspectives on LLMs and deeper understanding of their mechanisms. Code is available at https://github.com/bytetriper/LM4LV.

cs.CV

DSDRNet: Disentangling Representation and Reconstruct Network for Domain Generalization

Domain generalization faces challenges due to the distribution shift between training and testing sets, and the presence of unseen target domains. Common solutions include domain alignment, meta-learning, data augmentation, or ensemble learning, all of which rely on domain labels or domain adversarial techniques. In this paper, we propose a Dual-Stream Separation and Reconstruction Network, dubbed DSDRNet. It is a disentanglement-reconstruction approach that integrates features of both inter-instance and intra-instance through dual-stream fusion. The method introduces novel supervised signals by combining inter-instance semantic distance and intra-instance similarity. Incorporating Adaptive Instance Normalization (AdaIN) into a two-stage cyclic reconstruction process enhances self-disentangled reconstruction signals to facilitate model convergence. Extensive experiments on four benchmark datasets demonstrate that DSDRNet outperforms other popular methods in terms of domain generalization capabilities.

cs.CV

Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models

Adapter-based parameter-efficient transfer learning has achieved exciting results in vision-language models. Traditional adapter methods often require training or fine-tuning, facing challenges such as insufficient samples or resource limitations. While some methods overcome the need for training by leveraging image modality cache and retrieval, they overlook the text modality's importance and cross-modal cues for the efficient adaptation of parameters in visual-language models. This work introduces a cross-modal parameter-efficient approach named XMAdapter. XMAdapter establishes cache models for both text and image modalities. It then leverages retrieval through visual-language bimodal information to gather clues for inference. By dynamically adjusting the affinity ratio, it achieves cross-modal fusion, decoupling different modal similarities to assess their respective contributions. Additionally, it explores hard samples based on differences in cross-modal affinity and enhances model performance through adaptive adjustment of sample learning intensity. Extensive experimental results on benchmark datasets demonstrate that XMAdapter outperforms previous adapter-based methods significantly regarding accuracy, generalization, and efficiency.

cs.CV

Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning

The chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple aspects simultaneously and employ dynamic adjustment and updating mechanisms. Therefore, we propose a novel Aggregation-Graph-of-Thought (AGoT) mechanism for soft-prompt tuning in multi-modal representation learning. The proposed AGoT models the human thought process not only as a chain but also models each step as a reasoning aggregation graph to cope with the overlooked multiple aspects of thinking in single-step reasoning. This turns the entire reasoning process into prompt aggregation and prompt flow operations. Experiments show that our multi-modal model enhanced with AGoT soft-prompting achieves good results in several tasks such as text-image retrieval, visual question answering, and image recognition. In addition, we demonstrate that it has good domain generalization performance due to better reasoning.

cs.AI

CIRS: Bursting Filter Bubbles by Counterfactual Interactive Recommender System

While personalization increases the utility of recommender systems, it also brings the issue of filter bubbles. E.g., if the system keeps exposing and recommending the items that the user is interested in, it may also make the user feel bored and less satisfied. Existing work studies filter bubbles in static recommendation, where the effect of overexposure is hard to capture. In contrast, we believe it is more meaningful to study the issue in interactive recommendation and optimize long-term user satisfaction. Nevertheless, it is unrealistic to train the model online due to the high cost. As such, we have to leverage offline training data and disentangle the causal effect on user satisfaction. To achieve this goal, we propose a counterfactual interactive recommender system (CIRS) that augments offline reinforcement learning (offline RL) with causal inference. The basic idea is to first learn a causal user model on historical data to capture the overexposure effect of items on user satisfaction. It then uses the learned causal user model to help the planning of the RL policy. To conduct evaluation offline, we innovatively create an authentic RL environment (KuaiEnv) based on a real-world fully observed user rating dataset. The experiments show the effectiveness of CIRS in bursting filter bubbles and achieving long-term success in interactive recommendation. The implementation of CIRS is available via https://github.com/chongminggao/CIRS-codes.

cs.IR