Searcharxiv⌕ Search

arXiv subjects

Bo Li

Publications and source records attributed to Bo Li.

At least 37 records · Page 2Linked to original sources

Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs

Test-time adaptation (TTA) on graphs aims to adapt a graph neural network (GNN) that is well-trained on the training graph to the test graph, which involves potential distribution shifts that may harm model generalization and test-time inference. While recent efforts have investigated TTA on static graphs, there is still a research gap on dynamic graphs learned with dynamic GNN (DGNN) models, where both structural connectivity and node semantics evolve continuously over time. This makes adapting a DGNN model for reliable test-time performance substantially challenging. To fill this gap, in this work, we propose a novel framework of temporal memory-aware Online Test-Time Adaptation on Dynamic Graphs, named DGOTTA, to effectively adapt well-trained DGNNs during test time. Specifically, the proposed DGOTTA contains three modules: (1) temporal-aware augmentation, to extend the diversity of test dynamic graphs for addressing complex temporal and spatial shifts; (2) memory-aware model prediction, to alleviate catastrophic forgetting; (3) consistency-guided online adaptation, to enforce temporal alignment and memory smoothness. Extensive experiments on three real-world datasets and four DGNN backbones demonstrate that DGOTTA significantly improves generalization under diverse distribution shifts and multiple model architectures.

cs.LG↗

How Well Can Strategyproof Tournament Rules Resist Pairwise Manipulation?

A tournament rule maps the outcomes of all pairwise matches among $n$ teams to a possibly randomized winner. Desirable rules should be Condorcet consistent and monotone, yet also resistant to manipulation among coalition. Prior work mostly measures such manipulation additively through $k$-strongly non-manipulable at $α$ ($k$-SNM-$α$), meaning that no coalition of size $k$ can fix the matches among themselves to increase their total winning probability by $α$. Very recently, two new notions of non-manipulability were introduced. Multiplicative non-manipulability ($k$-MNM-$δ$) is defined analogously, using the multiplicative factor instead. Non-manipulability for $λ$ ($k$-NM$_λ$) characterizes the selfishness of a team, which restricts a coalition's gain to be less than $λ$ times the winning probability sacrificed by its members. In this work, we begin with a strict hierarchy among these three notions: NM$_λ$ is stronger than MNM, which is then stronger than SNM. This motivates us to consider those two notions that are stronger but less studied: pairwise multiplicative non-manipulability and $2$-non-manipulability for $λ$. We show that Randomized Death Match is $2$-MNM-$3/2$ and optimally matches the lower bound. Then, we introduce the BlockBonusedWinStrengths rule, which is Condorcet consistent, monotone, and $2$-NM$_2$. This rule substantially improves the previous upper bound of $λ=11$ and comes within a factor of two of the lower bound $λ=1$.

cs.GT↗

When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning

Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observed views, but also label-aware exploitation of view-specific evidence. Existing approaches usually emphasize either shared representation learning or decision-level fusion. The former improves robustness against missing views, yet tends to compress label-discriminative view-specific cues into a single latent representation. The latter preserves individual view predictions, but often relies on fixed or globally learned fusion weights, ignoring that different labels of different instances may require different views. To address these limitations, this paper presents V2L, a unified representation-decision framework for incomplete multi-view multi-label classification. On the representation side, V2L constructs semantically consistent variational posteriors from incomplete views through a perturbation-aware encoding mechanism, which provides a stable shared semantic basis. On the decision side, V2L introduces an active view-label relevance modeling strategy that estimates instance-wise and label-wise view contributions, allowing each label prediction to adaptively select useful view-specific evidence. From the perspective of model architecture, these two important strategies are integrated into a unified framework through a hybrid fusion architecture, simultaneously meeting the requirements of cross-view semantic consistency and representational complementarity. Extensive experiments under both incomplete and complete settings show that V2L achieves leading performance on five benchmarks. Code is available at: https://github.com/justsmart/V2L.

cs.CV↗

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is strict enough for one model may over-block another, and a policy that transfers across domains may miss application-specific safety relations. We present EvoSafeHarness, a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain. It jointly searches a natural-language policy and executable code logic, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules. Across four agent benchmark families, EvoSafeHarness achieves a stronger safety-utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost and achieves the best score in 14 of 15 cells. On AgentDojo, it reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at the same operating point, and transfers unchanged to unseen AgentDyn suites. It also achieves the best score on Agent-SafetyBench for every victim and keeps mean ASR below 20% under adaptive PAIR attacks with a refinement budget of 16. Analysis shows that domain semantics determine which safety relations and trajectory state are needed, while model and runtime behavior determine how and where those relations should be enforced.

cs.CR↗

Idea and Theory of Particle Access

Aiming at some problems existing in the current quality of service (QoS) mechanism of large-scale networks (i.e. poor scalability, coarse granularity for provided service levels, poor fairness between different service levels, and improving delay performance at the expense of sacrificing some resource utilization), the paper puts forward the idea and thoery of particle access. In the proposed particle access mechanism, the network first granulates the information flow (that is, the information flow is subdivided into information particles, each of which is given its corresponding attributes), and allocates access resources to the information particle group which is composed of all the information particles to be transmitted, so as to ensure that the occupied bandwidth resources is minimized on the premise of meeting the delay requirements of each information particle. Moreover, in the paper, the concepts of both information particle and information particle group are defined; Basic properties of the minimum reachable access bandwidth of an information particle group are analyzed; The influences of time attribute and attribute of bearing capacity of an information particle group on the minimum reachable access bandwidth are analyzed; Finally, an effective method for the calculation of the minimum reachable access bandwidth of an information particle group is given, and a particle access algorithm based on dynamically adjusting the minimum reachable access bandwidth is proposed. The research of the paper pave a new way for further improving QoS mechanisms of large-scale networks, and lay the corresponding theoretical foundation.

cs.NI↗

COSTA: Covariance-Optimized Design and Causal Inference under Network-Temporal Interference

Experiments on networks observed over time face network spillovers, temporal carryover, and dependence deliberately introduced by the design. We propose COSTA---Covariance-Optimized Spatiotemporal Treatment Allocation---a joint Bernoulli design for unit--time assignments. Under common treatment marginals and a nonnegative linear network--temporal exposure model, Horvitz--Thompson bias for the sustained all-treated versus all-control contrast is exactly the negative expected weight of an assignment cut. A covariance-level variance envelope yields an MSE bound that can be optimized directly over assignment covariance. To scale this design, we introduce a thresholded-Gaussian Kronecker parameterization that mirrors the network and temporal exposure operators while preserving valid Bernoulli marginals. We next develop inference theory for the joint effects of designed treatment dependence and interference-induced outcome dependence. Canonical correlations between latent blocks generating separated HT contributions supply the coefficients required by graph-$ψ$ central limit and network-HAC theory; a spectral-floor and far-row-mass condition gives a primitive sufficient check. The framework covers sparse, block, Kronecker, locally factored, and other structured covariance sequences satisfying these conditions. Semi-synthetic RetailRocket and MovieLens experiments show substantial default-setting RMSE reductions and well-calibrated model-assisted design-centered intervals across linear, nonlinear, and demand-substitution outcome surfaces.

stat.ME↗

Theory of Network Wave

Aiming at the disorder problem (i.e. uncertainty problem) of the utilization of network resources commonly existing in multi-hop transmission networks, the paper proposes the idea and the corresponding supporting theory, i.e. theory of network wave, by constructing volatility information transmission mechanism between the sending nodes and their corresponding receiving nodes of a pair of paths (composed of two primary paths), so as to improve the orderliness of the utilization of network resources. It is proved that the maximum asymptotic throughput of a primary path depends on its intrinsic period, which in itself is equal to the intrinsic interference intensity of a primary path. Based on the proposed theory of network wave, an algorithm for the transmission of information blocks based on the intrinsic period of a primary path is proposed, which can maximize the asymptotic throughput of a primary path. In the cases of traversals with equal opportunities, an algorithm for the cooperative volatility transmission of information blocks in a pair of paths based on the set of maximum supporting elements is proposed. It is proved that the algorithm can maximize the asymptotic joint throughput of a pair of paths. As for the cases of traversals with unequal opportunities, an algorithm for the cooperative volatility transmission of information blocks in a pair of paths based on the set of maximum supporting elements is also proposed. The research results of the paper lay an ideological and theoretical foundation for further exploring more general methods that can improve the orderly utilization of network resources.

cs.NI↗

DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models

Mixture-of-Experts (MoE) has become a prominent paradigm for scaling Large Language Models (LLMs). Parameter-efficient fine-tuning methods, such as LoRA, are widely adopted to adapt pretrained MoE LLMs to downstream tasks. However, existing approaches typically assign identical LoRA ranks to all expert modules, ignoring the heterogeneous specialization of pretrained experts. This uniform allocation leads to a resource mismatch: task-relevant experts are under-provisioned, while less relevant ones receive redundant parameters. To address this, we propose DR-LoRA, a Dynamic Rank LoRA framework for fine-tuning pretrained MoE models. Specifically, DR-LoRA initializes all expert LoRA modules with a small active rank and uses an expert saliency score, which combines routing frequency and gradient-based rank importance, to identify which experts would benefit most from additional capacity. It then periodically expands the active ranks of the task-critical expert LoRA, progressively constructing a heterogeneous rank distribution tailored to the target task. Experiments on three MoE models across six tasks show that DR-LoRA consistently outperforms LoRA and other strong baselines, demonstrating that task-adaptive heterogeneous rank allocation is an effective strategy to improve active capacity utilization in MoE fine-tuning.

cs.AI↗

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.

cs.CV↗

Kinetics of an Expanding Bacterial Colony: Continuum Modeling and Analysis

We study the spatiotemporal dynamics of the expansion of a bacterial colony on a hard substrate. Nutrient from the substrate diffuses into the colony and is taken up by the bacterial cells for them to grow and divide, expanding the initial monolayer and then pancake-shaped colony. The concentration of nutrient determines the local cell growth rate in the colony. Mass conservation relates such local growth rate with velocity which is approximated to be proportional to the pressure gradient by Darcy's law. Altogether, the growing colony is modeled as a moving-boundary problem with the nutrient concentration and pressure solving a reaction-diffusion equation and Laplace's equation, respectively. We analyze the self-consistent moving-boundary model with respect to different geometrical setting. For a general three-dimensional cylindrically symmetric model, we study the steady-state nutrient concentration. We construct and analyze a one-dimensional model for vertical expansion and a two-dimensional disk model for radial expansion of the colony. Our analysis finds that the nutrient depletes into the colony and slows down the vertical expansion of the colony. The vertical level where the nutrient concentration reaches the Monod constant, a threshold below which individual bacteria hardly grow, lowers down exponentially fast. We also estimate the asymptotic radial expansion rate. Moreover, we establish that the region where the nutrient concentration is above the threshold, allowing bacteria to grow and the colony to expand radially, is a ring-shaped peripheral region of fixed thickness. All these are consistent with experiment and agent-based simulations reported in literature.

math.AP↗

Differentially Private Continual Release with Relative Error

This work investigates several fundamental tasks, including $\mathsf{MaxSum}$, $\mathsf{MinSum}$, $\mathsf{MaxSelect}$, and $\mathsf{MinSelect}$, in the continual release model under differential privacy. Previous research has demonstrated that any algorithm for these tasks must admit a large purely additive error. We show that the error can be substantially reduced if a relative error term is allowed, provided that the input stream is generated non-adaptively. However, when input data records can be selected adaptively, we prove that a large error is inevitable for the task of selecting an attribute with a small cumulative sum, whereas small error bounds remain achievable for other tasks. This reveals a significant separation between non-adaptive and adaptive streams. We also complement our algorithms with nearly matching lower bounds.

cs.DS↗

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable Walkthrough induction from sparse-reward trajectories: extracting compact, state-conditioned, and verifiable procedures. Our key observation is that delayed credit identifies actions associated with progress but cannot determine whether they produce facts required by later actions. We propose Trace, a credit-guided, dependency-grounded framework that compiles noisy trajectories into executable Walkthrough Memory. It detects progress anchors from rewards and persistent state changes, propagates credit to identify valuable transitions, and estimates action prerequisites from cross-episode success and failure evidence. Backward dependency slicing then traces required facts to their producers, extracting dependency-consistent action chains while removing irrelevant loops and detours. The resulting Walkthroughs encode entry conditions, ordered state--action--effect steps, and completion and failure predicates, supporting reuse, intermediate-state resumption, and programmatic verification. Experiments on J-TTL, WebShop, and ScienceWorld with three open-source LLMs show that Trace consistently outperforms eight test-time learning and memory baselines. Compared with the strongest baseline, it improves average AUC and Final-$3$ by $30.0%$ and $40.5%$, respectively, while using fewer inference tokens. These results show that long-horizon interaction benefits more from state-conditioned executable procedures than from complete trajectories or abstract summaries.

cs.LG↗

`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs

As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vulnerability to audio-based adversarial attacks. These systems typically follow two architectural paradigms: cascaded pipelines, where automatic speech recognition converts audio inputs into text before LLM processing, and end-to-end large audio-language models (LALMs), which directly interpret raw audio signals. Beyond architectural differences, cascaded pipelines are primarily vulnerable to text-level jailbreak strategies delivered through speech, whereas end-to-end LALMs introduce additional acoustic-semantic attack vectors. However, existing studies often focus on a single paradigm and provide limited coverage of the broader audio attack space. To bridge this gap, we propose an adaptive jailbreak attack framework for systematic evaluation of both cascaded pipelines and LALMs under a unified experimental setting. At its core, the framework uses a feedback-guided mutation engine to automatically generate and refine jailbreak candidates across both textual prompts and audio perturbations, thereby expanding attack diversity and coverage. Experiments on six representative audio-based systems demonstrate that both paradigms remain substantially vulnerable to audio jailbreak attacks. Compared with state-of-the-art methods, our framework achieves consistently higher attack success rates across diverse audio-based LLM systems.

cs.CR↗

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $ρ= +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at https://github.com/Zane-ZYQiu/entry-point-umm.

cs.CV↗

On the Incompatibility of Weighted PROPX and Pareto Optimality for Indivisible Chores

Proportionality (PROP) is one of the simplest fairness criteria for allocating items among agents with additive preferences. With indivisible chores, however, PROP is not always satisfiable. We study proportionality up to any item (PROPX), which requires every agent to satisfy proportionality after any chore is removed from her bundle. Under strictly positive costs, we settle the weighted compatibility question negatively: weighted PROPX and Pareto optimality are incompatible already for two agents and four chores. Moreover, for every $n\geq3$, we give an $n$-agent, $(n+1)$-chore counterexample whose shares can be arbitrarily close to equal. These counterexamples are item-minimal: under strictly positive costs, weighted PROPX and Pareto optimality are always compatible when the number of chores is at most the number of agents, and they are compatible for two agents with at most three chores. Our impossibility result contrasts with the compatibility theorem of Mahara (2026) for weighted envy-freeness up to one item (EF1) and Pareto optimality .

cs.GT↗

PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy

Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirely risks over-correction against valid opinions. Effective control must therefore both reduce and increase sycophancy with predictable and gradual effect. Yet, existing methods fail to ensure a bidirectional and monotonic relationship between steering strength and behavioral outcome across models and datasets. We introduce PCA-guided Activation Scaling (PAS), an activation steering framework that decomposes residual stream activations into a PCA-identified sycophancy-honesty subspace and an orthogonal residual, then applies distinct scaling exponents to achieve monotonic, bidirectional control. Across three LLMs and three datasets, PAS achieves strong monotonicity (Spearman $ρ$ = +0.92) and an average shift of 15.4% per direction, compared with 8.7% for the baselines. Ablation studies confirm that the decomposition, asymmetric exponents, and layer selection are each essential for maintaining monotonic control. The data and code are available at https://github.com/Bellafc/PCS.

cs.CL↗

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.

cs.OS↗

Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization

Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of great practical importance. However, directly applying existing causal inference methods to clicked samples introduces sample selection bias and increased variance due to the exclusion of non-click data. Recent studies on CVR prediction introduce "ideal loss", which optimizes model parameters using an unbiased estimate of the loss over the full sample. Nevertheless, there is no guarantee that unbiasedness of the loss implies unbiasedness of the final estimator. We revisit this challenge from the perspective of semiparametric theory. Specifically, we develop a new doubly robust causal effect estimator for chain-structured outcomes such as CVR, and derive its theoretical properties in detail. It achieves a faster convergence rate compared to nuisance parameters estimation and is therefore more robust when using flexible nonparametric estimators, including neural networks. Based on these theoretical findings, we further design a framework based on targeted regularization to improve numerical stability and practical applicability. Extensive experiments on synthetic and real-world data demonstrate the effectiveness and robustness of our method. In addition, we find that naively combining loss debiasing with standard causal estimators underperforms our method, highlighting the necessity of developing the new estimator tailored to this CVR-style objective with solid theoretical guarantees.

cs.LG↗