SearcharxivSearch

arXiv subjects

Bin Guo

Publications and source records attributed to Bin Guo.

At least 19 recordsLinked to original sources

Analytic Construction of Rational Curves on Fano Manifolds

Inspired by methods for constructing entire curves in Oka geometry, we give an analytic construction of rational curves on a complex Fano manifold $X$. Yau's theorem provides a K\"ahler metric with positive Ricci curvature. Using this curvature to guide deformations of holomorphic discs, we construct maps from discs of radii tending to infinity with uniformly bounded area. A central point is to preserve the derivative normalization through the limiting process. This yields a nonconstant entire map $f:\mathbb C\rightarrow X$ of finite area. This map extends across infinity to a nonconstant holomorphic map $\mathbb P^1\to X$. Combined with algebraic arguments in characteristic zero, the construction yields proofs of the rational connectedness of Fano manifolds and of Hartshorne's conjecture on ample tangent bundles.

math.CV

Analytic and Algebraic Oka-1 Approximation for Smooth Projective Morphisms with Rationally Connected Fibers

Let $\pi:Z\rightarrow Y$ be a smooth projective morphism of complex manifolds with connected rationally connected fibers. We prove holomorphic approximation on arbitrary compact sets and finite-jet interpolation on arbitrary closed discrete sets for continuous liftings defined on open Riemann surfaces and holomorphic near those sets. For smooth projective morphisms of smooth complex algebraic varieties and algebraic base maps from smooth affine curves, the approximating liftings can be chosen algebraic, with interpolation on any finite set. In both cases the resulting lifting is homotopic to the initial one through continuous liftings of the fixed base map. For connected smooth projective complex manifolds, this gives the equivalence between the algebraic Oka-1 property and rational connectedness. Every rationally connected smooth projective complex manifold is also Oka-1.

math.CV

AdaSprite: Resource-efficient Online Co-Adaptation for V2I Systems Under Large-scale Data Drifts

The rise of vehicle-infrastructure (V2I) collaboration enables safer and broader perception. To process large-scale V2I video streams, vision-language models (VLMs) are promising as they unify multi-view vision into end-to-end task grounding, reducing handcrafted design. We use Vision Mixture-of-Experts (V-MoE) as the distributed visual backbone of VLMs, leveraging sparse expert routing to enable conditional computation across diverse viewpoints under resource constraints. Yet, V-MoEs face a critical challenge: large-scale data shifts over minutes to hours in V2I systems, amplified by agnostic participants and biased features propagating through experts. To maintain accuracy efficiently, we find it beneficial to co-adapt multiple V-MoEs on edge servers, avoiding the latency and privacy risks of cloud offloading and the accuracy sacrifices of on-device methods. However, the resource-constrained edge poses challenges for efficient co-adaptation: i) DRAM fragmentation and imbalance limit expert parallelism, ii) memory-I/O bottlenecks restrict computation reuse, and iii) asynchronous adaptation increases task-switch overhead. Also, prior work rarely explores the upper bound of concurrent tasks under limited edge resources, a critical factor for practical V2I deployment. To address these, we present AdaSprite. By combining cooperative elastic scaling with multi-level multiplexing, AdaSprite optimizes expert lifespans to reduce DRAM fragmentation, exploits predictable activation patterns for efficient I/O reuse, and employs twin-buffer scheduling to leverage sparsity. On a weak edge, AdaSprite supports up to 17 concurrent V2I tasks (vs. up to 6 for baselines), improving SLO attainment by 1.6x and throughput by 2.1x. Also, it allows users to trade accuracy and concurrency for second-level adaptation.

cs.OS

XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and Identification

Wireless sensing has emerged as a promising approach for tracking and identification using commodity Internet of Things devices. However, the features derived from a single wireless modality are often fragile to variations in environmental layouts and walking trajectories. Furthermore, most existing studies are based on datasets collected in specific scenarios with limited trajectory diversity and sensing modalities, preventing a robust evaluation of system generalization. \textcolor{blue}{To address this gap, we introduce \textbf{XGait}, a multi-modality wireless sensing dataset that synchronously captures human walking using Wi-Fi and acoustic transceivers across three indoor scenarios, with vision-based measurements serving as ground truth. Specifically, XGait contains more than 22K walking samples from 27 participants, covering diverse directions and trajectories to support both indoor tracking and identity recognition. To bridge the heterogeneity of wireless sensing modalities, we propose a unified Doppler spectrogram representation that maps Wi-Fi and acoustic signals into a shared time--frequency space, along with a standardized benchmark pipeline for pre-processing, temporal alignment, and feature construction, enabling reproducible evaluation and systematic cross-modal analysis. Extensive evaluations demonstrate that Wi-Fi and acoustic sensing exhibit complementary strengths, particularly under complex trajectories and challenging propagation conditions, thereby paving the way for novel research in the field of multi-modality wireless sensing.} The dataset and code are available at https://github.com/warrior-087/XGait.

cs.HC

What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents

Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounded. This survey separates these two questions. We define five non-exclusive functional roles for language: Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling. For each role, we trace the path from linguistic content to its embodied consumer and identify the observations or interventions that can test the claimed responsibility. Applying this framework to the reviewed literature reveals a recurring gap between functional use and evidential support. Interpretable or revised linguistic intermediates may be incorrect, go unused, or fail to affect later behavior. Even when actions are directly conditioned on language, system-level success does not by itself isolate language's contribution. We therefore evaluate grounding claim by claim, asking whether the reported evidence supports the specific responsibility assigned to language. Using role claims rather than architectures as the unit of comparison allows us to compare modular and end-to-end embodied agents without extending conclusions beyond the reported evidence.

cs.CL

Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression

Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student). Owing to its flexibility and broad applicability, KD has been extensively applied in the compression of server-side models to meet the Quality of Service (QoS) requirements of client users. Despite significant advancements, the performance of distillation is substantially compromised when a large disparity exists between the capabilities of the server and the requirements of the client. To alleviate this problem, we propose a novel distillation approach, named Progressive$^2$, which operates through the combination of a progressively stronger teacher and a progressively smaller student. On the side of the teacher, rather than involving all layers simultaneously, we progressively select additional layers for distillation following a raw-to-rich semantic progression, establishing a systematic learning curriculum. Furthermore, we design a teacher-side multi-feature fusion adapter for the teacher to improve training stability, which is theoretically supported by the framework of Lipschitz continuity. On the side of the student, rather than directly training a tiny model, we gradually reduce the size of the network to facilitate an iterative co-evolution with the teacher. Progressive$^2$ serves as a flexible framework; the progressive strategy of the teacher can be deployed independently to achieve an optimal balance between accuracy and training efficiency, while the joint integration of the teacher and the student yields further improvements in overall performance.

cs.LG

Task-Oriented Sensing and Covert Transmissions for Collaborative Multi-AUV Systems

In underwater covert cooperative missions, autonomous underwater vehicles (AUVs) often cannot rely on active sonar to continuously obtain complete information, since active sensing and frequent communications increase the risk of exposure. As a result, AUVs primarily rely on passive observation, an approach that yields incomplete local perception and limited task efficiency. Although underwater acoustic communications can mitigate this limitation through information sharing, they are simultaneously constrained by long delays, severe interference, low reliability, and the risk of covert exposure. Existing communications-oriented multi-agent reinforcement learning (MARL) studies often model communication as an ideal information flow, whereas traditional communication optimization primarily focuses on link-level performance. However, both are insufficient to characterize the actual contribution of perceptual information to cooperative tasks under realistic conditions of covert physical communications. This paper proposes a Sensed Information Value Realization Multi-Agent Reinforcement Learning (SVR-MARL) framework that leverages practical information to characterize the utility of information for cooperative tasks and learns distributed cooperative policies under realistic communication and covert constraints. Through a case study of covert multi-AUV cooperative localization and tracking, the potential of the proposed framework to improve collaborative task efficiency while reducing unnecessary communication and exposure risks is demonstrated.

cs.LG

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents

As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including beliefs, desires, intentions, and emotions (BDI/E), serves as an intermediate signal connecting agent behaviors with interaction outcomes and reflects how conversational strategies shape users during multi-turn interactions. However, existing evaluation paradigms primarily focus on surface-level responses or final outcomes, providing limited insight into the underlying cognitive processes. This limitation makes it difficult to diagnose why agents succeed or fail and to optimize their interaction strategies. To address this challenge, we propose Cognitive World Model (CogWM), an LLM-based cognitive user model that jointly models users' BDI/E states and corresponding responses, enabling explicit cognitive trajectory tracking. Trained on 150K user-turn samples with Qwen3-14B, CogWM achieves superior performance over existing user simulation baselines in both response fidelity and cognitive state understanding. Interactions with six state-of-the-art LLMs demonstrate that CogWM enables progressive comparison of agents through cognitive trajectories, revealing distinct agent patterns and complementary relationships between cognitive evolution and behavioral outcomes.

cs.AI

Optimizing the Sensitivity-Noise Trade-off in Non-Hermitian Sensing via Off-Exceptional-Deficiency Operation

A central challenge in non-Hermitian sensing is that spectral singularities simultaneously amplify both the signal and environmental noise. We address this predicament in a double-chain Hatano-Nelson model featuring unidirectional interlayer coupling. At the exceptional deficiency (ED) limit, the system exhibits a macroscopically degenerate complex spectrum and a pronounced non-Hermitian skin effect (NHSE), yielding a sensitivity that scales exponentially with lattice size $N$ while remaining robust across a six-order-of-magnitude detuning range. By introducing diagonal spatial disorder, we demonstrate that the NHSE is progressively suppressed, whith eigenspace cosine similarity analysis quantifying a well-defined fault-tolerance threshold. To reconcile the sensitivity-noise trade-off, we delineate "At-ED" and "Off-ED" operating regimes. While the At-ED configuration imposes fractional-order noise amplification (SNR $\propto \delta^{-1/2}$) that saturates at a suboptimal plateau, migrating to the Off-ED regime eliminates this geometric singularity and restores a linear scaling law (SNR $\propto \delta^{-1}$), achieving an SNR enhancement of several orders of magnitude. Crucially, this improvement is achieved while fully preserving the exponential sensitivity scaling, albeit at a slightly reduced absolute sensitivity compared to the strict At-ED limit. Our findings establish the Off-ED framework as a concrete paradigm for next-generation topological sensors that reconcile extreme sensitivity with robust noise immunity.

quant-ph

Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents

Persuasive dialogue requires reasoning about others' latent mental states, a capability known as Theory of Mind (ToM). However, due to reliance on simple prompting strategies and insufficient ToM knowledge, existing LLMs often fail to capture the intrinsic dependencies among mental states, leading to fragmented representations and unstable reasoning. To address these challenges, we introduce the ToM-based Persuasive Dialogue (ToM-PD) task, grounded in the Belief-Desire-Intention (BDI) framework, which explicitly models the sequential dependencies among mental states in multi-turn dialogues. To facilitate research on this task, we construct a large-scale annotated dataset, ToM-based Broad Persuasive Dialogues (ToM-BPD), capturing fine-grained mental states and corresponding persuasive strategies. We further propose Think Thrice Before You Speak (TTBYS), a knowledge-enhanced stepwise reasoning framework that leverages both explicit and implicit prior experiences to improve LLMs' inference of desires, beliefs, and persuasive strategies. Experimental results demonstrate that Qwen3-8B equipped with TTBYS outperforms GPT-5 by 1.20%, 22.80%, and 16.97% in predicting desires, beliefs, and persuasive strategies, respectively. Case studies further show that our approach enhances interpretability and consistency in reasoning.

cs.AI

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, query-agnostic modeling of visual evidence. We introduce Response-G1, a novel framework that establishes explicit, structured alignment between the accumulated video evidence and the query's expected response conditions via scene graphs. The framework operates in three fine-tuning-free stages: (1) online query-guided scene graph generation from streaming clips; (2) memory-based retrieval of the most semantically relevant historical scene graphs; and (3) retrieval-augmented trigger prompting for per-frame "silence/response" decisions. By grounding both evidence and conditions in a shared graph representation, Response-G1 achieves more interpretable and accurate response timing decisions. Experimental results on established benchmarks demonstrate the superiority of our method in both proactive and reactive tasks, validating the advantage of explicit scene graph modeling and retrieval in streaming video understanding.

cs.CV

Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents

In-group favoritism refers to the phenomena of favoring members of one's in-group over out-group members and is widely observed in numerous social cooperative behaviors. Recently, in-group favoritism biases have also been identified in generative language models. However, whether the in-group favoritism exists when persona agents are faced with contradicting information (e.g., misinformation), and how to mitigate the adverse effects of in-group favoritism biases in persona agents have been understudied. To address these problems, we propose a Truth or Tribe simulation framework to study the agent cooperation within the spread of contradicting information through a triadic interaction paradigm, and conduct controlled trials to evaluate the primary moderating factors. Extensive results showcase that persona agents display strong in-group favoritism, accepting incorrect answers from identity-similar peers at much higher rates than from dissimilar peers. In-group favoritism continues to emerge in defeasible reasoning contexts where no absolute truth exists, and it intensifies as cognitive complexity increases. Furthermore, three intervention strategies--Identity-Blind Instruction, Structured Counterfactual Reasoning, and Heterogeneous Perspective Ensemble--are proposed to mitigate the in-group favoritism.

cs.AI

Stochastic Momentum Tracking Push-Pull for Decentralized Optimization over Directed Graphs

Decentralized optimization over directed networks is frequently challenged by asymmetric communication and the inherent high variance of stochastic gradients, which collectively cause severe oscillations and hinder algorithmic convergence. To address these challenges, we propose the Stochastic Momentum Tracking Push-Pull (SMTPP) algorithm, which tracks the momentum term rather than raw stochastic gradients within the Push-Pull architecture. This design successfully decouples the variance reduction capacity from the algebraic connectivity of the graph.Although the inherent topology mismatch of directed graphs precludes exact convergence under persistent stochastic noise, SMTPP rigorously compresses this unavoidable steady-state error floor into a minimal neighborhood determined by network connectivity and gradient variance. Furthermore, SMTPP guarantees convergence on any strongly connected directed graph. Extensive experiments on non-convex logistic regression demonstrate that the algorithm is highly robust to network connectivity. By effectively dampening topology-induced oscillations, SMTPP achieves convergence rates and overall performance that closely match those of centralized baselines, regardless of whether the network is sparse or dense.

math.OC

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook

Sensor-based Human Activity Recognition (HAR) underpins many ubiquitous and wearable computing applications, yet current models remain limited by scarce labels, sensor heterogeneity, and weak generalization across users, devices, and contexts. Foundation models, which are generally pretrained at scale using self-supervised and multimodal learning, offer a unifying paradigm to address these challenges by learning reusable, adaptable representations for activity understanding. This survey synthesizes emerging foundation models for sensor-based HAR. We first clarify foundational concepts, definitions, and evaluation criteria, then organize existing work using a lifecycle-oriented taxonomy spanning input design, pretraining, adaptation, and utilization. Rather than enumerating individual models, we analyze recurring design patterns and trade-offs across nine technical axes, including modality scope, tokenization, architectures, learning paradigms, adaptation mechanisms, and deployment settings. From this synthesis, we identify three dominant development trajectories: (1) HAR-specific foundation models trained from scratch on large sensor corpora, (2) adaptation of general time-series or multimodal foundation models to sensor-based HAR, and (3) integration of large language models for reasoning, annotation, and human-AI interaction. We conclude by highlighting open challenges in data curation, multimodal alignment, personalization, privacy, and responsible deployment, and outline directions toward general-purpose, interpretable, and human-centered foundation models for activity understanding. A complete, continuously updated index of papers and models is available in our companion repository: https://github.com/zhaxidele/Foundation-Models-Defining-A-New-Era-In-Human-Activity-Recognition.

eess.SP

AppFlow: Memory Scheduling for Cold Launch of Large Apps on Mobile and Vehicle Systems

GB-scale large apps like on-device LLMs and rich media editors are becoming the next-generation trend, but their heavy memory and I/O demands, especially during multitasking, cause devices to reclaim or kill processes, turning warm apps into cold launches. The challenge lies not in storing them, but in fast, accurate launching. For users, 1s is the usability cliff, yet our measurements show 86.6\% of GB-scale cold launches exceed it. Also, Android Vitals flags only $\geq$ 5s as slow, exposing a large satisfaction gap. Existing optimizations are designed in isolation and conflict. For example, preloading reduces I/O stalls but consumes scarce memory and is undone by reclamation, while reclamation and killing free memory but sacrifice background survivability, leading to repeated cold relaunches. Our key insight is that, although multitasking makes runtime behavior complex, each app's file access pattern remains predictable. The challenge lies in exploiting this predictability, i.e., preloading without exhausting memory, reclaiming without undoing gains, and killing selectively to preserve background survivability. We introduce AppFlow, a prediction-based system-wide scheduler that integrates a Selective File Preloader, an Adaptive Memory Reclaimer, and a Context-Aware Process Killer. Implemented across the Android framework and Linux kernel without app changes, AppFlow cuts GB-scale cold-launch latency by 66.5\% (e.g., 2s$\rightarrow$690ms) and sustains 95\% of launches within 1s over a 100-day test, significantly improving responsiveness and multitasking experience.

cs.OS

Large-data solutions in multi-dimensional thermoviscoelasticity with temperature-dependent viscosities

This paper investigates a quasilinear parabolic system arising in thermoviscoelasticity of Kelvin-Voigt type with temperature-dependent viscosity and coupled terms. The system, given by \begin{equation*} \begin{cases} u_{tt}=\nabla\cdot\big(\gamma(\Theta)\nabla u_t\big)+a\Delta u-\nabla\cdot f(\Theta), & x \in \Omega,\ t > 0, \Theta_t=\Delta\Theta+\gamma(\Theta)|\nabla u_t|^2-f(\Theta)\nabla u_t, & x \in \Omega,\ t > 0, u=0,\quad\frac{\partial\Theta}{\partial\nu}=0, & x \in \partial\Omega,\ t > 0, u(x,0)=u_0(x),\; u_t(x,0)=u_{0t}(x),\;\Theta(x,0)=\Theta_0(x), & x \in \Omega, \end{cases} \end{equation*} models heat generation by acoustic waves in solid materials and can be derived as a scalar simplification of more complex piezoelectric-thermoviscoelastic model. Under the assumptions that $u_0\in H_0^1(\Omega)$, $u_{0t}\in L^2(\Omega)$, $\Theta_0\in L^1(\Omega)$ with $\Theta_0\geqslant0$ a.e.~in $\Omega$, that $\gamma,f\in C^0([0,\infty))$ satisfy $f(0)=0$, and that there exist constants $k_\gamma,K_\gamma,K_f>0$ and $0<\alpha<\frac{N+2}{2N}$ such that $$k_\gamma\leqslant\gamma(\xi)\leqslant K_\gamma\quad\text{and}\quad |f(\xi)|\leqslant K_f(1+\xi)^\alpha\qquad\forall~\xi\geqslant0,$$ we establish the global existence of weak solutions for arbitrarily large initial data in bounded domains $\Omega\subset\mathbb{R}^N$ ($N\geqslant1$). The result extends recent one-dimensional finding \cite{WinklerZAMP} to the multi-dimensional setting without requiring any smallness condition on the data.

math.AP

Asymptotic behavior of the solution with positive temperature in nonlinear 3D thermoelasticity

In this paper, we study a hyperbolic-parabolic coupled system arising in nonlinear three-dimensional thermoelasticity. We establish the global well-posedness and asymptotic behavior of solutions. Our main result shows that, a thermoelastic body asymptotically converges to an equilibrium state with a uniform temperature distribution for every initial data, determined by energy conservation. The proof of the global well-posedness is divided into some steps. To begin with, we introduce an approximate problem and derive its solvability. Next, we establish a time-independent upper bound for the temperature via Moser iteration technique. Together with an estimate of gradient of entropy, we use a functional involving the Fisher information of the temperature, which enables us to handle a delicate Gronwall-type inequality, to obtain required estimates of the higher-order derivatives. Further, we prove the strict positivity of temperature by applying Moser iteration again on the negative part of the logarithm of the temperature, followed by a uniqueness argument for the weak solution. Finally, we define a dynamical system on a proper functional phase space and analyze the $\omega$-limit set for every initial data. This work provides a complete proof of the global well-posedness and the long-time behavior in the nonlinear three-dimensional thermoelasticity system.

math.AP

Detecting Fake Reviewer Groups in Dynamic Networks: An Adaptive Graph Learning Method

The proliferation of fake reviews, often produced by organized groups, undermines consumer trust and fair competition on online platforms. These groups employ sophisticated strategies that evade traditional detection methods, particularly in cold-start scenarios involving newly launched products with sparse data. To address this, we propose the \underline{D}iversity- and \underline{S}imilarity-aware \underline{D}ynamic \underline{G}raph \underline{A}ttention-enhanced \underline{G}raph \underline{C}onvolutional \underline{N}etwork (DS-DGA-GCN), a new graph learning model for detecting fake reviewer groups. DS-DGA-GCN achieves robust detection since it focuses on the joint relationships among products, reviews, and reviewers by modeling product-review-reviewer networks. DS-DGA-GCN also achieves adaptive detection by integrating a Network Feature Scoring (NFS) system and a new dynamic graph attention mechanism. The NFS system quantifies network attributes, including neighbor diversity, network self-similarity, as a unified feature score. The dynamic graph attention mechanism improves the adaptability and computational efficiency by captures features related to temporal information, node importance, and global network structure. Extensive experiments conducted on two real-world datasets derived from Amazon and Xiaohongshu demonstrate that DS-DGA-GCN significantly outperforms state-of-the-art baselines, achieving accuracies of up to \textbf{89.8\% and 88.3\%}, respectively.

cs.SI