Searcharxiv⌕ Search

arXiv subjects

Yu Guo

Publications and source records attributed to Yu Guo.

At least 55 records · Page 3Linked to original sources

General Junction Condition and Casimir Effect for (1+1)-Dimensional Scalar Network CFT

Recently, BCFT and ICFT have been generalized to the CFT on networks (NCFT). A key aspect of NCFT is how we connect the CFTs across different edges at the network nodes. Previous research has primarily concentrated on a specific junction condition (JC) that requires the field to be continuous at the nodes. In this paper, we investigate the most general junction conditions for $(1+1)$-dimensional free scalars that are consistent with the variational principle and energy conservation. These general junction conditions are characterized by an $O(p)$ group, where $p$ represents the number of edges connected at a node. We provide exact realizations of two typical JCs in real physical systems. Additionally, we derive both the lower and upper bounds on the network Casimir energy for $(1+1)$-dimensional free scalar fields and extend the lower bound to encompass general NCFTs. Finally, we analyze the Casimir effect in networks composed of regular polyhedra and examine the binding energy required to construct such networks from individual components.

hep-th↗

Chirped-pulse engineering for robust control of single-molecule orientation in a cavity

We present a theoretical investigation of coherent control over the orientation of an individual molecule strongly coupled with a cavity using chirped-pulse driving. Specifically, we explore the dynamics of carbonyl sulfide (OCS) molecules under the influence of two chirped pulses with different spectral phases. We compare two pulse configurations: one with equal chirp rates ($β_{+} = β_{-}$) and another with unequal chirp rates ($β_{+} \neq β_{-}$). Numerical simulations reveal that chirped pulses enable precise control of the molecular orientation, achieving a maximum orientation degree of 0.5773. By analyzing the distribution of molecular polariton states, we show that chirped pulses can activate multiphoton processes, leading to deviations from the predictions of first-order Magnus expansion methods. Additionally, we demonstrate the robustness of the maximum orientation with respect to chirp amplitude and detuning, providing insights into the role of pulse parameters in optimizing control. This work introduces a new strategy for controlling molecular orientation in cavity-based systems and offers valuable perspectives for future experimental applications.

quant-ph↗

Unified entropy entanglement

The unified entropy as a promotion of the von Neumann entropy exhibits distinct diversity which contains the Tsallis entropy, the Rényi entropy, the von Neumann entropy as special cases. The unified-($r,t$) entropy entanglement with $0 1$ and $qs\geq1$ and show that it is also an entanglement monotone and that both of them are monogamous. Going further, we present two kinds of global multipartite entanglement measures (GlMEMs) based on the unified entropy and each kind has two subclasses which are classified by the parameters $(q,s)$ and $(r,t)$. Consequently, from the view of the complete multipartite entanglement measure theory, we show that one of them is a complete multipartite entanglement monotone and is not only completely monogamous but also tightly completely monogamous, but the other three are even not complete. We also explore the genuine entanglement measures induced by the unified entropy and the relations with the bipartite entanglement and the global entanglement are discussed, respectively.

quant-ph↗

Computable lower bound of the parameterized entanglement monotone

Although numerous measures of entanglement have been proposed so far, the calculation of a given faithful entanglement measure is a hard work since it is always involved in some optimization process. It is, therefore, important to estimate the lower bound of a given entanglement measure for an arbitrary quantum state. This results in a subject of intensive mathematical research. In particular, along this line, the lower bounds of concurrence or other measures that are induced from concurrence have been explored a lot. Here, we investigate the lower bounds of two kinds of entanglement monotones, i.e., $q$-concurrence ($q>1$) and $α$-concurrence ($0<α<1$), or termed the parameterized entanglement monotone together. We obtain, in the light of the informationally complete ($N$, $M$)-positive operator-valued measure [($N$, $M$)-POVM], the lower bounds for the case of $\frac12<α<1$, $1<q<2$ for two-qudit states, and the case of $2\leqslant q<3$ for two-qubit states. We list several examples which show that the lower bounds based on ($N$, $M$)-POVM outperform that of GSIC-POVM and SIC-POVM, and all these measurement based bounds are better then the ones induced by positive partial transpose (PPT) and realignment criteria in literature. In addition, we obtain an analytical formula of the parameterized entanglement monotone with $\frac12<α<1$ and $1<q<2$ for the isotropic state.

quant-ph↗

When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale Tables

Table question answering (TableQA) is a fundamental task in natural language processing (NLP). The strong reasoning capabilities of large language models (LLMs) have brought significant advances in this field. However, as real-world applications involve increasingly complex questions and larger tables, substantial noisy data is introduced, which severely degrades reasoning performance. To address this challenge, we focus on improving two core capabilities: Relevance Filtering, which identifies and retains information truly relevant to reasoning, and Table Pruning, which reduces table size while preserving essential content. Based on these principles, we propose EnoTab, a dual denoising framework for complex questions and large-scale tables. Specifically, we first perform Evidence-based Question Denoising by decomposing the question into minimal semantic units and filtering out those irrelevant to answer reasoning based on consistency and usability criteria. Then, we propose Evidence Tree-guided Table Denoising, which constructs an explicit and transparent table pruning path to remove irrelevant data step by step. At each pruning step, we observe the intermediate state of the table and apply a post-order node rollback mechanism to handle abnormal table states, ultimately producing a highly reliable sub-table for final answer reasoning. Finally, extensive experiments show that EnoTab achieves outstanding performance on TableQA tasks with complex questions and large-scale tables, confirming its effectiveness.

cs.CL↗

Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search

Table Question Answering (TableQA) benefits significantly from table pruning, which extracts compact sub-tables by eliminating redundant cells to streamline downstream reasoning. However, existing pruning methods typically rely on sequential revisions driven by unreliable critique signals, often failing to detect the loss of answer-critical data. To address this limitation, we propose TabTrim, a novel table pruning framework which transforms table pruning from sequential revisions to gold trajectory-supervised parallel search. TabTrim derives a gold pruning trajectory using the intermediate sub-tables in the execution process of gold SQL queries, and trains a pruner and a verifier to make the step-wise pruning result align with the gold pruning trajectory. During inference, TabTrim performs parallel search to explore multiple candidate pruning trajectories and identify the optimal sub-table. Extensive experiments demonstrate that TabTrim achieves state-of-the-art performance across diverse tabular reasoning tasks: TabTrim-8B reaches 73.5% average accuracy, outperforming the strongest baseline by 3.2%, including 79.4% on WikiTQ and 61.2% on TableBench.

cs.CL↗

Experimental observation of entropic-singularity-induced nonadditive quantum communication in a qutrit platypus channel

The nonadditivity of channel capacity is a defining feature that distinguishes quantum communication from classical communication. In the quantum realm, the channel capacity is determined by coherent information, which is defined through the von Neumann entropies of the output and its environment. Despite its fundamental importance, experimental evidence of such nonadditive quantum communication has been elusive because of the complexity of the required quantum channel. Here, we experimentally observe entropic-singularity-induced coherent-information nonadditivity using the qutrit platypus channel implemented on a photonic platform. By preparing six-dimensional photonic entanglement, we directly measure the coherent information of a platypus channel, a qubit amplitude damping channel, and their joint uses, revealing a clear violation of additivity. Quantum process tomography further reveals the entropic singularity responsible for this effect, demonstrating how singular entropy landscapes in low-dimensional channels can enhance quantum communication beyond additive limits.

quant-ph↗

Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities

Multimodal magnetic resonance imaging (MRI) is crucial for brain tumor segmentation, with many methods leveraging its four key modalities to capture complementary information for effective sub-region analysis. However, the absence of several modalities is very common in practice, leading to severe performance degradation in existing full-modality segmentation methods. Limited by the structured data model, recent works often adopt a multi-stage training strategy for full-modality and missing-modality scenarios, which increases training costs and inadequately addresses the interference of miss. In this work, we propose a graph-based one-stage framework for robust brain tumor segmentation with missing modalities. Specifically, we introduce modality-specific virtual nodes that serve as supplementary information sources to compensate for missing modalities. To enhance model robustness against arbitrary modality combinations, we leverage the inherent flexibility of graph networks to devise a dynamic connection strategy. This mechanism dynamically adjusts the adjacency matrix based on modality availability, preserving beneficial information flow while mitigating interference effects caused by missing modalities. Furthermore, we enhance the graph network through heterogeneous weight matrices, enhancing its adaptability to multimodal scenarios. Extensive experiments on the BRATS-2018 and BRATS-2020 datasets demonstrate that our method outperforms the state-of-the-art methods on almost all subsets of incomplete modalities.

cs.AI↗

Agent-Centric Observation Adaptation for Robust Visual Control under Dynamic Perturbations

Real-world visual systems face time-varying perturbations, including weather, sensor noise, compression artifacts, and background distractions. Existing image restoration methods are typically designed for fixed corruption types and optimized for pixel-level fidelity, leaving open two questions: how restoration behaves under non-stationary corruption switching, and whether pixel-level fidelity preserves the task-relevant information needed by downstream models. To study this setting, we introduce the Visual Degraded Control Suite (VDCS), a benchmark that injects Markov-switching physical degradations into rendered scenes. We further identify a fundamental failure mode of reconstruction-based representations: faithfully reconstructing corrupted observations forces the latent state to encode corruption-specific nuisance information, thereby contaminating downstream models. From an information-bottleneck perspective, anchoring the representation to the clean foreground eliminates this contamination. Motivated by this analysis, we propose \emph{Agent-Centric Observations with Mixture-of-Experts} (ACO-MoE), a frozen, plug-and-play observation adapter that combines a routed bank of restoration experts with a foreground-mask branch. ACO-MoE is pretrained entirely offline on synthetic rendered data with automatically generated degradation pairs and simulation-derived foreground masks, requiring no manual annotation. At inference time, it takes only corrupted RGB as input without corruption labels, clean reference frames, or foreground masks. Across VDCS, DMC-GB, and RoboSuite, ACO-MoE consistently improves downstream control with both model-free and model-based backbones, recovering 95.3\% of clean-input performance under challenging Markov-switching corruptions. It also generalizes zero-shot to unseen visual perturbations excluded from adapter pretraining.

cs.RO↗

Inference-Time Budget Control for LLM Search Agents

LLM search agents increasingly rely on tools at inference time, but their trajectories are often constrained by hard limits on both tool calls and generated tokens. Under such dual budgets, better answers require not only stronger models, but also explicit control over which search action should receive the next budget unit and when the accumulated evidence is sufficient to commit a final answer. We study this problem in multi-hop question answering (QA) and formulate it as two-stage inference-time budget control. At search time, our controller assigns each feasible action a task-level Value-of-Information (VOI) score, defined as an operational estimate of marginal task value per unit budget under the current search state and remaining dual budget, and uses this score to choose among retrieval, decomposition, and answer commitment. After search, a selective evidence-grounded finalizer compares the trajectory answer with a refined candidate and rewrites only when the residual error appears to be a low-risk answer-form error. Across four multi-hop QA benchmarks, three LLM backbones, and four budget levels, the method yields positive aggregate gains over four audited baselines under the same hard dual-budget protocol. Ablations show that search-time budget control, especially budget-dependent penalty, provides the main performance gain, while answer-time control helps mainly when the retrieval path is already adequate. These results suggest that inference-time budget control for LLM search agents should govern both how budget is spent during search and how the final answer is committed.

cs.AI↗

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

When a chest X-ray shows consolidation but the question asks which finding is present, a medical vision-language model may answer "No consolidation." This is more than an incorrect choice: it is a polarity reversal that emits a clinical statement contradicting the image. We study this failure as negated-option attraction, where a model is drawn to a negated answer option even when it conflicts with both the visual evidence and the question. We introduce CXR-ContraBench (Chest X-Ray Contradiction Benchmark), a diagnostic benchmark spanning internal ReXVQA slices and external OpenI and CheXpert protocols. The benchmark centers on present-finding questions, where selecting "No X" despite visible X creates the main clinical risk, and uses absent-finding questions as secondary tests of whether models copy negated wording. Across CheXpert protocols, the failure is substantial and persistent. On a strict direct presence probe, MedGemma and Qwen2.5-VL reach only 31.49% and 30.21% accuracy, respectively; on a matched 135,754-record CheXpert training-split protocol, both models select negated options on over 62% of presence questions. Chain-of-thought prompting reduces some presence-side reversals but does not eliminate them and can amplify absence-side contradictions. Finally, QCCV-Neg (Question-Conditioned Consistency Verifier for Negation) deterministically repairs the measured polarity-confused subset without retraining, raising MedGemma and Qwen2.5-VL to 96.60% and 95.32% accuracy on the direct presence probe. These results show that standard accuracy can hide a clinically meaningful inference-time polarity failure. Source code and benchmark construction scripts are available at https://github.com/fangzr/cxr-contrabench-code.

cs.CV↗

Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem

Large language models(LLMs) are increasingly integrated with external systems through the Model Context Protocol(MCP),which standardizes tool invocation and has rapidly become a backbone for LLM-powered applications. While this paradigm enhances functionality,it also introduces a fundamental security shift:LLMs transition from passive information processors to autonomous orchestrators of task-oriented toolchains,expanding the attack surface,elevating adversarial goals from manipulating single outputs to hijacking entire execution flows. In this paper,we identify and characterize a systematic privacy-leakage attack pattern,termed Parasitic Toolchain Attacks,instantiated as MCP Unintended Privacy Disclosure(MCP-UPD). These attacks require no direct victim interaction;instead,adversaries embed malicious instructions into external data sources that LLMs access during legitimate tasks. Unlike traditional prompt injection and tool poisoning attacks,our attack targets the interconnected toolchain itself,assembling multiple legitimate tools into a coordinated workflow whose combined behavior accomplishes malicious objectives. In MCP-UPD,the malicious logic infiltrates the toolchain and unfolds in three phases:Parasitic Ingestion,Privacy Collection,and Privacy Disclosure,culminating in stealthy exfiltration of private data. Our root cause analysis reveals that MCP lacks both context-tool isolation and least-privilege enforcement,enabling adversarial instructions to propagate unchecked into sensitive tool invocations. To assess the severity,we design MCP-SEC and conduct the first large-scale security census of the MCP ecosystem,analyzing 12230 tools across 1360 servers. Our findings show that the MCP ecosystem is rife with real-world exploitable gadgets and diverse attack methods,underscoring systemic risks in MCP platforms and the urgent need for defense mechanisms in LLM-integrated environments.

cs.CR↗

One-shot Compositional 3D Head Avatars with Deformable Hair

We propose a compositional method for constructing a complete 3D head avatar from a single image. Prior one-shot holistic approaches frequently fail to produce realistic hair dynamics during animation, largely due to inadequate decoupling of hair from the facial region, resulting in entangled geometry and unnatural deformations. Our method explicitly decouples hair from the face, modeling these components using distinct deformation paradigms while integrating them into a unified rendering pipeline. Furthermore, by leveraging image-to-3D lifting techniques, we preserve fine-grained textures from the input image to the greatest extent possible, effectively mitigating the common issue of high-frequency information loss in generalized models. Specifically, given a frontal portrait image, we first perform hair removal to obtain a bald image. Both the original image and the bald image are then lifted to dense, detail-rich 3D Gaussian Splatting (3DGS) representations. For the bald 3DGS, we rig it to a FLAME mesh via non-rigid registration with a prior model, enabling natural deformation that follows the mesh triangles during animation. For the hair component, we employ semantic label supervision combined with a boundary-aware reassignment strategy to extract a clean and isolated set of hair Gaussians. To control hair deformation, we introduce a cage structure that supports Position-Based Dynamics (PBD) simulation, allowing realistic and physically plausible transformations of the hair Gaussian primitives under head motion, gravity, and inertial effects. Striking qualitative results, including dynamic animations under diverse head motions, gravity effects, and expressions, showcase substantially more realistic hair behavior alongside faithfully preserved facial details, outperforming state-of-the-art one-shot methods in perceptual realism.

cs.CV↗

Birdcast: Interest-aware BEV Multicasting for Infrastructure-assisted Collaborative Perception

Vehicle-to-infrastructure collaborative perception (V2I-CP) leverages a high-vantage node to transmit supplementary information, i.e., bird's-eye-view (BEV) feature maps, to vehicles, effectively overcoming line-of-sight limitations. However, the downlink V2I transmission introduces a significant communication bottleneck. Moreover, vehicles in V2I-CP require \textit{heterogeneous yet overlapping} information tailored to their unique occlusions and locations, rendering standard unicast/broadcast protocols inefficient. To address this limitation, we propose \textit{Birdcast}, a novel multicasting framework for V2I-CP. By accounting for individual maps of interest, we formulate a joint feature selection and multicast grouping problem to maximize network-wide utility under communication constraints. Since this formulation is a mixed-integer nonlinear program and is NP-hard, we develop an accelerated greedy algorithm with a theoretical $(1 - 1/\sqrt{e})$ approximation guarantee. While motivated by CP, Birdcast provides a general framework applicable to a wide range of multicasting systems where users possess heterogeneous interests and varying channel conditions. Extensive simulations on the V2X-Sim dataset demonstrate that Birdcast significantly outperforms state-of-the-art baselines in both system utility and perception quality, achieving up to 27\% improvement in total utility and a 3.2\% increase in mean average precision (mAP).

cs.NI↗

Shared Spatial Memory Through Predictive Coding

Constructing a consistent shared spatial memory is a critical challenge in multi-agent systems, where partial observability and limited bandwidth often lead to catastrophic failures in coordination. We introduce a multi-agent predictive coding framework that formulates coordination as the minimization of mutual uncertainty among agents. Through an information bottleneck objective, this framework prompts agents to learn not only who and what to communicate but also when. At the foundation of this framework lies a grid-cell-like metric as internal spatial coding for self-localization, emerging spontaneously from self-supervised motion prediction. Building upon this internal spatial code, agents gradually develop a bandwidth-efficient communication mechanism and specialized neural populations that encode partners' locations-an artificial analogue of hippocampal social place cells (SPCs). These social representations are further utilized by a hierarchical reinforcement learning policy that actively explores to reduce joint uncertainty. On the Memory-Maze benchmark, our approach shows exceptional resilience to bandwidth constraints: success degrades gracefully from 73.5% to 64.4% as bandwidth shrinks from 128 to 4 bits/step, whereas a full-broadcast baseline collapses from 67.6% to 28.6%. Our findings establish a theoretically principled and biologically plausible basis for how complex social representations emerge from a unified predictive drive, leading to collective intelligence.

cs.AI↗

Gravity Dual of Networks

The network has been attracting increasing attention for its role in driving the artificial intelligence revolution and enabling profound insights into gravity. This paper investigates the gravity dual of the conformal field theory defined on a network (AdS/NCFT). A typical network, consisting of edges and nodes, is dual to a spacetime with branches and connecting branes, which we refer to as Net-branes. We demonstrate that the junction condition on the Net-brane results in energy conservation at the network node, providing strong support for our proposal of AdS/NCFT. We find that the spectrum of gravitational Kaluza-Klein modes on the Net-brane is a combination of the spectra from the AdS/BCFT with Neumann boundary conditions and Dirichlet/Conformal boundary conditions, corresponding to the isolated and transparent modes, respectively. We study two-point functions for NCFTs and provide examples, such as free fields and AdS/NCFT with tensionless Net-branes. We propose that the RT surfaces intersect at the same point on the Net-brane for connected subsystems within the network and verify this with the strong additivity and monotonicity of entanglement entropy. We establish that the network entropy, defined as the difference in entanglement between NCFT and BCFT, is always non-negative and effectively illustrates the network's complexity. Finally, we briefly discuss the holographic perspective of the shortest path problem and reveal its relation to the shortest geodesic in bulk and the holographic two-point correlators of massive operators.

hep-th↗

Task-Oriented Semantic Compression for Localization at the Network Edge

Achieving precise visual localization in GPS-limited urban environments poses significant challenges for resource-constrained mobile platforms, particularly under strict bandwidth, memory, and processing limitations. Inspired by mammalian spatial cognition, we propose a task-oriented communication framework in which bandwidth-limited endpoints equipped with multi-camera systems extract compact multi-view features and offload localization tasks to collaborative edge servers. We introduce the Orthogonally-constrained Variational Information Bottleneck encoder (O-VIB), which incorporates automatic relevance determination (ARD) to prune non-informative features while enforcing orthogonality to minimize redundancy. This enables efficient and accurate localization with minimal transmission overhead. Extensive evaluation on a real-world urban localization dataset demonstrates that O-VIB achieves high-precision localization under stringent bandwidth budgets, outperforming existing methods across diverse communication constraints.

cs.CV↗

Boundary-Aware Multi-Behavior Dynamic Graph Transformer for Sequential Recommendation

In the landscape of contemporary recommender systems, user-item interactions are inherently dynamic and sequential, often characterized by various behaviors. Prior research has explored the modeling of user preferences through sequential interactions and the user-item interaction graph, utilizing advanced techniques such as graph neural networks and transformer-based architectures. However, these methods typically fall short in simultaneously accounting for the dynamic nature of graph topologies and the sequential pattern of interactions in user preference models. Moreover, they often fail to adequately capture the multiple user behavior boundaries during model optimization. To tackle these challenges, we introduce a boundary-aware Multi-Behavioral Dynamic Graph Transformer (MB-DGT) model that dynamically refines the graph structure to reflect the evolving patterns of user behaviors and interactions. Our model involves a transformer-based dynamic graph aggregator for user preference modeling, which assimilates the changing graph structure and the sequence of user behaviors. This integration yields a more comprehensive and dynamic representation of user preferences. For model optimization, we implement a user-specific multi-behavior loss function that delineates the interest boundaries among different behaviors, thereby enriching the personalized learning of user preferences. Comprehensive experiments across three datasets indicate that our model consistently delivers remarkable recommendation performance.

cs.IR↗