SearcharxivSearch

arXiv subjects

Xin Liao

Publications and source records attributed to Xin Liao.

At least 19 recordsLinked to original sources

Screen-Conditioned Watermarking Against Multi-Screen Collusion Attacks

Screen-shooting poses a significant threat to confidential information protection. While existing screen-shooting watermarking methods enable copyright verification, the copyrighted images carrying the same copyright watermark across different screens often exhibit highly similar and estimable watermark patterns. These shared patterns can be exploited for watermark removal and forgery, a threat we term the multi-screen collusion attack. To mitigate this threat, we propose CoMSMark, a collusion-resistant image-agnostic watermarking framework for multi-screen shooting, which reduces shared residual components across screens to resist multi-screen collusion attacks. Specifically, we incorporate screen ID through a style modulation mechanism, enabling the encoder to generate screen-specific watermark residuals for reliable source attribution. We further introduce a collusion suppression loss that reduces shared residual components and encourages high-entropy predictions for forged samples, improving resistance to collusion attacks. Finally, to enable efficient large-scale distribution, CoMSMark employs an image-agnostic encoding paradigm that generates watermark residuals independently of image content. Extensive experiments demonstrate that CoMSMark effectively resists both collusion-based watermark removal and forgery. It maintains an average watermark accuracy above 90% under removal attacks while keeping forged-watermark accuracy near 50%. Moreover, CoMSMark achieves competitive robustness under diverse screen-shooting conditions, including varying capture distances and angles.

cs.CR

Efficient Cross-Scale Invertible Hiding Network with Spatial-Frequency Collaboration and Non-Invertible Mechanism

Image hiding aims to conceal image-level messages within cover images at the same resolution. Invertible neural networks (INN)-based image hiding has emerged as an important branch. It treats concealing and revealing as a pair of inverse problems on image domain transformation and uses INN's forward and backward processes to address them. Due to architectural constraints, existing INN-based methods suffer from single-scale and single-domain feature extraction and limited nonlinear representation capability, resulting in inferior image quality. To mitigate these limitations, we propose an efficient cross-scale invertible hiding network with the spatial-frequency collaboration and the non-invertible mechanism, termed CrosInv. CrosInv exploits cross-scale and spatial-frequency collaborative features while enhancing nonlinear representation. Specifically, we introduce a cross-scale invertible module that bijectively maps inputs to cross-scale representations. To effectively integrate spatial and frequency information, the cross-scale invertible module employs pixel shuffle, Haar wavelet transformation, and their inverse operations for scale transformation. Furthermore, a non-invertible cross dense module is integrated to enhance the nonlinearity. Comprehensive experiments verify the effectiveness and superiority of the proposed CrosInv.

cs.CV

CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback

Recent LLM search agents use reinforcement learning with verifiable rewards (RLVR) to learn search-augmented reasoning from outcome rewards. On hard problems, these agents rarely sample end-to-end successful rollouts, leaving outcome-only RLVR with few positive-reward trajectories. We argue that improving learning on such problems requires additional guidance during training, and RLVR already contains verifier-side information that can provide it. This information can identify errors or omissions in the agent's submitted answer and guide revision within the rollout. We propose a training-time mechanism called \textbf{Credit-Attenuated Privileged Feedback} (CAPF), which makes this verifier-side information available through a Privileged Feedback call during training. CAPF lets the policy revise zero-reward attempts into positive-reward repair trajectories and attenuates credit for the feedback call and earlier actions to accommodate deployment without this call. Empirical research demonstrates that CAPF improves Qwen3-4B's average exact-match score from 44.7% under outcome-only RLVR to 48.5% on seven open-domain QA benchmarks.

cs.AI

TopoClaw: A Human-Centric and Topology-Aware Agent Operating System

Large language models (LLMs) have evolved AI assistants into autonomous reasoning engines that maintain context, invoke tools, and pursue long-horizon tasks. This has spurred Agent Operating Systems (Agent OS) as kernel-like layers for lifecycle management, memory, scheduling, and access control. Yet most designs remain agent-centric, treating the OS as a single-host runtime for internal reasoning and tool use, leaving open how autonomous actions integrate with distributed, collaborative, permission-sensitive workflows. TopoClaw is an open-source, human-centric, topology-aware Agent OS modeling the user's ecosystem as two coupled structures: a physical device topology of heterogeneous surfaces and a social relationship topology of shared spaces, teams, and delegated roles. It unifies device operation, messaging, and skills around accountable cross-boundary execution, with three core contributions: (1) cross-device action placement, decoupling intent from actuation and routing distributed actions across the device cluster based on hardware affordances and user context; (2) cross-user identity attribution, treating agents as socially situated "Digital Twins" that coordinate in multi-user spaces while preserving provenance, role-aware permissions, and human accountability; (3) cross-context authority governance, pairing broad capability with distributed, context-aware policy enforcement across physical and social trust boundaries to bound proactive autonomy at the OS layer. This report presents TopoClaw as an engineering-oriented reference architecture, covering its design principles, runtime, cross-device execution, collaboration mechanisms, security model, and deployment outlook.

cs.HC

Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection

Detecting AI-generated images across unseen architectures remains challenging, as existing models often overfit to generator-specific fingerprints and semantic content rather than learning universal forgery traces. We attribute this failure to feature entanglement: detectors learn these factors as a single entangled representation, where universal forgery traces are inextricably confounded with both generator-specific fingerprints and semantic content. Crucially, our spectral analysis reveals that this entanglement is avoidable: distinct generator-specific fingerprints (e.g., GAN stripes vs. Diffusion Model spots) occupy disjoint frequency subspaces and coexist as independent superpositions. Leveraging this physical orthogonality, we propose the Orthogonal Decomposition and Purification Network (ODP-Net) to structurally disentangle these factors. Specifically, ODP-Net employs (1) Instance-aware Orthogonal Decomposition to project features into mutually exclusive subspaces: universal forgery traces, generator-specific fingerprints, and semantic content; (2) Perturbation-based Purification to enforce semantic invariance via cross-sample feature injection; and (3) Manifold Alignment to bridge domain gaps. By explicitly decoupling universal forgery traces from generator-specific fingerprints and semantic content, ODP-Net achieves state-of-the-art performance on unseen architectures (e.g., Stable Diffusion 3), validating that structural disentanglement is key to generalization.

cs.CV

Injectable Thermochemical Micro-Explosion for Prompt Thrombolysis via Liquid Alkali Metal

Thrombotic vascular diseases contribute to significant global mortality, yet current therapeutic strategies face persistent challenges including bleeding risks, suboptimal efficiency, and procedural complexity. Here, we report a micro-explosive thermochemical thrombolysis (METCT) therapy via injectable liquid alkali metal (LAM) encapsulated in dimethyl silicone (LAM@oil), which enables prompt, efficient and safe vascular recanalization within an ultrafast timeframe (< 90 seconds). This LAM@oil system effectively disrupts thrombus tissue through a synergistic triple-action mechanism: Mechanical micro-explosions forces, alkaline ablation due to highly localized exothermic chemical reactions, and thermal thrombolysis mediated by elevated temperature. Upon thrombolysis completion, the non-toxic reaction byproducts (sodium and potassium ions) exhibit physiologically biocompatible and metabolizable effects. Critically, the LAM@oil demonstrates significantly higher thrombolytic efficacy compared to clinically available thrombolytic drugs (residual thrombus area percent 10.87%+-7.16% for LAM@oil vs. 80.86%+-13.32% for urokinase), with no associated bleeding risks. This strategy opens a byproduct-free, cost-effective, and high-efficiency alternative to conventional thrombolytics, holding big potential for clinical translation in acute thrombosis management.

physics.med-ph

NiMark: A Non-intrusive Watermarking Framework against Screen-shooting Attacks

Unauthorized screen-shooting poses a critical data leakage risk. Resisting screen-shooting attacks typically requires high-strength watermark embedding, inevitably degrading the cover image. To resolve the robustness-fidelity conflict, non-intrusive watermarking has emerged as a solution by constructing logical verification keys without altering the original content. However, existing non-intrusive schemes lack the capacity to withstand screen-shooting noise. While deep learning offers a potential remedy, we observe that directly applying it leads to a previously underexplored failure mode, the Structural Shortcut: networks tend to learn trivial identity mappings and neglect the image-watermark binding. Furthermore, even when logical binding is enforced, standard training strategies cannot fully bridge the noise gap, yielding suboptimal robustness against physical distortions. In this paper, we propose NiMark, an end-to-end framework addressing these challenges. First, to eliminate the structural shortcut, we introduce the Sigmoid-Gated XOR (SG-XOR) estimator to enable gradient propagation for the logical operation, effectively enforcing rigid image-watermark binding. Second, to overcome the robustness bottleneck, we devise a two-stage training strategy integrating a restorer to bridge the domain gap caused by screen-shooting noise. Experiments demonstrate that NiMark consistently outperforms representative state-of-the-art methods against both digital attacks and screen-shooting noise, while maintaining zero visual distortion.

eess.IV

DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization

The rapid evolution of AIGC technology enables misleading viewers by tampering mere small segments within a video, rendering video-level detection inaccurate and unpersuasive. Consequently, temporal forgery localization (TFL), which aims to precisely pinpoint tampered segments, becomes critical. However, existing methods are often constrained by \emph{local view}, failing to capture global anomalies. To address this, we propose a \underline{d}ual-stream graph learning and \underline{d}isentanglement framework for temporal forgery localization (DDNet). By coordinating a \emph{Temporal Distance Stream} for local artifacts and a \emph{Semantic Content Stream} for long-range connections, DDNet prevents global cues from being drowned out by local smoothness. Furthermore, we introduce Trace Disentanglement and Adaptation (TDA) to isolate generic forgery fingerprints, alongside Cross-Level Feature Embedding (CLFE) to construct a robust feature foundation via deep fusion of hierarchical features. Experiments on ForgeryNet and TVIL benchmarks demonstrate that our method outperforms state-of-the-art approaches by approximately 9\% in AP@0.95, with significant improvements in cross-domain robustness.

cs.CV

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for pattern discovery. A core difficulty of qualitative-data clustering lies in measuring similarity among attribute values that carry no inherent ordering or distance. To recover such relationships, existing studies typically rely on within-dataset co-occurrence statistics. This statistical route, however, becomes unreliable once the sample size is small, and the semantic context of each value is therefore left underexploited. Motivated by this limitation, this paper proposes BREVE (Balanced Representation via External Value Enrichment), a clustering framework that enriches each qualitative value with extra semantic dimensions drawn from an external knowledge base. That is, every unique value is expanded by a dense embedding that encodes its semantic content. To prevent the original value identity from being diluted by the added dimensions, a lightweight one-hot component is further appended. An adaptive weight, guided by cluster compactness, then determines how strongly the enrichment dimensions enter the final representation. With this design, experiments on eight benchmark datasets yield an average ARI rank of 1.3 against seven representative competitors.

cs.LG

Neural Tucker Convolutional Network for Water Quality Analysis

Water quality monitoring is a core component of ecological environmental protection. However, due to sensor failure or other inevitable factors, data missing often exists in long-term monitoring, posing great challenges in water quality analysis. This paper proposes a Neural Tucker Convolutional Network (NTCN) model for water quality data imputation, which features the following key components: a) Encode different mode entities into respective embedding vectors, and construct a Tucker interaction tensor by outer product operations to capture the complex mode-wise feature interactions; b) Use 3D convolution to extract fine-grained spatiotemporal features from the interaction tensor. Experiments on three real-world water quality datasets show that the proposed NTCN model outperforms several state-of-the-art imputation models in terms of accuracy.

cs.LG

Quantitative Stability of Two Weakly Interacting Kinks in the Stationary phi^6 Model

We study the stationary phi^6 model given by the equation -phi''(x) + 2 phi(x) - 8 phi(x)^3 + 6 phi(x)^5 = 0 for x in R, and establish sharp quantitative stability estimates for configurations close to two weakly interacting kinks. More precisely, there exist constants a > 0 and epsilon > 0 such that, for any function u in L-infinity satisfying || u - H_{0,1}(x + x1) - H_{-1,0}(x + x2) ||_{H1} < epsilon with x2 - x1 > a, there exist constants y1, y2 such that || u - H_{0,1}(x + y1) - H_{-1,0}(x + y2) ||_{H1} + exp(-sqrt(2) (y2 - y1)) <= C * || u'' - 2 u + 8 u^3 - 6 u^5 ||_{L2}.

math.AP

Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking Robustness

Unauthorized screen capturing and dissemination pose severe security threats such as data leakage and information theft. Several studies propose robust watermarking methods to track the copyright of Screen-Camera (SC) images, facilitating post-hoc certification against infringement. These techniques typically employ heuristic mathematical modeling or supervised neural network fitting as the noise layer, to enhance watermarking robustness against SC. However, both strategies cannot fundamentally achieve an effective approximation of SC noise. Mathematical simulation suffers from biased approximations due to the incomplete decomposition of the noise and the absence of interdependence among the noise components. Supervised networks require paired data to train the noise-fitting model, and it is difficult for the model to learn all the features of the noise. To address the above issues, we propose Simulation-to-Real (S2R). Specifically, an unsupervised noise layer employs unpaired data to learn the discrepancy between the modeled simulated noise distribution and the real-world SC noise distribution, rather than directly learning the mapping from sharp images to real-world images. Learning this transformation from simulation to reality is inherently simpler, as it primarily involves bridging the gap in noise distributions, instead of the complex task of reconstructing fine-grained image details. Extensive experimental results validate the efficacy of the proposed method, demonstrating superior watermark robustness and generalization compared to state-of-the-art methods.

cs.CV

ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem

With the rapid development of (multimodal) large language model-based agents, the landscape of agentic service management has evolved from single-agent systems to multi-agent systems, and now to massive-agent ecosystems. Current massive-agent ecosystems face growing challenges, including impersonal service experiences, a lack of standardization, and untrustworthy behavior. To address these issues, we propose ColorEcosystem, a novel blueprint designed to enable personalized, standardized, and trustworthy agentic service at scale. Concretely, ColorEcosystem consists of three key components: agent carrier, agent store, and agent audit. The agent carrier provides personalized service experiences by utilizing user-specific data and creating a digital twin, while the agent store serves as a centralized, standardized platform for managing diverse agentic services. The agent audit, based on the supervision of developer and user activities, ensures the integrity and credibility of both service providers and users. Through the analysis of challenges, transitional forms, and practical considerations, the ColorEcosystem is poised to power personalized, standardized, and trustworthy agentic service across massive-agent ecosystems. Meanwhile, we have also implemented part of ColorEcosystem's functionality, and the relevant code is open-sourced at https://github.com/opas-lab/color-ecosystem.

cs.MA

From nonlinear Schrödinger equation to interacting particle system: 1 < p < 2

We investigate the limiting behavior of solutions with infinitely many peaks to nonlinear Schrödinger equations [-epsilon^2 Delta u_epsilon + u_epsilon = u_epsilon^p, u_epsilon > 0 in R^n,] as epsilon -> 0, where p is Sobolev subcritical. We derive the interaction law among the limiting peak points and complete the analysis for the previously unresolved range 1 < p < 2, extending the work of Ao, Lv, and Wang (J. Differential Equations, 2025).

math.AP

Multiple sign-changing solutions for semilinear subelliptic Dirichlet problem

We study the following perturbation from symmetry problem for the semilinear subelliptic equation \[ \left\{ \begin{array}{cc} -\triangle_{X} u=f(x,u)+g(x,u) & \mbox{in}~Ω, \\[2mm] u\in H_{X,0}^{1}(Ω),\hfill \end{array} \right. \] where $\triangle_{X}=-\sum_{i=1}^{m}X_{i}^{*}X_{i}$ is the self-adjoint sub-elliptic operator associated with Hörmander vector fields $X=(X_{1},X_{2},\ldots,X_{m})$, $Ω$ is an open bounded subset in $\mathbb{R}^n$, and $H_{X,0}^{1}(Ω)$ denotes the weighted Sobolev space. We establish multiplicity results for sign-changing solutions using a perturbation method alongside refined techniques for invariant sets. The pivotal aspect lies in the estimation of the lower bounds of min-max values associated with sign-changing critical points. In this paper, we construct two distinct lower bounds of these min-max values. The first one is derived from the lower bound of Dirichlet eigenvalues of $-\triangle_{X}$, while the second one is based on the Morse-type estimates and Cwikel-Lieb-Rozenblum type inequality in degenerate cases. These lower bounds provide different sufficient conditions for multiplicity results, each with unique advantages and are not mutually inclusive, particularly in the general non-equiregular case. This novel observation suggests that in some sense, the situation for sub-elliptic equations would have essential difference from the classical elliptic framework.

math.AP

A Nonlinear Low-rank Representation Model with Convolutional Neural Network for Imputing Water Quality Data

The integrity of Water Quality Data (WQD) is critical in environmental monitoring for scientific decision-making and ecological protection. However, water quality monitoring systems are often challenged by large amounts of missing data due to unavoidable problems such as sensor failures and communication delays, which further lead to water quality data becoming High-Dimensional and Sparse (HDS). Traditional data imputation methods are difficult to depict the potential dynamics and fail to capture the deep data features, resulting in unsatisfactory imputation performance. To effectively address the above issues, this paper proposes a Nonlinear Low-rank Representation model (NLR) with Convolutional Neural Networks (CNN) for imputing missing WQD, which utilizes CNNs to implement two ideas: a) fusing temporal features to model the temporal dependence of data between time slots, and b) Extracting nonlinear interactions and local patterns to mine higher-order relationships features and achieve deep fusion of multidimensional information. Experimental studies on three real water quality datasets demonstrate that the proposed model significantly outperforms existing state-of-the-art data imputation models in terms of estimation accuracy. It provides an effective approach for handling water quality monitoring data in complex dynamic environments.

cs.LG

MetaSSL: A General Heterogeneous Loss for Semi-Supervised Medical Image Segmentation

Semi-Supervised Learning (SSL) is important for reducing the annotation cost for medical image segmentation models. State-of-the-art SSL methods such as Mean Teacher, FixMatch and Cross Pseudo Supervision (CPS) are mainly based on consistency regularization or pseudo-label supervision between a reference prediction and a supervised prediction. Despite the effectiveness, they have overlooked the potential noise in the labeled data, and mainly focus on strategies to generate the reference prediction, while ignoring the heterogeneous values of different unlabeled pixels. We argue that effectively mining the rich information contained by the two predictions in the loss function, instead of the specific strategy to obtain a reference prediction, is more essential for SSL, and propose a universal framework MetaSSL based on a spatially heterogeneous loss that assigns different weights to pixels by simultaneously leveraging the uncertainty and consistency information between the reference and supervised predictions. Specifically, we split the predictions on unlabeled data into four regions with decreasing weights in the loss: Unanimous and Confident (UC), Unanimous and Suspicious (US), Discrepant and Confident (DC), and Discrepant and Suspicious (DS), where an adaptive threshold is proposed to distinguish confident predictions from suspicious ones. The heterogeneous loss is also applied to labeled images for robust learning considering the potential annotation noise. Our method is plug-and-play and general to most existing SSL methods. The experimental results showed that it improved the segmentation performance significantly when integrated with existing SSL frameworks on different datasets. Code is available at https://github.com/HiLab-git/MetaSSL.

cs.CV

Radio-Frequency Quantum Rectification in Kagome Superconductor CsV3Sb5

Rectification of electromagnetic fields into direct current (DC) is pivotal for energy harvesting, wireless charging, and next-generation communication technologies. The superconducting diode effect, which exploits the nonreciprocal transport of dissipationless superconducting currents, offers ultra-low power consumption and high rectification ratios. Combining the superconducting diode effect with the AC Josephson effect holds promise for converting radio-frequency (rf) irradiation into a quantized DC output. However, experimental realization has been hindered by challenges in achieving the necessary symmetry breaking and fabricating high-performance Josephson junctions. Here we demonstrate the quantum rectification in kagome superconductor CsV3Sb5, which hosts emergent Josephson effects and a zero-field Josephson diode. Under rf irradiation, a DC voltage emerges without applied bias, scaling linearly with frequency as V = hf/2e, where h is Planck's constant, f is the microwave frequency, and e is the electron charge. Furthermore, the rectified voltage exhibits quantized steps with increasing rf power, consistent with Shapiro step quantization. Our work establishes CsV3Sb5 as a versatile platform for wireless quantum power supplies and charging, and underscores the intertwined order parameters as a promising pathway for precise quantum matter control.

cond-mat.supr-con