SearcharxivSearch

arXiv subjects

Sheng Liu

Publications and source records attributed to Sheng Liu.

At least 19 recordsLinked to original sources

Multi-Tool Image Editing Attribution in Facial Forgery

As generative AI tools become increasingly powerful and easy to use, people can easily edit portrait images with a prompt, necessitating the task of image editing attribution, which predicts the involved editing tools from the given image. Existing attribution methods hold the single-tool assumption and can only attribute a specific editing tool, but struggle to handle the more complex and increasingly common multi-tool editing scenarios, where artifacts left by different editing tools are composite and overlapped. To address this gap, we explore Multi-Tool Image Editing Attribution (MIEA), which aims to identify multiple editing tools involved in a multi-tool edited facial image. To simulate the real-life editing operations on facial images, we then construct a new dataset, MultiEdit, which contains 500k+ edited facial images and covers six types of editing tools that support face swapping (Deepfake) and various facial enhancements. Inspired by the findings from data analysis, we design DPEC, a multi-tool attribution method that can capture distinguishable, locality-aware editing tool traces from both spatial and frequency domains with the support of an error-based curriculum learning strategy. Experiments show \Method\ outperforms nine methods for facial images edited in at most five steps.

cs.CV

Descent and Brauer-Manin Obstructions on Deligne-Mumford Stacks

We generalize and compare local-global obstructions for algebraic stacks over number fields. For smooth separated Deligne-Mumford stacks of finite type with a quasi-projective coarse moduli space, we prove that the descent obstruction is contained in the Brauer--Manin obstruction. By lifting $\mathbb{G}_m$-gerbes, we obtain an inclusion between the corresponding composite obstructions. We also show that in this setting the descent obstruction coincides with both the \'etale Brauer-Manin obstruction and the iterated descent obstruction. The Brauer-Manin obstruction also coincides with the iterated Brauer-Manin obstruction.

math.AG

Let AEDs Move: Urban Mobility Enhanced Defibrillator Deployment

Timely access to automated external defibrillators (AEDs) remains a major limitation of public-access defibrillation systems because stationary AED networks cannot adapt to spatial and temporal variation in out-of-hospital cardiac arrest (OHCA) demand. We study a platform-enabled mobile AED model in which AEDs are carried by existing urban ride-hailing fleets that can be dispatched to OHCA locations. Rather than assuming additional AED capacity, we hold the total AED supply fixed and compare alternative allocations between stationary and mobile deployment, including hybrid policies that retain both forms of coverage. Our analysis leverages real-world urban data from New York City and Toronto, integrating georeferenced public AED inventories, EMS-reported OHCA incident data, and ride-hailing mobility data. The results show that mobile AEDs can substantially improve both response times and reliability, but the benefits depend on how AED capacity is divided between stationary and mobile deployment. In New York City, average response-time gains exceed 2.5 minutes and peak when approximately 55 to 70 percent of AED capacity is mobile; beyond this range, both gains and reliability decline as stationary coverage is reduced. In Toronto, average gains approach 4 minutes and then plateau as the system moves toward a fully mobile configuration. These contrasting patterns show that mobile AEDs can materially improve access, but the best mix of stationary and mobile AEDs must be tailored to local coverage and mobility conditions.

math.OC

ProjFormer: Point Cloud Completion via Geometric-Projective Transformer and Cross-Modal Semantic Constraints

Point cloud completion is inherently ill-posed due to severe sparsity and ambiguity in partial observations. Existing multi-view methods alleviate this by incorporating 2D semantics, but often rely on learned attention and fixed fusion, which lack geometric consistency and adaptability. We propose ProjFormer, a cross-modal framework that enforces geometry-consistent 2D-3D interaction through explicit projection and adaptive feature routing. A Projective Guided View Attention module aligns 3D points with multi-view features via deterministic projection, enabling efficient and geometrically consistent aggregation. Building on this, a geometry-aware routing network performs point-wise adaptive fusion of structural and observation-driven features for progressive refinement. Experiments show that, under a lightweight design, ProjFormer delivers competitive performance with improved structural completeness.

cs.CV

Spin-canting-induced Giant Nonlinear Optical Magnetochirality in a 2D Ferrotoroid

Achieving magnetically switchable chiral light emission is an important goal for 2D opto-spintronics. However, conventional strategies face a fundamental trade-off between dynamic tunability and polarization contrast. Nonlinear optics, particularly the emerging mechanism of chiral second-harmonic generation (SHG), offers a distinct strategy to bypass this restriction, yet its experimental realization remains elusive due to stringent symmetry requirements. Here, we report giant nonlinear optical magnetochirality in a centrosymmetric 2D ferrotoroid, bilayer (2L) CrSBr. We reveal that a field-induced spin-canting state breaks the parity-time (PT) symmetry of the unperturbed antiferromagnetic (AFM) ground state, activating a spin-chirality-driven i-type susceptibility. The coherent interference between this emergent i-type and intrinsic c-type SHG susceptibilities generates a macroscopic circularly polarized SHG signal whose helicity is magnetically switchable. Leveraging this sensitive mechanism, we uncover remanent magnetic states after field saturation that evade conventional linear probes. By exploiting the non-volatility of these states, we demonstrate magneto-optical memory and logic operations. Our work establishes a general symmetry-driven strategy for tailoring nonlinear magnetochirality, while providing a sensitive optical probe for subtle spin textures in the 2D limit.

physics.optics

Entanglement-based quantum key distribution with data in hollow-core fiber

The coexistence of quantum information and classical signals in a single fiber is essential for future quantum networks that leverage the well-established optical fiber infrastructure. Although multiplexing technologies can separate quantum and classical signals, pure silica core fibers (PSCFs) remain fundamentally limited by the high nonlinearity, which generates substantial Raman scattering and four-wave mixing noise. Hollow-core fibers (HCFs), guiding light predominantly in air, offer an attractive solution with intrinsically ultra-low nonlinearity and strongly suppressed nonlinear noise. In this work, we demonstrate the entanglement-based key coexisting with data over an 18-km HCF link. We achieve time-encoded high-dimensional quantum key distribution (HD-QKD) carrying 0 dBm of bidirectional received power, corresponding to a theoretical data capacity of up to 2.3 Tbps. During 24 hours of continuous operation, an average secret key rate (SKR) of 10.56 kbps is obtained. Theoretical analysis further predicts SKRs above 135 kbps over transmission distances exceeding 200 km using state-of-the-art low-loss HCFs. These results show significantly improved performance compared with PSCF-based systems and highlight the potential of HCFs for scalable quantum-classical coexistence compatible with the architectures of established fiber-optic networks.

quant-ph

ASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy Segmentation

Quantitative drug-induced sleep endoscopy (DISE) requires reliable airway boundaries at specific anatomical levels. Pixel-level DISE annotations are scarce, and manual contouring limits the scalability of quantitative assessment. To address this limitation, we developed ASTRA-Net for known-plane DISE segmentation with limited real annotations. Stage 1 aligned intermediate ConvNeXt-Base representations from 14,250 unlabeled virtual endoscopy frames derived from computed tomography and real DISE frames. Virtual images were used only for feature alignment. Stage 2 fine-tuned four independent UNet++ decoders on 401 real annotated frames. Structured zero-mask supervision constrained incompatible plane outputs and invalid frames. Six alignment configurations used maximum mean discrepancy, domain adversarial learning, or both objectives. On a hold-out evaluation set of 100 frames, the five-model MMD-only segmentation ensemble achieved a mean Dice of 0.8927, with a 95% image-level bootstrap interval of 0.8631 to 0.9160. The mean intersection over union was 0.8239. A classification- enabled variant of the same alignment configuration reached a restricted four-plane top-1 accuracy of 0.92 on the same hold-out frames. These results indicate that ASTRA-Net can support frame-level, plane-specific DISE boundary delineation when real annotations are limited.

cs.CV

SPECTRA: Context-Conditioned Spectral Movement Primitives for Robot Skill Generalization

Robot imitation learning for manipulation should preserve demonstrated task geometry while producing dynamically admissible robot motions. Existing pipelines often learn task-dependent trajectories and impose execution limits afterward through filtering, smoothing, clipping, or time scaling, which may distort task-critical end-effector paths. We propose the Spectral Movement Primitive (SMP), a frequency-domain imitation learning framework that couples task-space skill generation with joint-space execution regulation. Demonstrations are represented by truncated finite-horizon Fourier coefficients. An empirically selected low-frequency task band captures the dominant motion geometry, while higher harmonics contribute disproportionately to derivative growth. A frame-aware context-conditioned GMM/GMR prior predicts the task-band coefficients in a canonical task frame, and the resulting Cartesian trajectory is mapped to joint space through sequential inverse kinematics. A phase-coupled regulator then limits the requested phase progression without modifying the spectral coefficients, thereby enforcing joint velocity and acceleration limits while preserving the represented path. Experiments evaluate task-band reconstruction, robustness to composite demonstration corruption, out-of-distribution cross-board generalization, joint-space dynamic admissibility, end-effector path preservation, and deployment on a Franka Panda robot. Results show compact geometric reconstruction, consistent transfer across unseen task frames, substantial reductions in dynamic violations and jerk, and preservation of the intended end-effector path during phase regulation.

cs.RO

Traveling Salesman Tardiness

How fragile is the routing time window of delivery systems against spatial distributional uncertainty? We study the tardiness risk of Traveling Salesman Problem (TSP) solutions with respect to a service deadline (target) over the routing time. Using the robust satisficing model, we introduce the TSP tardiness index to quantify the target's fragility under distributional uncertainty in customer locations. Assuming there are m potential customer locations from historical samples on a service region D (of area |D|), we prove that the TSP tardiness index is {\Theta}(n * sqrt(|D|m) / {\tau}) for n realized locations with respect to the routing time target {\tau}, under non-boundary conditions. This result establishes a new scaling law that extends beyond the existing deterministic and probabilistic TSP bounds. We further extend it to a multi-vehicle case and derive simple partition rules for managing delivery systems. Our numerical experiments using synthetic and real-world routing data validate the value of the TSP tardiness index in characterizing and managing the overtime risk of routing systems.

math.OC

Deviations beyond the Kibble-Zurek mechanism in a Spin-Orbit-Coupled Bose-Einstein Condensate with phenomenological damping

We investigate the quench dynamics in a one-dimensional spin-orbit-coupled Bose-Einstein condensate (SOC-BEC) across the phase transition from plane-wave (PW) to stripe (ST), incorporating phenomenological damping. In the dissipation-free case, a state stagnation phenomenon emerges during the PW-ST quench: for slow quenches, the system remains trapped in the PW phase due to the energy gap induced by critical slowing down, which prevents spontaneous relaxation to the stripe ground state. To explore this phenomenon and examine the universal scaling predicted by the Kibble-Zurek mechanism (KZM) in open systems, we introduce a dissipative Gross-Pitaevskii equation with a phenomenological damping term. Numerical simulations reveal that weak dissipation preserves the expected KZM power-law scaling for the freeze-out time and defect density, whereas strong dissipation or long quench times lead to significant deviations. Our results demonstrate that the KZM remains applicable in dissipative quantum systems under appropriate conditions, providing insights into nonequilibrium dynamics in open quantum systems.

cond-mat.quant-gas

Electric-field-driven magnetic switching and tightly bound interlayer excitons in bilayer CrSBr

Electric field control of magnetic order in two-dimensional (2D) van der Waals magnets is a central goal for low-power spin-based technologies. In the ambient-stable antiferromagnet CrSBr, strong magnetic anisotropy and robust exciton-spin coupling provide a favorable platform, yet deterministic electric field control of its magnetic phases has not been achieved. Here we demonstrate electric-field-driven reversible switching between antiferromagnetic and ferromagnetic states in dual-gated bilayer CrSBr without intentional carrier doping. In parallel, photoluminescence measurements resolve a tightly bound interlayer exciton with an intrinsic dipole moment of only ~1 e angstrom. The electric field dependence of the magnetic phase transition reveals two coexisting mechanisms: a linear magnetoelectric effect in the antiferromagnetic state and an electric-field-modulated interlayer exchange coupling. Their interplay accounts for the asymmetric evolution of the critical magnetic field. Our results establish bilayer CrSBr as a promising 2D material for electrically controlled spin-optoelectronic functionalities.

cond-mat.mes-hall

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stale, and surface the right evidence at decision time. Existing benchmarks measure recall over static dialogue, collapse memory into a single end-of-task accuracy, and reduce visual observations to captions, leaving us unable to localize failures to writing, maintenance, retrieval, or use. The rise of agent harnesses that author their own memory sharpens this gap, since we have no principled way to compare hand-designed pipelines with self-managing alternatives. To close these gaps, we formulate multimodal agent memory as an Action-World Interaction Loop with an observable four-stage lifecycle, and instantiate it in WorldMemArena: 400 multi-session multimodal tasks spanning Lifelong Evolution (evolving personal and task states) and Agentic Execution (memory from real observations, actions, and feedback), annotated with gold memory points, updates, distractors, and evidence chains for stage-level diagnosis. This enables the first head-to-head comparison of long-context, manually designed (RAG and external memory systems), and harness-based memory agents. Results show that: (1) better memory writing and storage do not guarantee better performance; (2) multimodal memory still struggles to fully use visual evidence; (3) systems are unstable across domains and degrade on realistic agentic trajectories; and (4) harness memory is more flexible but remains costly and less reliable.

cs.CV

Anderson Transition and Mobility Edges in a Family of 3D Fractal Lattices

Anderson localization is fundamentally controlled by dimensionality, yet the nature of the Anderson transition in continuously tunable noninteger dimensions remains largely unexplored. Here, we introduce a family of three-dimensional fractal lattices with continuously tunable spectral dimension $d_s\in[2,3]$, providing a controlled platform for studying localization physics beyond integer dimensions and across the lower critical dimension $d_s=2$. Using large-scale finite-size scaling analysis, we systematically investigate the Anderson transition and identify mobility edges throughout the fractal family. The critical disorder strength evolves continuously from $0$ to $16.6$ as the spectral dimension increases from $2$ to $3$. We show that the spectral dimension predominantly governs the universality class of the transition, while the precise critical point is additionally influenced by microscopic geometric details of the underlying fractal lattice. The critical exponent exhibits an approximate inverse dependence on $d_s$, providing quantitative insight into scaling theory in noninteger dimensions. Our results establish tunable fractal lattices as a versatile framework for exploring localization and quantum critical phenomena beyond conventional integer-dimensional systems.

cond-mat.dis-nn

Auditing Agent Harness Safety

LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. However, a harness can return a correct, benign answer over a trajectory that accesses unauthorized resources or leaks context to the wrong agent. Output-level evaluation cannot see these failures, yet most safety benchmarks score only final outputs or terminal states, even though many violations occur mid-trajectory rather than at termination. The central question is whether the harness respects user intent, permission boundaries, and information-flow constraints throughout execution. To address this gap, we propose HarnessAudit, a framework that audits full execution trajectories across boundary compliance, execution fidelity, and system stability, with a focus on multi-agent harnesses where these risks are most pronounced. We further introduce HarnessAudit-Bench, a benchmark of 210 tasks across eight real-world domains, instantiated in both single-agent and multi-agent configurations with embedded safety constraints. Evaluating ten harness configurations across frontier models and three multi-agent frameworks, we find that: (i) task completion is misaligned with safe execution, and violations accumulate with trajectory length; (ii) safety risks vary across domains, task types, and agent roles; (iii) most violations concentrate in resource access and inter-agent information transfer; and (iv) multi-agent collaboration expands the safety risk surface, while harness design sets the upper bound of safe deployment.

cs.CL

CT-Guided Spatially-varying Regularization for Voxel-Wise Deformable Whole-Body PET Registration

Whole-body Positron Emission Tomography (PET) registration is essential for multi-parametric tumor characterization and assessment of metastatic disease progression. In deep learning-based deformable registration, the dense displacement field (DDF) regularizer is crucial for stabilizing optimization and preventing unrealistic deformations in large 3D volumes. A key challenge in whole-body deformable registration is anatomical heterogeneity, rigid structures (e.g., bones) should undergo stronger regularization, whereas soft tissues require more flexible deformation and weaker constraints. In this work, we propose a simple yet effective CT-guided spatially-varying regularization strategy for whole-body cross-tracer deformable PET registration. The key idea is to use the paired CT volume from the PET/CT acquisition to construct a voxel-wise regularization map for the DDF, replacing the conventional single global regularization weight. This yields anatomy-adaptive regularization strength across rigid and soft tissues. The proposed method is evaluated on a real clinical cross-tracer PET/CT dataset of 296 patients involving 18F-PSMA and 18F-FDG, showing that the proposed method achieves statistically significant improvements over weakly-supervised registration baseline in both whole-body registration performance and organ-wise alignment.

eess.IV

ViLL-E: Video LLM Embeddings for Retrieval

Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Captioning. However, they underperform specialized embedding-based models in Retrieval tasks, such as Text-toVideo Retrieval and Moment Retrieval. We introduce ViLL-E (Video-LLM-Embed), a unified VideoLLM architecture endowed with a novel embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones. We train this model with a three-stage training methodology combining generative and contrastive learning: initial large-scale pre-training with video-caption pairs; followed by continual training on a smaller, detailed-caption dataset; and concluding with task-specific fine-tuning on a novel multi-task dataset covering Video QA, Temporal Localization, Video Retrieval, and Video-Text Matching. Our model significantly improves temporal localization (on avg. 7% over other VideoLLMs) and video retrieval (up to 4% over dual encoder models), achieving performance comparable to state-of-the-art specialized embedding models while remaining competitive on VideoQA tasks. Furthermore, our joint contrastive-generative training unlocks new zero-shot capabilities, significantly outperforming state-of-the-art methods in composed video retrieval (+5% over SotA) and retrieval from long text (+2% over SotA).

cs.CV

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

Modern clinical practice increasingly depends on reasoning over heterogeneous, evolving, and incomplete patient data. Although recent advances in multimodal foundation models have improved performance on various clinical tasks, most existing models remain static, opaque, and poorly aligned with real-world clinical workflows. We present Cerebra, an interactive multi-agent AI team that coordinates specialized agents for EHR, clinical notes, and medical imaging analysis. These outputs are synthesized into a clinician-facing dashboard that combines visual analytics with a conversational interface, enabling clinicians to interrogate predictions and contextualize risk at the point of care. Cerebra supports privacy-preserving deployment by operating on structured representations and remains robust when modalities are incomplete. We evaluated Cerebra using a massive multi-institutional dataset spanning 3 million patients from four independent healthcare systems. Cerebra consistently outperformed both state-of-the-art single-modality models and large multimodal language model baselines. In dementia risk prediction, it achieved AUROCs up to 0.80, compared with 0.74 for the strongest single-modality model and 0.68 for language model baselines. For dementia diagnosis, it achieved an AUROC of 0.86, and for survival prediction, a C-index of 0.81. In a reader study with experienced physicians, Cerebra significantly improved expert performance, increasing accuracy by 17.5 percentage points in prospective dementia risk estimation. These results demonstrate Cerebra's potential for interpretable, robust decision support in clinical care.

cs.AI

FedTrident: Resilient Road Condition Classification Against Poisoning Attacks in Federated Learning

FL has emerged as a transformative paradigm for ITS, notably camera-based Road Condition Classification (RCC). However, by enabling collaboration, FL-based RCC exposes the system to adversarial participants launching Targeted Label-Flipping Attacks (TLFAs). Malicious clients (vehicles) can relabel their local training data (e.g., from an actual uneven road to a wrong smooth road), consequently compromising global model predictions and jeopardizing transportation safety. Existing countermeasures against such poisoning attacks fail to maintain resilient model performance near the necessary attack-free levels in various attack scenarios due to: 1) not tailoring poisoned local model detection to TLFAs, 2) not excluding malicious vehicular clients based on historical behavior, and 3) not remedying the already-corrupted global model after exclusion. To close this research gap, we propose FedTrident, which introduces: 1) neuron-wise analysis for local model misbehavior detection (notably including attack goal identification, critical feature extraction, and GMM-based model clustering and filtering); 2) adaptive client rating for client exclusion according to the local model detection results in each FL round; and 3) machine unlearning for corrupted global model remediation once malicious clients are excluded during FL. Extensive evaluation across diverse FL-RCC models, tasks, and configurations demonstrates that FedTrident can effectively thwart TLFAs, achieving performance comparable to that in attack-free scenarios and outperforming eight baseline countermeasures by 9.49% and 4.47% for the two most critical metrics. Moreover, FedTrident is resilient to various malicious client rates, data heterogeneity levels, complicated multi-task, and dynamic attacks.

cs.CR