Searcharxiv⌕ Search

arXiv subjects

Yu Shi

Publications and source records attributed to Yu Shi.

At least 37 records · Page 2Linked to original sources

Holographic Timelike Entanglement and Subregion Complexity in Localized AdS3*S3*T4 Black Holes

We study timelike entanglement entropy and timelike subregion complexity in localized black holes with asymptotic AdS3*S3*T4 geometry, focusing on the black-pole solution. Unlike the BTZ solution, the black pole exhibits a nontrivial dependence on the internal sphere through the functions $K_y(r,θ)$ and $G(r,θ)$. Both observables are constructed from spacelike and timelike Lorentzian branches, but they probe the geometry in different ways: timelike entanglement yields a complex lifted area, while timelike complexity gives a real, finite renormalized volume. We employ a localized timelike prescription in which the branch profile is built at an angular label $θ_0$ and subsequently lifted over the physical internal angle $θ$. In the large-$r$ regime, the leading angular dependence drops out, recovering the expected short-interval behaviour. In the exact black-pole geometry, the temporal families become non-monotonic, making a fixed-boundary-interval selection essential. As the boundary interval increases, the selected branches move inward and become sensitive to the localized cap-horizon transition region. These results demonstrate that timelike Lorentzian observables probe localized-geometry effects that are absent in BTZ and in the leading large-$r$ description.

hep-th↗

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabilities, yet often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework that unifies generalist vision-language learning with specialist objectives, enabling both broad generalization and specialist-level diagnostic capability. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that captures lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision.

cs.CV↗

Multiparameter Quantum Estimation and Degeneracy Structure in Three-Flavor Neutrino Oscillations

Achieving precision measurements of neutrino oscillation parameters and resolving parameter degeneracies remain central challenges in neutrino physics. This work presents a systematic investigation of three-flavor neutrino oscillations within the framework of quantum estimation theory using the quantum Fisher information matrix (QFIM). The behavior of all six independent elements of the QFIM associated with the parameters theta23, deltaCP, and Delta(m31)^2 is analyzed, and the impact of parameter correlations on the quantum Cramér-Rao bound is studied. Furthermore, we demonstrate that parameter degeneracies in neutrino oscillation probabilities do not necessarily imply indistinguishability of the underlying quantum states. By employing quantum fidelity and the QFIM, we show that degenerate parameter sets can exhibit distinct quantum-information characteristics that remain hidden at the probability level, revealing quantum-state differences between probability-degenerate solutions.

hep-ph↗

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic by enforcing uniform thresholds across all tokens independently. This pointwise treatment conflicts with autoregressive generation in two critical ways. First, uniform thresholds ignore autoregressive asymmetry. Early-stage deviations produce compounding sequence-level drift, causing static thresholds to under-regulate early divergence and excessively constrain late-stage exploration. Second, evaluating token-level divergence in isolation overlooks cumulative prefix drift, granting the same divergence allowance regardless of how far the conditioning history has already deviated from the rollout policy. To address this limitation, we propose CPPO (Cumulative Prefix-divergence Policy Optimization), a token-level masking rule that aligns updates with a finite-horizon policy-improvement bound via two coupled mechanisms. First, a position-weighted threshold imposes stricter limits at early positions whose effects persist longer, relaxing constraints for late-stage tokens. Second, a cumulative prefix budget tracks historical deviations, dynamically restricting further token-level deviation to prevent compounding errors along the prefix. Empirically, CPPO enhances training stability and significantly improves reasoning accuracy across various model scales.

cs.LG↗

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical workflow that spans \textit{screening}, \textit{diagnosis} and \textit{treatment planning}, where each stage involves distinct imaging modalities, task objectives, and reasoning patterns. However, constrained by data scarcity and model versatility, existing medical MLLMs are typically evaluated on isolated modalities or narrow task families, limiting their ability to support workflow-level clinical reasoning. In this work, we first introduce \textbf{BreastStage}, a workflow-aligned breast imaging instruction corpus comprising 1.86M instruction-following pairs curated from 17 sub-datasets across 5 imaging modalities and 136 task templates. Its held-out split, \textbf{BreastStage-Bench}, provides a comprehensive benchmark for evaluating multimodal reasoning across the breast cancer care continuum. Building on this corpus, we propose \textbf{BreastGPT}, a unified MLLM equipped with a dual-branch visual encoder and concept-preserving token compression to bridge the scale gap between standard radiology and gigapixel pathology. On BreastStage-Bench, BreastGPT achieves 75.66\% closed-ended accuracy and 89.92\% open-ended score, outperforming both general-purpose and medical-specific MLLMs across clinical stages and task formats. These results suggest that workflow-aligned data and cross-scale visual modeling are critical for clinically grounded medical MLLMs. All data, code, and model checkpoints are released at https://yangyy-liu.github.io/BreastGPT.io.

cs.CV↗

Relativity from the Perspectives of Observers

This paper reviews the role of observers in the development of relativity theory, from special relativity to general relativity, emphasizing that observer-dependent descriptions are as fundamental as the covariance of physical laws. This paper reviews the role of observers in the development of relativity theory, from special relativity to general relativity, emphasizing that observer-dependent descriptions are as fundamental as the covariance of physical laws. After the introduction of a geometric framework for observers using timelike worldlines, Frenet-Serret formulas, projection operators, and the Frobenius condition for hypersurface-orthogonal families, the paper revisits key problems in early relativistic mechanics, such as the transformation of velocity and acceleration, the variational principle for particle motion, and the Ehrenfest paradox concerning rigid rotation. It shows that while early physicists often conflated coordinate systems with reference frames, their results remain valid because the underlying geometric objects are observer-independent. The historical analysis, from Einstein's 1905 work to the development of general relativity and later advances such as Hawking radiation, demonstrates that clarifying the concept of observers not only resolved paradoxes but also paved the way toward a field-theoretic formulation of gravity. The paper concludes that observer dependence, far from being a nuisance, is an essential ingredient for understanding spacetime physics.

physics.hist-ph↗

Azimuthal decorrelation in diffractive dijet production

We calculate the azimuthal angular decorrelation of diffractive dijets in ultra-peripheral heavy-ion, $ep$, and $eA$ collisions to probe non-perturbative diffractive transverse momentum-dependent distributions. Focusing on the dominant semi-inclusive channel with an unobserved semi-hard gluon, we perform an all-order resummation of soft gluon emissions for the transverse energy-energy correlator observable, accounting for both initial and final state radiation. We also analyze heavy-quark pair production and demonstrate the sensitivity of the decorrelation to the jet axis definition. Finally, we provide numerical predictions for relevant kinematics at LHC UPCs, HERA, and the future EIC. Our results demonstrate that the acoplanarity of diffractive dijet production could serve as a promising probe of diffractive transverse momentum-dependent distributions.

hep-ph↗

MatterSim-MT: A multi-task foundation model for in silico materials characterization

Accurate property characterization is a major bottleneck in materials design. While first-principles methods and task-specific machine-learning models have driven important progress, they remain fundamentally limited in scalability and generalizability across the vast space of structures and properties relevant to real-world materials design. We present MatterSim-MT, a multi-task foundation model for in silico materials simulation and property characterization. The model is pretrained on over 35 million first-principles-labeled structures covering 89 elements, temperatures up to 5000 K and pressures up to 1000 GPa, and is fine-tuned on various properties including Bader charges, magnetic moments, Born effective charges, and dielectric matrices. Out of the box, MatterSim-MT not only serves as a foundation model for predicting material structure, dynamics and thermodynamics, its multi-task architecture also enables a wide range of complex simulations that cannot be captured by potential energy surfaces alone. For example, we demonstrate pressure-dependent LO-TO phonon splitting in SiC with close agreement with experiment, electric hysteresis in ferroelectric BaTiO3, and the cationic-to-anionic redox transition during delithiation of a Li-rich cathode material. Finally, we show that MatterSim-MT scales well with more data and parameters, can be efficiently fine-tuned to higher levels of theory, and can be efficiently extended to new systems via active learning. Overall, we believe this approach provides a scalable route to accurate in silico materials characterization.

cond-mat.mtrl-sci↗

Training-Free Quantum Generative Paradigm via Local Parent Hamiltonians

We propose a training-free quantum generative paradigm, which is fundamentally different from current generative models, which demand substantial computational power, face practical scalability limits, and often function as opaque black boxes, despite their remarkable success. We enable image and text generation without parameter training, by constructing a local parent Hamiltonian whose ground state encodes the target distribution and then solving the global Hamiltonian. Rooted directly in quantum mechanical principles, this approach establishes a new pathway for generative modeling that leverages superposition and entanglement to maintain global consistency.

quant-ph↗

Generative structure search for efficient and diverse discovery of molecular and crystal structures

Predicting stable and metastable structures is central to molecular and materials discovery, but remains limited by the cost of searching high-dimensional energy landscapes. Deep generative models offer efficient structure sampling, yet their outputs remain shaped by training data and can underexplore minima that are rare but physically relevant. We introduce generative structure search (GSS), a unified framework that formulates diffusion-based generation and random structure search (RSS) as limiting regimes of a common sampling process driven by learned score fields and physical forces. Coupling these drivers lets GSS use data priors to accelerate sampling while retaining energy-guided exploration of local minima. Across molecular and crystalline systems, GSS recovers diverse metastable structures with more than tenfold lower sampling cost than RSS for broad coverage and remains effective for compositions outside the training distribution. The results establish a physically grounded generative search strategy for discovering structures beyond the reach of data-driven sampling alone.

cs.AI↗

FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards

Live future prediction refers to the task of making predictions about real-world events before they unfold. This task is increasingly studied using large language model-based agent systems, and it is important for building agents that can continually learn from the real world. It can provide a large number of prediction questions grounded in diverse real-world events, while preventing answer leakage. To leverage the advantages of future prediction, we present FutureWorld, a live agentic reinforcement learning environment that closes the training loop between prediction, outcome realization, and parameter updates. Specifically, we modify and extend verl-tool, resulting in a new framework that we call verl-tool-future. Unlike standard reinforcement learning training frameworks that rely on immediate rewards, verl-tool-future stores prediction-time rollouts, backfills rewards after real-world outcomes become available, and then replays the completed trajectories for policy update. Across three open-source agents, successive FutureWorld training rounds lead to consistent improvements in prediction accuracy, probabilistic scoring, and calibration, demonstrating that delayed real-world outcome feedback can serve as an effective reinforcement learning signal.

cs.AI↗

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents

As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locations, and behavioral traces. While existing benchmarks primarily focus on task completion, grounding, or defenses against third-party attacks, current visual privacy datasets remain largely restricted to static natural images, limiting their ability to capture the contextual dependence and task relevance of privacy risks in GUI task trajectories. To bridge this gap, we introduce \textbf{GUIGuard-Bench}, a first-step benchmark for studying privacy-preserving GUI agents in trajectory-based GUI workflows. GUIGuard-Bench contains 241 real GUI-agent trajectories with 4,080 screenshots across Android and PC environments. Each screenshot is annotated at the region level with privacy bounding boxes, semantic privacy categories, risk levels, and whether the private information is necessary for completing the task. Built on these annotations, GUIGuard-Bench supports three complementary evaluations: privacy recognition, offline planning fidelity under protected screenshots, and the utility impact of different protection strategies. Our results show that current models can often detect whether a screenshot contains private information, but they struggle with fine-grained localization, category recognition, risk assessment, and task-necessity judgment. We also find that closed-source models, exemplified by Claude Sonnet 4.6, can maintain largely consistent planner semantics in Android environments after privacy protection is applied. Our results highlight privacy recognition as a critical bottleneck for practical GUI agents. Project: https://futuresis.github.io/GUIGuard-page/

cs.CR↗

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored, requiring comprehensive e-commercial dataset and elaborative utilization of multi-modal condition. In this paper, we propose a brand-new e-commerce dataset, named Fashion130k, with various occasions, models, and garment types. For the consistent generation of garment, we design a framework with Unified Multi-modal Condition (UMC) to align and integrate the text and visual prompts into generation model. Specifically, we explore an embedding refiner to extract the unified embeddings of multi-modal prompts, within which a Fusion Transformer is proposed to align the multi-modal embeddings by adjusting the modality gap between text and image. Based on unified embeddings, the attention in generation model is redesigned to emphasis the correlations between prompts and noise image, inducing that the noise image can select the pivotal tokens of prompts for consistent outfit generation. Our dataset and proposed framework offer a general and nuanced exploration of multi-modal prompts for generation models. Extensive experiments on real-world applications and benchmark demonstrate the effectiveness of UMC in visual consistency, achieving promising result than that of SoTA methods.

cs.CV↗

Quantumness of top quark pairs produced at LHC within SMEFT framework

Top and anti-top quark pair production at LHC provides a unique setting to probe non-classical correlations at the TeV scale. We study quantum information (QI) properties of the $t\bar{t}$ spin state in $pp$ collisions at $\sqrt{s}=13$ TeV within the Standard Model Effective Field Theory (SMEFT), focusing on dimension-6 operators that induce anomalous chromo- and weak dipole moments of the top quark within their current experimental bounds. The $t\bar{t}$ spin density matrix is reconstructed from the joint angular distribution of the final state charged leptons in the $k$-$r$-$n$ helicity basis. We analyze three complementary QI quantities: concurrence-based quantum entanglement (QE), geometric quantum discord (GQD), and the Bell parameter,across four $t\bar{t}$ invariant-mass bins. Within the Standard Model (SM), non-vanishing QE appears only near threshold ($m_{t\bar{t}}\lesssim 400$ GeV), while GQD remains nonzero across the full phase space, indicating persistent non-classical correlations even for separable states. Anomalous chromo-dipole interactions modify these observables primarily near threshold: $\hatμ_t$ induces asymmetric shifts, whereas $\hat{d}_t$ produces a mild symmetric response without Bell inequality violation. Among weak dipole operators, the CP-even coupling $C_2^V$ generates the largest deformation of the QI observables, while $ΔC_1^{A,V}$ leave them unchanged. These results demonstrate that QI observables derived from the $t\bar{t}$ spin density matrix provide a complementary probe of anomalous top-quark interactions with distinct sensitivity to CP-even and CP-odd operator structures.

hep-ph↗

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the image. Previous methods often involve complex steps of specifying the text content and attributes, such as font size, color, and layout, without considering the stylistic consistency with the reference image. To address this, we propose UM-Text, a unified multimodal model for context understanding and visual text editing by natural language instructions. Specifically, we introduce a Visual Language Model (VLM) to process the instruction and reference image, so that the text content and layout can be elaborately designed according to the context information. To generate an accurate and harmonious visual text image, we further propose the UM-Encoder to combine the embeddings of various condition information, where the combination is automatically configured by VLM according to the input instruction. During training, we propose a regional consistency loss to offer more effective supervision for glyph generation on both latent and RGB space, and design a tailored three-stage training strategy to further enhance model performance. In addition, we contribute the UM-DATA-200K, a large-scale visual text image dataset on diverse scenes for model training. Extensive qualitative and quantitative results on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance.

cs.CV↗

Harnessing Pre-Resolution Signals for Future Prediction Agents

Many high-stakes decisions depend on forecasts made before outcomes are known. In this future prediction setting, the central challenge is that public evidence evolves over time, while the main supervision signal arrives only after resolution: the realized outcome mainly assesses final correctness, offering only coarse guidance on what to track, what to verify, and which judgments to leave uncertain along the way. Our key observation is that revisiting the same unresolved question over time creates informative temporal contrasts across evolving evidence and repeated forecasts, exposing what earlier attempts missed before resolution and yielding a diagnostic signal we call the pre-resolution signal. We instantiate this idea in Milkyway, a future prediction agent with a persistent future prediction harness, an editable external state that stores reusable procedural guidance across revisits to the same unresolved question. As the same unresolved question is revisited, Milkyway extracts pre-resolution signals from evolving evidence and repeated forecasts, uses them to update the harness, and improves later forecasts on that question before resolution. After resolution, the realized outcome serves as a post-resolution check of provisional updates. On the FutureX and FutureWorld benchmarks, Milkyway achieves strong performance against competitive baselines, and a mechanism study suggests that the gains stem from harness evolution driven by pre-resolution signals rather than repeated prediction alone.

cs.AI↗

AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification

Identifying Bug-Inducing Commits (BICs) is fundamental for understanding software defects and enabling downstream tasks such as defect prediction and automated program repair. Yet existing SZZ-based approaches rely on git blame, restricting the search space to commits that directly modified the fixed lines. Our preliminary study on 2,102 validated bug-fixing commits reveals this limitation is significant: 28% of BICs require traversing commit history beyond blame results and 14% are blameless. We present AgenticSZZ, the first approach to apply Temporal Knowledge Graphs (TKGs) to software evolution analysis. AgenticSZZ reframes BIC identification from ranking blame commits into a graph search problem, where temporal ordering is fundamental to causal reasoning about bug introduction. The approach operates in two phases: (1) constructing a TKG that encodes commits with temporal and structural relationships, expanding the search space by traversing file history backward from blame commits and the bug-fixing commit; and (2) leveraging an LLM agent to navigate the graph using specialized tools for candidate exploration and causal analysis. Evaluation on three datasets shows that AgenticSZZ achieves F1-scores of 0.47 to 0.79, with statistically significant F1 improvements over state-of-the-art by up to 34%. Ablation confirms that both components and context expansion each contribute: the TKG and agent form an exploration-exploitation synergy, while context expansion unlocks ancestor BIC discovery, yielding 60 additional true positives. A sensitivity analysis across five open-weight LLMs reveals that effective TKG navigation requires sufficiently capable models, and that the TKG architecture amplifies stronger LLMs, widening the advantage. By transforming BIC identification into graph search, we open a new direction for temporal and causal reasoning in software evolution analysis.

cs.SE↗

Probing Saturation Effect in Heavy Meson Pair Correlation in Forward $pA$ Collisions

Forward two-particle angular correlations in $pA$ collisions have long been recognized as a particularly sensitive observable for exploring gluon saturation effects. In the back-to-back regime, two-particle correlations receive substantial contributions from both soft-gluon radiation and saturation effects. In this work, we study heavy meson pair correlation in forward proton-nucleus collisions by incorporating a unified Sudakov resummation for heavy meson pair correlations in the Color Glass Condensate effect theory. Our results are in good agreement with the $Δϕ$ data measured by the LHCb Collaboration for $D^0 \bar D^0$ pairs in forward $pp$ and $pA$ collisions, as well as $J/ψ$ pairs from $b\bar b$ decays in forward $pp$ collisions. Furthermore, we present predictions for $D\bar D$ and $B\bar B$ correlations in the forward rapidity regions at the Large Hadron Collider. A pronounced mass-hierarchy is observed in the nuclear modification factor, $R_{pA}\big|_{m_b}<R_{pA}\big|_{m_c}$, indicating stronger sensitivity to saturation effects at small $x$. As the rapidity increases, the suppression becomes more pronounced while the mass hierarchy remains robust. This study will help us to search for the saturation signal via heavy-meson pair correlations in forward $pA$ collisions.

hep-ph↗