SearcharxivSearch

arXiv subjects

Yuchen Guo

Publications and source records attributed to Yuchen Guo.

At least 19 recordsLinked to original sources

Observation of electron spin interactions between Rydberg atoms

We report the observation of electron spin interactions between Rydberg atoms, which are driven by spin-orbit coupling through second-order dipole perturbation and exhibit a spatial anisotropy governed by the atomic configuration. Specifically, we observe coherent electron spin exchange dynamics, with the measured coupling strength agreeing well with both numerical calculations and theoretical models. Furthermore, we show that global microwave dressing enables active engineering and dynamical freezing of the spin exchange by introducing a differential AC Stark shift between the participating states. Additionally, we achieve tunability of the interaction by applying a stronger magnetic field, which effectively modifies the energy contributions of the underlying spin-orbit coupling channels. Finally, measuring spin dynamics in one-dimensional multi-atom chains aligned parallel or perpendicular to the magnetic field provides a self-consistent validation of the anisotropic XXZ framework. This electron spin interaction natively features spin-position coupling and enables a natural mapping onto the Heisenberg-Kitaev model. Our findings reveal a new class of electron spin-spin interactions among Rydberg atoms, expanding the scope of quantum simulation with Rydberg atom arrays.

cond-mat.quant-gas

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.

cs.AI

Rigorous Statements and Proofs of the Lemmas in Simon's Algorithm for the Dihedral Coset Problem and Their Underlying Hypothesis

In a recent preprint, Simon proposed a polynomial-time quantum algorithm for the Dihedral Coset Problem and rested the analysis on four lemmas. Three of them carry only proof sketches, and this paper gives each of those three a statement that admits a single reading together with a complete proof. Lemma 1 follows from an exact second-moment computation for the subset-sum counts, and it holds with probability tending to one in place of the constant originally claimed. The amplitude bound of Lemma 3 follows from an exact Parseval identity on the cube of measurement outcomes and holds at every threshold with no well-behavedness hypothesis, so that predicate leaves the argument entirely. For Lemma 4, we compute both balls-in-bins covariances exactly and find that the second carries a term a fixed ball count leaves out. The assumption that the distinguished group contains no faulty samples can also be dropped. The two branch amplitudes share a signed prefactor, so the counting estimates control their difference and not the ratio the lemma states. We prove the additive form and show that the closing argument consumes nothing more than that. A single hypothesis survives all of this. It asks that the partition into the two sides be fixed independently of the measured string, and the rule the algorithm gives for choosing that partition does not supply it. Establishing these four lemmas therefore does not by itself establish the correctness of the algorithm.

cs.CR

Parent Hamiltonian and intrinsic phase transition in non-Hermitian photonic systems

Non-Hermitian systems host phenomena absent in Hermitian physics, but realizing Hamiltonians with intrinsic non-Hermitian properties remains challenging. The theoretical method of non-Hermitian parent Hamiltonian (NH-PH) enables the construction of a non-Hermitian system from a pair of matrix product states (MPSs) with tailored properties. Here, we report the first experimental generation of NH-PHs. This generation starts from MPSs that represent asymmetric Affleck--Kennedy--Lieb--Tasaki (AKLT) states. The construction is validated with single photons via imaginary-time evolution of the generated NH-PH to obtain its left and right ground states. We then characterize the properties of the system by measuring four different order parameters that probe non-reciprocal correlations, chiral imbalance, and conventional antiferromagnetic correlations. Furthermore, extending the framework to a larger system with a different model, we observe an intrinsic non-Hermitian phase transition, manifested by abrupt jumps of an order parameter when the designated zero-energy modes cease to be the globally lowest-energy states. Our work provides the first experimental realization and characterization of non-Hermitian Hamiltonians with controllable and customizable properties, opening new avenues for exploring intrinsic non-Hermitian phenomena across diverse physical platforms.

quant-ph

Tensor network characterization and mitigation of readout errors

Readout errors are a major bottleneck to extracting reliable information from near-term quantum processors, especially when spatial correlations are non-negligible. We present a unified tensor-network framework that models the readout process as a matrix product operator (MPO), enabling efficient characterization and mitigation beyond uncorrelated approximations. The MPO model is trained via likelihood optimization on calibration data and applies to multiple tasks, including nonlocal observable estimation, random circuit sampling, and random-measurement protocols, such as classical shadows and learning-based tomography. Experiments on a superconducting processor and numerical simulations up to 20 qubits show that the MPO model captures correlated readout errors that uncorrelated models miss, with a sample cost that grows only near-linearly with system size. When extended to two-dimensional systems, the framework can also be integrated with tensor-network quantum error-correction decoders by performing joint inference over data and readout errors. These results establish tensor-network readout error mitigation as a scalable and versatile approach for noise-aware quantum data processing.

quant-ph

Discovery of a Barred-Spiral Galaxy at $z_{spec}$ = 3.16 II. The Star Formation History

We present a detailed analysis of a massive barred galaxy at $z_{spec}=3.1591$ using deep multi-band imaging from HST and JWST. For the first time, we resolve its morphology and stellar structures thanks to the JWST/NIRCam NIR and MIR photometry. The galaxy possesses two distinct components with significantly different colors. Through careful image decomposition and masking, we isolate and characterize the flux contribution from each component. The galaxy exhibits a clear spiral morphology, and in a separate companion paper, we present evidence suggesting the presence of a stellar bar. Based on spatially resolved spectral energy distribution modeling with Prospector, we derive the star formation history and other physical properties of the bar and the surrounding regions. The total stellar mass of the galaxy is constrained as $\log(M_*/M_{\odot}) = 10.63\pm0.13$. We find that the bar region contains around 30% of the total stellar mass, but only accounts for around 8% of the recent star formation rate. The region containing the potential bar shows a significantly older mass-weighted stellar age, supporting the inside-out scenario for galaxy formation, and providing tentative evidence for bar quenching in the early stage. The quick onset of a stellar bar at this redshift requires a low dark matter fraction, suggesting the baryon-dominated nature of high-$z$ massive galaxies, and offering rare insight into galaxy evolution at around two billion years after the Big Bang.

astro-ph.GA

Discovery of a Barred-Spiral Galaxy at $z_{spec}$ = 3.16 I: Bar Identification and Properties

The formation of stellar bars is an important milestone in the secular evolution of spiral galaxies, which typically indicates the presence of a massive rotationally supported disk. Determining when these structures first appeared in the early universe is crucial to constraining the timeline of galactic disk assembly. Here, we report the discovery of COSMOS-74706, a barred spiral galaxy at $z_{spec} = 3.159$. Imaging of COSMOS-74706 with JWST/NIRCam indicates a disk-like morphology and spiral structure with an elongated central feature aligned between the spiral arms, most conspicuously visible in the F200W, F277W, and F356W filters. Three independent methods all support the presence of a bar: visual inspection of residuals from S\'ersic-profile fitting shows a linear structure, isophotal ellipse-fitting displays characteristic profiles of ellipticity and position angle consistent with a bar signature, and Fourier decomposition of the galaxy produces a central bisymmetric mode above a threshold strength calibrated to $z=1-3$ barred spirals. Leveraging archival Keck/MOSDEF spectroscopy overlapping with a blue clump on the edge of the galaxy, a robust redshift is inferred, with photometric constraints indicating that this structure lies at the same redshift as the main spiral. This spectroscopic evidence, placing an unlensed barred spiral at $z>3$ supports the idea that galaxies with rotationally supported disks and disk-halo properties that are conducive to bar formation were already in place within 2 Gyr after the Big Bang.

astro-ph.GA

Model Merging to Evolution: Parameter Space Exploration for Expert Models

Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby reducing computational resource requirements. However, existing methods operate within the convex combination space of expert models, failing to explore high-performance regions outside this space. This paper proposes the MERGEvolve framework, which unifies model merging and evolution within an evolution strategy by treating the merged model as the initialization for evolutionary exploration of the parameter space. During the merging phase, expert models act as deterministic sources to build a strong initial point. The evolution phase then explores the parameter space using random noise. Theoretical analysis shows that MERGEvolve explores regions outside the convex combination space. Extensive experiments on single-task and multi-task benchmarks demonstrate that MERGEvolve consistently achieves performance competitive with advanced model merging baselines. Ablation studies confirm that a high-quality initial point is critical for efficient exploration of the parameter space.

cs.NE

Exploring the Relationship Between Bars, Star Formation Activity, and Host Galaxy Properties from $\mathbf{z \sim 0}$ to $\mathbf{z \sim 2}$

We present the most comprehensive study to date of the relationship between bars, star formation, and galaxy properties from $z \sim$ 0 to $z \sim$ 2. We use a mass-complete sample of 1,171 galaxies from the JWST CEERS survey with $M_\star > 10^{10} M_\odot$ and repeat the analysis using COSMOS-Web data. Our results are: 1) At high redshift ($z \sim$ $1-2$) barred galaxies tend to have high sSFRs and low S\'ersic indices ($n \leq 2$), while at low redshifts barred galaxies emerge with both low sSFR and higher $n$, suggestive of quiescent galaxies with bulges. 2) The fractional contribution of barred quiescent galaxies to the bar fraction rises steeply from $z \sim$ 2 to $z \sim$ 0, while that of barred actively star-forming galaxies falls. 3) The fraction of quiescent galaxies that are barred rises steeply over the last 10 Gyr. 4) Our empirical results show good agreement with the TNG50-1 simulations for bars with $a_{\mathrm{bar}}$ $>$ 1.5 kpc. Our results allow for the possibility that bar-driven secular evolution may lead to quiescence and/or that bars are more likely to persist and grow in gas-poor, quiescent galaxies. The steep rise in the quiescent bar fraction over 10 Gyr may represent an evolutionary sequence whereby gas-rich disks at high redshift first develop short, dynamically young bars and over time, repeated bar-driven gas inflows lead to central starbursts and declining gas fractions that strengthen the bar as the galaxy transitions toward quiescence.

astro-ph.GA

PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-resident skills internalized through experience. PEAM pairs a slow deliberative LLM for open-ended reasoning with a fast parametric module for reflexive execution of consolidated skills. The fast module is a multimodal Mixture-of-Experts LoRA architecture with per-category physically isolated adapters, enabling parameter-level continual learning without catastrophic forgetting. We treat failure as a first-class training signal: failure--correction trajectory pairs are internalized through a joint behavioral-cloning and contrastive objective, so the agent learns not only what succeeds but also how corrected actions differ from failed ones. To govern consolidation, PEAM introduces a parameterization-worthiness score for deciding which experience should be internalized, and a scale-free self-triggered consolidation mechanism for deciding when to internalize without task-specific hand-tuned thresholds, making the agent self-evolving as the trigger transfers across task distributions without re-tuning. Experiments in Minecraft show that PEAM improves long-horizon task performance, mitigates forgetting on previously consolidated skills, and improves parametric-versus-retrieval efficiency over retrieval-based embodied agents and parametric memory variants.

cs.AI

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instructions remain target-referential: they describe, constrain, or imply the region to be segmented. However, in real-world embodied interaction, human instructions are often at the intent-level, which includes the desired outcome without naming the region that enables it. To bridge this gap, we introduce SegWorld, where the model reasons about the scene through a multi-level visual chain-of-thought (CoT) before committing to a mask. Before receiving any instructions, it proactively observes the scene, describing visible objects and inferring plausible events they may support. Given an instruction, it continues the chain: from the object relevant to the intent, through the action that satisfies it, to the physical interaction site, the object part that affords the action. We formalize SegWorld as probabilistic inference, in which proactive observation supplies a linguistic scene context that improves mask prediction when instructions are given at the level of intent. We construct an intent-to-part benchmark for evaluating affordance-bearing part segmentation from high-level goals. Experiments show SegWorld matches instruction-driven baselines on target-referential instructions and improves substantially on intent-level ones.

cs.CV

Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more holistic understanding that jointly captures procedural context, semantic reasoning, and precise visual grounding. However, existing approaches typically address these components in isolation, leading to fragmented representations and limited semantic consistency. To address this limitation, we propose SurgMLLM, a unified surgical scene understanding framework that bridges high-level reasoning and low-level visual grounding within a single model. Given surgical videos, SurgMLLM fine-tunes a multimodal large language model (MLLM) to support structured interpretability reasoning, which is used to jointly model phases, instrument-verb-target (IVT) triplets, and triplet-entity segmentation tokens. These tokens are then temporally aggregated and serve as prompts for a segmentation network, enabling accurate pixel-wise grounding of triplet instruments and targets. The entire framework is trained end-to-end with a unified objective that couples language-based reasoning supervision with visual grounding losses, promoting coherent cross-task learning and clinically consistent scene representations. To facilitate unified evaluation, we introduce CholecT45-Scene, extending CholecT45 dataset with 64,299 frames of pixel-level mask annotations for instruments and targets, aligned with existing triplet labels. Extensive experiments show that SurgMLLM significantly advances surgical scene understanding, improving the primary triplet recognition metric AP_IVT from 40.7% to 46.0% and consistently outperforming prior methods in phase recognition and segmentation. These results highlight the effectiveness of unified reasoning-and-grounding for reliable, context-aware surgical assistance.

cs.CV

Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical complementary information, but existing high-fidelity fusion methods suffer from prohibitive latency, rendering them impractical for real-time edge deployment. To address this, we propose FusionProxy, a real-time image fusion module designed as a fully independent, plug-and-play component with diffusion level quality. FusionProxy exploits two complementary statistics of a teacher sample ensemble: per-pixel variance in raw image space, used to weight pixel-level supervision, and per-pixel variance inside frozen foundation backbones, used to route feature-level alignment spatially. Once trained, FusionProxy can be directly integrated into any visual perception system without joint optimization. Extensive experiments demonstrate that our method achieves superior performance on static recognition tasks and significantly enhances robustness in dynamic tasks, including closed-loop autonomous driving. Crucially, FusionProxy achieves real-time inference speeds on diverse platforms, from high-end GPUs to commodity hardware, providing a flexible and generalizable solution for all-day perception.

cs.CV

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance perception. However, existing evaluation approaches tend to over-optimize both hand-crafted no-reference statistics and full-reference metrics that treat the source images as pseudo ground truths. Recent IVIF reward-modelling efforts learn from human ratings but use scalar regression on aggregated scores, neither leveraging the reasoning of Multimodal Large Language Models (MLLMs) nor encoding per-image perceptual ambiguity in their supervision, but naively introducing MLLMs with discrete one-hot supervision likewise collapses fused images of similar quality into different rating levels. To address this, we introduce FuScore, which utilizes an MLLM to mimic human visual perception by producing continuous quality score, rather than discrete level predictions, enabling fine-grained discrimination among fused images of similar quality. We exploit the agreement among four IVIF-specific sub-dimensions to construct a per-image soft label whose sharpness reflects how consensual the overall judgment is. We further introduce a tripartite objective combining per-image distributional supervision, within-source-pair Thurstone fidelity for method-level ordering, and cross-source-pair Thurstone fidelity for scene-level ordering across scenes. Extensive experiments demonstrate that FuScore achieves state-of-the-art correlation with human visual preferences.

cs.CV

Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis

The evolution of artificial intelligence (AI) has rendered the boundary between humanity and computational machinery increasingly ambiguous. In the presence of more interwoven relationships within human-machine symbiosis, the very notion of AI-generated information becomes difficult to define, as such information arises not from either humans or machines in isolation, but from their mutual shaping. At times AI acts in place of the human, automating the task; at others it extends what the human can do, augmenting their capability. Therefore, a more pertinent question lies not merely in whether AI has participated, but in how it has participated. In general, the role assumed by AI is often specified, either implicitly or explicitly, in the input prompt, yet becomes less apparent or altogether unobservable when the generated content alone is available. Once detached from the dialogue context, the functional role may no longer be traceable. This study considers the problem of tracing the functional role played by AI in natural language generation. A methodology is proposed to infer the latent role specified by the prompt, embed this role into the content during the probabilistic generation process and subsequently recover the nature of AI participation from the resulting text. Experimentation is conducted under a representative scenario in which AI acts either as an assistive agent that edits human-written content or as a creative agent that generates new content from a brief concept. The experimental results support the validity of the proposed methodology in terms of discrimination between roles, robustness against perturbations and preservation of linguistic quality. We envision that this study may contribute to future research on the ethics of AI with regard to whether AI has been used fairly, transparently and appropriately.

cs.AI

LumiVideo: An Intelligent Agentic System for Video Color Grading

Video color grading is a critical post-production process that transforms flat, log-encoded raw footage into emotionally resonant cinematic visuals. Existing automated methods act as static, black-box executors that directly output edited pixels, lacking both interpretability and the iterative control required by professionals. We introduce LumiVideo, an agentic system that mimics the cognitive workflow of professional colorists through four stages: Perception, Reasoning, Execution, and Reflection. Given only raw log video, LumiVideo autonomously produces a cinematic base grade by analyzing the scene's physical lighting and semantic content. Its Reasoning engine synergizes an LLM's internalized cinematic knowledge with a Retrieval-Augmented Generation (RAG) framework via a Tree of Thoughts (ToT) search to navigate the non-linear color parameter space. Rather than generating pixels, the system compiles the deduced parameters into industry-standard ASC-CDL configurations and a globally consistent 3D LUT, analytically guaranteeing temporal consistency. An optional Reflection loop then allows creators to refine the result via natural language feedback. We further introduce LumiGrade, the first log-encoded video benchmark for evaluating automated grading. Experiments show that LumiVideo approaches human expert quality in fully automatic mode while enabling precise iterative control when directed.

cs.CV

All-Mem: Agentic Lifelong Memory via Dynamic Topology Evolution

Lifelong interactive agents are expected to assist users over months or years, which requires continually writing long term memories while retrieving the right evidence for each new query under fixed context and latency budgets. Existing memory systems often degrade as histories grow, yielding redundant, outdated, or noisy retrieved contexts. We present \textbf{All-Mem}, an online/offline lifelong memory framework that maintains a topology structured memory bank via explicit, non destructive consolidation, avoiding the irreversible information loss typical of summarization based compression. In online operation, it anchors retrieval on a bounded visible surface to keep coarse search cost bounded. Periodically offline, an LLM diagnoser proposes confidence scored topology edits executed with gating using three operators: Split, Merge, and Update, while preserving immutable evidence for traceability. At query time, typed links enable hop bounded, budgeted expansion from active anchors to archived evidence when needed. Experiments on \textbf{LoCoMo} and \textbf{LongMemEval-s} show improved retrieval and QA over representative baselines. The code is available at https://github.com/LvCan926/All-Mem.

cs.IR

Quantum criticality in open quantum systems from the purification perspective

Open quantum systems host mixed-state phases that go beyond the symmetry-protected topological and spontaneous symmetry-breaking paradigms established for closed, pure-state systems. Developing a unified and physically transparent classification of such phases remains a central challenge. In this work, we introduce a purification-based framework that systematically characterizes all mixed-state phases in one-dimensional systems with $\mathbb{Z}_2^{\sigma} \times \mathbb{Z}_2^{\tau}$ symmetry. By introducing an ancillary $\kappa$ chain and employing decorated domain-wall constructions, we derive eight purified fixed-point Hamiltonians labeled by topological indices $(\mu_{\sigma\tau},\mu_{\tau\kappa},\mu_{\kappa\sigma}) \in \{\pm1\}^3$. Tracing out the ancilla recovers the full structure of mixed-state phases, including symmetric, strong-to-weak spontaneous symmetry breaking, average symmetry-protected topological phases, and their nontrivial combinations. Interpolations between the eight fixed points naturally define a three-dimensional phase diagram with a cube geometry. The edges correspond to elementary transitions associated with single topological indices, while the faces host intermediate phases arising from competing domain-wall decorations. Along the edges, we identify a class of critical behavior that connects distinct strong-to-weak symmetry-breaking patterns associated with distinct strong subgroups, highlighting a mechanism unique to mixed-state settings. Large-scale tensor-network simulations reveal a rich phase structure, including pyramid-shaped symmetry-breaking regions and a fully symmetry-broken phase at the cube center. Overall, our purification approach provides a geometrically transparent and physically complete classification of mixed-state phases, unified with a single $\mathbb{Z}_2^{\sigma} \times \mathbb{Z}_2^{\tau} \times \mathbb{Z}_2^{\kappa}$ model.

quant-ph