SearcharxivSearch

arXiv subjects

Qin Guo

Publications and source records attributed to Qin Guo.

At least 19 recordsLinked to original sources

TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning

Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps, routinely accumulating hundreds of thousands of context tokens per session. Existing compression strategies typically allocate retention budgets without accounting for the downstream consequences of removing individual tool outputs. Aggressive compression may therefore trigger costly tool re-invocations that offset the initial savings. We call this the compression--consequence gap. To close it, we propose TRACER, which formulates compression as a sequential per-tool decision problem. A lightweight REINFORCE policy assigns query-conditioned retention ratios using only information available at each compression event. Its consequence-aware objective jointly accounts for task success, total token consumption, and post-compression tool re-invocations. To improve credit assignment, TRACER uses a learned outcome model to compare the predicted consequences of the selected retention ratio with those of fully retaining each tool output. On held-out production queries across three compressor backends, TRACER reduces total token consumption by 29--46% relative to keeping all context while maintaining comparable or higher task success. Compared with a tool-type-conditional static policy, TRACER provides an additional 15--18% of token savings. Interventional rollouts show that the learned per-tool credit scores correlate with measured single-tool consequences. The learned policy also yields positive savings when transferred across agent backbones and compressor architectures, and reduces token consumption by 18--25% on five held-out LOCA-bench environments. These results demonstrate the value of consequence-aware, per-tool context retention for improving the efficiency of long-horizon language agents.

cs.AI

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions. In this work, we introduce UniGP, a framework built upon MMDiT, which unifies controllable generation and dense prediction through simple joint training, without the need for complex task-specific designs or losses, while preserving the backbone's versatile priors. By learning controllable generation and prediction under different conditions, our model effectively captures the joint distribution of image-geometry pairs. UniGP is capable of versatile controllable generation, dense prediction, and joint generation. Specifically, the proposed UniGP consists of DUGP and a unified dataset training strategy. The former, following the principle of Occam's razor, uses only a copied image branch of MMDiT to model dense distributions beyond RGB, while the latter integrates heterogeneous datasets into a unified training framework to jointly model generation and perception tasks. Extensive experiments demonstrate that our unified model surpasses prior unified approaches and performs on par with specialized methods. Furthermore, we demonstrate that multi-task joint training provides complementary benefits: generative priors enrich perceptual details, while perceptual learning improves structural alignment in generation.

cs.CV

WildActor: Unconstrained Identity-Preserving Video Generation

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level consistency, or produce copy-paste artifacts where subjects appear rigid due to pose locking. We present Actor-18M, a large-scale human video dataset designed to capture identity consistency under unconstrained viewpoints and environments. Actor-18M comprises 1.6M videos with 18M corresponding human images, covering both arbitrary views and canonical three-view representations. Leveraging Actor-18M, we propose WildActor, a framework for any-view conditioned human video generation. We introduce an Asymmetric Identity-Preserving Attention mechanism coupled with a Viewpoint-Adaptive Monte Carlo Sampling strategy that iteratively re-weights reference conditions by marginal utility for balanced manifold coverage. Evaluated on the proposed Actor-Bench, WildActor consistently preserves body identity under diverse shot compositions, large viewpoint transitions, and substantial motions, surpassing existing methods in these challenging settings.

cs.CV

A Secure Semantic Communication System Based on Knowledge Graph

This study proposes a novel approach to ensure the security of textual data transmission in a semantic communication system. In the proposed system, a sender transmits textual information to a receiver, while a potential eavesdropper attempts to intercept the information. At the sender side, the text is initially preprocessed, where each sentence is annotated with its corresponding topic, and subsequently extracted into a knowledge graph. To achieve the secure transmission of the knowledge graph, we propose a channel encryption scheme that integrates constellation diagonal transformation with multi-parameter weighted fractional Fourier transform (MP-WFRFT). At the receiver side, the textual data is first decrypted, and then recovered via a transformer model. Experimental results demonstrate that the proposed method reduces the probability of information compromise. The legitimate receiver achieves a Bilingual Evaluation Understudy (BLEU) score of 0.9, whereas the BLEU score of the eavesdropper remains below 0.3. Compared to the baselines, the proposed method can improve the security by up to 20%.

cs.CR

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

Although significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is difficult to generate multiple overlapping humans and animals based on keypoint controls solely. These challenges arise from two main aspects: the inherent limitations of existing controllable methods and the lack of suitable datasets. First, we design a DiT-based framework, named UniMC, to explore unifying controllable multi-class image generation. UniMC integrates instance- and keypoint-level conditions into compact tokens, incorporating attributes such as class, bounding box, and keypoint coordinates. This approach overcomes the limitations of previous methods that struggled to distinguish instances and classes due to their reliance on skeleton images as conditions. Second, we propose HAIG-2.9M, a large-scale, high-quality, and diverse dataset designed for keypoint-guided human and animal image generation. HAIG-2.9M includes 786K images with 2.9M instances. This dataset features extensive annotations such as keypoints, bounding boxes, and fine-grained captions for both humans and animals, along with rigorous manual inspection to ensure annotation accuracy. Extensive experiments demonstrate the high quality of HAIG-2.9M and the effectiveness of UniMC, particularly in heavy occlusions and multi-class scenarios.

cs.CV

VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction

Objective: Ulcerative colitis (UC), characterized by chronic inflammation with alternating remission-relapse cycles, requires precise histological healing (HH) evaluation to improve clinical outcomes. To overcome the limitations of annotation-intensive deep learning methods and suboptimal multi-instance learning (MIL) in HH prediction, we propose VIGIL, the first vision-language guided MIL framework integrating white light endoscopy (WLE) and endocytoscopy (EC). Methods:VIGIL begins with a dual-branch MIL module KS-MIL based on top-K typical frames selection and similarity metric adaptive learning to learn relationships among frame features effectively. By integrating the diagnostic report text and specially designed multi-level alignment and supervision between image-text pairs, VIGIL establishes joint image-text guidance during training to capture richer disease-related semantic information. Furthermore, VIGIL employs a multi-modal masked relation fusion (MMRF) strategy to uncover the latent diagnostic correlations of two endoscopic image representations. Results:Comprehensive experiments on a real-world clinical dataset demonstrate VIGIL's superior performance, achieving 92.69\% accuracy and 94.79\% AUC, outperforming existing state-of-the-art methods. Conclusion: The proposed VIGIL framework successfully establishes an effective vision-language guided MIL paradigm for UC HH prediction, reducing annotation burdens while improving prediction reliability. Significance: The research outcomes provide new insights for non-invasive UC diagnosis and hold theoretical significance and clinical value for advancing intelligent healthcare development.

q-bio.QM

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the \textbf{M}otion-priors \textbf{C}onditional \textbf{D}iffusion \textbf{M}odel (\textbf{MCDM}), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the \textbf{TalkingFace-Wild} dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Code, models, and datasets will be publicly available.

cs.CV

Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention Modulation

Recently, diffusion-based methods, like InstructPix2Pix (IP2P), have achieved effective instruction-based image editing, requiring only natural language instructions from the user. However, these methods often inadvertently alter unintended areas and struggle with multi-instruction editing, resulting in compromised outcomes. To address these issues, we introduce the Focus on Your Instruction (FoI), a method designed to ensure precise and harmonious editing across multiple instructions without extra training or test-time optimization. In the FoI, we primarily emphasize two aspects: (1) precisely extracting regions of interest for each instruction and (2) guiding the denoising process to concentrate within these regions of interest. For the first objective, we identify the implicit grounding capability of IP2P from the cross-attention between instruction and image, then develop an effective mask extraction method. For the second objective, we introduce a cross attention modulation module for rough isolation of target editing regions and unrelated regions. Additionally, we introduce a mask-guided disentangle sampling strategy to further ensure clear region isolation. Experimental results demonstrate that FoI surpasses existing methods in both quantitative and qualitative evaluations, especially excelling in multi-instruction editing task.

cs.CV

Neural Network-Based Histologic Remission Prediction In Ulcerative Colitis

BACKGROUND & AIMS: Histological remission (HR) is advocated and considered as a new therapeutic target in ulcerative colitis (UC). Diagnosis of histologic remission currently relies on biopsy; during this process, patients are at risk for bleeding, infection, and post-biopsy fibrosis. In addition, histologic response scoring is complex and time-consuming, and there is heterogeneity among pathologists. Endocytoscopy (EC) is a novel ultra-high magnification endoscopic technique that can provide excellent in vivo assessment of glands. Based on the EC technique, we propose a neural network model that can assess histological disease activity in UC using EC images to address the above issues. The experiment results demonstrate that the proposed method can assist patients in precise treatment and prognostic assessment. METHODS: We construct a neural network model for UC evaluation. A total of 5105 images of 154 intestinal segments from 87 patients undergoing EC treatment at a center in China between March 2022 and March 2023 are scored according to the Geboes score. Subsequently, 103 intestinal segments are used as the training set, 16 intestinal segments are used as the validation set for neural network training, and the remaining 35 intestinal segments are used as the test set to measure the model performance together with the validation set. RESULTS: By treating HR as a negative category and histologic activity as a positive category, the proposed neural network model can achieve an accuracy of 0.9, a specificity of 0.95, a sensitivity of 0.75, and an area under the curve (AUC) of 0.81. CONCLUSION: We develop a specific neural network model that can distinguish histologic remission/activity in EC images of UC, which helps to accelerate clinical histological diagnosis. keywords: ulcerative colitis; Endocytoscopy; Geboes score; neural network.

cs.CV

ChatFace: Chat-Guided Real Face Editing via Diffusion Latent Space Manipulation

Editing real facial images is a crucial task in computer vision with significant demand in various real-world applications. While GAN-based methods have showed potential in manipulating images especially when combined with CLIP, these methods are limited in their ability to reconstruct real images due to challenging GAN inversion capability. Despite the successful image reconstruction achieved by diffusion-based methods, there are still challenges in effectively manipulating fine-gained facial attributes with textual instructions.To address these issues and facilitate convenient manipulation of real facial images, we propose a novel approach that conduct text-driven image editing in the semantic latent space of diffusion model. By aligning the temporal feature of the diffusion model with the semantic condition at generative process, we introduce a stable manipulation strategy, which perform precise zero-shot manipulation effectively. Furthermore, we develop an interactive system named ChatFace, which combines the zero-shot reasoning ability of large language models to perform efficient manipulations in diffusion semantic latent space. This system enables users to perform complex multi-attribute manipulations through dialogue, opening up new possibilities for interactive image editing. Extensive experiments confirmed that our approach outperforms previous methods and enables precise editing of real facial images, making it a promising candidate for real-world applications. Project page: https://dongxuyue.github.io/chatface/

cs.CV

Operator transpose within normal ordering and its applications for quantifying entanglement

Partial transpose is an important operation for quantifying the entanglement, here we study the (partial) transpose of any single (two-mode) operators. Using the Fock-basis expansion, it is found that the transposed operator of an arbitrary operator can be obtained by replacement of a^{†}(a) by a(a^{†}) instead of c-number within normal ordering form. The transpose of displacement operator and Wigner operator are studied, from which the relation of Wigner function, characteristics function and average values such as covariance matrix are constructed between density operator and transposed density operator. These observations can be further extended to multi-mode cases. As applications, the partial transpose of two-mode squeezed operator and the entanglement of two-mode squeezed vacuum through a laser channel are considered.

quant-ph

Photon blockade in a bi-mode nonlinear nano-cavity embedded with a quantum-dot

We study the interaction between a quantum-dot and a bi-mode micro/nano-optical cavity composed of second-order nonlinear materials. Compared with the Jaynes-Cummings (J-C) model, except for a coherent weak driving field, a strong pump light illuminates the two-mode optical cavity. Analytical results indicate that the model exhibits abundant non-classical optical phenomena, such as conventional photon blockade induced by the nonlinear interaction between polaritons. It constitutes unconventional photon blockade induced by quantum interference due to parametric driving. We compare the photon statistical properties and average photon number of the proposed model, J-C model, and double-mode driven optical cavity under the same parameters and the proposed model can obtain stronger antibunching photons and higher average photon number.

quant-ph

Conventional and unconventional photon blockade effects in an atom-cavity system

A two-level system interacting with a cavity field is an important model for investigating the photon blockade (PB) effect. Most work on this topic has been based on the assumption that the atomic transition frequency is resonant with the fundamental mode frequency of the cavity. We relax this constraint and reexamine PB in a more general atom--cavity system with arbitrary atomic and cavity detunings from a driving field. The results show that when the signs of the atomic and cavity detunings are the same, PB occurs only in the strong-coupling regime, but for opposite signs of the atomic and cavity detunings, strong photon antibunching is observed in both the weak- and strong-coupling regimes and a better PB effect is achieved compared with the case when the signs are the same. More interestingly, we find that this PB arises from quantum interference for both weak and strong nonlinearities. These results deepen our understanding of the underlying mechanism of PB and may be help in the construction of single-photon sources with higher purity and better flexibility using atom--cavity systems.

quant-ph

Hierarchically porous Ni monolith@branch-structured NiCo2O4 for high energy density supercapacitors

NiCo2O4 of varying nanostrucutures ranging from nanowires, nanoplates to nano-plates@nanowires were successfully grown on microporous (MP) Ni foams via one-step hydrothermal process. The investigation of electrochemical capacitance favors Ni-Co2O4 of nanoplates@nanowires microstructures which possesses specific capacitance of 1380.3 F/g and 1033F/g at 5A/g and 50A/g respectively and 86.7% capacitance re-tention after 5000 cycles at 30A/g. The relationship between morphology and specific capacitance was further explored by the model of surface roughness factor (RF), which is indicative of the active electrode-electrolyte interface areas. The RF of porous Ni@NiCo2O4 was remarkably improved by employing hierarchically porous (HP) Ni monoliths as substrates, which illustrates the model of high energy density (12.6 F/cm2) electrodes for super-capacitors.

cond-mat.mtrl-sci

Intermediate Coherent-entangled State Representation: Generation and its applications

By combining the beam splitter and the Fresnel transform, a protocol is proposed to generate a new entangled state representation, called the intermediate coherent-entangled state (ICES) representation. The properties, such as eigenvalue equation, completeness relation and orthogonal relation, are investigated. The conjugate state representation of the ICES and the Schmidt decomposing of the ICES are also discussed. As applications, a new squeezing operator and some operator identities by using the ICES are obtained.

quant-ph

M Times Photon Subtraction-Addition Coherent Superposition Operated Odd-Schrődinger-cat State: Nonclassicality and Decoherence

We introduce a new non-Gaussian state, generated by m times coherent superposition operation $a\cos θ+a^{\dagger }e^{iφ}\sin θ$ (MCSO) on odd-Schrodinger-cat state (OSCS). Its normalized constant is turned out to be related with the Hermite polynomial. We further investigate the nonclassical properties of the MCSO-OSCS through Mandel's Q-parameter, quadrature squeezing, the photocount distribution and Wigner function (WF). It is shown that the nonclassicality of the MCSO-OSCS is influenced by the number of times (m) of coherent superpositon operation, the angle $θ$ and the amplitude of the coherent state (|$α_{0}$|). Especially the volume of negative region of WF increases with the increment of parameters m, $θ$ and $α_{0}$. We also investigate the decoherence of the MCSO-OSCS in terms of the fadeaway of the negativity of WF in a thermal environment.

quant-ph

Comments on "Asking Photons Where They Have Been"

By using the nonuniform discrete Fourier transform on the center positions of the symmetric intensity distribution of the output beam, we recover the vibration information of two mirrors, which is lost in the analysis of Danan \emph{et al.} in the work [Phys. Rev. Lett. \textbf{111}, 240402 (2013), arXiv:1304.7469]. We believe a photon always follows continuous trajectories, and a photon has to enter an interferometer at first and then leave it, if this photon has ever been inside the interferometer.

quant-ph