SearcharxivSearch

arXiv subjects

Qiang Gao

Publications and source records attributed to Qiang Gao.

At least 19 recordsLinked to original sources

Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

Traffic sign detection faces a long-tailed data distribution. Many rare signs matter as much as common ones from a regulatory standpoint, yet they have very few samples. Generative data augmentation is one way out. General-purpose inpainting models, however, distort digits, deform geometry and perspective, and shift colours when applied directly to sign regions. We trace this to a single gap: the conditioning signal is too abstract for the physical composition of a sign. We propose a structured-prior-guided diffusion inpainting framework with physical consistency. It injects the semantic, appearance and geometric priors of a sign through three orthogonal pathways: a JSON-formatted text prompt, a front-view vector template rendered with measured dominant colours (via IP-Adapter), and an affine-aligned vector template (via ControlNet). Two physical consistency losses constrain colour with a CIELAB chromaticity $L_1$ term and edge structure with a Sobel gradient term. We train by self-supervised reconstruction on a large set of images collected in-house at AMAP, then evaluate zero-shot on the public TT100K-2021 dataset, a different source. Our method uses a Stable Diffusion 1.5 backbone of about 1.4B parameters. It beats seven representative competitors on every metric of reconstruction fidelity, physical consistency and semantic controllability. Its OCR exact-match rate reaches 91.1\%, against 44.2\% for the 12B industrial model FLUX.1 Fill [dev], and it needs only $1/14$ of that model's inference time. Leave-one-out ablations confirm that each of the three prior pathways and both loss terms contribute on their own. In downstream detection, the synthetic data raises the group-pooled AP50 of rare classes by $1.23\times$ to $7.40\times$ over a real-data-only baseline. Code and pre-trained models are available at https://github.com/52hz-whale/TrafficSignInpaint.

cs.CV

A quantum geometric mechanism for chiral domain wall metastability: Application to twisted transition-metal dichalcogenides

Band topology can have an imprint on the excitations of a ferromagnet. A known example is quantum Hall ferromagnets and their lattice analogs; when both flavors have the same Chern number $C$, a smooth skyrmion texture binds charge $-eC$ per unit winding. Here, we consider instead the case of conjugate Chern bands related by time-reversal. We show that, despite the vanishing net charge response, a smooth texture can be associated with a dipole response---a domain wall (DW) with an in-plane winding along its length can bind a nonzero dipole density transverse to the wall. The strength of this dipole density is controlled by a dimensionless coefficient $c_G$. Although not quantized, the geometric dipole coefficient $c_G$ is a moment of the second Chern form of the occupied projector in mixed (momentum and order-parameter) space and is generally nonzero. The dipole-decorated DW can thus become metastable at a finite radius due to the competition between dipolar repulsion and the usual surface tension, even at a finite Zeeman field. In a realistic model of twisted MoTe$_2$, we find that $c_G$ drops sharply across a transition within the valley-polarized (VP) phase from a $C=1$ to a $C=0$ ferromagnet. This naturally explains recent pump-probe experiments~\cite{exp} at hole filling $\nu=1$, in which a long-lived excitation survives reverse fields far exceeding the saturation field but disappears at an intermediate displacement field despite only weak changes in conventional magnetic diagnostics. Metastable spin textures thus serve as a sensitive probe of band quantum geometry, and as an intrinsic bottleneck for fast optical control of moir\'e ferromagnets in Chern-conjugate bands.

cond-mat.str-el

Observation of metastable chiral domain walls in a topological magnet

The interplay between topology and correlation can give rise to exotic collective excitations. The integer and fractional quantum anomalous Hall (QAH) magnets recently discovered in two-dimensional (2D) flatband systems are predicted to host spin excitations distinct from those in conventional magnets. Experimentally, nevertheless, these new excitations remain largely unexplored. Here we investigate spin-valley excitations in a twisted MoTe2 moir\'e superlattice using resonant ultrafast pump-probe spectroscopy. We observe a metastable spin-valley excitation in the QAH magnet below T ~ 3.7 K that survives reverse magnetic field several times larger than the saturation field. The behavior of this excitation is sharply distinct from ordinary domain walls and magnons, indicating a new type of spin-valley textures unique to topological magnets. We propose that these textures are chiral domain walls with an in-plane winding of the pseudospin order parameter along the domain wall. Their metastability arises from the interplay between the topological winding in real space and the quantum geometry of the parent bands in momentum space through a universal mechanism. These chiral domain walls govern the nonequilibrium dynamics of QAH magnets and may play a central role in their stability. Our study highlights intrinsic quantum geometry effects on spin excitations in topological magnets; and provides key insights into the fundamental mechanism limiting stability of topological protection.

cond-mat.mes-hall

KAT-Coder-V2.5 Technical Report

We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator. Its capability is bottlenecked less by model scale than by the scarcity of reproducible environments, verifiable rewards, and high-value trajectories, which we address with an end-to-end agentic post-training framework. AutoBuilder reconstructs multilingual repositories into sandboxed environments with fail-to-pass and pass-to-pass verification at scale, from which we regenerate self-contained task specifications, recover near-miss trajectories, and distill supervision through process-aware filtering, while KwaiClawEnv synthesizes large-scale tool-use trajectories from executable services and real task seeds. We further scale reinforcement learning with harness randomization, a reliability-hardened sandbox, an asymmetric actor--critic PPO with hindsight-augmented value estimation, and a harness-oriented reward framework, and unify SWE, Agent-Claw, and WebCoding experts via Multi-Teacher On-Policy Distillation. Across six software-engineering and agentic benchmarks, KAT-Coder-V2.5 delivers the best agentic tool-use result on PinchBench and ranks second only to the frontier Opus 4.8 on repository-level software engineering. Our service is available at https://streamlake.com/product/kat-coder.

cs.SE

Normalized solutions of quasilinear Schr\"odinger equations in the general $L^2$-supercritical case

This paper is devoted to studying the existence of normalized solutions for the following quasilinear Schr\"odinger equation \begin{equation*} \begin{aligned} -\Delta u-u\Delta u^2 +\lambda u=h(u) \quad\mathrm{in}\ \mathbb{R}^{3}, \end{aligned} \end{equation*} where $\lambda$ appears as a Lagrange multiplier, $h$ is a $L^2$-supercritical and Sobolev subcritical nonlinearity. The solutions correspond to critical points of the energy functional subject to the $L^2$-norm constraint $\int_{\mathbb{R}^3}|u|^2dx=a^2>0$. Taking into account the Pohozaev manifold and perturbation method, we obtain the existence of ground state normalized solutions and infinitely many normalized solutions. Moreover, our results cover several relevant existing results in \cite{LZ2023}. And in the end, we get the asymptotic properties of energy as $a$ tends to $+\infty$ and $a$ tends to $0^+$.

math.AP

Microsopic Theory of Spin Polarons in Chern Ferromagnets

We develop a microscopic theory of charged excitations in an SU(2) Chern ferromagnet and obtain closed-form wavefunctions for a hierarchy of charge-$e$ spin polaron states binding an arbitrary number of spin flips. In an ideal Chern-$1$ band with a normal-ordered contact interaction, we show that these polarons are exact eigenstates of the Hamiltonian with the same energy as single-hole excitations. Away from this ideal limit, we promote these states to a variational family by introducing a single size parameter and a geometry-informed single-particle dressing. Our momentum-space wavefunctions admit two equivalent representations: a ratio of Jastrow factors of Weierstrass functions of relative momenta or an antisymmetrized geminal product of particle-hole wavefunctions. The latter enables efficient evaluation of overlaps and expectation values for large system sizes and many spin flips. Benchmarking in the lowest Landau level, the single-spin-flip ansatz achieves $\gtrsim 99\%$ overlap with exact diagonalization and accurately captures binding energies, while the multi-spin-flip energies interpolate smoothly toward the large-texture (skyrmion) regime. For Chern bands with tunable quantum geometry, we find that interaction-generated single particle dispersion quickly destabilizes the spin polarons once quantum geometry becomes sufficiently non-uniform. When such dispersion is suppressed, however, the bound states persist deeper into the non-uniform regime, with the binding energy slowly decreasing and the bound state becoming larger as the quantum geometry becomes more concentrated. Our results provide a microscopic foundation for analyzing doped Chern ferromagnets in moir\'e platforms and lay the groundwork for variational wavefunctions of multi-polaron excitations and phases.

cond-mat.str-el

Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding

Large language models (LLMs) often produce content that contradicts or overlooks information provided in the input context, a phenomenon known as faithfulness hallucination. In this paper, we propose Context-Fidelity Boosting (CFB), a lightweight and general decoding-time framework that reduces such hallucinations by increasing the generation probability of source-supported tokens. Motivated by logit-shaping principles from watermarking techniques, CFB applies additive token-level logit adjustments based on a token's degree of support from the input context. Specifically, we develop three boosting strategies: static boosting, which applies a fixed bias to source-supported tokens; context-aware boosting, which scales this bias using the divergence between next-token distributions with and without context; and token-aware boosting, which further redistributes the adaptive bias according to local relevance estimated from source-position attention and source-scoped semantic similarity. CFB requires no retraining or architectural changes, making it compatible with a wide range of LLMs. Experiments on summarization and question answering tasks across multiple open-source LLMs show that CFB consistently improves faithfulness metrics with minimal generation overhead. Our implementation is fully open-sourced.

cs.CL

Generalizable CT-Free PET Attenuation and Scatter Correction for Pediatric Patients

Computed tomography (CT)-based attenuation and scatter correction improves quantitative PET but adds radiation exposure that is particularly undesirable in pediatric imaging. Existing CT-free methods are commonly trained in homogeneous settings and often degrade under scanner or radiotracer shifts, which limits their clinical utility. We propose the Generalizable PET Correction Network (GPCN), a dual-domain network for domain-robust CT-free PET attenuation and scatter correction. GPCN combines a multi-band contextual refinement module, which models pediatric anatomical variability through wavelet-based multiscale decomposition and long-range spatial context modeling, with a frequency-aware spectral decoupling module, which performs coordinate-conditioned amplitude/phase refinement in the Fourier domain. By synergizing multi-band spatial contextual modeling with asymmetric frequency-spectrum decoupling, the network explicitly separates invariant topological structures from domain-specific noise, thereby achieving precise quantitative recovery of both anatomical organs and focal lesions. This design aims to separate anatomy-dominant structures from domain-sensitive spectral residuals and to improve robustness across heterogeneous imaging conditions. We train and evaluate the method on 1085 pediatric whole-body PET scans acquired with two scanners and five radiotracers. In both joint training and zero-shot cross-domain evaluation, GPCN outperforms representative baselines and maintains stable quantitative accuracy on unseen scanner-tracer combinations. The method is further supported by ablation, region-wise quantitative analysis, and downstream segmentation experiments. In our cohort, the CT component of the conventional protocol corresponded to an average effective dose of 10.8 mSv, indicating the potential clinical value of reliable CT-free correction for pediatric PET.

eess.IV

SemanticAgent: A Semantics-Aware Framework for Text-to-SQL Data Synthesis

Existing text-to-SQL synthesis pipelines still conflate executability with semantic validity: syntactic checks and execution-based validation can retain queries that execute successfully while violating database semantics. To address these limitations, we propose SemanticAgent, a semantic-aware synthesis framework. SemanticAgent organizes synthesis around three specialized modules: an analyzer, a synthesizer, and a verifier. Through a three-stage protocol of semantic analysis, stepwise synthesis, and diagnostic refinement, SemanticAgent transforms execution-based validation alone into a traceable reasoning process. Our framework generates synthetic data that consistently outperforms prior synthesis methods under semantic-quality evaluation, leading to stronger downstream fine-tuning performance, especially on semantically demanding benchmarks.

cs.AI

The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview

This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze recent advances in the field. To reflect the evolving objectives of image super-resolution, the challenge includes two tracks: (1) a restoration track, which emphasizes pixel-wise fidelity and ranks submissions based on PSNR; and (2) a perceptual track, which focuses on visual realism and evaluates results using a perceptual score. A total of 194 participants registered for the challenge, with 31 teams submitting valid entries. This report summarizes the challenge design, datasets, evaluation protocol, main results, and methods of participating teams. The challenge provides a unified benchmark and offers insights into current progress and future directions in image super-resolution.

cs.CV

TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation

Medical image segmentation remains challenging due to limited fine-grained annotations, complex anatomical structures, and image degradation from noise, low contrast, or illumination variation. We propose TAMISeg, a text-guided segmentation framework that incorporates clinical language prompts and semantic distillation as auxiliary semantic cues to enhance visual understanding and reduce reliance on pixel-level fine-grained annotations. TAMISeg integrates three core components: a consistency-aware encoder pretrained with strong perturbations for robust feature extraction, a semantic encoder distillation module with supervision from a frozen DINOv3 teacher to enhance semantic discriminability, and a scale-adaptive decoder that segments anatomical structures across different spatial scales. Experiments on the Kvasir-SEG, MosMedData+, and QaTa-COV19 datasets demonstrate that TAMISeg consistently outperforms existing uni-modal and multi-modal methods in both qualitative and quantitative evaluations. Code will be made publicly available at https://github.com/qczggaoqiang/TAMISeg.

cs.CV

SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction

Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching. A shared set-aware encoder and conditional DiT backbone learn latent transport, making endpoint supervision directly delta-aligned without an auxiliary delta objective. Across genetic, chemical, developmental and immune perturbations, SCALE recovered gene-expression changes, response directions and population structure. In CRISPR data with dominant cell-line effects, SCALE outperformed competing methods across seven metrics and maintained separation among gene-target representations rather than collapsing them into a shared region. SCALE further prioritized cytokines predicted to produce distinct immune activation and inflammatory responses. Experiments using matched PBMC samples from three donors confirmed these predicted differences. Together, SCALE enables perturbation-specific prediction from unpaired populations and supports experimental prioritization.

cs.LG

Millimeter-Scale, Atomically Controlled 2D Topological Insulators Revealed by Multimodal Spectroscopy

Quantum spin Hall insulators, or synonymously known as 2D topological insulators, are crucial 2D systems hosting topologically protected edge states. The working temperature of this topological quantum phase is dictated by the inverted bandgap. However, the previously identified large-gap 2D topological insulators are either extremely chemically unstable, or cannot be made with atomistic precision over macroscopic scales. Here, we establish two-quintuple-layer Bi2Te3 and MnBi2Te4/Bi2Te3 heterostructures as atomically controlled, millimeter-scale 2D topological insulators, enabled by precision layer-by-layer growth that yields a carpet-like morphology extending coherently over macroscopic distances. This carpet-like growth mode renders the films amenable to mechanical exfoliation and subsequent wet or dry transfer. Multimodal spectroscopies and microscopies reveal the integer-layer tuned electronic structure of (Bi2Te3)n with excellent agreement to theory. Photon-energy-dependent photoemission and time-resolved photoemission identify band inversion and band dynamics, respectively, while scanning tunneling spectroscopy resolves topological edge states, characteristic of the 2D topological insulator phase. Thickness- and photon-energy-dependent photoemission further validates MnBi2Te4/Bi2Te3 as a robust 2D topological insulator. The large inverted gaps of ~100 meV in (Bi2Te3)2 and ~150 meV in MnBi2Te4/Bi2Te3 suggest operation near ambient temperature. These results define a scalable materials platform for next-generation, low-loss quantum and energy-efficient devices.

cond-mat.mtrl-sci

Scaling DPPs for RAG: Density Meets Diversity

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding generation in external knowledge, yielding relevance responses that are aligned with factual evidence and evolving corpora. Standard RAG pipelines construct context through relevance ranking, performing point-wise scoring between the user query and each corpora chunk. This formulation, however, ignores interactions among retrieved candidates, leading to redundant contexts that dilute density and fail to surface complementary evidence. We argue that effective retrieval should optimize jointly for both density and diversity, ensuring the grounding evidence that is dense in information yet diverse in coverage. In this study, we propose ScalDPP, a diversity-aware retrieval mechanism for RAG that incorporates Determinantal Point Processes (DPPs) through a lightweight P-Adapter, enabling scalable modeling of inter-chunk dependencies and complementary context selection. In addition, we develop a novel set-level objective, Diverse Margin Loss (DML), that enforces ground-truth complementary evidence chains to dominate any equally sized redundant alternatives under DPP geometry. Experimental results demonstrate the superiority of ScalDPP, substantiating our core statement in practice.

cs.LG

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

Hate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular and ambiguous forms to evade censorship, resulting in substantial semantic drift and rendering previously trained models ineffective. Test-Time Adaptation (TTA) offers a solution by adapting models during inference to narrow the cross-domain gap, while conventional TTA methods target mild distribution shifts and struggle with the severe semantic drift in HVD. To tackle these challenges, we propose SCANNER, the first TTA framework tailored for HVD. Motivated by the insight that, despite the evolving nature of hateful manifestations, their underlying cores remain largely invariant (i.e., targeting is still based on characteristics like gender, race, etc), we leverage these stable cores as a bridge to connect the source and target domains. Specifically, SCANNER initially reveals the stable cores from the ambiguous layout in evolving hateful content via a principled centroid-guided alignment mechanism. To alleviate the impact of outlier-like samples that are weakly correlated with centroids during the alignment process, SCANNER enhances the prior by incorporating a sample-level adaptive centroid alignment strategy, promoting more stable adaptation. Furthermore, to mitigate semantic collapse from overly uniform outputs within clusters, SCANNER introduces an intra-cluster diversity regularization that encourages the cluster-wise semantic richness. Experiments show that SCANNER outperforms all baselines, with an average gain of 4.69% in Macro-F1 over the best.

cs.CV

M-RAG: Semantic Key-Value Indexing for Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) turns external documents into evidence for large language models. In practice, this is also a data access problem: a system must decide what to index, what to retrieve, and what evidence to place in the context under a token budget. Most RAG pipelines use text chunks for both lookup and generation. This couples two different objectives. Retrieval benefits from compact and discriminative records, while generation needs contextual and faithful evidence. As a result, small chunks may fragment answer-bearing information, whereas large chunks may introduce noise and waste the context budget. We propose M-RAG, a semantic key-value indexing layer for budget-constrained RAG query processing. M-RAG extracts meta-markers from complete documents, where each record contains a retrieval key, an information value, and provenance pointers. Online retrieval operates over the key field, which can be searched by dense vector retrieval or sparse lexical retrieval; the paired values are returned as generation payloads and assembled under the token budget. Provenance pointers further support coverage validation and position-aware context ordering. This design separates the physical index entry from the evidence payload without changing the underlying retriever or generator. Experiments on LongBench QA subtasks show that M-RAG achieves competitive or better accuracy than representative chunk-based baselines, especially under tight token budgets. Further analyses show high document coverage, stronger robustness under expanding candidate corpora, and lower online retrieval latency. These results suggest that semantic key-value indexing is a practical access method for RAG workloads.

cs.IR

From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement

The proliferation of harmful memes on online media poses significant risks to public health and stability. Existing detection methods heavily rely on large-scale labeled data for training, which necessitates substantial manual annotation efforts and limits their adaptability to the continually evolving nature of harmful content. To address these challenges, we present ALARM, the first lAbeL-free hARmful Meme detection framework powered by Large Multimodal Model (LMM) agent self-improvement. The core innovation of ALARM lies in exploiting the expressive information from "shallow" memes to iteratively enhance its ability to tackle more complex and subtle ones. ALARM consists of a novel Confidence-based Explicit Meme Identification mechanism that isolates the explicit memes from the original dataset and assigns them pseudo-labels. Besides, a new Pairwise Learning Guided Agent Self-Improvement paradigm is introduced, where the explicit memes are reorganized into contrastive pairs (positive vs. negative) to refine a learner LMM agent. This agent autonomously derives high-level detection cues from these pairs, which in turn empower the agent itself to handle complex and challenging memes effectively. Experiments on three diverse datasets demonstrate the superior performance and strong adaptability of ALARM to newly evolved memes. Notably, our method even outperforms label-driven methods. These results highlight the potential of label-free frameworks as a scalable and promising solution for adapting to novel forms and topics of harmful memes in dynamic online environments.

cs.CV

Impact of Electron Correlations on Infinite-Layer Cuprates and Nickelates

Optimization of unconventional superconductivity involves a balance of interaction strengths. Precise determination of correlation strength across different material families is therefore important. Here, we present a combined X-ray absorption spectroscopy (XAS) and resonant inelastic X-ray scattering (RIXS) study of infinite-layer PrNiO$_2$ and SrCuO$_2$ that enables fair comparison of their interaction strengths. For both compounds, we study the orbital and magnetic excitations and extract their dispersions along high-symmetry directions. Using a single-band Hubbard model and including higher-order exchange interactions, we derive the correlation factor $U/t$ for both compounds. A key finding is that despite a smaller Coulomb repulsion $U$, PrNiO$_2$ exhibits a correlation strength that is 20% stronger than that of its isostructural cuprate counterpart SrCuO$_2$. This indicates that a moderation of the correlation strength may further optimize superconductivity in nickelates.

cond-mat.supr-con