SearcharxivSearch

arXiv subjects

Xu Jiang

Publications and source records attributed to Xu Jiang.

At least 19 recordsLinked to original sources

A scattering correction method for CT reconstruction based on the Wavelet Adaptive Material-dependent Boltzmann Transport Equation (WAM-BTE)

X-ray computed tomography (CT) is an essential imaging technology in clinical diagnosis. However, scattered photons can reduce image contrast and introduce CT value bias, which severely degrades image quality. Recently, scatter correction methods based on the Boltzmann transport equation (BTE) have attracted increasing attention due to their high physical accuracy and flexibility. Nevertheless, existing BTE-based methods usually employ single-material models, which cannot accurately describe the nonlinear energy dependence of photon interaction cross-sections in different materials. In this work, a scatter correction method based on the wavelet adaptive material-dependent BTE (WAM-BTE) is proposed. The conventional single-material model is extended to a multi-material model by introducing material-dependent scattering distributions. Furthermore, an adaptive multi-scale framework is established through wavelet decomposition. The low-frequency wavelet coefficients are used for coarse-scale scatter estimation to reduce computational complexity, while the high-frequency wavelet energy of high attenuation materials is utilized for adaptive local refinement. Theoretical analysis demonstrates that, when using the Haar basis function, the low-frequency wavelet coefficients at the $w$-th level are mathematically equivalent to block-average downsampling with a scale factor of $2^w$, except for a deterministic normalization factor. Experimental results show that the proposed WAM-BTE method achieves comparable accuracy to the Monte Carlo method while preserving the computational efficiency of coarse-scale estimation. The scatter calculation time for a single projection view is reduced to the millisecond level.

physics.med-ph

Topologically Configurable Nonlinear Vortex Generation at van der Waals Heterostructures

van der Waals (vdW) materials offer a highly tunable and efficient platform at nanoscale for nonlinear and quantum optics. Twist-stacked vdW heterostructures enable elegant control of symmetry and interlayer coupling. Prior studies mainly focus on planar twisted interfaces, while neglecting the naturally formed and mandatory defects in such vdW heterostructures. Here, we demonstrate nonlinear singular optics with topologically configurable nonlinear vortex generation at the corner singularity of vdW heterostructures. By tailoring azimuthally discrete second-harmonic phase gradients at each interface, we obtain programmable nonlinear vortex emitters with dominant target OAM components. Nonlinear OAM beams with topological charge $\ell = 1$ and $\ell = -2$ are experimentally realized, respectively. Our work unlocks the untapped potentials of nonlinear singular optics in twisted vdW materials as a reconfigurable and lithography-free platform for nonlinear structured light generation, important in quantum nonlinear optics and related fields.

physics.optics

SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction

In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark bands, and structural distortions that compromise clinical diagnosis and quantitative analysis. Existing metal artifact reduction (MAR) methods remain limited: optimization-based methods may leave residual artifacts or blur structures, regression networks may generalize poorly across scenarios, and generative models without sample-specific structural guidance and physical constraints may produce anatomically inconsistent structures. Flow Matching learns a continuous-time velocity field that deterministically transports a source distribution to a target distribution, providing a flexible MAR prior. However, standard unconditional Flow Matching does not exploit sample-specific structure, spatially nonuniform metal-induced degradation, or measured projections. To address these limitations, we propose SCMA, a structure-conditioned and metal-aware Flow Matching framework. First, a linear-interpolation-corrected image is fed into the velocity network with the intermediate state as a sample-specific structural condition, guiding inference toward artifact-free CT images while preserving anatomy. Second, time-varying spatial weights from the metal mask and its distance transform are incorporated into the Flow Matching loss to emphasize severe degradation within and around metal regions. Finally, conditional Flow Matching updates alternate with projection-consistency correction during inference, allowing reliable measurements outside metal traces to constrain predictions. Experiments on simulated and real CT data demonstrate that SCMA more effectively suppresses metal artifacts, preserves local anatomical structures, and reduces hallucination-like structures inconsistent with projection measurements than representative MAR methods.

cs.CV

OctoT2I: A Self-Evolving Agentic Text-to-Image Router

The explosive growth of Text-to-Image (T2I) models, from large-scale versions to lightweight, real-time ones, now faces diminishing marginal returns from single-model scaling. Agentic T2I methods emerged to alleviate this bottleneck by using multiple models. However, existing agentic T2I methods suffer from three key challenges: reliance on expensive handcrafted priors or human annotations, rigid single-path decision mechanisms, and a neglect of inference efficiency. To address these challenges, we introduce OctoT2I, a novel agentic framework that reformulates the T2I task as a joint optimization of generation quality and inference efficiency. OctoT2I implements a stateful, multi-round routing strategy that adaptively selects the most suitable tool based on its knowledge and memory. This strategy is enabled by a knowledge base built from scratch by our novel Self-Evolving Mechanism. This mechanism, which requires no human supervision, first autonomously defines foundational Conceptual Dimensions (eg, style, color, count) and then intelligently explores their combinations via an iterative" Propose--Solve--Evaluate--Learn"(PSEL) loop. The PSEL loop efficiently discovers each tool's capability frontier, driving continuous improvement without external guidance. Extensive experiments demonstrate that OctoT2I achieves competitive performance (0.96) on GenEval while delivering a 90.3% inference speedup and a 56.6% energy-efficiency gain over the leading baseline (Flow-GRPO), striking an exceptional balance between performance and efficiency. Code and models will be made available.

cs.AI

ContextGuard: Structured Self-Auditing for Context Learning in Language Models

Recent benchmarks reveal that despite strong reasoning capabilities, large language models (LLMs) still struggle to faithfully apply complex contextual knowledge. These failures are often not wholesale reasoning collapses: in context-rich tasks, models may follow the central reasoning path while missing peripheral, persistent, or format-sensitive requirements.

cs.CL

Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis

While LLMs excel at reasoning over prompts using static pretrained knowledge, they struggle significantly with context learning-the ability to dynamically extract, internalize, and apply new knowledge from complex, task-specific contexts. Recent evaluations on the CL-Bench reveal a critical capability gap: frontier models solve only 17.2% of context-dependent tasks on average.

cs.AI

VISD: Enhancing Video Reasoning via Structured Self-Distillation

Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token level contributions, leading to inefficient learning. Conversely, existing self distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token level supervision. To stably integrate dense supervision with RL, we introduce a direction magnitude decoupling mechanism, where rollout level advantages computed from rewards determine update direction, while structured privileged signals modulate token level update magnitudes. This design enables semantically aligned and fine grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates curriculum scheduling and EMA based teacher stabilization to support robust optimization over long video sequences. Experiments on diverse benchmarks show that VISD consistently outperforms strong baselines, improving answer accuracy and spatio temporal grounding quality. Notably, VISD reaches these gains with nearly 2x faster convergence in optimization steps, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.

cs.CV

DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment

Reinforcement learning is crucial for aligning large language models to perform complex reasoning tasks. However, current algorithms such as Group Relative Policy Optimization suffer from coarse grained, sequence level credit assignment, which severely struggles to isolate pivotal reasoning steps within long Chain of Thought generations. Furthermore, the standard unbounded Kullback Leibler divergence penalty induces severe gradient instability and mode seeking conservatism, ultimately stifling the discovery of novel reasoning trajectories. To overcome these limitations, we introduce Distribution Guided Policy Optimization, a novel critic free reinforcement learning framework that reinterprets distribution deviation as a guiding signal rather than a rigid penalty. DGPO replaces the volatile KL divergence with the bounded Hellinger distance to safely quantify token level exploration without the risk of gradient explosion. To effectively distinguish genuine reasoning breakthroughs from hallucinatory noise, we propose an entropy gating mechanism that scales this deviation by the policy`s epistemic uncertainty. By dynamically redistributing the coarse sequence-level advantage to individual tokens based on these gated scores, DGPO heavily incentivizes critical exploratory steps while suppressing unwarranted, low-entropy deviations. Consequently, DGPO completely eliminates the traditional token-level KL penalty and achieves fine-grained credit reallocation without the computational overhead of an additional value network. Extensive empirical evaluations demonstrate that DGPO sets a new state-of-the-art for critic free alignment. Notably, on the Qwen2.5-32B architecture, DGPO achieves 60.0% Avg@32 accuracy and 46.0% Avg@32 accuracy on the challenging AIME2024 and AIME2025 benchmarks respectively, substantially outperforming competitive baselines like DAPO.

cs.LG

A Theory of Covenant Accounting Adjustment

We develop an incomplete-contracting model with accounting-based covenants to study how covenant accounting adjustments are made and what properties they exhibit. Standard accounting rules (e.g., GAAP) can generate false-alarm errors or undue-optimism errors. The manager can exert costly effort to privately identify these errors and propose adjustments. If errors are not corrected, control rights may be inefficiently allocated, leading to costly renegotiation. We show that (1) adjustments always correct false-alarm errors, but correct undue-optimism errors only when their magnitude is small; and (2) the manager may expend socially wasteful effort to identify these errors. The model yields testable empirical predictions and policy implications.

econ.GN

Beyond Fixed Inference: Quantitative Flow Matching for Adaptive Image Denoising

Diffusion and flow-based generative models have shown strong potential for image restoration. However, image denoising under unknown and varying noise conditions remains challenging, because the learned vector fields may become inconsistent across different noise levels, leading to degraded restoration quality under mismatch between training and inference. To address this issue, we propose a quantitative flow matching framework for adaptive image denoising. The method first estimates the input noise level from local pixel statistics, and then uses this quantitative estimate to adapt the inference trajectory, including the starting point, the number of integration steps, and the step-size schedule. In this way, the denoising process is better aligned with the actual corruption level of each input, reducing unnecessary computation for lightly corrupted images while providing sufficient refinement for heavily degraded ones. By coupling quantitative noise estimation with noise-adaptive flow inference, the proposed method improves both restoration accuracy and inference efficiency. Extensive experiments on natural, medical, and microscopy images demonstrate its robustness and strong generalization across diverse noise levels and imaging conditions.

cs.CV

Polarization Engineering of Second-Harmonic Generation in 3R-MoS$_2$ Waveguides

Chip-scale nonlinear optics enables strong light-matter interactions within compact devices, serving as a fundamental platform for multifunctional integrated photonics from classical optical signal processing to quantum information technologies. Transition metal dichalcogenide (TMDC) waveguides have recently emerged as a highly promising platform owing to their giant material nonlinearity and extended interaction lengths. To date, however, research has predominantly focused on conversion efficiency, leaving the mechanisms governing the polarization state of nonlinear signal largely unexplored. Here, we establish a comprehensive framework for engineering the polarization of second-harmonic generation (SHG) in 3R-MoS$_2$ waveguides. By synergizing polarization-resolved measurements with theoretical modeling, we reveal that the SHG polarization is determined by guided-mode interactions constrained by waveguide geometry and crystal symmetry, and further reshaped during propagation. We demonstrate that thickness-dependent guided-mode confinement and in-plane crystal symmetry provide robust, static control over SHG polarization, while propagation length offers a dynamic tuning knob for continuously tailoring the nonlinear output. Our findings provide a deterministic approach for on-chip polarization engineering, opening opportunities for reconfigurable nonlinear light sources and quantum photonic circuits.

physics.optics

LogicGraph : Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification

Evaluations of large language models (LLMs) primarily emphasize convergent logical reasoning, where success is defined by producing a single correct proof. However, many real-world reasoning problems admit multiple valid derivations, requiring models to explore diverse logical paths rather than committing to one route. To address this limitation, we introduce LogicGraph, the first benchmark aimed to systematically evaluate multi-path logical reasoning, constructed via a neuro-symbolic framework that leverages backward logic generation and semantic instantiation. This pipeline yields solver-verified reasoning problems formalized by high-depth multi-path reasoning and inherent logical distractions, where each instance is associated with an exhaustive set of minimal proofs. We further propose a reference-free evaluation framework to rigorously assess model performance in both convergent and divergent regimes. Experiments on state-of-the-art language models reveal a common limitation: models tend to commit early to a single route and fail to explore alternatives, and the coverage gap grows substantially with reasoning depth. LogicGraph exposes this divergence gap and provides actionable insights to motivate future improvements. Our code and data will be released at https://github.com/kkkkarry/LogicGraph.

cs.AI

T2MBench: A Benchmark for Out-of-Distribution Text-to-Motion Generation

Most existing evaluations of text-to-motion generation focus on in-distribution textual inputs and a limited set of evaluation criteria, which restricts their ability to systematically assess model generalization and motion generation capabilities under complex out-of-distribution (OOD) textual conditions. To address this limitation, we propose a benchmark specifically designed for OOD text-to-motion evaluation, which includes a comprehensive analysis of 14 representative baseline models and the two datasets derived from evaluation results. Specifically, we construct an OOD prompt dataset consisting of 1,025 textual descriptions. Based on this prompt dataset, we introduce a unified evaluation framework that integrates LLM-based Evaluation, Multi-factor Motion evaluation, and Fine-grained Accuracy Evaluation. Our experimental results reveal that while different baseline models demonstrate strengths in areas such as text-to-motion semantic alignment, motion generalizability, and physical quality, most models struggle to achieve strong performance with Fine-grained Accuracy Evaluation. These findings highlight the limitations of existing methods in OOD scenarios and offer practical guidance for the design and evaluation of future production-level text-to-motion models.

cs.CV

TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing

Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Language Models (MLLMs) to promote information exchange between instructions and images, ensuring the controllability and flexibility of image editing. However, these frameworks often build a multi-instruction dataset to train the model to handle multiple editing tasks, which is not only time-consuming and labor-intensive but also fails to achieve satisfactory results. In this paper, we present TalkPhoto, a versatile training-free image editing framework that facilitates precise image manipulation through conversational interaction. We instruct the open-source LLM with a specially designed prompt template to analyze user needs after receiving instructions and hierarchically invoke existing advanced editing methods, all without additional training. Moreover, we implement a plug-and-play and efficient invocation of image editing methods, allowing complex and unseen editing tasks to be integrated into the current framework, achieving stable and high-quality editing results. Extensive experiments demonstrate that our method not only provides more accurate invocation with fewer token consumption but also achieves higher editing quality across various image editing tasks.

cs.CV

Neural Collapse in Test-Time Adaptation

Test-Time Adaptation (TTA) enhances model robustness to out-of-distribution (OOD) data by updating the model online during inference, yet existing methods lack theoretical insights into the fundamental causes of performance degradation under domain shifts. Recently, Neural Collapse (NC) has been proposed as an emergent geometric property of deep neural networks (DNNs), providing valuable insights for TTA. In this work, we extend NC to the sample-wise level and discover a novel phenomenon termed Sample-wise Alignment Collapse (NC3+), demonstrating that a sample's feature embedding, obtained by a trained model, aligns closely with the corresponding classifier weight. Building on NC3+, we identify that the performance degradation stems from sample-wise misalignment in adaptation which exacerbates under larger distribution shifts. This indicates the necessity of realigning the feature embeddings with their corresponding classifier weights. However, the misalignment makes pseudo-labels unreliable under domain shifts. To address this challenge, we propose NCTTA, a novel feature-classifier alignment method with hybrid targets to mitigate the impact of unreliable pseudo-labels, which blends geometric proximity with predictive confidence. Extensive experiments demonstrate the effectiveness of NCTTA in enhancing robustness to domain shifts. For example, NCTTA outperforms Tent by 14.52% on ImageNet-C. Project page is publicly available at https://github.com/Cevaaa/NCTTA.

cs.CV

Direct Mapping of Intrinsic Topology of Bound States in the Continuum via Nonlinear Emission

The direct mapping of the intrinsic topology in a leaky photonic band is crucial and challenging in topological photonics. For instance, observables in bound states in the continuum (BICs) feature complex topological textures such as a polarization vortex in momentum space, which nonetheless is difficult to be characterized in far-field scattering, especially considering the dominant direct channel. Here, we propose and experimentally demonstrate a hybrid nonlinear metasurface that enables a direct visualization of the intrinsic topology in BICs via second-harmonic generation (SHG). The enhanced local-source of SHG from the ultrathin indium tin oxide can effectively excite the emissions from the eigenmodes of a TiO2 photonics crystal slab, achieving three-order enhancement of SHG magnitudes. Importantly, these enhanced SH emissions carry topological polarization textures of BICs to the far field. With this, we can directly construct polarization vector maps of symmetry-protected BICs and chiral symmetry-broken quasi-BICs, clearly visualizing the winding structure around V points, the generation and evolution of chiral C points. This work provides a universal approach for characterizing topological photonic systems via coherent nonlinearity processes, opening new avenues for studying topological phenomena in non-Hermitian photonic systems.

physics.optics

Physics-Inspired Gaussian Kolmogorov-Arnold Networks for X-ray Scatter Correction in Cone-Beam CT

Cone-beam CT (CBCT) employs a flat-panel detector to achieve three-dimensional imaging with high spatial resolution. However, CBCT is susceptible to scatter during data acquisition, which introduces CT value bias and reduced tissue contrast in the reconstructed images, ultimately degrading diagnostic accuracy. To address this issue, we propose a deep learning-based scatter artifact correction method inspired by physical prior knowledge. Leveraging the fact that the observed point scatter probability density distribution exhibits rotational symmetry in the projection domain. The method uses Gaussian Radial Basis Functions (RBF) to model the point scatter function and embeds it into the Kolmogorov-Arnold Networks (KAN) layer, which provides efficient nonlinear mapping capabilities for learning high-dimensional scatter features. By incorporating the physical characteristics of the scattered photon distribution together with the complex function mapping capacity of KAN, the model improves its ability to accurately represent scatter. The effectiveness of the method is validated through both synthetic and real-scan experiments. Experimental results show that the model can effectively correct the scatter artifacts in the reconstructed images and is superior to the current methods in terms of quantitative metrics.

cs.CV

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines feeding Large Language Models (LLMs) or native Multimodal LLMs (MLLMs), face key limitations: the former loses structural detail, while the latter struggles with context modeling. Retrieval-Augmented Generation (RAG) helps ground models in external data, but documents' multimodal nature, i.e., combining text, tables, charts, and layout, demands a more advanced paradigm: Multimodal RAG. This approach enables holistic retrieval and reasoning across all modalities, unlocking comprehensive document intelligence. Recognizing its importance, this paper presents a systematic survey of Multimodal RAG for document understanding. We propose a taxonomy based on domain, retrieval modality, and granularity, and review advances involving graph structures and agentic frameworks. We also summarize key datasets, benchmarks, applications and industry deployment, and highlight open challenges in efficiency, fine-grained representation, and robustness, providing a roadmap for future progress in document AI.

cs.CL