SearcharxivSearch

arXiv subjects

Peng Yuan

Publications and source records attributed to Peng Yuan.

At least 19 recordsLinked to original sources

An Ambiguity-Function-Assisted Newtonized Channel Estimation Method for Pulse-Shaped AFDM Under Fractional Delay and Doppler

Accurate channel estimation for pulse-shaped AFDM systems over doubly selective channels with fractional normalized delay and Doppler remains challenging. This paper proposes a low-complexity ambiguity-function-assisted newtonized channel (AFNC) estimation method. Specifically, we first present a closed-form input-output relation for pulse-shaped affine frequency division multiplexing (AFDM) under fractional normalized delay and Doppler. As a further step, we demonstrate that the input-output relation admits a low-complexity representation by offline precomputing and storing the discretized ambiguity function of the shaping pulse, followed by tailored cyclic-shift and stacking operations. Building on this representation, AFNC performs fractional delay-Doppler channel estimation through Newtonized refinement, where the required Jacobian and Hessian updates are computed efficiently using the low-complexity input-output representation. Simulation results confirm the effectiveness of the proposed approach.

eess.SP

Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents

Large Language Models (LLMs) have been widely adopted in conversational applications. However, their reliance on parametric knowledge limits reliability in real-world scenarios that require dynamic or domain-specific information. Retrieval-Augmented Generation (RAG) addresses this limitation by incorporating external knowledge during generation, but existing text-based and graph-based RAG methods often struggle with noisy or irrelevant contexts. In this work, we propose Structure-aware Retrieval Augmented Generation (SA-RAG), which uses tables as an intermediate structured representation to provide a compact and controllable interface that reduces noise while preserving essential information. We introduce a quality-aware table metadata generation framework that models metadata normalization and effectiveness, improving metadata quality and downstream performance. Furthermore, we explore both training-free and training-based table generation methods. Generation validation and direct preference optimization further improve table quality while maintaining semantic and structural consistency. Experiments on two noisy real-world datasets show that SA-RAG significantly outperforms existing RAG baselines. Our code is publicly available at a public repository.

cs.CL

WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark

Existing browser agent benchmarks face a fundamental trilemma: real-website benchmarks lack reproducibility due to content drift, controlled environments sacrifice realism by omitting real-web noise, and both require costly manual curation that limits scalability. We present WebForge, the first fully automated framework that resolves this trilemma through a four-agent pipeline -- Plan, Generate, Refine, and Validate -- that produces interactive, self-contained web environments end-to-end without human annotation. A seven-dimensional difficulty control framework structures task design along navigation depth, visual complexity, reasoning difficulty, and more, enabling systematic capability profiling beyond single aggregate scores. Using WebForge, we construct WebForge-Bench, a benchmark of 934 tasks spanning 7 domains and 3 difficulty levels. Multi-model experiments show that difficulty stratification effectively differentiates model capabilities, while cross-domain analysis exposes capability biases invisible to aggregate metrics. Together, these results confirm that multi-dimensional evaluation reveals distinct capability profiles that a single aggregate score cannot capture. Code and benchmark are publicly available at https://github.com/yuandaxia2001/WebForge.

cs.AI

VCC-DSA: A Novel Vascular Consistency Constrained DSA Imaging Model for Motion Artifact Suppression

Digital Subtraction Angiography (DSA) is a clinically significant imaging technique for diagnosing cerebrovascular disease, as gold-standard. However, the artifacts caused by motion of high-attenuation tissues such as bones, teeth, and catheters, seriously reduce the visibility of blood vessels. This paper presents a novel Vascular Consistency Constrained DSA Imaging Model (VCC-DSA) for robust motion suppression and precise vascular imaging with the following designs: 1) We specially design a Learning-based Subtraction Mapping Paradigm, so that the ill-posed problem of existing learning-based methods can be solved to enhance the stability of the algorithm. 2) Our model effectively develops Residual Dense Blocks and details-shortcut to improve the performance under complex structures, such as moving bones overlapping with blood vessels, and small features, like peripheral vessels. 3) An innovative Vascular Consistency Strategy is proposed to extract intrinsically consistency from the various relative motions in mask-live images, so that spontaneously distils the vascular structure with contrast-agent development and robustly suppress motion artifacts, and also naturally alleviates the high matching requirements of data. 4) We creatively design a Mixup-based Data Self-evolution Strategy for data-intra self-enhancement in training loop, so that the training data gains dynamically optimized to promote model better learning the vascular features, and excluding the irrelevant structures in live/mask image and even the inevitable-artifacts/fake-structure in label. Prospectively, to further evaluate practical value, an actual general anesthesia animal experiment is specially conducted, besides the assessment on human clinical data. Compared with other method, our model improves the PSNR and SSIM by 73.4% and 8.56%, respectively.

eess.IV

FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data

Composed Image Retrieval (CIR) retrieves target images using a reference image paired with modification text. Despite rapid advances, all existing methods and datasets operate at the image level -- a single reference image plus modification text in, a single target image out -- while real e-commerce users reason about products shown from multiple viewpoints. We term this mismatch View Incompleteness and formally define a new Multi-View CIR task that generalizes standard CIR from image-level to product-level retrieval. To support this task, we construct FashionMV, the first large-scale multi-view fashion dataset for product-level CIR, comprising 127K products, 472K multi-view images, and over 220K CIR triplets, built through a fully automated pipeline leveraging large multimodal models. We further propose ProCIR (Product-level Composed Image Retrieval), a modeling framework built upon a multimodal large language model that employs three complementary mechanisms -- two-stage dialogue, caption-based alignment, and chain-of-thought guidance -- together with an optional supervised fine-tuning (SFT) stage that injects structured product knowledge prior to contrastive training. Systematic ablation across 16 configurations on three fashion benchmarks reveals that: (1) alignment is the single most critical mechanism; (2) the two-stage dialogue architecture is a prerequisite for effective alignment; and (3) SFT and chain-of-thought serve as partially redundant knowledge injection paths. Our best 0.8B-parameter model outperforms all baselines, including general-purpose embedding models 10x its size. The dataset, model, and code are publicly available at https://github.com/yuandaxia2001/FashionMV.

cs.CV

Yunque DeepResearch Technical Report

Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full potential is hindered by critical limitations, including escalating contextual noise in long-horizon tasks, fragility leading to cascading errors, and a lack of modular extensibility. To address these challenges, we introduce Yunque DeepResearch, a hierarchical, modular, and robust framework. The architecture is characterized by three key components: (1) a centralized Multi-Agent Orchestration System that routes subtasks to an Atomic Capability Pool of tools and specialized sub-agents; (2) a Dynamic Context Management mechanism that structures completed sub-goals into semantic summaries to mitigate information overload; and (3) a proactive Supervisor Module that ensures resilience through active anomaly detection and context pruning. Yunque DeepResearch achieves state-of-the-art performance across a range of agentic deep research benchmarks, including GAIA, BrowseComp, BrowseComp-ZH, and Humanity's Last Exam. We open-source the framework, reproducible implementations, and application cases to empower the community.

cs.CL

An Anti-Interference AFDM System: Interference Impacts Analyses and Parameter Optimization

This paper proposes an anti-interference affine frequency division multiplexing (AFDM) system to ensure reliability and resource efficiency under malicious high-power interference originating from adversarial devices in high-mobility scenarios. Closed-form expressions of interferences in the discrete affine Fourier transform (DAFT) domain are derived by utilizing the stationary phase principle and the Affine Fourier transform convolution theorem, which indicates that interference impacts can be classified into stationary and non-stationary categories. On this basis, we reveal the analytical relationship between packet throughput and the paramerters of spread spectrum and error correction coding in our proposed anti-interference system, which enables the design of a parameter optimization algorithm that maximizes packet throughput. For reception, by jointly utilizing the autocorrelation function of spreading sequence and the cyclic-shift property of AFDM input-output relation, we design a linear-complexity correlation-based DAFT domain detector (CDD) capable of achieving full diversity gain, which performs correlation-based equalization to avoid matrix inversion. Numerical results validate the accuracy of the derived closed-form expressions and verify that the proposed anti-interference AFDM system could achieve high packet throughput under interference in high-mobility scenarios.

cs.IT

Coherently Enhanced Axion-Photon Conversion via Seeded Photons for Short-Pulse Axion Detection

We propose a seeded axion-photon conversion scheme to enhance the sensitivity of light-shining-through-a-wall (LSW) experiments for axion detection, where the axions are generated from short pulse lasers and the usual resonant cavity is not applicable. By injecting a weak, coherent seed electromagnetic (EM) field into the axion-photon conversion region, the axion-induced EM field can constructively interfere with the seed field, amplifying the number of regenerated photons to a level exceeding that of the unseeded scenario. We evaluate the expected signal enhancement, statistical limits from Poisson counting with seed fluctuations and background, and the potential improvement in coupling sensitivity. Compared to a standard LSW setup, the seeded scheme can achieve orders-of-magnitude higher photon yield per axion, potentially surpassing resonance-enhanced experiments in certain parameter regimes. This approach presents a promising pathway to extend the reach of laboratory axion searches, particularly in scenarios where the resonant cavities are impractical.

hep-ph

Affine-Doppler Division Multiplexing for High-Mobility Wireless Communications Systems

Affine Frequency Division Multiplexing (AFDM) has been regarded as a candidate integrated sensing and communications (ISAC) waveform owing to its superior communication performance, outperforming the Orthogonal Time-Frequency Space (OTFS) that has been researched for a longer time. However, since the above two waveforms are incompatible with each other, the state-of-the-art methods well-designed for OTFS may not be directly applicable to AFDM. This paper introduces a new orthogonal multicarrier waveform, namely Affine-Doppler Division Multiplexing (ADDM), which can provide a generic framework and subsume the existing OTFS and AFDM as a particular case. ADDM modulating information symbols in the Affine-Doppler (A-D) domain based on a two-dimensional (2D) transform can enjoy both excellent unambiguous Doppler and Doppler resolution, which is the same as AFDM but outperforms OTFS. Moreover, benefiting from the 2D transform, the symbols block of ADDM in the A-D domain undergoes a 2D cyclic shift produced by the delay and the Doppler of the channel, similar to the 2D cyclic shift in the delay-Doppler domain of cyclic prefix (CP)-OTFS. This offers a potential to directly apply the state-of-the-art methods well-designed for OTFS and AFDM to ADDM. Numerical results show that ADDM achieves comparable BER performance with AFDM but outperforms OTFS in high-mobility scenarios.

eess.SP

4D Virtual Imaging Platform for Dynamic Joint Assessment via Uni-Plane X-ray and 2D-3D Registration

Conventional computed tomography (CT) lacks the ability to capture dynamic, weight-bearing joint motion. Functional evaluation, particularly after surgical intervention, requires four-dimensional (4D) imaging, but current methods are limited by excessive radiation exposure or incomplete spatial information from 2D techniques. We propose an integrated 4D joint analysis platform that combines: (1) a dual robotic arm cone-beam CT (CBCT) system with a programmable, gantry-free trajectory optimized for upright scanning; (2) a hybrid imaging pipeline that fuses static 3D CBCT with dynamic 2D X-rays using deep learning-based preprocessing, 3D-2D projection, and iterative optimization; and (3) a clinically validated framework for quantitative kinematic assessment. In simulation studies, the method achieved sub-voxel accuracy (0.235 mm) with a 99.18 percent success rate, outperforming conventional and state-of-the-art registration approaches. Clinical evaluation further demonstrated accurate quantification of tibial plateau motion and medial-lateral variance in post-total knee arthroplasty (TKA) patients. This 4D CBCT platform enables fast, accurate, and low-dose dynamic joint imaging, offering new opportunities for biomechanical research, precision diagnostics, and personalized orthopedic care.

cs.CV

Stability of Large-Amplitude Viscous Shock Under Periodic Perturbation for 1-d Viscoelasticity with Non-Convex Constitutive Relations

This paper investigates the large-time behavior of the viscous shock profile for the one-dimensional system of viscoelasticity, subject to initial perturbations that approach space-periodic functions at far fields. We specifically address the case with non-convex constitutive stress relations and non-degenerate Lax's shock. Under the assumptions of suitably small initial perturbations satisfying a zero-mass type condition, we prove that the solution of the system converges to a viscous shock profile with a shift, which is partially determined by the space-periodic perturbation. Notably, our result imposes no amplitude restrictions on the viscous shock waves. This work extends the result of Kawashima-Matsumura (\textit{Commun. Pure Appl. Math.} \textbf{47}, 1994) by simultaneously handling both large-amplitude shocks and space-periodic perturbations, while also generalizing the result of Huang-Yuan (\textit{Commun. Math. Phys.} \textbf{387}, 2021) by allowing for a non-convex constitutive relation. The key ingredient of proof is decomposing the large-amplitude shock wave into small-amplitude shocks and, for each, introducing suitable transform and weight functions to counteract the adverse effects of non-convex constitutive relations encountered during weighted energy estimates on the system in effective velocity and deformation gradient variables.

math.AP

The non-linear multiple stopping problem: between the discrete and the continuous time

We consider the non-linear optimal multiple stopping problem under general conditions on the non-linear evaluation operators, which might depend on two time indices: the time of evaluation/assessment and the horizon (when the reward or loss is incurred). We do not assume convexity/concavity or cash-invariance. We focus on the case where the agent's stopping strategies are what we call Bermudan stopping strategies, a framework which can be seen as lying between the discrete and the continuous time. We first study the non-linear double optimal stopping problem by using a reduction approach. We provide a necessary and a sufficient condition for optimal pairs, and a result on existence of optimal pairs. We then generalize the results to the non-linear $d$-optimal stopping problem. We treat the symmetric case (of additive and multiplicative reward families) as examples.

math.OC

Dynamical phases of short-term memory mechanisms in RNNs

Short-term memory is essential for cognitive processing, yet our understanding of its neural mechanisms remains unclear. Neuroscience has long focused on how sequential activity patterns, where neurons fire one after another within large networks, can explain how information is maintained. While recurrent connections were shown to drive sequential dynamics, a mechanistic understanding of this process still remains unknown. In this work, we introduce two unique mechanisms that can support this form of short-term memory: slow-point manifolds generating direct sequences or limit cycles providing temporally localized approximations. Using analytical models, we identify fundamental properties that govern the selection of each mechanism. Precisely, on short-term memory tasks (delayed cue-discrimination tasks), we derive theoretical scaling laws for critical learning rates as a function of the delay period length, beyond which no learning is possible. We empirically verify these results by training and evaluating approximately 80,000 recurrent neural networks (RNNs), which are publicly available for further analysis. Overall, our work provides new insights into short-term memory mechanisms and proposes experimentally testable predictions for systems neuroscience.

q-bio.NC

Latent computing by biological neural networks: A dynamical systems framework

Although individual neurons and neural populations exhibit the phenomenon of representational drift, perceptual and behavioral outputs of many neural circuits can remain stable across time scales over which representational drift is substantial. These observations motivate a dynamical systems framework for neural network activity that focuses on the concept of \emph{latent processing units,} core elements for robust coding and computation embedded in collective neural dynamics. Our theoretical treatment of these latent processing units yields five key attributes of computing through neural network dynamics. First, neural computations that are low-dimensional can nevertheless generate high-dimensional neural dynamics. Second, the manifolds defined by neural dynamical trajectories exhibit an inherent coding redundancy as a direct consequence of the universal computing capabilities of the underlying dynamical system. Third, linear readouts or decoders of neural population activity can suffice to optimally subserve downstream circuits controlling behavioral outputs. Fourth, whereas recordings from thousands of neurons may suffice for near optimal decoding from instantaneous neural activity patterns, experimental access to millions of neurons may be necessary to predict neural ensemble dynamical trajectories across timescales of seconds. Fifth, despite the variable activity of single cells, neural networks can maintain stable representations of the variables computed by the latent processing units, thereby making computations robust to representational drift. Overall, our framework for latent computation provides an analytic description and empirically testable predictions regarding how large systems of neurons perform robust computations via their collective dynamics.

q-bio.NC

An Integrated Sensing and Communications System Based on Affine Frequency Division Multiplexing

This paper proposes an integrated sensing and communications (ISAC) system based on affine frequency division multiplexing (AFDM) waveform. To this end, a metric set is designed according to not only the maximum tolerable delay/Doppler, but also the weighted spectral efficiency as well as the outage/error probability of sensing and communications. This enables the analytical investigation of the performance trade-offs of AFDM-ISAC system using the derived analytical relation among metrics and AFDM waveform parameters. Moreover, by revealing that delay and the integral/fractional parts of normalized Doppler can be decoupled in the affine Fourier transform-Doppler domain, an efficient estimation method is proposed for our AFDM-ISAC system, whose unambiguous Doppler can break through the limitation of subcarrier spacing. Theoretical analyses and numerical results verify that our proposed AFDM-ISAC system may significantly enlarge unambiguous delay/Doppler while possessing good spectral efficiency and peak-to-sidelobe level ratio in high-mobility scenarios.

eess.SP

The Impact of Ionic Anharmonicity on Superconductivity in Metal-Stuffed B-C Clathrates

Metal-stuffed B$-$C compounds with sodalite clathrate structure have captured increasing attention due to their predicted exceptional superconductivity above liquid nitrogen temperature at ambient pressure. However, by neglecting the quantum lattice anharmonicity, the existing studies may result in an incomplete understanding of such a lightweight system. Here, using state-of-the-art ab initio methods incorporating quantum effects and machine learning potentials, we revisit the properties of a series of $XY$$\text{B}_{6}\text{C}_{6}$ clathrates where $X$ and $Y$ are metals. Our findings show that ionic quantum and anharmonic effects can harden the $E_g$ and $E_u$ vibrational modes, enabling the dynamical stability of 15 materials previously considered unstable in the harmonic approximation, including materials with previously unreported ($XY$)$^{1+}$ state, which is demonstrated here to be crucial to reach high critical temperatures. Further calculations based on the anisotropic Migdal-Eliashberg equation demonstrate that the $T_\text{c}$ values for KRb$\text{B}_{6}\text{C}_{6}$ and Rb$\text{B}_{3}\text{C}_{3}$ among these stabilized compounds are 102 and 115 K at 0 and 15 GPa, respectively, both being higher than $T_\text{c}$ of 92 K of KPb$\text{B}_{6}\text{C}_{6}$ at the anharmonic level. These record-high $T_\text{c}$ values, surpassing liquid nitrogen temperatures, emphasize the importance of anharmonic effects in stabilizing B-C clathrates with large electron-phonon coupling strength and advancing the search for high-$T_\text{c}$ superconductivity at (near) ambient pressure.

cond-mat.supr-con

Terahertz-driven Two-Dimensional Mapping for Electron Temporal Profile Measurement

The precision measurement of real-time electron temporal profiles is crucial for advancing electron and X-ray devices used in ultrafast imaging and spectroscopy. While high temporal resolution and large temporal window can be achieved separately using different technologies, real-time measurement enabling simultaneous high resolution and large window remains challenging. Here, we present the first THz-driven sampling electron oscilloscope capable of measuring electron pulses with high temporal resolution and a scalable, large temporal window simultaneously. The transient THz electric field induces temporal electron streaking in the vertical axis, while extended interaction along the horizontal axis leads to a propagation-induced time delay, enabling electron beam sampling with sub-cycle THz wave. This allows real-time femtosecond electron measurement with a tens-of-picosecond window, surpassing previous THz-based techniques by an order of magnitude. The measurement capability is further enhanced through projection imaging, deflection cavity tilting, and shorted antenna utilization, resulting in signal spatial magnification, extended temporal window, and increased field strength. The technique holds promise for a wide range of applications and opens new opportunities in ultrafast science and accelerator technologies.

physics.optics

MDS-GNN: A Mutual Dual-Stream Graph Neural Network on Graphs with Incomplete Features and Structure

Graph Neural Networks (GNNs) have emerged as powerful tools for analyzing and learning representations from graph-structured data. A crucial prerequisite for the outstanding performance of GNNs is the availability of complete graph information, i.e., node features and graph structure, which is frequently unmet in real-world scenarios since graphs are often incomplete due to various uncontrollable factors. Existing approaches only focus on dealing with either incomplete features or incomplete structure, which leads to performance loss inevitably. To address this issue, this study proposes a mutual dual-stream graph neural network (MDS-GNN), which implements a mutual benefit learning between features and structure. Its main ideas are as follows: a) reconstructing the missing node features based on the initial incomplete graph structure; b) generating an augmented global graph based on the reconstructed node features, and propagating the incomplete node features on this global graph; and c) utilizing contrastive learning to make the dual-stream process mutually benefit from each other. Extensive experiments on six real-world datasets demonstrate the effectiveness of our proposed MDS-GNN on incomplete graphs.

cs.LG