SearcharxivSearch

arXiv subjects

Xiang Yang

Publications and source records attributed to Xiang Yang.

At least 19 recordsLinked to original sources

GeoFF3D: Coordinate-Anchored Feed-Forward Reconstruction for Large-Scale UAV Mapping

Existing feed-forward 3D reconstruction methods typically process a bounded number of images and recover cameras and geometry in local or internally normalized frames. Extending them to large-scale UAV mapping requires scalable multi-chunk processing and reliable aggregation, while full Sim(3) alignment can become unstable for near collinear trajectories. We present GeoFF3D, which combines a coordinate-anchored model with a spatial large-scale reconstruction framework (SLRF). The model uses georeferenced camera translations and optional geometric priors to predict camera poses and dense point maps directly in a gravity-aligned Z-up metric frame. SLRF partitions images into spatially overlapping chunks, propagates shared-view priors, and aggregates local reconstructions hierarchically, while remaining applicable to different bounded-view models. Across nine aerial mapping blocks, GeoFF3D achieves the best average reconstruction quality, improving F@5 from 0.829 for Pi3X + SLRF to 0.877. On long UAVScenes sequences, it reaches 0.848, compared with 0.687 for Pi3X + SLRF and 0.451 for the strongest evaluated SLAM/streaming baseline. GeoFF3D reconstructs 2,000 images in approximately five minutes, demonstrating scalable and robust large-scale UAV reconstruction.The code is available at https://github.com/yanxian-ll/GeoFF3D.

cs.CV

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful content remains a critical challenge, particularly in image-to-image (I2I) editing tasks. Existing safety mechanisms are primarily designed for text-to-image (T2I) synthesis or U-Net-based architectures, which limits their effectiveness for unified safety mitigation in DiT-based frameworks. To bridge this gap, we propose Unified Visual Safety Regulator (UVR), a training-free safe generation framework that regulates unsafe semantics in generated images. UVR is grounded in an analysis of attention dynamics from the perspective of information flow in MM-Attn. We identify a task-independent start-up stage, during which unsafe semantics in output patches rapidly emerge and can be accurately localized, followed by task-specific semantic amplification and interference stages, where harmful signals are further propagated and entangled with benign content. Based on these observations, UVR mitigates unsafe generation through unified, targeted attention modulation and explicit restriction of harmful information flow over the identified unsafe output patches. Experiments across various concepts show that UVR achieves state-of-the-art safety performance by achieving 91% and 77% erase rate in image synthesis and editing tasks, while preserving visual quality and fidelity with minimal degradation. Code is available at https://github.com/deng12yx/UVR.

cs.CV

QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples

Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text. Existing retrieval-augmented generation (RAG) systems are optimized primarily for semantic relevance, but retrieving plausible passages does not guarantee correct query execution. We introduce QO-Bench, a diagnostic benchmark for query-operator question answering over typed event tuples. The benchmark covers 22,984 news articles and 614 corporate events, with 18 query templates instantiating 785 questions. Each gold answer is deterministically computed from typed event tuples and scored by recall, with answers matched to the gold tuples by exact match rather than an LLM judge. This design enables operator-level diagnosis such as joins and intersection. We evaluate RAG, ReAct RAG, GraphRAG, and information-extraction-to-SQL under matched conditions, with a long-context oracle ceiling to isolate retrieval failure. A two-axis framework -- index-time preservation versus query-time execution -- predicts where each paradigm fails, and the results bear it out: systems retrieve relevant text but discard the typed values operators need, and the deployable paradigm ranking inverts across operators, with similarity retrieval leading on filter/project and extraction-to-SQL on intersection and counting. Even given the gold evidence, a long-context oracle stays far from saturated, so operator execution -- not retrieval alone -- is a core bottleneck that a stronger answer model does not remove. QO-Bench reframes the goal from passage relevance to query-operator-preserving retrieval. The benchmark, predictions, and code are released at https://github.com/ZHANG-MENGAO/qo-bench.

cs.CL

UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction

Feed-forward 3D reconstruction has advanced rapidly, but current models remain unreliable in UAV photogrammetric acquisition. We argue that this failure is caused not only by appearance-domain shift, but also by UAV-specific camera-geometry variations, especially oblique views and HFOV-height ambiguity. Existing UAV datasets mainly emphasize scene diversity and provide limited coverage of camera configurations, which restricts robustness evaluation and UAV-domain adaptation. To address this gap, we introduce UAVFF3D, a geometry-aware real-synthetic benchmark for feed-forward UAV 3D reconstruction. UAVFF3D contains more than 170k real UAV images and more than 370k synthetic images rendered from high-quality textured 3D models, covering diverse HFOVs, flight altitudes, viewing directions, and acquisition patterns. It also includes a controlled HFOV-height test subset for diagnosing projection-geometry ambiguity. We further propose an evaluation protocol that jointly assesses camera-geometry estimation and dense scene reconstruction under a shared global alignment, avoiding the bias caused by separate camera and geometry alignments. Experiments on representative feed-forward reconstruction models show that UAVFF3D-based domain adaptation consistently improves camera and geometry estimation, reducing Ray Error by up to 84.2%, Pose ATE by up to 76.0%, and Chamfer Distance by up to 41.1%. In oblique scenes, adaptation reduces the oblique-nadir rotation gap by up to 90.7%. Under HFOV-height ambiguity, it improves robustness across HFOV-height configurations and yields more stable performance across HFOV settings. Incorporating camera priors further improves reconstruction under UAV-specific acquisition geometries. The dataset and evaluation code are available at https://github.com/yanxian-ll/UAVFF3D .

cs.CV

From centrality to productivity: How firms reconfigure technological search in innovation networks?

Firms' positions in innovation networks determine their access to external knowledge, yet how these positions shape technological search behavior and influence productivity remains underexplored. We propose that central network positions systematically reconfigure firms' innovation strategies by promoting exploratory search across emerging technological domains while sustaining broader technological portfolios. This behavioral reorientation allows central firms to diversify their innovation efforts and leverage knowledge spillovers more effectively, translating network advantages into higher productivity. Using panel data on Chinese listed firms and patent-based measures of innovation networks, we construct a dynamic patent citation network to track changes in firms' network centrality and technological search patterns over time. Our findings show that firms with greater centrality are more likely to enter novel technological fields and expand their technological scope, leading to measurable gains in total factor productivity. We further demonstrate that the impact of network centrality on exploratory search is amplified by scientific embeddedness, whereas the productivity returns from exploration depend on technological distance. By connecting structural network positions with behavioral adaptations in technological search, this study uncovers a direct micro-level mechanism through which innovation networks drive firm performance. These results highlight the strategic value of network centrality in shaping not just access to knowledge, but also the direction and efficiency of innovation activities.

physics.soc-ph

SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers

Recent Text-to-Image (T2I) models based on rectified-flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when triggered by multi-token interactions. Existing mitigation methods largely rely on fine-tuning or attention modulation for concept unlearning; however, their expensive computational overhead and design tailored to U-Net-based denoisers hinder direct adaptation to transformer-based diffusion models (e.g., MMDiT). In this paper, we conduct an in-depth analysis of the attention mechanism in MMDiT and find that unsafe semantics concentrate within interpretable, low-dimensional subspaces at head level, where a finite set of safety-critical heads is responsible for unsafe feature extraction. We further observe that perturbing the Rotary Positional Embedding (RoPE) applied to the query and key vectors can effectively modify some specific concepts in the generated images. Motivated by these insights, we propose SafeRoPE, a lightweight and fine-grained safe generation framework for MMDiT. Specifically, SafeRoPE first constructs head-wise unsafe subspaces by decomposing unsafe embeddings within safety-critical heads, and computes a Latent Risk Score (LRS) for each input vector via projection onto these subspaces. We then introduce head-wise RoPE perturbations that can suppress unsafe semantics without degrading benign content or image quality. SafeRoPE combines both head-wise LRS and RoPE perturbations to perform risk-specific head-wise rotation on query and key vector embeddings, enabling precise suppression of unsafe outputs while maintaining generation fidelity. Extensive experiments demonstrate that SafeRoPE achieves SOTA performance in balancing effective harmful content mitigation and utility preservation for safe generation of MMDiT. Codes are available at https://github.com/deng12yx/SafeRoPE.

cs.CV

ShadowGS: Shadow-Aware 3D Gaussian Splatting for Satellite Imagery

3D Gaussian Splatting (3DGS) has emerged as a novel paradigm for 3D reconstruction from satellite imagery. However, in multi-temporal satellite images, prevalent shadows exhibit significant inconsistencies due to varying illumination conditions. To address this, we propose ShadowGS, a novel framework based on 3DGS. It leverages a physics-based rendering equation from remote sensing, combined with an efficient ray marching technique, to precisely model geometrically consistent shadows while maintaining efficient rendering. Additionally, it effectively disentangles different illumination components and apparent attributes in the scene. Furthermore, we introduce a shadow consistency constraint that significantly enhances the geometric accuracy of 3D reconstruction. We also incorporate a novel shadow map prior to improve performance with sparse-view inputs. Extensive experiments demonstrate that ShadowGS outperforms current state-of-the-art methods in shadow decoupling accuracy, 3D reconstruction precision, and novel view synthesis quality, with only a few minutes of training. ShadowGS exhibits robust performance across various settings, including RGB, pansharpened, and sparse-view satellite inputs.

cs.CV

Constraining the Nanohertz Gravitational Wave Background with an X-ray Pulsar Timing Array from NICER observations

We present constraints on the nanohertz gravitational wave background (GWB) using X-ray pulsar timing data from the Neutron Star Interior Composition Explorer(\textit{NICER}). By analyzing six millisecond pulsars over a six-year observational baseline, we employed a Bayesian framework to model noise components and search for a common red signal consistent with a GWB from supermassive black hole binaries (assuming a spectral index $\gamma_{\rm gwb}=13/3$). Our results show no significant evidence for a GWB, yielding a 95\% upper limit of $\log_{10}(A_{\rm gwb})<-13.4$. Weak evidence for Hellings-Downs spatial correlations was found (S=2.5), though the signal remains statistically inconclusive. Compared to radio and $\gamma$-ray pulsar timing arrays, the \textit{NICER} constraint is currently less stringent but demonstrates the feasibility of X-ray timing with \textit{NICER} for GWB studies and highlights the potential for improved sensitivity with future X-ray missions.

astro-ph.HE

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware heterogeneity across bimanual robotic platforms. To bridge this gap, we introduce RoboCOIN, a large-scale multi-embodiment bimanual manipulation dataset comprising over 180,000 demonstrations collected from 15 distinct robotic platforms. Spanning 16 diverse environments-including residential, commercial, and industrial settings-the dataset features 421 bimanual tasks systematically categorized by 39 bimanual collaboration actions and 432 objects. A key innovation of our work is the hierarchical capability pyramid, which provides granular annotations ranging from trajectory-level concepts to segment-level subtasks and frame-level kinematics. Furthermore, we present CoRobot, an efficient data processing pipeline powered by the Robot Trajectory Markup Language (RTML), designed to facilitate quality assessment, automated annotation, and unified multi-embodiment and data management. Extensive experiments demonstrate the effectiveness of RoboCOIN in enhancing the performance of various bimanual manipulation models across a wide spectrum of robotic embodiments. The entire dataset and codebase are fully open-sourced, providing a valuable resource for advancing research in bimanual and multi-embodiment manipulation.

cs.RO

GLOMIA-Pro: A Generalizable Longitudinal Medical Image Analysis Framework for Disease Progression Prediction

Longitudinal medical images are essential for monitoring disease progression by capturing spatiotemporal changes associated with dynamic biological processes. While current methods have made progress in modeling spatiotemporal patterns, they face three key limitations: (1) lack of generalizable framework applicable to diverse disease progression prediction tasks; (2) frequent overlook of the ordinal nature inherent in disease staging; (3) susceptibility to representation collapse due to structural similarities between adjacent time points, which can obscure subtle but discriminative progression biomarkers. To address these limitations, we propose a Generalizable LOngitudinal Medical Image Analysis framework for disease Progression prediction (GLOMIA-Pro). GLOMIA-Pro consists of two core components: progression representation extraction and progression-aware fusion. The progression representation extraction module introduces a piecewise orthogonal attention mechanism and employs a novel ordinal progression constraint to disentangle finegrained temporal imaging variations relevant to disease progression. The progression-aware fusion module incorporates a redesigned skip connection architecture which integrates the learned progression representation with current imaging representation, effectively mitigating representation collapse during cross-temporal fusion. Validated on two distinct clinical applications: knee osteoarthritis severity prediction and esophageal cancer treatment response assessment, GLOMIA-Pro consistently outperforms seven state-of-the-art longitudinal analysis methods. Ablation studies further confirm the contribution of individual components, demonstrating the robustness and generalizability of GLOMIA-Pro across diverse clinical scenarios.

q-bio.QM

A data-driven convergence booster for accelerating and stabilizing pseudo time-stepping

This paper introduces a novel data-driven convergence booster that not only accelerates convergence but also stabilizes solutions in cases where obtaining a steady-state solution is otherwise challenging. The method constructs a reduced-order model (ROM) of the solution residual using intermediate solutions and periodically solves a least-square problem in the low-dimensional ROM subspace. The second-order approximation of the residual and the use of normal equation distinguish this work from similar approaches in the literature from the methodology perspective. From the application perspective, in contrast to prior studies that focus on linear systems or idealized problems, we rigorously assess the method's performance on realistic computational fluid dynamics (CFD) applications. In addition to reducing the time complexity of point-iterative solvers for linear systems, we demonstrate substantial reductions in the number of pseudo-time steps required for implicit schemes solving the nonlinear Navier-Stokes equations. Across a range of two- and three-dimensional flows-including subsonic inviscid and transonic turbulent cases-the method consistently achieves a 3 to 4 times speedup in wall-clock time. Lastly, the proposed method acts as a robust stabilizer, capable of converging to steady solutions in flows that would otherwise exhibit persistent unsteadiness-such as vortex shedding or transonic buffet-without relying on symmetry boundary conditions.

physics.flu-dyn

Electroweak baryogenesis and electron EDM in the TNMSSM

We have studied the impact of CP-violating (CPV) effects on electroweak baryogenesis (EWBG) and the electric dipole moment (EDM) of electron ($d_e$), the electric dipole moment (EDM) of mercury neutron ($d_n$) and the electric dipole moment (EDM) of ($d_{Hg}$) in an extension of the Minimal Supersymmetric Standard Model (MSSM). The model incorporating a $SU(2)$ triplet with hypercharges of $\pm1$ and a gauge singlet from the standard model that is mutually coupled, is collectively referred to as the next-to-minimal supersymmetric standard model with triplets (TNMSSM). Furthermore, we discuss the strong first-order phase transition (PT) achieved by this model via tree-level effects. The numerical results indicate that the TNMSSM can account for the observed baryon asymmetry. Additionally, the regions favored by EWBG can be compatible with the corresponding electron EDM bounds when different contributions to the $d_e$ cancel each other out.

hep-ph

3rd Workshop on Maritime Computer Vision (MaCVi) 2025: Challenge Results

The 3rd Workshop on Maritime Computer Vision (MaCVi) 2025 addresses maritime computer vision for Unmanned Surface Vehicles (USV) and underwater. This report offers a comprehensive overview of the findings from the challenges. We provide both statistical and qualitative analyses, evaluating trends from over 700 submissions. All datasets, evaluation code, and the leaderboard are available to the public at https://macvi.org/workshop/macvi25.

cs.CV

Electric dipole moments from the perspective of a scalar triplet and singlet extension of the MSSM: A study of neutrons, electrons, mercury, and b and c quarks

In the framework of the minimal supersymmetric model extension with new scalar triplets and singlet (TNMSSM), we analyze the electric dipole moment (EDM) of neutron ($d_n$), electron EDM($d_e$), the mercury EDM($d_{Hg}$), $b$ quark ($d_b$) and $c$ quark ($d_c$) by considering the contributions from the one-loop diagrams, some two-loop diagrams and the Weinberg operators. The effects of TNMSSM specific CPV sources $\chi_d,\;\chi_t$ on $d_n$, $d_e$, $d_{Hg}$, $d_b$, $d_c$ are specialized, it is found that they have significant contributions to these EDMs, and the current upper bounds on $d_n$ impose strict constraints on $\chi_d,\;\chi_t$. The theoretical predictions on $d_b$, $d_c$ can reach about $10^{-22}$ e$\cdot$cm and $10^{-23}$ e$\cdot$cm respectively by taking the upper bounds on $d_n$, $d_e$, $d_{Hg}$ into account, which have great potential to be observed in future.

hep-ph

Variantional autoencoder with decremental information bottleneck for disentanglement

One major challenge of disentanglement learning with variational autoencoders is the trade-off between disentanglement and reconstruction fidelity. Previous studies, which increase the information bottleneck during training, tend to lose the constraint of disentanglement, leading to the information diffusion problem. In this paper, we present a novel framework for disentangled representation learning, DeVAE, which utilizes hierarchical latent spaces with decreasing information bottlenecks across these spaces. The key innovation of our approach lies in connecting the hierarchical latent spaces through disentanglement-invariant transformations, allowing the sharing of disentanglement properties among spaces while maintaining an acceptable level of reconstruction performance. We demonstrate the effectiveness of DeVAE in achieving a balance between disentanglement and reconstruction through a series of experiments and ablation studies on dSprites and Shapes3D datasets. Code is available at https://github.com/erow/disentanglement_lib/tree/pytorch#devae.

cs.LG

Robust experimental data assimilation for the Spalart-Allmaras turbulence model

This study presents a methodology focusing on the use of computational model and experimental data fusion to improve the Spalart-Allmaras (SA) closure model for Reynolds-averaged Navier-Stokes solutions. In particular, our goal is to develop a technique that not only assimilates sparse experimental data to improve turbulence model performance, but also preserves generalization for unseen cases by recovering classical SA behavior. We achieve our goals using data assimilation, namely the Ensemble Kalman filtering approach (EnKF), to calibrate the coefficients of the SA model for separated flows. A holistic calibration strategy is implemented via the parameterization of the production, diffusion, and destruction terms. This calibration relies on the assimilation of experimental data collected in the form of velocity profiles, skin friction, and pressure coefficients. Despite using observational data from a single flow condition around a backward-facing step (BFS), the recalibrated SA model demonstrates generalization to other separated flows, including cases such as the 2D NASA wall mounted hump (2D-WMH) and modified BFS. Significant improvement is observed in the quantities of interest, i.e., skin friction coefficient ($C_f$) and pressure coefficient ($C_p$) for each flow tested. Finally, it is also demonstrated that the newly proposed model recovers SA proficiency for flows, such as a NACA-0012 airfoil and axisymmetric jet (ASJ), and that the individually calibrated terms in the SA model target specific flow-physics wherein the calibrated production term improves the re-circulation zone while destruction improves the recovery zone.

physics.flu-dyn

Adaptive DNN Surgery for Selfish Inference Acceleration with On-demand Edge Resource

Deep Neural Networks (DNNs) have significantly improved the accuracy of intelligent applications on mobile devices. DNN surgery, which partitions DNN processing between mobile devices and multi-access edge computing (MEC) servers, can enable real-time inference despite the computational limitations of mobile devices. However, DNN surgery faces a critical challenge: determining the optimal computing resource demand from the server and the corresponding partition strategy, while considering both inference latency and MEC server usage costs. This problem is compounded by two factors: (1) the finite computing capacity of the MEC server, which is shared among multiple devices, leading to inter-dependent demands, and (2) the shift in modern DNN architecture from chains to directed acyclic graphs (DAGs), which complicates potential solutions. In this paper, we introduce a novel Decentralized DNN Surgery (DDS) framework. We formulate the partition strategy as a min-cut and propose a resource allocation game to adaptively schedule the demands of mobile devices in an MEC environment. We prove the existence of a Nash Equilibrium (NE), and develop an iterative algorithm to efficiently reach the NE for each device. Our extensive experiments demonstrate that DDS can effectively handle varying MEC scenarios, achieving up to 1.25$\times$ acceleration compared to the state-of-the-art algorithm.

cs.GT

Theoretical and computational analysis of the electrophoretic polymer mobility inversion induced by charge correlations

Electrophoretic (EP) mobility reversal is commonly observed for strongly charged macromolecules in multivalent salt solutions. This curious effect takes place, e.g., when a charged polymer, such as DNA, adsorbs excess counterions so that the counterion-dressed surface charge reverses its sign, leading to the inversion of the polymer drift driven by an external electric field. In order to characterize this seemingly counterintuitive phenomenon that cannot be captured by electrostatic mean-field theories, we adapt here a previously developed strong-coupling-dressed Poisson-Boltzmann approach to the cylindrical geometry of the polyelectrolyte-salt system. Within the framework of this formalism, we derive an analytical polymer mobility formula dressed by charge correlations. In qualitative agreement with polymer transport experiments, this mobility formula predicts that the increment of the monovalent salt, the decrease of the multivalent counterion valency, and the increase of the dielectric permittivity of the background solvent, suppress charge correlations and increase the multivalent bulk counterion concentration required for EP mobility reversal. These results are corroborated by coarse-grained molecular dynamics simulations showing how multivalent counterions induce mobility inversion at dilute concentrations and suppress the inversion effect at large concentrations. This re-entrant behavior, previously observed in the aggregation of like-charged polymer solutions, calls for verification by polymer transport experiments.

cond-mat.soft