SearcharxivSearch

arXiv subjects

Yuan Hu

Publications and source records attributed to Yuan Hu.

At least 19 recordsLinked to original sources

CoCoNav: Conformal Control for Safe Robot Navigation in Crowds

Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors. Existing reactive methods can produce oscillatory behavior, while predictive planners often treat forecasts as exact or rely on restrictive error models. Incorporating conservative uncertainty sets as hard constraints can also render model predictive control (MPC) infeasible. We propose \textit{CoCoNav}, a crowd-navigation framework that combines online conformal calibration with runtime-certified planning. A horizon-specific conformal proportional--integral controller adapts trajectory-error bounds to regulate long-run empirical coverage, enabling the framework to respond to changing prediction errors. A \textit{relax-then-verify} planner preserves solver feasibility by generating nominal trajectories with soft-constrained MPC and separately certifying them, together with contingency maneuvers, against the calibrated bounds before execution. Simulations and quadruped experiments show that CoCoNav achieves a favorable balance among collision avoidance, task success, and navigation efficiency relative to the evaluated baselines.

cs.RO

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

Unified remote sensing multimodal models exhibit a pronounced spatial reversal curse: Although they can accurately recognize and describe object locations in images, they often fail to faithfully execute the same spatial relations during text-to-image generation, where such relations constitute core semantic information in remote sensing. Motivated by this observation, we propose Uni-RS, the first unified multimodal model tailored for remote sensing, to explicitly address the spatial asymmetry between understanding and generation. Specifically, we first introduce explicit Spatial-Layout Planning to transform textual instructions into spatial layout plans, decoupling geometric planning from visual synthesis. We then impose Spatial-Aware Query Supervision to bias learnable queries toward spatial relations explicitly specified in the instruction. Finally, we develop Image-Caption Spatial Layout Variation to expose the model to systematic geometry-consistent spatial transformations. Extensive experiments across multiple benchmarks show that our approach substantially improves spatial faithfulness in text-to-image generation, while maintaining strong performance on multimodal understanding tasks like image captioning, visual grounding, and VQA tasks.

cs.CV

ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval

With the rise of multimodal learning, image retrieval plays a crucial role in connecting visual information with natural language queries. Existing image retrievers struggle with processing long texts and handling unclear user expressions. To address these issues, we introduce the conversational query rewriting (CQR) task into the image retrieval domain and construct a dedicated multi-turn dialogue query rewriting dataset. Built on full dialogue histories, CQR rewrites users' final queries into concise, semantically complete ones that are better suited for retrieval. Specifically, We first leverage Large Language Models (LLMs) to generate rewritten candidates at scale and employ an LLM-as-Judge mechanism combined with manual review to curate approximately 7,000 high-quality multimodal dialogues, forming the ReCQR dataset. Then We benchmark several SOTA multimodal models on the ReCQR dataset to assess their performance on image retrieval. Experimental results demonstrate that CQR not only significantly enhances the accuracy of traditional image retrieval models, but also provides new directions and insights for modeling user queries in multimodal systems.

cs.IR

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visual question answering, and visual grounding. However, existing RS MLLMs lack the pixel-level dialogue capability, which involves responding to user instructions with segmentation masks for specific instances. In this paper, we propose GeoPix, a RS MLLM that extends image understanding capabilities to the pixel level. This is achieved by equipping the MLLM with a mask predictor, which transforms visual features from the vision encoder into masks conditioned on the LLM's segmentation token embeddings. To facilitate the segmentation of multi-scale objects in RS imagery, a class-wise learnable memory module is integrated into the mask predictor to capture and store class-wise geo-context at the instance level across the entire dataset. In addition, to address the absence of large-scale datasets for training pixel-level RS MLLMs, we construct the GeoPixInstruct dataset, comprising 65,463 images and 140,412 instances, with each instance annotated with text descriptions, bounding boxes, and masks. Furthermore, we develop a two-stage training strategy to balance the distinct requirements of text generation and masks prediction in multi-modal multi-task optimization. Extensive experiments verify the effectiveness and superiority of GeoPix in pixel-level segmentation tasks, while also maintaining competitive performance in image- and region-level benchmarks.

cs.CV

A Universal Method to Transform Aromatic Hydrocarbon Molecules into Confined Carbyne inside Single-Walled Carbon Nanotubes

Carbyne, a sp1-hybridized allotrope of carbon, is a linear carbon chain with exceptional theoretically predicted properties that surpass those of sp2-hybridized graphene and carbon nanotubes (CNTs). However, the existence of carbyne has been debated due to its instability caused by Peierls distortion, which limits its practical development. The only successful synthesis of carbyne has been achieved inside CNTs, resulting in a form known as confined carbyne (CC). However, CC can only be synthesized inside multi-walled CNTs, limiting its property-tuning capabilities to the inner tubes of the CNTs. Here, we present a universal method for synthesizing CC inside single-walled carbon nanotubes (SWCNTs) with diameter of 0.9-1.3 nm. Aromatic hydrocarbon molecules are filled inside SWCNTs and subsequently transformed into CC under low-temperature annealing. A variety of aromatic hydrocarbon molecules are confirmed as effective precursors for formation of CC, with Raman frequencies centered around 1861 cm-1. Enriched (6,5) and (7,6) SWCNTs with diameters less than 0.8 nm are less effective than the SWCNTs with diameter of 0.9-1.3 nm for CC formation. Furthermore, resonance Raman spectroscopy reveals that optical band gap of the CC at 1861 cm-1 is 2.353 eV, which is consistent with the result obtained using a linear relationship between the Raman signal and optical band gap. This newly developed approach provides a versatile route for synthesizing CC from various precursor molecules inside diverse templates, which is not limited to SWCNTs but could extend to any templates with appropriate size, including molecular sieves, zeolites, boron nitride nanotubes, and metal-organic frameworks.

cond-mat.mtrl-sci

An Empirical Implementation of the Shadow Riskless Rate

We address the problem of asset pricing in a market where there is no risky asset. Previous work developed a theoretical model for a shadow riskless rate (SRR) for such a market in terms of the drift component of the state-price deflator for that asset universe. Assuming asset prices are modeled by correlated geometric Brownian motion, in this work we develop a computational approach to estimate the SRR from empirical datasets. The approach employs: principal component analysis to model the effects of the individual Brownian motions; singular value decomposition to capture the abrupt changes in condition number of the linear system whose solution provides the SRR values; and a regularization to control the rate of change of the condition number. Among other uses (e.g., for option pricing, developing a term structure of interest rate), the SRR can be employed as an investment discriminator between asset classes. We apply the computational procedure to markets consisting of groups of stocks, varying asset type and number. The theoretical and computational analysis provides not only the drift, but also the total volatility of the state-price deflator. We investigate the time trajectory of these two descriptive components of the state-price deflator for the empirical datasets.

q-fin.MF

Hall effect on the joint cascades of magnetic energy and helicity in helical magnetohydrodynamic turbulence

Helical magnetohydrodynamic turbulence with Hall effects is ubiquitous in heliophysics and plasma physics, such as star formation and solar activities, and its intrinsic mechanisms are still not clearly explained. Direct numerical simulations reveal that when the forcing scale is comparable to the ion inertial scale, Hall effects induce remarkable cross helicity. It then suppresses the inverse cascade efficiency, leading to the accumulation of large-scale magnetic energy and helicity. The process is accompanied by the breaking of current sheets via filaments along magnetic fields. Using the Ulysses data, the numerical findings are separately confirmed. These results suggest a novel mechanism wherein small-scale Hall effects could strongly affect large-scale magnetic fields through cross helicity.

physics.plasm-ph

Vision-Language Models in Remote Sensing: Current Progress and Future Trends

The remarkable achievements of ChatGPT and GPT-4 have sparked a wave of interest and research in the field of large language models for Artificial General Intelligence (AGI). These models provide intelligent solutions close to human thinking, enabling us to use general artificial intelligence to solve problems in various applications. However, in remote sensing (RS), the scientific literature on the implementation of AGI remains relatively scant. Existing AI-related research in remote sensing primarily focuses on visual understanding tasks while neglecting the semantic understanding of the objects and their relationships. This is where vision-language models excel, as they enable reasoning about images and their associated textual descriptions, allowing for a deeper understanding of the underlying semantics. Vision-language models can go beyond visual recognition of RS images, model semantic relationships, and generate natural language descriptions of the image. This makes them better suited for tasks requiring visual and textual understanding, such as image captioning, and visual question answering. This paper provides a comprehensive review of the research on vision-language models in remote sensing, summarizing the latest progress, highlighting challenges, and identifying potential research opportunities.

cs.CV

Unifying Market Microstructure and Dynamic Asset Pricing

We introduce a discrete binary tree for pricing contingent claims with the underlying security prices exhibiting history dependence characteristic of that induced by market microstructure phenomena. Example dependencies considered include moving average or autoregressive behavior. Our model is market-complete, arbitrage-free, and preserves all of the parameters governing the historical (natural world) price dynamics when passing to an equivalent martingale (risk-neutral) measure. Specifically, this includes the instantaneous mean and variance of the asset return and the instantaneous probabilities for the direction of asset price movement. We believe this is the first paper to demonstrate the ability to include market microstructure effects in dynamic asset/option pricing in a market-complete, no-arbitrage, format.

q-fin.MF

The Implied Views of Bond Traders on the Spot Equity Market

This study delves into the temporal dynamics within the equity market through the lens of bond traders. Recognizing that the riskless interest rate fluctuates over time, we leverage the Black-Derman-Toy model to trace its temporal evolution. To gain insights from a bond trader's perspective, we focus on a specific type of bond: the zero-coupon bond. This paper introduces a pricing algorithm for this bond and presents a formula that can be used to ascertain its real value. By crafting an equation that juxtaposes the theoretical value of a zero-coupon bond with its actual value, we can deduce the risk-neutral probability. It is noteworthy that the risk-neutral probability correlates with variables like the instantaneous mean return, instantaneous volatility, and inherent upturn probability in the equity market. Examining these relationships enables us to discern the temporal shifts in these parameters. Our findings suggest that the mean starts at a negative value, eventually plateauing at a consistent level. The volatility, on the other hand, initially has a minimal positive value, peaks swiftly, and then stabilizes. Lastly, the upturn probability is initially significantly high, plunges rapidly, and ultimately reaches equilibrium.

q-fin.TR

A rigorous model reduction for the anisotropic-scattering transport process

In this letter, we propose a reduced-order model to bridge the particle transport mechanics and the macroscopic fluid dynamics in the highly scattered regime. A rigorous mathematical derivation and a concise physical interpretation are presented for an anisotropic-scattering transport process with arbitrary order of scattering kernel. The prediction of the theoretical model perfectly agrees with the numerical experiments. A clear picture of the diffusion physics is revealed for the neutral particle transport in the asymptotic optically thick regime.

math-ph

Viewpoint Integration and Registration with Vision Language Foundation Model for Image Change Understanding

Recently, the development of pre-trained vision language foundation models (VLFMs) has led to remarkable performance in many tasks. However, these models tend to have strong single-image understanding capability but lack the ability to understand multiple images. Therefore, they cannot be directly applied to cope with image change understanding (ICU), which requires models to capture actual changes between multiple images and describe them in language. In this paper, we discover that existing VLFMs perform poorly when applied directly to ICU because of the following problems: (1) VLFMs generally learn the global representation of a single image, while ICU requires capturing nuances between multiple images. (2) The ICU performance of VLFMs is significantly affected by viewpoint variations, which is caused by the altered relationships between objects when viewpoint changes. To address these problems, we propose a Viewpoint Integration and Registration method. Concretely, we introduce a fused adapter image encoder that fine-tunes pre-trained encoders by inserting designed trainable adapters and fused adapters, to effectively capture nuances between image pairs. Additionally, a viewpoint registration flow and a semantic emphasizing module are designed to reduce the performance degradation caused by viewpoint variations in the visual and semantic space, respectively. Experimental results on CLEVR-Change and Spot-the-Diff demonstrate that our method achieves state-of-the-art performance in all metrics.

cs.CV

RSGPT: A Remote Sensing Vision Language Model and Benchmark

The emergence of large-scale large language models, with GPT-4 as a prominent example, has significantly propelled the rapid advancement of artificial general intelligence and sparked the revolution of Artificial Intelligence 2.0. In the realm of remote sensing (RS), there is a growing interest in developing large vision language models (VLMs) specifically tailored for data analysis in this domain. However, current research predominantly revolves around visual recognition tasks, lacking comprehensive, large-scale image-text datasets that are aligned and suitable for training large VLMs, which poses significant challenges to effectively training such models for RS applications. In computer vision, recent research has demonstrated that fine-tuning large vision language models on small-scale, high-quality datasets can yield impressive performance in visual and language understanding. These results are comparable to state-of-the-art VLMs trained from scratch on massive amounts of data, such as GPT-4. Inspired by this captivating idea, in this work, we build a high-quality Remote Sensing Image Captioning dataset (RSICap) that facilitates the development of large VLMs in the RS field. Unlike previous RS datasets that either employ model-generated captions or short descriptions, RSICap comprises 2,585 human-annotated captions with rich and high-quality information. This dataset offers detailed descriptions for each image, encompassing scene descriptions (e.g., residential area, airport, or farmland) as well as object information (e.g., color, shape, quantity, absolute position, etc). To facilitate the evaluation of VLMs in the field of RS, we also provide a benchmark evaluation dataset called RSIEval. This dataset consists of human-annotated captions and visual question-answer pairs, allowing for a comprehensive assessment of VLMs in the context of RS.

cs.CV

SwinRDM: Integrate SwinRNN with Diffusion Model towards High-Resolution and High-Quality Weather Forecasting

Data-driven medium-range weather forecasting has attracted much attention in recent years. However, the forecasting accuracy at high resolution is unsatisfactory currently. Pursuing high-resolution and high-quality weather forecasting, we develop a data-driven model SwinRDM which integrates an improved version of SwinRNN with a diffusion model. SwinRDM performs predictions at 0.25-degree resolution and achieves superior forecasting accuracy to IFS (Integrated Forecast System), the state-of-the-art operational NWP model, on representative atmospheric variables including 500 hPa geopotential (Z500), 850 hPa temperature (T850), 2-m temperature (T2M), and total precipitation (TP), at lead times of up to 5 days. We propose to leverage a two-step strategy to achieve high-resolution predictions at 0.25-degree considering the trade-off between computation memory and forecasting accuracy. Recurrent predictions for future atmospheric fields are firstly performed at 1.40625-degree resolution, and then a diffusion-based super-resolution model is leveraged to recover the high spatial resolution and finer-scale atmospheric details. SwinRDM pushes forward the performance and potential of data-driven models for a large margin towards operational applications.

cs.AI

Option pricing using a skew random walk pricing tree

Motivated by the Corns-Satchell, continuous time, option pricing model, we develop a binary tree pricing model with underlying asset price dynamics following Itô-Mckean skew Brownian motion. While the Corns-Satchell market model is incomplete, our discrete time market model is defined in the natural world; extended to the risk neutral world under the no-arbitrage condition where derivatives are priced under uniquely determined risk-neutral probabilities; and is complete. The skewness introduced in the natural world is preserved in the risk neutral world. Furthermore, we show that the model preserves skewness under the continuous-time limit. We provide numerical applications of our model to the valuation of European put and call options on exchange-traded funds tracking the S&P Global 1200 index.

q-fin.MF

A unified gas-kinetic particle method for frequency-dependent radiative transfer equations with isotropic scattering process on unstructured mesh

In this paper, we extend the unified kinetic particle (UGKP) method to the frequency-dependent radiative transfer equation with both absorption-emission and scattering processes. The extended UGKP method could not only capture the diffusion and free transport limit, but also provide a smooth transition in the physical and frequency space in the regime between the above two limits. The proposed scheme has the properties of asymptotic-preserving, regime-adaptive, and entropy-preserving, which make it an accurate and efficient scheme in the simulation of multiscale photon transport problems. The methodology of scheme construction is a coupled evolution of macroscopic energy equation and the microscopic radiant intensity equation, where the numerical flux in macroscopic energy equation and the closure in microscopic radiant intensity equation are constructed based on the integral solution. Both numerical dissipation and computational complexity are well controlled especially in the optical thick regime. A 2D multi-thread code on a general unstructured mesh has been developed. Several numerical tests have been simulated to verify the numerical scheme and code, covering a wide range of flow regimes. The numerical scheme and code that we developed are highly demanded and widely applicable in the high energy density engineering applications.

physics.comp-ph

Microwave heating as a universal method to transform confined molecules into armchair graphene nanoribbons

Armchair graphene nanoribbons (AGNRs) with sub-nanometer width are potential materials for fabrication of novel nanodevices thanks to their moderate direct band gaps. AGNRs are usually synthesized by polymerizing precursor molecules on substrate surface. However, it is time-consuming and not suitable for large-scale production. AGNRs can also be grown by transforming precursor molecules inside single-walled carbon nanotubes via furnace annealing, but the obtained AGNRs are normally twisted. In this work, microwave heating is applied for transforming precursor molecules into AGNRs. The fast heating process allows synthesizing the AGNRs in seconds. Several different molecules were successfully transformed into AGNRs, suggesting that it is a universal method. More importantly, as demonstrated by Raman spectroscopy, aberration-corrected high-resolution transmission electron microscopy and theoretical calculations, less twisted AGNRs are synthesized by the microwave heating than the furnace annealing. Our results reveal a route for rapid production of AGNRs in large scale, which would benefit future applications in novel AGNRs-based semiconductor devices.

cond-mat.mtrl-sci

PolyBuilding: Polygon Transformer for End-to-End Building Extraction

We present PolyBuilding, a fully end-to-end polygon Transformer for building extraction. PolyBuilding direct predicts vector representation of buildings from remote sensing images. It builds upon an encoder-decoder transformer architecture and simultaneously outputs building bounding boxes and polygons. Given a set of polygon queries, the model learns the relations among them and encodes context information from the image to predict the final set of building polygons with fixed vertex numbers. Corner classification is performed to distinguish the building corners from the sampled points, which can be used to remove redundant vertices along the building walls during inference. A 1-d non-maximum suppression (NMS) is further applied to reduce vertex redundancy near the building corners. With the refinement operations, polygons with regular shapes and low complexity can be effectively obtained. Comprehensive experiments are conducted on the CrowdAI dataset. Quantitative and qualitative results show that our approach outperforms prior polygonal building extraction methods by a large margin. It also achieves a new state-of-the-art in terms of pixel-level coverage, instance-level precision and recall, and geometry-level properties (including contour regularity and polygon complexity).

cs.CV