SearcharxivSearch

arXiv subjects

Hao Tian

Publications and source records attributed to Hao Tian.

At least 19 recordsLinked to original sources

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding human activity and urban function only weakly. To address this limitation, we propose BEACON, a tri-modal contrastive learning framework that aligns three complementary views of urban space: physical representations from AE embeddings, semantic representations from point-of-interest (POI) text, and human behavioral representations from hourly POI visitation, while keeping the deployed representation image-only. Using the Houston Metropolitan Area as a case study area, we evaluated the performance of the BEACON framework on nine downstream tasks, including seven regression and two classification tasks against six baselines (raw coordinates, Space2Vec, SatCLIP, TESSERA, Clay and AlphaEarth), using frozen linear and MLP probes over five seeds. Under a linear probe, BEACON improves relative R^2 over AlphaEarth by up to 43% for obesity prevalence, 34% for poor mental health, and 22% for median household income, while remaining competitive in the prediction of physical and environmental variables. These findings highlight the value of augmenting geospatial foundation models with semantic and behavioral signals, extending their applicability from physical Earth observation to human-centered urban analytics.

cs.LG

Critical Microwave Mach-Zehnder-Type Interferometry with Dual-LO Rydberg Atoms

High-precision phase measurement of microwave fields underpins a wide range of applications, including wireless communications, distributed radar, plasma diagnostics, and antenna metrology. Existing Rydberg-atom-based approaches, however, often face trade-offs among phase resolution, measurement range, and system complexity. Here we demonstrate a Rydberg-atom-based microwave Mach-Zehnder-type interferometer using a dual-local-oscillator configuration. The two local oscillators establish two coherent interferometric pathways in the Rydberg medium. Their coherent mixing with the signal field produces an interferometric intermediate-frequency output governed by a phase-to-intensity transfer characteristic that enables critical-point enhancement. This scheme supports direct phase retrieval with a resolution exceeding $0.1^\circ$ and unambiguous full $360^\circ$ phase coverage with the reconfigurable dual-LO architecture. Moreover, near the critical interference point, the system exhibits a sharply enhanced phase-to-amplitude transduction, where weak amplitude variations are converted into pronounced phase responses, yielding a sensitivity enhancement exceeding 25 dB. Besides, the same interferometric transfer mechanism enables microwave propagation-distance and polarization metrology, achieving a propagation-distance precision below 20 $μ$m at 5.7 GHz together with a polarization-angle resolution exceeding $0.1^\circ$. This approach eliminates the need for complex optical configurations and lock-in detection, providing a simple, scalable, and reconfigurable Mach-Zehnder-type quantum microwave interferometry framework for multifunctional high-precision microwave metrology.

physics.atom-ph

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

cs.AI

Correlation between the two-armed $V_R$ spiral in the $Z$--$V_Z$ plane and moving groups

We use a cross-matched sample of 3.7 million stars from Gaia DR3 and LAMOST DR7 to investigate the velocity substructures in the Milky Way disk. The median radial velocity $V_R$ as a function of guiding-center radius $R_g$ exhibits alternating positive and negative stripes, which are strongly correlated with known moving groups. By examining the $V_R$ distribution in the $Z$--$V_Z$ phase space, we find that the D1, P2, D2, and P3 $V_R$ stripes display clear two-armed spirals. Among the moving groups embedded in these $V_R$ stripes, the Coma Berenices moving group in the P3 stripe exhibits the most pronounced two-armed spiral and serves as the primary contributor to the left arm of the overall P3 spiral. The angular momentum, eccentricities, orbital frequencies, and frequency ratios of its stars are consistent with either the corotation resonance of the spiral arms or the $m=4$ inner Lindblad resonance of the bar. Test-particle simulations confirm that a bar with a varying pattern speed, together with static or transient spiral arms, can produce such two-armed spirals.

astro-ph.GA

Long Tidal Tails of NGC 5024 Hidden in LMS-1 and NGC 5053 Tidal Streams

We report the discovery of long tidal tails associated with the globular cluster (GC) NGC 5024. A modified matched filter applied to Gaia DR3 data reveals a broad stellar stream spanning $α\approx 230^{\circ}-175^{\circ}$. The stellar stream overlaps on the sky with the LMS-1 and with the simulated stream of the GC NGC 5053, and all three share similar proper motions and metallicity. Our member candidates may therefore be a mixture of stars from NGC 5024, NGC 5053, and LMS-1. Nevertheless, the trailing tail extends roughly $20^{\circ}$ beyond the known LMS-1 stream. Furthermore, the radial velocity (RV) as a function of $α$ is used to distinguish the genuine stream candidates of NGC 5024 from the streams of NGC 5053 and LMS-1. Among the sources in common with DESI (Dark Energy Spectroscopic Instrument) DR1, we found two distinct sequences in RV-$α$ plane, corresponding to the stellar streams of NGC 5024 and NGC 5053 (or LMS-1), respectively. This constitutes the first strong evidence for the existence of extensive tidal tails around NGC 5024.

astro-ph.GA

Compressive Spectrum Sensing via Spectral Multiplexing in Rydberg Atomic Receiver

Rydberg-atomic receivers exhibit exceptional sensitivity yet are fundamentally constrained by the narrow instantaneous bandwidth, limiting their practical deployment in broadband scenarios. Prior approaches typically expand the bandwidth by physically broadening the atomic response, which usually requires auxiliary electromagnetic fields or stringent parameter tuning, thereby increasing overall system complexity. Here, we propose a compressive spectral multiplexing framework implemented in a waveguide-coupled Rydberg atomic receiver using a frequency-modulated local oscillator (FMLO). The FMLO creates multiple parallel sensing channels that collectively constitute a physical compressive sensing matrix, generating multiple narrowband intermediate-frequency replicas of the input signal. Thus, a broadband microwave spectrum is projected onto a set of narrowband atomic responses. It is demonstrated that spectral information spanning a bandwidth of over 640 MHz can be effectively compressed into the intrinsic atomic bandwidth of 126 kHz, achieving a spectrum compression ratio exceeding 1000. Furthermore, these output replicas offer intrinsic measurement redundancy and facilitate signal-to-noise ratio enhancement. An approximate 10 dB gain is achieved in the required bit-energy-to-noise-power-density ratio for multi-channel communication via maximal-ratio combining. This approach requires no auxiliary fields or broadband electronics, providing a simple and scalable pathway for chip-scale quantum receivers, latency-critical sensing, and next-generation wireless communications.

quant-ph

Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction

This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object priors. We follow the trend of dynamic 3D Gaussian Splatting based methods, and address several significant challenges. To model complex hand-object interaction with mutual occlusion and edge blur, we present interaction-aware hand-object Gaussians with newly introduced optimizable parameters aiming to adopt piecewise linear hypothesis for clearer structural representation. Moreover, considering the complementarity and tightness of hand shape and object shape during interaction dynamics, we incorporate hand information into object deformation field, constructing interaction-aware dynamic fields to model flexible motions. To further address difficulties in the optimization process, we propose a progressive strategy that handles dynamic regions and static background step by step. Correspondingly, explicit regularizations are designed to stabilize the hand-object representations for smooth motion transition, physical interaction reality, and coherent lighting. Experiments show that our approach surpasses existing dynamic 3D-GS-based methods and achieves state-of-the-art performance in reconstructing dynamic hand-object interaction.

cs.CV

DamageArbiter: A Multimodal Arbitration Framework for Disaster Damage Assessment from Street-View Imagery

Analyzing street-view imagery with computer vision models offers a promising approach for rapid, hyperlocal disaster damage assessment, but existing approaches typically rely on black-box pre-trained vision models, which lack interpretability and reliability. This study proposes DamageArbiter, a multimodal disagreement-driven arbitration framework designed to improve the accuracy and reliability of street-view-based damage assessment. DamageArbiter leverages the complementary strengths of unimodal and multimodal models and employs a lightweight logistic regression meta-classifier to arbitrate cases in which model predictions disagree. Using 2,556 post-disaster street-view images, paired with manually generated or large language model (LLM)-generated text descriptions, we systematically compared DamageArbiter with fine-tuned unimodal (image-only and text-only) models and CLIP-based multimodal models in terms of classification performance and overconfidence errors. Results show that DamageArbiter improved accuracy to 75.85% and the Matthews correlation coefficient (MCC) to 0.6188, compared with the best-performing text-only baseline (63.07% accuracy, 0.4126 MCC), image-only baseline (74.33% accuracy, 0.5947 MCC), and CLIP baseline (74.22% accuracy, 0.5915 MCC). The overconfidence analysis further reveals that DamageArbiter substantially reduced the overconfidence error from 70.58% for the best-performing baseline, the image-only ViT model, to 16.45%. Overall, this study demonstrates that accuracy alone is insufficient for evaluating disaster damage classification models and highlights the importance of measuring overconfidence errors as part of model reliability assessment. DamageArbiter thus offers a more reliable framework for rapid, hyperlocal disaster damage assessment from street-view imagery.

cs.CV

Tracing the kinematic perturbations of the Milky Way spiral arms with APOGEE DR17 and Gaia DR3

Aims. We constrain the dynamical perturbations of the spiral arms in the Milky Way disk, based on the non-axisymmetric streaming motions of RGB stars revealed by APOGEE and \textit{Gaia}. Methods. We develop a revised steady-state radial-velocity response model that incorporates both the \(V_{R,\sin}\) and the dynamically important \(V_{R,\cos}\) components for a two-armed logarithmic spiral potential. The model is validated using orbit integrations with \texttt{AGAMA} and Bayesian parameter recovery with \texttt{dynesty}, and is applied to the smoothed two-dimensional radial-velocity field of RGB stars while accounting for Lindblad and corotation resonances. Results. The revised model reproduces the phase and amplitude of the mock radial-velocity field to the \(\sim2\%\) level, substantially improving upon earlier \(V_{R,\sin}\)-only formulations. Applied to the observational data, it yields a robust pitch angle of \(p \simeq 10^\circ\) and a local surface density contrast of \(ξ\simeq 5\)--\(18\%\) at the solar radius. The radial scale length is less well-constrained (\(h_{R,1} \simeq 40\)--\(50\,\mathrm{kpc}\)) due to intrinsic parameter covariance. Resonance effects strongly shape the velocity field, thus affecting the fitting: the radial velocity becomes extremely large near the Lindblad resonances, whereas it vanishes close to the corotation resonance. Conclusions. Our results demonstrate that including both the \(V_{R,\sin}\) and \(V_{R,\cos}\) terms is essential for a physically consistent interpretation of stellar streaming motions induced by a spiral potential. The observed kinematics constrain the spiral pattern speed to \(Ω_{p} \approx 10\)--\(20\,\mathrm{km\,s}^{-1}\mathrm{kpc}^{-1}\).

astro-ph.GA

Piezoelectric resonators in thin-film barium titanate from room temperature to millikelvin

Ferroelectric materials, with their strong nonlinearities, underpin key technologies across radio-frequency (RF) signal processing, optical communications, and emerging quantum systems. Barium titanate (BTO) is a notable example, combining strong piezoelectric and electro-optic responses. While bulk BTO has been studied for decades, the piezoelectric properties of its recently available thin films, and their behavior at the millikelvin temperatures relevant to quantum hardware, remain largely unexplored. Here, we fabricate and characterize surface acoustic wave (SAW) resonators on thin-film BTO. The measured devices exhibit high electromechanical coupling (k2eff 0.14 at 5.2 GHz) and operate up to 7.8 GHz. From these measurements, combined with finite-element modeling of the multi-domain microstructure, we extract an effective piezoelectric coefficient d33eff of 53 pC/N, comparable to bulk BTO. Exploiting the intrinsic ferroelectricity, we further demonstrate low-voltage switching with a fast (100 ns) response, attractive for reconfigurable RF front-ends and parametric amplifiers. Extending these measurements to millikelvin temperatures, we find that the piezoelectric response persists, with d33eff 19 pC/N, pointing to the potential of BTO for piezoelectric coupling in superconducting quantum circuits. These results position thin-film BTO as a promising piezoelectric platform for both classical and quantum information technologies.

physics.app-ph

Ultra-broadband Anti-Jamming Communication via a Rydberg Atomic Receiver

Ultra-broadband anti-jamming communication represents a promising approach to secure and robust information transfer through spread-spectrum techniques, effectively combatting malicious interference and eavesdropping. Rydberg atoms, enhanced by waveguide coupling, facilitate ultra-broadband spectrum sensing without traditional RF components. This framework provides an experimental platform for ultra-wide anti-jamming communication. Here, we demonstrate real-time signal demodulation based on frequency-hopping spread spectrum (FHSS) in a waveguide-coupled Rydberg receiver, achieving ultra-broad frequency-hopping covering 100 kHz to 20 GHz and a hopping rate of 100 khop/s. When confined to a standard operational band (e.g., the 2.4 GHz ISM band), our system achieves a high channel density of 8 channels per MHz. Beyond this, by leveraging its ultra-broad and continuous bandwidth, the system supports over 150,000 channels. Experimental results reveal a 51 dB enhancement in narrowband interference tolerance compared with single-frequency systems, confirming its outstanding anti-jamming capability. The reported system demonstrates significant potential for secure communications based on quantum technology, especially communication in complex electromagnetic environments.

physics.atom-ph

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence. Here, overconfidence is defined behaviorally as the tendency to produce confident, assertive, and well-formatted outputs even when the underlying knowledge is incomplete or unverifiable, rather than as a calibration gap between stated confidence and accuracy. To examine this issue, we introduce GIScholarBench, a benchmark built from 10,865 papers published in 25 core GIScience journals between 2020 and 2025. The benchmark covers three tasks with increasing cognitive complexity: metadata retrieval, literature linking, and research direction generation. We evaluate Claude Sonnet 4.5, Gemini 3, and ChatGPT 5.3 through their native web interfaces under real-world user-facing conditions. Results show consistent overconfidence across all tasks. In metadata retrieval, ChatGPT 5.3 achieves the highest accuracy, but all models still generate definitive titles and DOIs when predictions are wrong. In literature linking, Claude Sonnet 4.5 recovers the most references, but all models show a clear gap between top-ranked retrieval and longer citation lists, suggesting that references are extended beyond reliable retrieval capacity. In research direction generation, AI-generated directions show lower topic coverage, higher novel miss rates, and lower semantic diversity than real future-citing papers. These findings suggest that LLM overconfidence is task-invariant but takes different forms: factual overgeneration in retrieval, unreliable citation expansion in literature linking, and overconfidence in output completeness during research ideation.

cs.IR

Multiple populations detection with the Chinese Space Station Survey Telescope main survey camera

Multiple stellar populations (MPs), characterized by star-to-star light-element abundance variations, are ubiquitous in globular clusters (GCs). Spectroscopy directly reveals these anomalies, while photometric studies, especially with the \textit{Hubble Space Telescope} (\textit{HST}), have been essential for tracing MP sequences in colour-magnitude diagrams (CMDs). However, the limited field of view of \textit{HST} confines most studies to cluster centres. The upcoming \textit{Chinese Space Station Survey Telescope} (CSST), with its wide field of view and UV-optical coverage, will enable systematic MP studies over entire clusters. We assess the capability of the CSST wide-field camera to detect and characterize MPs in GCs using realistic simulations. Synthetic stellar population models with different helium abundances ($ΔY$) and CNO variations were used to simulate CSST observations of GCs at distances of 9.6 and 20~kpc under different exposure times. MP detectability was evaluated using CMDs in seven CSST bands and UV-optical pseudo-colour diagrams. For a GC at 9.6~kpc, the $NUV-u$ colour is highly sensitive to $ΔY$ and CNO variations, with separations of $Δ(NUV-u)\approx0.16$ mag for red giants and up to 0.44 mag for dwarfs. MPs can be resolved when the total UV exposure exceeds $\sim1000$~s and the optical exposure exceeds $\sim300$~s. At 20~kpc, encompassing $\sim80\%$ of Galactic GCs, CSST still retains strong diagnostic power, resolving populations with $ΔY\geq0.06$ and $δ[\mathrm{N/Fe}]\geq0.64$, and separating MPs down to $i\sim19.5$ mag in clusters with large chemical spreads. The $NUV$-$u$-$g$ combination provides diagnostic performance comparable to the \textit{HST} F275W--F336W--F438W system. CSST will enable homogeneous MP surveys across the full spatial extent of star clusters in the Milky Way and nearby galaxies.

astro-ph.SR

Stellar Density Classification and Regression for CSST Multi-color Imaging Using Deep Learning

The Chinese Space Station Survey Telescope (CSST) aims to map the universe across an unprecedented dynamic range of stellar densities, spanning from extragalactic voids to the crowded Galactic center (e.g. a few stars and galaxies in the voids and $>10^5$ stars per detector in Galactic center). However, processing such heterogeneous data with a general source extraction pipeline introduces significant systematic uncertainties, standard algorithms exhibit poor accuracy in crowded fields and suffer from increased astrometric uncertainty in void regions. To mitigate these systematics, we propose a hierarchical, two-stage deep learning model for adaptive data reduction. The first stage ('classification') employs a ResNet-34 model to classify images into six discrete density categories, achieving $98.83\%$ in global accuracy. This classification acts as a critical decision gate, ensuring high calibration accuracy in the crowded fields. In the second stage ('regression'), a ResNet-50 regression model predicts the bright stars ($<23.5$ mag) in the field, which is essential for astrometric calibration, achieving a mean absolute error (MAE) of 0.0824 dex. By decoupling density characterization from source extraction, our model ensures that photometric and astrometric algorithms are optimally matched to the stellar density environment, thereby enhancing the fidelity and homogeneity of CSST as well as future large sky survey data products.

astro-ph.IM

Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarcity of large-scale training data spanning diverse real-world applications. Existing datasets rely heavily on costly manual annotations and are typically confined to narrow domains. To address this challenge, we propose Video2GUI, a fully automated framework that extracts grounded GUI interaction trajectories directly from unlabeled Internet videos. Video2GUI employs a coarse-to-fine filtering strategy to identify high-quality GUI tutorial videos and convert them into structured agent trajectories. Applying this pipeline to 500 million video metadata entries, we construct WildGUI, a large-scale dataset containing 12 million interaction trajectories spanning over 1,500 applications and websites. Pre-training Qwen2.5-VL and Mimo-VL on WildGUI yields consistent improvements of 5-20% across multiple GUI grounding and action benchmarks, matching or surpassing state-of-the-art performance. We will release both the WildGUI dataset and the Video2GUI pipeline to support future research of GUI agents.

cs.CL

Comparative analysis of missing data imputation methods for CSST survey: Impact on photometric redshift estimation performance

Improving the accuracy of photometric redshifts (photo-$z$) is essential for reliable statistical studies of cosmology and galaxy evolution. However, missing photometric bands are a common observational challenge that can significantly degrade photo-$z$ estimation accuracy. In this work, we present a systematic evaluation of data imputation methods aimed at improving photo-$z$ performance. We benchmark a range of representative machine learning (ML) and deep learning (DL) architectures, identifying k-nearest neighbors (KNN) and the attention-based SAITS model as the leading performers. These models are then applied to China Space Station Survey Telescope (CSST) mock data to assess their performance under realistic observational conditions. Our results show that KNN yields the highest accuracy under idealized missing completely at random (MCAR) conditions with complete training sets, whereas robustness tests reveal that SAITS significantly outperforms KNN when training data is incomplete or when applied to realistic mixed-mechanism scenarios. We find that domain consistency between training and testing missingness patterns is a prerequisite for optimal performance, highlighting the risks of domain shift in supervised regression tasks. Furthermore, our analysis demonstrates that while general imputation models are highly effective for MCAR and missing at random (MAR) data, they are detrimental when applied to missing not at random (MNAR) data arising from flux limits, as statistical models fail to capture the physical information inherent in these non-detections. Consequently, we advocate for more sophisticated architectures capable of disentangling stochastic missingness from physical non-detections to address these distinct mechanisms individually.

astro-ph.GA

MiMo-Embodied: X-Embodied Foundation Model Technical Report

We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied AI. MiMo-Embodied sets new records across 17 embodied AI benchmarks in Task Planning, Affordance Prediction and Spatial Understanding, while also excelling in 12 autonomous driving benchmarks across Environmental Perception, Status Prediction, and Driving Planning. Across these tasks, MiMo-Embodied significantly outperforms existing open-source, closed-source, and specialized baselines. Our results indicate that through multi-stage learning, curated data construction, and CoT/RL fine-tuning, these two domains exhibit strong positive transfer and mutually reinforce one another. We provide a detailed analysis of our model design and training methodologies to facilitate further research. Code and models are available at https://github.com/XiaomiMiMo/MiMo-Embodied.

cs.RO

GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly visually ground with human-like sophistication, or are they merely pattern-matching on simplified datasets? Current benchmarks fail to capture real-world complexity where humans effortlessly navigate intricate references and recognize when grounding is impossible. To rigorously assess MLLMs' true capabilities, we introduce GroundingME, a benchmark that systematically challenges models across four critical dimensions: (1) Discriminative: distinguishing highly similar objects, (2) Spatial: understanding complex relational descriptions, (3) Limited: handling occlusions or tiny objects, and (4) Rejection: recognizing ungroundable queries. Through careful curation combining automated generation with human verification, we create 1,005 challenging examples mirroring real-world complexity. Evaluating 25 state-of-the-art MLLMs reveals a profound capability gap: the best model achieves only 45.1% accuracy, while most score 0% on rejection tasks. We explore two strategies for improvements: (1) test-time scaling selects optimal response by thinking trajectory to improve overall performance by up to 4.5%, and (2) data-mixture training boosts rejection accuracy from 0% to 27.9%. GroundingME thus serves as both a diagnostic tool revealing current limitations in MLLMs and a roadmap toward human-level visual grounding. Project page: https://groundingme.github.io

cs.CV