Searcharxiv⌕ Search

arXiv subjects

Zhiyong Liu

Publications and source records attributed to Zhiyong Liu.

At least 19 recordsLinked to original sources

Anisotropic dynamical reconstruction of quantum geometry in quenched Chern insulators

Under unitary dynamics, the Chern number of an evolving quantum state remains conserved even when a quench drives the Hamiltonian across a topological phase transition. In sharp contrast, we reveal an anisotropic dynamical reconstruction of the quantum metric. Following a sudden quench in a two-dimensional Chern insulator, the metric develops a principal frame in which one eigenvalue grows in time whereas the other remains nearly unchanged. At long times, the associated axes align with the energy-gradient direction and the tangent direction of the constant-energy contours of the post-quench Hamiltonian, respectively. This dynamically selected frame is distinct from that of the static post-quench ground-state metric. We identify momentum-dependent relative dynamical phases as its origin: they enhance the distinguishability of neighboring states separated along the energy-gradient direction, while states along an equal-energy contour remain nearly phase locked. Consequently, local and global metric observables acquire characteristic long-time signatures of the post-quench Hamiltonian. Our results establish a nonequilibrium mechanism by which coherent dynamics reorganizes quantum geometry, suggesting new possibilities for its dynamical control.

quant-ph↗

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.

cs.CV↗

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation

Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Language Models (MLLMs) for vision-language reasoning and generate intermediate target cues (e.g., points or boxes) to guide a segmentation model. However, compressing rich reasoning into sparse cues often introduces ambiguity and noise, preventing these cues from accurately preserving the reasoning intent. While multiple complementary cues can enrich target information, existing methods typically feed them jointly into a single segmentation process, allowing ambiguous or erroneous cues to affect the entire prediction. Therefore, we propose DGSeg, a reasoning segmentation framework that learns to fuse predictions guided by semantic and spatial cues. Specifically, the MLLM jointly reasons about both target identity and spatial location, producing complementary semantic and spatial cues that are fed into separate segmentation branches. Their predictions are adaptively integrated by a lightweight dynamic gating module trained with relative branch-quality supervision to suppress noisy or conflicting regions. Extensive experiments demonstrate that DGSeg consistently outperforms strong baselines on multiple benchmarks and achieves 69.6% and 67.3% gIoU on the challenging ReasonSeg validation and test splits. Code is available at https://github.com/RZZeng/DGSeg.

cs.CV↗

SMR: Scheduler with Multi-Channel Map-Encoded Reinforcement Learning for Radio Telescopes

Observation scheduling for large single-dish radio telescopes is a multi-objective optimization problem: schedulers must maximize on-source scientific return under strict mechanical and environmental constraints. Previous dynamic scheduling relies on expert-designed heuristics, while existing reinforcement-learning (RL) approaches often struggle with variable-length target lists and lack an intrinsic representation of sky geometry. We present SMR (Scheduler with Map-encoded Reinforcement Learning), which projects discrete targets onto an azimuth--elevation (Az--El) grid in the local horizon frame. The resulting aligned multi-channel sky maps encode target attributes together with direction-dependent cues such as satellite-interference risk and elevation-dependent receiver gain. This representation provides an explicit spatial inductive bias and enables SMR to learn directly from the sky state. Simulations based on real catalogs and site parameters show that, compared with a tuned look-ahead greedy baseline, SMR achieves about a 10\% relative improvement in time utilization by learning non-myopic scheduling strategies. In the full three-channel setting, SMR further achieves joint trade-off among efficiency, interference avoidance, and observation quality, with up to 17\% higher LIER and 54\% higher HGOR relative to an MLP baseline while maintaining higher utilization across both 12 h and 24 h horizons. Overall, SMR provides a simple and extensible way for data-driven single-dish scheduling.

astro-ph.IM↗

Globally Localizing Lunar Rover in Pixels via Graph Alignment

Precise rover localization is a prerequisite for autonomous lunar exploration, yet the absence of Global Navigation Satellite System (GNSS) signals and the cumulative drift of local localization methods severely constrain long-range missions. Cross-view localization provides a promising drift-free global solution by matching rover-view and satellite-view imagery. However, the lunar environment poses unique challenges for correspondence alignment, including inter-entity entanglement, inter-viewpoint divergence, and simulation-to-real domain shift. To address these challenges, we propose Warped Alignment of Reprojected Graphs (WARG), a framework that leverages unified graph learning and reprojected graph matching for robust cross-view alignment. Pretrained on the synthetic LuSNAR dataset, WARG achieves an average test error of 0.32 m and demonstrates robust zero-shot generalization to the synthetic lunar south pole region with an error of 3.63 m. More importantly, when validated on real-world data from the YuTu-2 rover, WARG achieves a localization error of 1.68 m within a 100 m x 100 m search area, corresponding to nearly one-pixel precision in low-resolution satellite imagery with a spatial resolution of 1.40 m/pixel. Beyond accuracy, WARG is computationally efficient, containing only 1.56M parameters, corresponding to 16.12% of previous lightweight models, and operating at 5.49 Hz on an NVIDIA RTX A6000 GPU, approaching GNSS-level update frequency. Finally, we observe that WARG naturally develops low-level spatial awareness, including semantic segmentation and structural reasoning, through cross-view localization learning, highlighting its potential as a promising paradigm for spatial intelligence with minimal annotation cost. The source code is available at https://github.com/maochen-casia/warg.

cs.CV↗

From flat to narrow bands: Engineering quantum emission in a one-dimensional Lieb lattice

We develop a comprehensive theoretical framework that unifies quantum emission dynamics in one-dimensional Lieb lattices, bridging the gap between ideal flat-band coherence and realistic narrow-band dissipation. By coupling an emitter to sublattices with finite flat-band wavefunction overlap, we activate a collective, size-independent interaction fundamentally distinct from dispersive-band processes. Controllably breaking lattice symmetry transforms the flat band into a narrow dispersive band, enabling a continuous crossover from non-Markovian to Markovian dynamics governed by the competition between coupling strength and engineered bandwidth. Crucially, we derive explicit scaling laws that provide a quantitative blueprint for tuning spontaneous emission from coherent trapping to Markovian decay. Our work provides a unified framework that connects idealized flat-band physics to emerging narrow-band platforms such as moir$\rm\acute{e}$ photonic crystals, offering a practical toolkit for interpreting experiments and engineering quantum emission in structured photonic environments.

physics.optics↗

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real-world scenarios, where queries may correspond to multiple or no moments. Thus, we formulate Generalized Moment Retrieval (GMR), a unified setting that requires retrieving the complete set of relevant moments or predicting an empty set. To enable systematic study of GMR, we introduce Soccer-GMR, a large-scale benchmark built on challenging soccer videos that reflect general GMR scenarios, with realistic negative and positive queries. The benchmark is constructed via a duration-flexible semi-automated pipeline with human verification, enabling scalable data generation while maintaining high annotation quality. We further design a unified evaluation protocol with complementary metrics tailored for null-set rejection, positive-query localization, and end-to-end GMR performance. Finally, we establish strong baselines across two modeling paradigms: a lightweight plug-and-play GMR adapter for discriminative VMR models, and a GMR-tailored GRPO reward for fine-tuning multimodal large language models (MLLMs). Extensive experiments show consistent gains across all metrics and expose key limitations of current methods, positioning GMR as a more realistic and challenging benchmark for video-language understanding.

cs.CV↗

Fly360: Omnidirectional Obstacle Avoidance within Drone View

Obstacle avoidance in unmanned aerial vehicles (UAVs), as a fundamental capability, has gained increasing attention with the growing focus on spatial intelligence. However, current obstacle-avoidance methods mainly depend on limited field-of-view sensors and are ill-suited for UAV scenarios which require full-spatial awareness when the movement direction differs from the UAV's heading. This limitation motivates us to explore omnidirectional obstacle avoidance for panoramic drones with full-view perception. We first study an under explored problem setting in which a UAV must generate collision-free motion in environments with obstacles from arbitrary directions, and then construct a benchmark that consists of three representative flight tasks. Based on such settings, we propose Fly360, a two-stage perception-decision pipeline with a fixed random-yaw training strategy. At the perception stage, panoramic RGB observations are input and converted into depth maps as a robust intermediate representation. For the policy network, it is lightweight and used to output body-frame velocity commands from depth inputs. Extensive simulation and real-world experiments demonstrate that Fly360 achieves stable omnidirectional obstacle avoidance and outperforms forward-view baselines across all tasks. Our model is available at https://zxkai.github.io/fly360/

cs.RO↗

GA-Field: Geometry-Aware Vehicle Aerodynamic Field Prediction

Accurate aerodynamic field prediction is crucial for vehicle drag evaluation, but the computational cost of high-fidelity CFD hinders its use in iterative design workflows. While learning-based methods enable fast and scalable inference, accurately aerodynamic fields modeling remains challenging, as it demands capturing both long-range geometric effects and fine-scale flow structures. Existing approaches typically encode geometry only once at the input and formulate prediction as a one-shot mapping, which often leads to diluted global shape awareness and insufficient resolution of sharp local flow variations. To address these issues, we propose GA-Field, a Geometry-Aware Field prediction network that introduces two complementary design components: (i) a global geometry injection mechanism that repeatedly conditions the network on a compact 3D geometry embedding at multiple stages to preserve long-range geometric consistency, and (ii) a coarse-to-fine field refinement strategy to recover sharp local aerodynamic details. GA-Field achieves new state-of-the-art performance on ShapeNet-Car and the large-scale DrivAerNet++ benchmark for surface pressure, wall shear stress, and 3D velocity prediction tasks, while exhibiting strong out-of-distribution generalization across different vehicle categories.

cs.CE↗

FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting

Time series forecasting is central to data analysis and web technologies. The recent success of Large Language Models (LLMs) offers significant potential for this field, especially from the cross-modality aspect. Most methods adopt an LLM-as-Predictor paradigm, using LLM as the forecasting backbone and designing modality alignment mechanisms to enable LLM to understand time series data. However, the semantic information in the two modalities of time series and text differs significantly, making it challenging for LLM to fully understand time series data. To mitigate this challenge, our work follows an LLM-as-Enhancer paradigm to fully utilize the advantage of LLM in text understanding, where LLM is only used to encode text modality to complement time series modality. Based on this paradigm, we propose FiCoTS, an LLM-enhanced fine-to-coarse framework for multimodal time series forecasting. Specifically, the framework facilitates progressive cross-modality interaction by three levels in a fine-to-coarse scheme: First, in the token-level modality alignment module, a dynamic heterogeneous graph is constructed to filter noise and align time series patches with text tokens; Second, in the feature-level modality interaction module, a global cross-attention mechanism is introduced to enable each time series variable to connect with relevant textual contexts; Third, in the decision-level modality fusion module, we design a gated network to adaptively fuse the results of the two modalities for robust predictions. These three modules work synergistically to let the two modalities interact comprehensively across three semantic levels, enabling textual information to effectively support temporal prediction. Extensive experiments on seven real-world benchmarks demonstrate that our model achieves state-of-the-art performance. The codes will be released publicly.

cs.LG↗

An Interpretable AI Framework to Disentangle Self-Interacting and Cold Dark Matter in Galaxy Clusters: The CKAN Approach

Convolutional neural networks have shown their ability to differentiate between self-interacting dark matter (SIDM) and cold dark matter (CDM) on galaxy cluster scales. However, their large parameter counts and ''black-box'' nature make it difficult to assess whether their decisions adhere to physical principles. To address this issue, we have built a Convolutional Kolmogorov-Arnold Network (CKAN) that reduces parameter count and enhances interpretability, and propose a novel analytical framework to understand the network's decision-making process. With this framework, we leverage our network to qualitatively assess the offset between the dark matter distribution center and the galaxy cluster center, as well as the size of heating regions in different models. These findings are consistent with current theoretical predictions and show the reliability and interpretability of our network. By combining network interpretability with unseen test results, we also estimate that for SIDM in galaxy clusters, the minimum cross-section $(σ/m)_{\mathrm{th}}$ required to reliably identify its collisional nature falls between $0.1\,\mathrm{cm}^2/\mathrm{g}$ and $0.3\,\mathrm{cm}^2/\mathrm{g}$. Moreover, CKAN maintains robust performance under simulated JWST and Euclid noise, highlighting its promise for application to forthcoming observational surveys.

astro-ph.IM↗

Hierarchical Image Matching for UAV Absolute Visual Localization via Semantic and Structural Constraints

Absolute localization, aiming to determine an agent's location with respect to a global reference, is crucial for unmanned aerial vehicles (UAVs) in various applications, but it becomes challenging when global navigation satellite system (GNSS) signals are unavailable. Vision-based absolute localization methods, which locate the current view of the UAV in a reference satellite map to estimate its position, have become popular in GNSS-denied scenarios. However, existing methods mostly rely on traditional and low-level image matching, suffering from difficulties due to significant differences introduced by cross-source discrepancies and temporal variations. To overcome these limitations, in this paper, we introduce a hierarchical cross-source image matching method designed for UAV absolute localization, which integrates a semantic-aware and structure-constrained coarse matching module with a lightweight fine-grained matching module. Specifically, in the coarse matching module, semantic features derived from a vision foundation model first establish region-level correspondences under semantic and structural constraints. Then, the fine-grained matching module is applied to extract fine features and establish pixel-level correspondences. Building upon this, a UAV absolute visual localization pipeline is constructed without any reliance on relative localization techniques, mainly by employing an image retrieval module before the proposed hierarchical image matching modules. Experimental evaluations on public benchmark datasets and a newly introduced CS-UAV dataset demonstrate superior accuracy and robustness of the proposed method under various challenging conditions, confirming its effectiveness.

cs.CV↗

Boosting Open-Vocabulary Object Detection by Handling Background Samples

Open-vocabulary object detection is the task of accurately detecting objects from a candidate vocabulary list that includes both base and novel categories. Currently, numerous open-vocabulary detectors have achieved success by leveraging the impressive zero-shot capabilities of CLIP. However, we observe that CLIP models struggle to effectively handle background images (i.e. images without corresponding labels) due to their language-image learning methodology. This limitation results in suboptimal performance for open-vocabulary detectors that rely on CLIP when processing background samples. In this paper, we propose Background Information Representation for open-vocabulary Detector (BIRDet), a novel approach to address the limitations of CLIP in handling background samples. Specifically, we design Background Information Modeling (BIM) to replace the single, fixed background embedding in mainstream open-vocabulary detectors with dynamic scene information, and prompt it into image-related background representations. This method effectively enhances the ability to classify oversized regions as background. Besides, we introduce Partial Object Suppression (POS), an algorithm that utilizes the ratio of overlap area to address the issue of misclassifying partial regions as foreground. Experiments on OV-COCO and OV-LVIS benchmarks demonstrate that our proposed model is capable of achieving performance enhancements across various open-vocabulary detectors.

cs.CV↗

Meshfree method for solving the elliptic Monge-Ampere equation with Dirichlet boundary

We prove the convergence of meshfree method for solving the elliptic Monge-Ampere equation with Dirichlet boundary on the bounded domain. L2 error is obtained based on the kernel-based trial spaces generated by the compactly supported radial basis functions. We obtain the convergence result when the testing discretization is finer than the trial discretization. The convergence rate depend on the regularity of the solution, the smoothness of the computing domain, and the approximation of scaled kernel-based spaces. The presented convergence theory covers a wide range of kernel-based trial spaces including stationary approximation and non-stationary approximation. An extension to non-Dirichlet boundary condition is in a forthcoming paper.

math.NA↗

Precessing jet nozzle connecting to a spinning black hole in M87

The nearby radio galaxy M87 offers a unique opportunity to explore the connections between the central supermassive black hole and relativistic jets. Previous studies of the inner region of M87 revealed a wide opening angle for the jet originating near the black hole. The Event Horizon Telescope resolved the central radio source and found an asymmetric ring structure consistent with expectations from General Relativity. With a baseline of 17 years of observations, there was a shift in the jet's transverse position, possibly arising from an eight to ten-year quasi-periodicity. However, the origin of this sideways shift remains unclear. Here we report an analysis of radio observations over 22 years that suggests a period of about 11 years in the position angle variation of the jet. We infer that we are seeing a spinning black hole that induces the Lense-Thirring precession of a misaligned accretion disk. Similar jet precession may commonly occur in other active galactic nuclei but has been challenging to detect owing to the small magnitude and long period of the variation.

astro-ph.HE↗

The Qitai Radio Telescope

This study presents a general outline of the Qitai radio telescope (QTT) project. Qitai, the site of the telescope, is a county of Xinjiang Uygur Autonomous Region of China, located in the east Tianshan Mountains at an elevation of about 1800 m. The QTT is a fully steerable, Gregorian type telescope with a standard parabolic main reflector of 110 m diameter. The QTT has adopted an um-brella support, homology-symmetric lightweight design. The main reflector is active so that the deformation caused by gravity can be corrected. The structural design aims to ultimately allow high-sensitivity observations from 150 MHz up to 115 GHz. To satisfy the requirements for early scientific goals, the QTT will be equipped with ultra-wideband receivers and large field-of-view mul-ti-beam receivers. A multi-function signal-processing system based on RFSoC and GPU processor chips will be developed. These will enable the QTT to operate in pulsar, spectral line, continuum and Very Long Baseline Interferometer (VLBI) observing modes. Electromagnetic compatibility (EMC) and radio frequency interference (RFI) control techniques are adopted throughout the system design. The QTT will form a world-class observational platform for the detection of low-frequency (nanoHertz) gravitational waves through pulsar timing array (PTA) techniques, pulsar surveys, the discovery of binary black-hole systems, and exploring dark matter and the origin of life in the universe.

astro-ph.IM↗

Fractional Quantum Zeno Effect Emerging from Non-Hermitian Physics

Exploring non-Hermitian phenomenology is an exciting frontier of modern physics. However, the demonstration of a non-Hermitian phenomenon that is quantum in nature has remained elusive. Here, we predict quantum non-Hermitian phenomena: the fractional quantum Zeno (FQZ) effect and FQZ-induced photon antibunching. We consider a quantum optics platform with reservoir engineering, where nonlinear emitters are coupled to a bath of decaying bosonic modes whose own decay rates form band structures. By engineering the dissipation band, the spontaneous emission of emitters can be suppressed by strong dissipation through an algebraic scaling with fractional exponents - the FQZ effect. This fractional scaling originates uniquely from the divergent dissipative density of states near the dissipation band edge, different from the traditional closed-bath context. We find FQZ-induced strong photon antibunching in the steady state of a driven emitter even for weak nonlinearities. Remarkably, we identify that the sub-Poissonian quantum statistics of photons, which has no classical analogs, stems here from the key role of non-Hermiticity. Our setup is experimentally feasible with the techniques used to design lattice models with dissipative couplings.

quant-ph↗

AIGC Empowering Telecom Sector White Paper_chinese

In the global craze of GPT, people have deeply realized that AI, as a transformative technology and key force in economic and social development, will bring great leaps and breakthroughs to the global industry and profoundly influence the future world competition pattern. As the builder and operator of information and communication infrastructure, the telecom sector provides infrastructure support for the development of AI, and even takes the lead in the implementation of AI applications. How to enable the application of AIGC (GPT) and implement AIGC in the telecom sector are questions that telecom practitioners must ponder and answer. Through the study of GPT, a typical representative of AIGC, the authors have analyzed how GPT empowers the telecom sector in the form of scenarios, discussed the gap between the current GPT general model and telecom services, proposed for the first time a Telco Augmented Cognition capability system, provided answers to how to construct a telecom service GPT in the telecom sector, and carried out various practices. Our counterparts in the industry are expected to focus on collaborative innovation around telecom and AI, build an open and shared innovation ecosystem, promote the deep integration of AI and telecom sector, and accelerate the construction of next-generation information infrastructure, in an effort to facilitate the digital transformation of the economy and society.

cs.AI↗