SearcharxivSearch

arXiv subjects

Jing Zhong

Publications and source records attributed to Jing Zhong.

At least 19 recordsLinked to original sources

Drive by Hindsight and Foresight: Tool-Grounded Synergistic Reasoning over Hierarchical Memory for Autonomous Driving

VLMs have shown promise for autonomous driving, yet still suffer from hallucination, weak spatio-temporal perception, and limited generalization. Recent methods improve reasoning and decision-making through CoT explanations, retrieval-augmented generation or the static injection of tool outputs. Although these mechanisms enrich the context, the model neither proactively perceives scene information nor accumulates experience after answering. To overcome these limitations, we present, to our knowledge, the first synergistic framework that tightly couples hierarchical memory with proactive tool invocation in a closed reasoning loop. Our contributions are threefold. (i) Hierarchical Driving Memory: a scene-level short-term memory maintains the dynamic scene state, and an evolving long-term memory retrieves reusable experience and tool strategies. (ii) Memory-Tool Synergistic Reasoning Framework: guided by the scene state and retrieved experience, the model adaptively invokes tools to refine its reasoning at inference time and consolidates reusable experience into a long-term memory pool offline. (iii) Data Generation and Two-stage Training Pipeline: verified memory-tool trajectories built by multi-step teacher rollout are used to train with SFT and GRPO. Our 7B model reaches an overall reasoning score of 80.03 and MCQ accuracy of 79.09% on DriveLMM-o1, surpassing the strongest baseline by 7.74 MCQ points and generalizes strongly across benchmarks. Notably, ablation and analysis studies validate the effectiveness of each component and further reveal the complementary roles of hierarchical memory. Short-term memory strengthens spatio-temporal understanding, improving STSBench accuracy by 24.2 points, while offline long-term memory consolidation yields an additional 3.57-point MCQ gain with all parameters frozen, demonstrating continual self-evolution through accumulated driving experience.

cs.CV

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

Two structural insights have been overlooked in automated residential floor plan generation. First, design is inherently progressive. Architects begin with rough strokes and refine them over time, whereas existing methods typically require their conditioning representation to be fully specified before generation, a fundamental mismatch with how design actually works. Second, the 2D floor plan is not an optional intermediate but an irreplaceable spatial contract. Once room boundaries, doors, and windows are fixed, furnishing reduces from open-ended spatial reasoning to bounded constraint satisfaction. Bypassing this contract, as existing 3D systems do by delegating layout to language models, yields overlapping rooms and implausible proportions; directly calling general-purpose language models likewise produces geometrically invalid layouts. Guided by these insights, we present PlanCraft. SketchPlan supplies the missing training signal by replaying the architect's drawing process on 80K real floor plans, producing partial sketches at every completeness level. PlanCraft-Diff progressively sharpens an incomplete sketch into a geometrically precise, vectorizable floor plan through a coarse-to-fine strategy. With the spatial contract established, PlanCraft-Agent then furnishes the scene within well-defined room boundaries. Experiments show that PlanCraft achieves a 61.1\% lower FID than the best existing 2D method and surpasses existing 3D systems by 15 points in expert-rated spatial rationality, with a sketch at only 25\% completion already outperforming all fully specified baselines.

cs.CV

Revisiting the Radial Velocities of Nearby Open Clusters using Gaia DR3

Open clusters (OCs) are essential laboratories for probing stellar dynamics and tracing the structure and evolution of the Milky Way. Accurate measurements of their average radial velocities (RVs) and RV dispersions are crucial for estimating dynamical masses, orbital evolution, and overall kinematic states. Gaia Data Release 3 (DR3) provides an unprecedented volume of high-precision RVs. However, when applied to OCs, Gaia DR3 RVs often yield unusually large and overestimated RV dispersions. This inflation is primarily driven by RV measurement systematics for hot and faint stars, as well as unresolved binary contamination. To mitigate this, we revisit the average RVs and RV dispersions of OCs within 500 pc of the Sun using Gaia DR3. We evaluate the reliability of RV measurements and employ a color-based filtering method. By selecting member stars within an intrinsic color range of $0.2 \le (BP-RP)_{0} \le 1.2$ mag, we exclude hot and cool stars with less reliable RVs. This filtering significantly reduces the inferred RV dispersions while maintaining average RVs consistent with previous literature. Specifically, the median RV dispersion decreases by 26%, dropping from 3.76 km s$^{-1}$ prior to filtering to 2.79 km s$^{-1}$ afterward. These filtered RV dispersions remain systematically larger than tangential velocity dispersions, likely due to residual biases and undetected binaries. However, RV dispersions measured exclusively from red clump giants (found in 6 clusters) are remarkably small ($\lesssim 1.6$ km s$^{-1}$) and align closely with tangential dispersions ($\lesssim 1$ km s$^{-1}$). Ultimately, we provide a practical, Gaia-only strategy to derive more realistic RV dispersions for OCs, identifying red clump giants as exceptionally high-fidelity kinematic tracers for robust cluster studies.

astro-ph.GA

UrbanGraph: Physics-Informed Spatio-Temporal Dynamic Heterogeneous Graphs for Urban Microclimate Prediction

With rapid urbanization, predicting urban microclimates has become critical, as it affects building energy demand and public health risks. However, existing generative and homogeneous graph approaches fall short in capturing physical consistency, spatial dependencies, and temporal variability. To address this, we introduce UrbanGraph, a framework founded on a novel structure-based inductive bias. Unlike implicit graph learning, UrbanGraph transforms physical first principles into a dynamic causal topology, explicitly encoding time-varying causalities (e.g., shading and convection) directly into the graph structure to ensure physical consistency and data efficiency. Results show that UrbanGraph achieves state-of-the-art performance across all baselines. Specifically, the use of explicit causal pruning significantly reduces the model's floating-point operations (FLOPs) by 73.8% and increases training speed by 21% compared to implicit graphs. Our contribution includes the first high-resolution benchmark for spatio-temporal microclimate modeling, and a generalizable explicit topological encoding paradigm applicable to urban spatio-temporal dynamics governed by known physical equations.

cs.LG

Antarctic TianMu Staring Observation Project I: Overview and Implementation of the Prototype Telescope

Wide-field rapid sky surveys serve as critical observational methods for time-domain astronomical research. The Antarctic region, with several months of continuous dark nights annually, is an ideal site for time-domain astronomical observations. The Antarctic TianMu Staring Observation Project aims to deploy a fleet of small telescopes, adopting an array observation model to conduct time-domain optical observations in Antarctica, featuring wide-sky coverage, high-cadence sampling, long-period staring, and simultaneous multi-band measurements. Considering the severe challenges optical telescopes face in Antarctica, including extremely low temperatures, unattended operation, and limited power supply and network transmission, we have designed and developed the Antarctic TianMu prototype telescope based on drift-scan charge-coupled device technology. In October 2022, our prototype (with an aperture of 18 cm), named AT-Proto was transported to Zhongshan Station in Antarctica aboard China's 39th Antarctic Research Expedition. It has since operated stably and reliably in the frigid environment for over two years, demonstrating the significant advantages of this technology in polar astronomical observations. The experimental observation results of AT-Proto provide a solid foundation for the subsequent construction of a time-domain astronomy observation array in Antarctica.

astro-ph.IM

Antarctic TianMu Staring Observation Project II: Data reduction and preliminary results

The Antarctic TianMu Staring Observation Program is a time-domain optical sky survey project carried out in Antarctica, capable of large sky coverage, high-cadence sampling, and long-period staring. It utilizes the exceptional observing conditions in Antarctica to conduct high-cadence time-domain sky surveys. At present, we have successfully developed an 18-cm aperture Antarctic TianMu prototype, which has been deployed at Zhongshan Station in Antarctica for two consecutive years of trouble-free observations, during which more than 300,000 original images were obtained. This paper systematically outlines the commissioning data of the prototype telescope in 2023, the primary data processing pipeline, and the preliminary data products. The core pipeline encompasses four key stages: Data preprocessing, instrumental effect correction, astrometric solution, and full-field stellar photometry. Here, we release the 2023 data products, which specifically include reduced image data and a photometric catalog, for which, preliminary analyses demonstrate robust performance. Using Gaia Data Release 3 as a reference catalog, the astrometric precision, quantified by the root mean square of positional errors, is determined to be better than approximately 2 arcseconds, validating the observational capabilities of the system. For a 30-second exposure, the detection limit in the G-band is achieved at 15.00~mag, with a detection threshold of 1.5~$σ$. The photometric errors are below 0.1~mag for the majority of stars brighter than 14.00~mag. Furthermore, it improves significantly, reaching better than 0.01~mag for most stars brighter than 11.00~mag and 12.00~mag when employing the adaptive aperture photometry and point spread function photometry methods, respectively.

astro-ph.IM

MCI: Multi-Channel Imager on the Chinese Space Station Survey Telescope

The Multi-Channel Imager (MCI) is a powerful near-ultraviolet (NUV) and visible imager onboard the Chinese Space Station Survey Telescope (CSST). The MCI provides three imaging channels, which are the NUV channel, the Blue channel and the Red channel, with the wavelength range of 255-430 nm, 430-700 nm, and 700-1000 nm, respectively. MCI's three channels can target the same field simultaneously, which is unique compared to other imagers onboard the Hubble Space Telescope (HST) or the James Webb Space Telescope (JWST). Each channel employs a CCD focal plane of 9216 x 9232 pixels and $\sim$7\arcmin.5 x 7\arcmin.5 field of view (FOV), which are about $\gtrsim 4$ times greater than the FOVs of HST imagers. The MCI's three channels feature unprecedented sensitivities and field of views complement the NUV and visible capabilities of the CSST for high-precision photometry and weak-signal detection, which would help build a new standard-star system and the deepest UV-Optical exposures for CSST. Rich filter sets of MCI would help explore other sciences such as local emission line mapping, high-z Ly$α$ emitters searching, etc. Here we present key design features, results of current ground tests, and suggested observing strategies of the MCI.

astro-ph.IM

Understanding the Planetary Formation and Evolution in Star Clusters(UPiC)-II: Catalog of planets/candidates in Open Clusters and Moving Groups

Detecting planets in open clusters offers a unique opportunity to test planet formation theories in clustered environments. The precisely determined ages of young open clusters make their planets particularly valuable for tracing the early evolution of planetary systems. As the second paper of the UPiC project, this study focuses on stars in stellar groups that host transiting planets or planetary candidates. We categorize these stellar groups into Open Clusters (OCs) and Moving Groups (MGs) based on the Jacobi radius to investigate potential differences in their planetary systems. By cross-matching the latest star cluster catalogs with catalogs of transiting planets and candidates, we have compiled the most extensive catalog to date, containing 106 confirmed planets and 168 candidates within OCs and MGs. We refitted the structural parameters of these stellar groups and identified substructures using the \texttt{HDBSCAN} and Gaussian Mixture Model (GMM) algorithms. Our analysis reveals the density evolution of both MGs and OCs during their first Gyr. We find that MGs consistently exhibit a significantly higher planet fraction than OCs, regardless of sample selection, particularly for Hot Jupiters. Furthermore, exoplanet radii show a clear dichotomy at early stages: most sub-Jupiters evolve into Neptune-sized planets within 100 Myr, while super-Jupiters undergo only minimal contraction. These results suggest that young sub-Jupiters (\textless 100 Myr) represent puffy, Neptune-mass planets undergoing vigorous photoevaporation, whereas Jupiter-mass planets can maintain their atmospheres. We also report evidence for the early emergence of the hot-Neptune desert at 100 Myr in both OCs and MGs.

astro-ph.EP

GreenPlanner: Practical Floorplan Layout Generation via an Energy-Aware and Function-Feasible Generative Framework

Building design directly affects human well-being and carbon emissions, yet generating spatial-functional and energy-compliant floorplans remains manual, costly, and non-scalable. Existing methods produce visually plausible layouts but frequently violate key constraints, yielding invalid results due to the absence of automated evaluation. We present GreenPlanner, an energy- and functionality-aware generative framework that unifies design evaluation and generation. It consists of a labeled Design Feasibility Dataset for learning constraint priors; a fast Practical Design Evaluator (PDE) for predicting energy performance and spatial-functional validity; a Green Plan Dataset (GreenPD) derived from PDE-guided filtering to pair user requirements with regulation-compliant layouts; and a GreenFlow generator trained on GreenPD with PDE feedback for controllable, regulation-aware generation. Experiments show that GreenPlanner accelerates evaluation by over $10^{5}\times$ with $>$99% accuracy, eliminates invalid samples, and boosts design efficiency by 87% over professional architects.

cs.AI

Mock Observations for the CSST Mission: Multi-Channel Imager--Instrument Simulation

The Chinese Space Station Survey Telescope (CSST), a two-meter aperture astronomical space telescope under China's manned space program, is equipped with multiple back-end scientific instruments. As an astronomical precision measurement module of the CSST, the Multi-Channel Imager (MCI) can cover a wide wavelength range from ultraviolet to near-infrared with three-color simultaneous high-precision photometry and imaging, which meets the scientific requirements for various fields. The diverse scientific objectives of MCI require not only a robust airborne platform, advanced optical systems, and observing facilities but also comprehensive software support for scientific operations and research. To this end, it is essential to develop realistic observational simulation software to thoroughly evaluate the MCI data stream and provide calibration tools for future scientific investigations. The MCI instrument simulation software will serve as a foundation for the development of the MCI data processing pipeline and will facilitate improvements in both hardware and software, as well as in the observational operation strategy, in alignment with the mission's scientific goals. In conclusion, we present a comprehensive overview of the MCI instrument simulation and some corresponding performances of the MCI data processing pipeline.

astro-ph.IM

Mock Observations for the CSST Mission: Multi-Channel Imager--The Cluster Field

The Multi-Channel Imager (MCI), one of the instruments aboard the China Survey Space Telescope (CSST), is designed to simultaneously observe the sky in three filters, covering wavelengths from the near-ultraviolet (NUV) to the near-infrared (NIR). With its large field of view ($7.5^{\prime}\times7.5^{\prime}$), MCI is particularly well-suited for observing galaxy clusters, providing a powerful tool for investigating galaxy evolution, dark matter and dark energy through gravitational lensing. Here we present a comprehensive simulation framework of a strong lensing cluster as observed by MCI, aiming to fully exploit its capabilities in capturing lensing features. The framework simulates a strong lensing cluster from the CosmoDC2 catalog, calculating the gravitational potential and performing ray-tracing to derive the true positions, shapes and light distribution of galaxies within the cluster field. Additionally, the simulation incorporates intra-cluster light (ICL) and spectral energy distributions (SEDs), enabling further strong lensing analyses, such as ICL seperation from galaxy light and mass reconstruction combining strong and weak lensing measurements. This framework provides a critical benchmark for testing the MCI data pipeline and maximizing its potential in galaxy cluster research.

astro-ph.IM

A catalog of new blue stragglers in open clusters with Gaia DR3

The high-precision {\it Gaia} data release 3 (DR3) enables the discovery of numerous open clusters in the Milky Way, providing an excellent opportunity to search for blue straggler stars in open clusters and investigate their formation and evolution in these environments. Using the member stars from literature open cluster catalogs, we visually inspected the color-magnitude diagram (CMD) of each cluster and selected cluster candidates that potentially host blue stragglers. We then reassessed cluster memberships using the {\tt pyUPMASK} algorithm with {\it Gaia} DR3 and performed isochrone fitting to derive physical parameters for each cluster, including age, distance modulus, mean reddening, and metallicity. Finally, we empirically identified straggler stars based on their positions relative to the best-fitting isochrone, zero-age main sequence (ZAMS), and equal-mass binary sequence on the CMD. In total, we identified 272 new straggler stars in 99 open clusters, comprising 153 blue stragglers, 98 probable blue stragglers, and 21 yellow stragglers. Compared to the reported blue straggler catalogs based on earlier {\it Gaia} data, our results increase the number of open clusters with stragglers in the Milky Way by 22.2\%, and the total number of blue stragglers by 11.2\%.

astro-ph.SR

Binary clusters in the Galactic disk I: Systematic identification and classification using Gaia DR3

Aims. We aim to identify and classify BCs using high-precision astrometric and kinematic data, and to investigate their physical properties, mutual gravitational interactions, and formation rates. Methods. We used a comprehensive star cluster catalog that contains 4,084 high-quality clusters. Based on spatial and kinematic proximity, we identified 400 cluster pairs involving 686 unique clusters. These pairs were classified into three types: primordial BCs, systems formed through tidal capture or resonant trapping, and hyperbolic encounter pairs. For each system, we calculated the tidal factor to quantify the strength of mutual tidal interaction. Additionally, we constructed multi-cluster systems by identifying transitive connections among cluster pairs. Results. Among the 400 identified cluster pairs, nearly 60.8% (243 pairs) are probably primordial BCs, exhibiting both similar ages and motions. This supports a scenario where they formed together in the same giant molecular cloud. We find that 82.5% of the cluster pairs have strong mutual tidal forces. In addition, 278 star clusters are identified as members of 82 multi-cluster systems, including 27 newly reported groups. Cross-matching with the literature confirms the recovery of previously reported systems and leads to the discovery of 268 new cluster pairs. In our sample, about 16.8% of star clusters are involved in some type of interaction with another cluster, and 9.94% of star clusters are likely born in primordial BCs. Conclusions. Our results provide a comprehensive, homogeneously identified sample of Galactic BCs. The high fraction of primordial BCs and their mutual tidal interaction suggest that cluster formation in pairs is a main outcome of star formation. This work offers new observational constraints on the formation and dynamical evolution of multiple star cluster systems.

astro-ph.GA

UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models

Urban cultures and architectural styles vary significantly across cities due to geographical, chronological, historical, and socio-political factors. Understanding these differences is essential for anticipating how cities may evolve in the future. As representative cases of historical continuity and modern innovation in China, Beijing and Shenzhen offer valuable perspectives for exploring the transformation of urban streetscapes. However, conventional approaches to urban cultural studies often rely on expert interpretation and historical documentation, which are difficult to standardize across different contexts. To address this, we propose a multimodal research framework based on vision-language models, enabling automated and scalable analysis of urban streetscape style differences. This approach enhances the objectivity and data-driven nature of urban form research. The contributions of this study are as follows: First, we construct UrbanDiffBench, a curated dataset of urban streetscapes containing architectural images from different periods and regions. Second, we develop UrbanSense, the first vision-language-model-based framework for urban streetscape analysis, enabling the quantitative generation and comparison of urban style representations. Third, experimental results show that Over 80% of generated descriptions pass the t-test (p less than 0.05). High Phi scores (0.912 for cities, 0.833 for periods) from subjective evaluations confirm the method's ability to capture subtle stylistic differences. These results highlight the method's potential to quantify and interpret urban style evolution, offering a scientifically grounded lens for future design.

cs.CV

ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models

Architectural cultures across regions are characterized by stylistic diversity, shaped by historical, social, and technological contexts in addition to geograph-ical conditions. Understanding architectural styles requires the ability to describe and analyze the stylistic features of different architects from various regions through visual observations of architectural imagery. However, traditional studies of architectural culture have largely relied on subjective expert interpretations and historical literature reviews, often suffering from regional biases and limited ex-planatory scope. To address these challenges, this study proposes three core contributions: (1) We construct a professional architectural style dataset named ArchDiffBench, which comprises 1,765 high-quality architectural images and their corresponding style annotations, collected from different regions and historical periods. (2) We propose ArchiLense, an analytical framework grounded in Vision-Language Models and constructed using the ArchDiffBench dataset. By integrating ad-vanced computer vision techniques, deep learning, and machine learning algo-rithms, ArchiLense enables automatic recognition, comparison, and precise classi-fication of architectural imagery, producing descriptive language outputs that ar-ticulate stylistic differences. (3) Extensive evaluations show that ArchiLense achieves strong performance in architectural style recognition, with a 92.4% con-sistency rate with expert annotations and 84.5% classification accuracy, effec-tively capturing stylistic distinctions across images. The proposed approach transcends the subjectivity inherent in traditional analyses and offers a more objective and accurate perspective for comparative studies of architectural culture.

cs.CV

FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs

In the architectural design process, floorplan design is often a dynamic and iterative process. Architects progressively draw various parts of the floorplan according to their ideas and requirements, continuously adjusting and refining throughout the design process. Therefore, the ability to predict a complete floorplan from a partial one holds significant value in the design process. Such prediction can help architects quickly generate preliminary designs, improve design efficiency, and reduce the workload associated with repeated modifications. To address this need, we propose FloorplanMAE, a self-supervised learning framework for restoring incomplete floor plans into complete ones. First, we developed a floor plan reconstruction dataset, FloorplanNet, specifically trained on architectural floor plans. Secondly, we propose a floor plan reconstruction method based on Masked Autoencoders (MAE), which reconstructs missing parts by masking sections of the floor plan and training a lightweight Vision Transformer (ViT). We evaluated the reconstruction accuracy of FloorplanMAE and compared it with state-of-the-art benchmarks. Additionally, we validated the model using real sketches from the early stages of architectural design. Experimental results show that the FloorplanMAE model can generate high-quality complete floor plans from incomplete partial plans. This framework provides a scalable solution for floor plan generation, with broad application prospects.

cs.AI

Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance

In the context of the digital development of architecture, the automatic segmentation of walls and windows is a key step in improving the efficiency of building information models and computer-aided design. This study proposes an automatic segmentation model for building facade walls and windows based on multimodal semantic guidance, called Segment Any Architectural Facades (SAAF). First, SAAF has a multimodal semantic collaborative feature extraction mechanism. By combining natural language processing technology, it can fuse the semantic information in text descriptions with image features, enhancing the semantic understanding of building facade components. Second, we developed an end-to-end training framework that enables the model to autonomously learn the mapping relationship from text descriptions to image segmentation, reducing the influence of manual intervention on the segmentation results and improving the automation and robustness of the model. Finally, we conducted extensive experiments on multiple facade datasets. The segmentation results of SAAF outperformed existing methods in the mIoU metric, indicating that the SAAF model can maintain high-precision segmentation ability when faced with diverse datasets. Our model has made certain progress in improving the accuracy and generalization ability of the wall and window segmentation task. It is expected to provide a reference for the development of architectural computer vision technology and also explore new ideas and technical paths for the application of multimodal learning in the architectural field.

cs.CV

FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction

In the architectural design process, floor plan generation is inherently progressive and iterative. However, existing generative models for floor plans are predominantly end-to-end generation that produce an entire pixel-based layout in a single pass. This paradigm is often incompatible with the incremental workflows observed in real-world architectural practice. To address this issue, we draw inspiration from the autoregressive 'next token prediction' mechanism commonly used in large language models, and propose a novel 'next room prediction' paradigm tailored to architectural floor plan modeling. Experimental evaluation indicates that FPDS demonstrates competitive performance in comparison to diffusion models and Tell2Design in the text-to-floorplan task, indicating its potential applicability in supporting future intelligent architectural design.

cs.CL