SearcharxivSearch

arXiv subjects

Xiaoqing Xu

Publications and source records attributed to Xiaoqing Xu.

At least 19 recordsLinked to original sources

Asynchronous Parallel Search for Exact Multi-Objective Shortest Paths with Versioned Frontier Snapshots and Indexed Dominance Pruning

Exact multi-objective shortest-path (MOSP) search computes the complete Pareto set between specified start and goal vertices, and its computational cost can grow rapidly with expanding nondominated label sets and frequent dominance tests over per-vertex Pareto frontiers. Efficiently parallelizing exact MOSP remains an open challenge. This paper presents SIP-MOSP (Snapshot-based Indexed-Pruning MOSP), an asynchronous exact framework that separates label expansion from frontier maintenance within a single cooperative search. SIP-MOSP combines immutable versioned frontier snapshots with indexed dominance pruning, enabling concurrent label processing without concurrent access to the same mutable frontier. Together, these mechanisms reduce synchronization overhead and accelerate dominance testing. We instantiate the framework with block-minimum (SIP-MOSP-BM) and segment-tree-minimum (SIP-MOSP-ST) indices and prove exactness. We evaluate both variants against four state-of-the-art exact MOSP baselines covering sequential and parallel search. Experiments across multiple objective dimensions on a road network, an Internet service provider topology, and an 180-vertex complete directed graph show that SIP-MOSP achieves speedups of up to 46.9* over the best-performing sequential baseline and up to 7.05* over the best-performing parallel baseline on mutually solved instances. In the 20-objective complete-graph setting, where many instances remain unsolved by the sequential baselines within one hour, SIP-MOSP-ST achieves a 3.34* speedup while reducing peak memory by a factor of 60.3 relative to the best-performing parallel baseline. These results demonstrate that SIP-MOSP is an efficient shared-memory framework for exact MOSP across structurally diverse graph topologies.

cs.DC

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of how general multimodal models can be developed into an industrial-grade foundation model for content ecosystems. The model adopts a compact InternViT-300M + MLP + Qwen3 1.7B architecture, balancing fine-grained visual perception, language-semantic alignment, and deployment cost within an approximately 2B-parameter budget. To balance business specialization with the retention of general capabilities, we developed a data iteration and curation mechanism and trained the model through a progressive three-stage pipeline: pre-training, mid-training, and post-training. Ablation studies and offline business evaluations show that Xuanwu VL-2B achieves an average score of 67.90 across seven OpenCompass multimodal metrics (vs. 64.27 for InternVL 3.5 2B), an average recall of 94.38% over seven independent business moderation tasks, and a weighted overall recall of 82.82% on policy-violating text in challenging adversarial OCR scenarios, outperforming Gemini-2.5-Pro (76.72%). These results show that, under a limited parameter budget, Xuanwu VL-2B achieves a practical balance among business alignment, visual perception, general capability retention, and deployment cost.

cs.AI

Democratizing planetary-scale analysis: An ultra-lightweight Earth embedding database for accurate and flexible global land monitoring

The rapid evolution of satellite-borne Earth Observation (EO) systems has revolutionized terrestrial monitoring, yielding petabyte-scale archives. However, the immense computational and storage requirements for global-scale analysis often preclude widespread use, hindering planetary-scale studies. To address these barriers, we present Embedded Seamless Data (ESD), an ultra-lightweight, 30-m global Earth embedding database spanning the 25-year period from 2000 to 2024. By transforming high-dimensional, multi-sensor observations from the Landsat series (5, 7, 8, and 9) and MODIS Terra into information-dense, quantized latent vectors, ESD distills essential geophysical and semantic features into a unified latent space. Utilizing the ESDNet architecture and Finite Scalar Quantization (FSQ), the dataset achieves a transformative ~340-fold reduction in data volume compared to raw archives. This compression allows the entire global land surface for a single year to be encapsulated within approximately 2.4 TB, enabling decadal-scale global analysis on standard local workstations. Rigorous validation demonstrates high reconstructive fidelity (MAE: 0.0130; RMSE: 0.0179; CC: 0.8543). By condensing the annual phenological cycle into 12 temporal steps, the embeddings provide inherent denoising and a semantically organized space that outperforms raw reflectance in land-cover classification, achieving 79.74% accuracy (vs. 76.92% for raw fusion). With robust few-shot learning capabilities and longitudinal consistency, ESD provides a versatile foundation for democratizing planetary-scale research and advancing next-generation geospatial artificial intelligence.

cs.CV

Wetland mapping from sparse annotations with satellite image time series and temporal-aware segment anything model

Accurate wetland mapping is essential for ecosystem monitoring, yet dense pixel-level annotation is prohibitively expensive and practical applications usually rely on sparse point labels, under which existing deep learning models perform poorly, while strong seasonal and inter-annual wetland dynamics further render single-date imagery inadequate and lead to significant mapping errors; although foundation models such as SAM show promising generalization from point prompts, they are inherently designed for static images and fail to model temporal information, resulting in fragmented masks in heterogeneous wetlands. To overcome these limitations, we propose WetSAM, a SAM-based framework that integrates satellite image time series for wetland mapping from sparse point supervision through a dual-branch design, where a temporally prompted branch extends SAM with hierarchical adapters and dynamic temporal aggregation to disentangle wetland characteristics from phenological variability, and a spatial branch employs a temporally constrained region-growing strategy to generate reliable dense pseudo-labels, while a bidirectional consistency regularization jointly optimizes both branches. Extensive experiments across eight global regions of approximately 5,000 km2 each demonstrate that WetSAM substantially outperforms state-of-the-art methods, achieving an average F1-score of 85.58%, and delivering accurate and structurally consistent wetland segmentation with minimal labeling effort, highlighting its strong generalization capability and potential for scalable, low-cost, high-resolution wetland mapping.

cs.CV

Subgraph Extraction-based Feedback-guided Iterative Scheduling for HLS

This paper proposes ISDC, a novel feedback-guided iterative system of difference constraints (SDC) scheduling algorithm for high-level synthesis (HLS). ISDC leverages subgraph extraction-based low-level feedback from downstream tools like logic synthesizers to iteratively refine HLS scheduling. Technical innovations include: (1) An enhanced SDC formulation that effectively integrates low-level feedback into the linear-programming (LP) problem; (2) A fanout and window-based subgraph extraction mechanism driving the feedback cycle; (3) A no-human-in-loop ISDC flow compatible with a wide range of downstream tools and process design kits (PDKs). Evaluation shows that ISDC reduces register usage by 28.5% against an industrial-strength open-source HLS tool.

cs.CL

How to hide your voice: Noise-cancelling bird photography blind

Getting close to birds is a great challenge in wildlife photography. Bird photography blinds may be the most effective and least intrusive way if properly designed. However, the acoustic design of the blinds has been overlooked so far. Herein, we present noise-cancelling blinds which allow photographing birds at close range. Firstly, we conduct a questionnaire in the eco-tourism centre located in Yunnan, China. Thus, we determine the birders' expectations of the indoor sound environment. We then identify diverse variables to examine the impact of architectural and acoustic decisions on noise propagation. Finally, we examine the acoustic performance of the blinds by considering the birds' hearing threshold. The numerical simulations are performed in the acoustics module of Comsol MultiPhysics. Our study demonstrated that photography blinds require a strong and thorough acoustic design for both human and bird well-being.

cs.HC

VisDrone-CC2020: The Vision Meets Drone Crowd Counting Challenge Results

Crowd counting on the drone platform is an interesting topic in computer vision, which brings new challenges such as small object inference, background clutter and wide viewpoint. However, there are few algorithms focusing on crowd counting on the drone-captured data due to the lack of comprehensive datasets. To this end, we collect a large-scale dataset and organize the Vision Meets Drone Crowd Counting Challenge (VisDrone-CC2020) in conjunction with the 16th European Conference on Computer Vision (ECCV 2020) to promote the developments in the related fields. The collected dataset is formed by $3,360$ images, including $2,460$ images for training, and $900$ images for testing. Specifically, we manually annotate persons with points in each video frame. There are $14$ algorithms from $15$ institutes submitted to the VisDrone-CC2020 Challenge. We provide a detailed analysis of the evaluation results and conclude the challenge. More information can be found at the website: \url{http://www.aiskyeye.com/}.

cs.CV

Net2: A Graph Attention Network Method Customized for Pre-Placement Net Length Estimation

Net length is a key proxy metric for optimizing timing and power across various stages of a standard digital design flow. However, the bulk of net length information is not available until cell placement, and hence it is a significant challenge to explicitly consider net length optimization in design stages prior to placement, such as logic synthesis. This work addresses this challenge by proposing a graph attention network method with customization, called Net2, to estimate individual net length before cell placement. Its accuracy-oriented version Net2a achieves about 15% better accuracy than several previous works in identifying both long nets and long critical paths. Its fast version Net2f is more than 1000 times faster than placement while still outperforms previous works and other neural network techniques in terms of various accuracy metrics.

cs.LG

Fast IR Drop Estimation with Machine Learning

IR drop constraint is a fundamental requirement enforced in almost all chip designs. However, its evaluation takes a long time, and mitigation techniques for fixing violations may require numerous iterations. As such, fast and accurate IR drop prediction becomes critical for reducing design turnaround time. Recently, machine learning (ML) techniques have been actively studied for fast IR drop estimation due to their promise and success in many fields. These studies target at various design stages with different emphasis, and accordingly, different ML algorithms are adopted and customized. This paper provides a review to the latest progress in ML-based IR drop estimation techniques. It also serves as a vehicle for discussing some general challenges faced by ML applications in electronics design automation (EDA), and demonstrating how to integrate ML models with conventional techniques for the better efficiency of EDA tools.

cs.LG

Thermal Analysis of a 3D Stacked High-Performance Commercial Microprocessor using Face-to-Face Wafer Bonding Technology

3D integration technologies are seeing widespread adoption in the semiconductor industry to offset the limitations and slowdown of two-dimensional scaling. High-density 3D integration techniques such as face-to-face wafer bonding with sub-10 $μ$m pitch can enable new ways of designing SoCs using all 3 dimensions, like folding a microprocessor design across multiple 3D tiers. However, overlapping thermal hotspots can be a challenge in such 3D stacked designs due to a general increase in power density. In this work, we perform a thorough thermal simulation study on sign-off quality physical design implementation of a state-of-the-art, high-performance, out-of-order microprocessor on a 7nm process technology. The physical design of the microprocessor is partitioned and implemented in a 2-tier, 3D stacked configuration with logic blocks and memory instances in separate tiers (logic-over-memory 3D). The thermal simulation model was calibrated to temperature measurement data from a high-performance, CPU-based 2D SoC chip fabricated on the same 7nm process technology. Thermal profiles of different 3D configurations under various workload conditions are simulated and compared. We find that stacking microprocessor designs in 3D without considering thermal implications can result in maximum die temperature up to 12°C higher than their 2D counterparts under the worst-case power-indicative workload. This increase in temperature would reduce the amount of time for which a power-intensive workload can be run before throttling is required. However, logic-over-memory partitioned 3D CPU implementation can mitigate this temperature increase by half, which makes the temperature of the 3D design only 6$^\circ$C higher than the 2D baseline. We conclude that using thermal aware design partitioning and improved cooling techniques can overcome the thermal challenges associated with 3D stacking.

cs.AR

Stack up your chips: Betting on 3D integration to augment Moore's Law scaling

3D integration, i.e., stacking of integrated circuit layers using parallel or sequential processing is gaining rapid industry adoption with the slowdown of Moore's law scaling. 3D stacking promises potential gains in performance, power and cost but the actual magnitude of gains varies depending on end-application, technology choices and design. In this talk, we will discuss some key challenges associated with 3D design and how design-for-3D will require us to break traditional silos of micro-architecture, circuit/physical design and manufacturing technology to work across abstractions to enable the gains promised by 3D technologies.

cs.AR

Analysis of the Mobility-Limiting Mechanisms of the Two-Dimensional Hole Gas on Hydrogen-Terminated Diamond

Here we present an analysis of the mobility-limiting mechanisms of a two-dimensional hole gas on hydrogen-terminated diamond surfaces. The scattering rates of surface impurities, surface roughness, non-polar optical phonons, and acoustic phonons are included. Using a Schrodinger/Poisson solver, the heavy hole, light hole, and split-off bands are treated separately. To compare the calculations with experimental data, Hall-effect structures were fabricated and measured at temperatures ranging from 25 to 700 K, with hole sheet densities ranging from 2 to 6$\times10^{12}\;\text{cm}^{-2}$ and typical mobilities measured from 60 to 100 cm$^{2}$/(V$\cdot$s) at room temperature. Existing data from literature was also used, which spans sheet densities above 1$\times10^{13}\;\text{cm}^{-2}$. Our analysis indicates that for low sheet densities, surface impurity scattering by charged acceptors and surface roughness are not sufficient to account for the low mobility. Moreover, the experimental data suggests that long-range potential fluctuations exist at the diamond surface, and are particularly enhanced at lower sheet densities. Thus, we propose a second type of surface impurity scattering which is caused by disorder related to the C-H dipoles.

physics.app-ph

Temperature Dependent Thermal Boundary Conductance of Monolayer MoS$_2$ by Raman Thermometry

The electrical and thermal behavior of nanoscale devices based on two-dimensional (2D) materials is often limited by their contacts and interfaces. Here we report the temperature-dependent thermal boundary conductance (TBC) of monolayer MoS$_2$ with AlN and SiO$_2$, using Raman thermometry with laser-induced heating. The temperature-dependent optical absorption of the 2D material is crucial in such experiments, which we characterize here for the first time above room temperature. We obtain TBC ~ 15 MWm$^-$$^2$K$^-$$^1$ near room temperature, increasing as ~ T$^0$$^.$$^6$$^5$ in the range 300 - 600 K. The similar TBC of MoS$_2$ with the two substrates indicates that MoS$_2$ is the "softer" material with weaker phonon irradiance, and the relatively low TBC signifies that such interfaces present a key bottleneck in energy dissipation from 2D devices. Our approach is needed to correctly perform Raman thermometry of 2D materials, and our findings are key for understanding energy coupling at the nanoscale.

cond-mat.mes-hall

Standard Cell Library Design and Optimization Methodology for ASAP7 PDK

Standard cell libraries are the foundation for the entire backend design and optimization flow in modern application-specific integrated circuit designs. At 7nm technology node and beyond, standard cell library design and optimization is becoming increasingly difficult due to extremely complex design constraints, as described in the ASAP7 process design kit (PDK). Notable complexities include discrete transistor sizing due to FinFETs, complicated design rules from lithography and restrictive layout space from modern standard cell architectures. The design methodology presented in this paper enables efficient and high-quality standard cell library design and optimization with the ASAP7 PDK. The key techniques include exhaustive transistor sizing for cell timing optimization, transistor placement with generalized Euler paths and back-end design prototyping for library-level explorations.

cs.AR

Imaging of objects through a thin scattering layer using a spectrally and spatially separated reference

Incoherently illuminated or luminescent objects give rise to a low-contrast speckle-like pattern when observed through a thin diffusive medium, as such a medium effectively convolves their shape with a speckle-like point spread function (PSF). This point spread function can be extracted in the presence of a reference object of known shape. Here it is shown that reference objects that are both spatially and spectrally separated from the object of interest can be used to obtain an approximation of the point spread function. The crucial observation, corroborated by analytical calculations, is that the spectrally shifted point spread function is strongly correlated to a spatially scaled one. With the approximate point spread function thus obtained, the speckle-like pattern is deconvolved to produce a clear and sharp image of the object on a speckle-like background of low intensity.

physics.optics

Tuning Electrical and Thermal Transport in AlGaN/GaN Heterostructures via Buffer Layer Engineering

Over the last decade, progress in wide bandgap, III-V materials systems based on gallium nitride (GaN) has been a major driver in the realization of high power and high frequency electronic devices. Since the highly conductive, two-dimensional electron gas (2DEG) at the AlGaN/GaN interface is based on built-in polarization fields (not doping) and is confined to very small thicknesses, its charge carriers exhibit much higher mobilities in comparison to their doped counterparts. In this study, we show that this heterostructured material also offers the unique ability to manipulate electrical transport separately from thermal transport through the examination of fully-suspended AlGaN/GaN diaphragms of varied GaN buffer layer thicknesses. Notably, we show that ~$100$ nm thin GaN layers can considerably impede heat flow without electrical transport degradation, and that a significant improvement (~$4$x) in the thermoelectric figure of merit ($\it zT$) over externally doped GaN is observed in 2DEG based heterostructures. We also observe state-of-the art thermoelectric power factors ($4-7\times$ $10^{-3}$$\,Wm^{-1}K^{-2}$) at room temperature) in the 2DEG of this material system. This remarkable tuning behavior and thermoelectric enhancement, elucidated here for the first time in a polarization-based heterostructure, is achieved since the electrons are at the heterostructured interface, while the phonons are within the material system. These results highlight the potential for using the 2DEG in III-V materials for on-chip thermal sensing and energy harvesting.

cond-mat.mtrl-sci

3D Object Imaging through Scattering Media

Human ability to visualize an image is usually hindered by optical scattering. Recent extensive studies have promoted imaging technique through turbid materials to a reality where color image can be restored behind scattering media in real time. The big challenge now is to recover a 3D object in a large field of view with depth resolving ability. Here, we reveal a new physical relationship between speckles generated from objects at different planes. With a single given point spread function, 3D imaging through scattering media is achieved even beyond the depth of field (DOF). Experimental testing of standard scattering media shows that the original DOF can be extended up to 5 times and the physical mechanism is depicted. This extended 3D imaging is expected to have important applications in science, technology, bio-medical, security and defense.

physics.optics

Imaging objects through scattering layers and around corners by retrieval of the scattered point spread function

We demonstrate a high-speed method to image objects through a thin scattering medium and around a corner. The method employs a reference object of known shape to retrieve the speckle-like point spread function of the scatterer. We extract the point spread function of the scatterer from a dynamic scene that includes a static reference object, and use this to image the dynamic objects. Sharp images are reconstructed from the transmission through a diffuser and from reflection off a rough surface. The sharp and clean reconstructed images from single shot data exemplify the robustness of the method.

physics.optics