SearcharxivSearch

arXiv subjects

Qiong Liu

Publications and source records attributed to Qiong Liu.

At least 19 recordsLinked to original sources

ByteAction: Byte-space Action Recognition Foundation Model

Byte-space Action Recognition (BAR) aims to recognize human actions directly from compressed image bitstreams without any pixel decoding. By operating entirely in byte space, BAR is inherently independent of file integrity and pixel-level reconstruction, making it naturally applicable to privacy-sensitive scenarios and robust against bitstream corruption. In this paper, we propose ByteAction, a BAR foundation model that achieves accurate action recognition on corrupted image bitstreams. ByteAction follows a dual-view byte-level recognition framework. It constructs weakly and strongly corrupted bitstream views, which are augmented by Bitstream Pattern Augmentation (BPA) and encoded with a shared ByteFormer backbone. The model is optimized with both classification and corruption consistency objectives. Specifically, we propose Bitstream Pattern Augmentation (BPA), which reshapes one-dimensional byte sequences into two-dimensional byte matrix and applies region-level erasure to encourage the model to learn robust cross-region byte dependencies. We further propose a Corruption Consistency Training strategy that constrains the model to maintain stable predictions across different corruption severities through bidirectional KL divergence. Experiments on the image bitstream from Stanford40, PPMI, and PASCAL VOC 2012 Action demonstrate that ByteAction achieves state-of-the-art corruption robustness across all scenarios while maintaining competitive intact bitstream performance.

cs.CV

Bitstream Action Recognition is Byte Modeling

Conventional action recognition typically relies on successful pixel decoding of the bitstream. However, bitstream corruption during storage or transmission may cause severe visual artifacts or even decoding failure, posing a significant challenge to reliable action recognition. Bitstream Action Recognition (BAR) aims to overcome the dependency on decoding and the vulnerability to corruption. In this paper, we propose a novel BAR framework, Bitstream Recognition via Anchoring Corrupted Embeddings (BRACE). BRACE is a dual-branch byte-modeling architecture that treats a corrupted bitstream and its intact counterpart as two byte realizations of the same action. This guides the generation of rich and stable representations for robustness to corruption through Intact-Anchored Representation Alignment (IARA). The intact representation serves as a stable anchor, and the corrupted one is aligned to it at the embedding and decision levels under Unreliable-Anchor Suppression (UAS), entirely in representation space and without repairing the bitstream. To address the scarcity of corrupted bitstreams in practice, we introduce the Real-world Bitstream Corruption Simulator (RBCS), a four-parameter simulator that reproduces bit-flip and byte-loss errors arising in transmission and storage. Building on RBCS, we construct the first large-scale BAR dataset (BAR-D), which comprises the BAR-Stanford40 and BAR-PPMI subsets and spans diverse corruption types and severity levels. Finally, we build a large benchmark on BAR-D involving 14 action recognition methods from the pixel, compressed, and bitstream domains. Extensive experiments demonstrate that BRACE has superior robustness to bitstream corruption than all comparison methods. Ablation studies further validate the effectiveness of the proposed RBCS augmentation and IARA.

cs.CV

Angle-I2P: Angle-Consistent-Aware Hierarchical Attention for Cross-Modality Outlier Rejection

Image-to-point-cloud registration (I2P) is a fundamental task in robotic applications such as manipulation,grasping, and localization. Existing deep learning-based I2P methods seek to align image and point cloud features in a learned representation space to establish correspondences, and have achieved promising results. However, when the inlier ratio of the initial matching pairs is low, conventional Perspective-n-Points (PnP) methods may struggle to achieve accurate results. To address this limitation, we propose Angle-I2P, an outlier rejection network that leverages angle-consistent geometric constraints and hierarchical attention. First, we design a scale-invariant, crossmodality geometric constraint based on angular consistency. This explicit geometric constraint guides the model in distinguishing inliers from outliers. Furthermore, we propose a global-tolocal hierarchical attention mechanism that effectively filters out geometrically inconsistent matches under rigid transformation, thereby improving the Inlier Ratio (IR) and Registration Recall (RR). Experimental results demonstrate that our method achieves state-of-the-art performance on the 7Scenes, RGBD Scenes V2, and a self-collected dataset, with consistent improvements across all benchmarks.

cs.CV

White Dwarfs with Infrared Excess from DESI EDR

Infrared (IR) excess emission around white dwarfs (WDs) is commonly attributed to circumstellar debris disks and/or low-mass companions, providing a unique window into the evolution of planetary systems and binary evolution after the main-sequence stage. Based on a spectroscopically confirmed WD sample from the DESI Early Data Release, we performed a systematic search for IR excess by combining multi-band photometry from SDSS, Pan-STARRS, UKIDSS, 2MASS, and WISE. Using spectral energy distribution (SED) fitting, we initially identified 72 IR-excess candidates and conducted a stringent contamination assessment based on higher-resolution imaging within 6 arcseconds of each target. After removing sources affected by blending or source confusion, we obtained a final sample of 62 reliable IR excess candidates. Among them, we identify three candidate WD+M dwarf binaries (two new systems), five candidate WD+brown dwarf (BD) binaries (all new), 38 candidate WD+dust disks (28 new), and 16 ambiguous systems that could be either WD+BD or WD+dust (15 new). Compared with previous samples, our catalog extends the parameter space of known dusty WDs toward older cooling ages. Due to the limited spatial resolution of WISE, follow-up high-resolution imaging and/or infrared spectroscopy is required to confirm the physical nature of all candidate systems and to further expand the parameter space of dust disks in terms of cooling age and other properties.

astro-ph.SR

Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition

Multi-modality of color and depth, i.e., RGB-D, is of great importance in recent research of indoor scene recognition. In this kind of data representation, depth map is able to describe the 3D structure of scenes and geometric relations among objects. Previous works showed that local features of both modalities are vital for promotion of recognition accuracy. However, the problem of adaptive selection and effective exploitation on these key local features remains open in this field. In this paper, a dynamic graph model is proposed with adaptive node selection mechanism to solve the above problem. In this model, a dynamic graph is built up to model the relations among objects and scene, and a method of adaptive node selection is proposed to take key local features from both modalities of RGB and depth for graph modeling. After that, these nodes are grouped by three different levels, representing near or far relations among objects. Moreover, the graph model is updated dynamically according to attention weights. Finally, the updated and optimized features of RGB and depth modalities are fused together for indoor scene recognition. Experiments are performed on public datasets including SUN RGB-D and NYU Depth v2. Extensive results demonstrate that our method has superior performance when comparing to state-of-the-arts methods, and show that the proposed method is able to exploit crucial local features from both modalities of RGB and depth.

cs.CV

White Dwarfs with Infrared Excess from LAMOST Data Release 11

Infrared (IR) excess observed around white dwarfs (WDs) is typically attributed to companions or debris disks. These systems are interesting because they offer a unique opportunity to study the late stages of stellar evolution and the interactions between WDs and surrounding material. The 11th data release (DR11) of the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) - one of the largest spectroscopic surveys to date - has recently provided spectra for 3092 WDs, many of which have yet to be systematically investigated for IR excess. In this study, we cross-correlated the LAMOST DR11 WD catalog with optical and IR surveys, including the Sloan Digital Sky Survey (SDSS), Two Micron All-Sky Survey (2MASS), UKIRT Infrared Deep Sky Survey (UKIDSS), and Wide-field Infrared Survey Explorer (WISE). We performed spectral energy distribution fitting using the VOSA tool for 1818 WDs and identified 167 IR excess WD candidates. After excluding 23 sources with potential contamination within 6" and five additional sources identified through WISE ccf flag analysis, we identified 139 objects with candidate IR excess. These include 30 candidate WD + M dwarf binaries (18 new systems), 19 candidate WD + brown dwarf (BD) binaries (eight new systems), 66 candidate WD + dust disks (38 new systems), and 24 candidate either WD + BD or WD + dust disks (19 new systems). Given the limited spatial resolution of WISE, all candidate systems require follow-up IR observations for confirmation, such as high spatial resolution imaging or IR spectroscopy. This will help expand the parameter space of dust disks, allowing us to explore a broader range of possibilities.

astro-ph.SR

Warm Debris Disk Candidates around Nearby FGK Stars from LAMOST DR12

Warm debris disks around main-sequence stars trace late-stage terrestrial planet formation. Motivated by the need for systematic searches of such systems, we identify debris disk candidates around FGK stars within 150 pc by combining a spectroscopically selected sample from LAMOST DR12 with Gaia astrometry and multi-band infrared photometry. Infrared excesses are identified through SED fitting and validated using conservative, source-by-source checks. This approach yields a final sample of 12 debris disk candidates including ten new detections. Stellar age research indicate that most of the host stars are several billion years old. NEOWISE monitoring reveals no significant W1/W2 variability, consistent with a circumstellar origin of the infrared excess. while a search for co-moving companions using Gaia DR3 reveals possible companions for only two candidates at very large projected separations ($\gtrsim 10^4$~au). Three candidates exhibit excess emission in both the W3 and W4 bands, allowing estimates of characteristic dust properties. This work establishes a small yet reliable sample of debris disk candidates anchored in homogeneous LAMOST spectroscopy, providing a foundation for future studies of debris disk evolution and stellar activity.

astro-ph.EP

Convolutional causal learning for aerodynamic flows

This study aims to capture aerodynamic causality from snapshot data with a time-varying mode decomposition technique referred to as information-theoretic machine learning. The current approach extracts time-dependent informative vortical structures, contributing to the future evolution of the aerodynamic coefficients. The present decomposition is employed with a convolutional neural network, enabling the identification of the spatial continuous mode. In addition, a low-order representation, characterizing the informative vortical structures and their corresponding aerodynamic coefficients, can also be identified by considering autoencoder-based data compression. The present technique is applied to a range of aerodynamic examples, including extreme vortex-gust airfoil interactions, experimentally measured transverse jet-wing interaction, and a turbulent separated wake across different Reynolds numbers. For the cases of gust-wing interaction, the time-varying gust effect on the lift response is extracted in an interpretable manner. With the example of a turbulent wake, the relationship between large-scale vortical motion and lift force is identified without any spatial length-scale information. The proposed approach could serve as a foundation for data-driven causal modeling and control for a range of unsteady flows.

physics.flu-dyn

Uplink RSMA Performance Analysis with Rate Adaptation: A Stochastic Geometry Approach

Rate-splitting multiple access (RSMA) has emerged as a promising technique for efficient interference management in next-generation wireless networks. While most existing studies focus on downlink and single-cell designs, the modeling and analysis of uplink RSMA under large-scale deployments remain largely unexplored. On the basis of stochastic geometry (SG), this paper introduces a unified analytical framework that integrates finite modulation and coding scheme (MCS)-based rate adaptation. This framework jointly captures spatial interference coupling and discrete rate behavior to bridge theoretical tractability and practical realism. Within this framework, we derive tractable expressions for the conditional received rate (CRR), its spatial average, and higher-order statistics via the meta distribution, thereby quantifying both the mean and user-specific rate performance. Results show that the proposed unified framework not only generalizes existing non-orthogonal multiple access (NOMA) and orthogonal multiple access (OMA) analyses but also provides new insights into how discrete rate adaptation reshapes interference dynamics and fairness in dense RSMA-enabled networks.

cs.IT

Fast Collaborative Inference via Distributed Speculative Decoding

Speculative decoding accelerates large language model (LLM) inference by allowing a small draft model to predict multiple future tokens for verification by a larger target model. In AI-native radio access networks (AI-RAN), this enables device-edge collaborative inference but introduces significant uplink overhead, as existing distributed speculative decoding schemes transmit full vocabulary logits at every step. We propose a sparsify-then-sample strategy, Truncated Sparse Logits Transmission (TSLT), which transmits only the logits and indices of a truncated candidate set. We provide theoretical guarantees showing that the acceptance rate is preserved under TSLT. TSLT is further extended to multi-candidate case, where multiple draft candidates per step increase acceptance probability. Experiments show that TSLT significantly reduces uplink communication while maintaining end-to-end inference latency and model quality, demonstrating its effectiveness for scalable, communication-efficient distributed LLM inference in future AI-RAN systems.

eess.SP

Anatomically and Metabolically Informed Diffusion for Unified Denoising and Segmentation in Low-Count PET Imaging

Positron emission tomography (PET) image denoising, along with lesion and organ segmentation, are critical steps in PET-aided diagnosis. However, existing methods typically treat these tasks independently, overlooking inherent synergies between them as correlated steps in the analysis pipeline. In this work, we present the anatomically and metabolically informed diffusion (AMDiff) model, a unified framework for denoising and lesion/organ segmentation in low-count PET imaging. By integrating multi-task functionality and exploiting the mutual benefits of these tasks, AMDiff enables direct quantification of clinical metrics, such as total lesion glycolysis (TLG), from low-count inputs. The AMDiff model incorporates a semantic-informed denoiser based on diffusion strategy and a denoising-informed segmenter utilizing nnMamba architecture. The segmenter constrains denoised outputs via a lesion-organ-specific regularizer, while the denoiser enhances the segmenter by providing enriched image information through a denoising revision module. These components are connected via a warming-up mechanism to optimize multi-task interactions. Experiments on multi-vendor, multi-center, and multi-noise-level datasets demonstrate the superior performance of AMDiff.

eess.IV

MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnP

Image-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings.

cs.CV

A survey on proximity monitoring and warning in construction

Various technologies have been applied to monitor the proximity between two construction entities, preventing struck-by accidents and thereby enhancing onsite safety. This study comprehensively reviews related efforts dedicated to proximity monitoring and warning (PMW) based on 97 relevant articles published between 2010 and 2024. The bibliometric analysis reveals the technical roadmap over time, as well as the five most influential leaders and the two largest research networks they have established. The qualitative review is then conducted from four perspectives: influencing factor study, hazard level definition and determination, proximity perception, and alarm issuing and receiving. Finally, the limitations and challenges of current proximity perception are discussed, along with corresponding future research directions, including end-to-end three-dimensional (3D) object detection, real-time 3D reconstruction and updating for dynamic construction scenes, and multimodal fusion. This review presents the current research status, limitations, and future directions of PMW, guiding the future development of PMW systems.

cs.DL

DRST: a Non-Intrusive Framework for Performance Analysis in Softwarized Networks

The last decade has witnessed the proliferation of network function virtualization (NFV) in the telco industry, thanks to its unparalleled flexibility, scalability, and cost-effectiveness. However, as the NFV infrastructure is shared by virtual network functions (VNFs), sporadic resource contentions are inevitable. Such contention makes it extremely challenging to guarantee the performance of the provisioned network services, especially in high-speed regimes (e.g., Gigabit Ethernet). Existing solutions typically rely on direct traffic analysis (e.g., packet- or flow-level measurements) to detect performance degradation and identify bottlenecks, which is not always applicable due to significant integration overhead and system-level constraints. This paper complements existing solutions with a lightweight, non-intrusive framework for online performance inference that easily adapts to drift (i.e., a change over time of the actual state of our system). Instead of direct data-plane collection, we reuse hardware features in the underlying NFV infrastructure, introducing negligible interference in the data-plane. Our Drift-Resilient and Self-Tuning (DRST) framework can be integrated into existing NFV systems with minimal engineering effort and operate without the need for predefined traffic models or VNF-specific customization. DRST is deployed via a lightweight MLOps pipeline that automates the adaptation under runtime drift. We show how DRST can deliver accurate performance inference or diagnose run-time bottlenecks, as demonstrated through a comprehensive evaluation across diverse NFV scenarios.

cs.NI

Dose-aware Diffusion Model for 3D PET Image Denoising: Multi-institutional Validation with Reader Study and Real Low-dose Data

Reducing scan times, radiation dose, and enhancing image quality for lower-performance scanners, are critical in low-dose PET imaging. Deep learning techniques have been investigated for PET image denoising. However, existing models have often resulted in compromised image quality when achieving low-count/low-dose PET and have limited generalizability to different image noise-levels, acquisition protocols, and patient populations. Recently, diffusion models have emerged as the new state-of-the-art generative model to generate high-quality samples and have demonstrated strong potential for medical imaging tasks. However, for low-dose PET imaging, existing diffusion models failed to generate consistent 3D reconstructions, unable to generalize across varying noise-levels, often produced visually-appealing but distorted image details, and produced images with biased tracer uptake. Here, we develop DDPET-3D, a dose-aware diffusion model for 3D low-dose PET imaging to address these challenges. Collected from 4 medical centers globally with different scanners and clinical protocols, we evaluated the proposed model using a total of 9,783 18F-FDG studies with low-dose levels ranging from 1% to 50%. With a cross-center, cross-scanner validation, the proposed DDPET-3D demonstrated its potential to generalize to different low-dose levels, different scanners, and different clinical protocols. As confirmed with reader studies performed by board-certified nuclear medicine physicians, experienced readers judged the images to be similar or superior to the full-dose images and previous DL baselines based on qualitative visual impression. Lesion-level quantitative accuracy was evaluated using a Monte Carlo simulation study and a lesion segmentation network. The presented results show the potential to achieve low-dose PET while maintaining image quality. Real low-dose scans was also included for evaluation.

eess.IV

WISE 12 micron search for exozodi candidates within 10 parsecs

The discovery of extra-terrestrial life is one of the ultimate goals for future exoplanet-seeking missions, with one major challenge being the presence of 'exozodiacal' dust near target stars or within their habitable zone. Therefore, it is critical to identify which stars possess exozodiacal dust and quantify their emission levels. In this study, we conducted a search for exozodi candidates within 10 parsecs using the Reyl'e sample. We performed proper motion calculations and cross-matched the sample with the WISE and 2MASS database, resulting in 339 preliminary target samples. We further analysed the infrared radiation characteristics of these targets, using spectral energy distribution (SED) fitting to predict photometric flux levels in the infrared and searching for 3sigma excesses in the WISE W3 band. During further selection processes, we applied various analysis methods to perform rigorous validation. We identified five exozodi candidates all of which are brown dwarfs (BDs). Given the clustering in candidate spectral types, we expect that these are not true exozodi candidates, rather the apparent excess arises from the inability of the BD photosphere models to accurately represent the SEDs of objects at the L-T transition. Indeed, for the object DENIS J025503.3-470049, excess is likely due to silicate clouds in the BD atmosphere. We suggest that a more stringent 5sigma excess is required to infer excess for this spectral type. The detection rate (0/339) in our sample shows that less than 1% M stars have exozodi above 21% excess levels. This is consistent with the rate of exozodi at similar level towards FGK stars in the Kennedy & Wyatt sample (25/24,174). We provide upper limits on the 12 micron exozodi emission for the sample, which is typically at 21% relative to the star. For most stars, in particular the low mass M stars, this is the first such upper limit in the literature.

astro-ph.EP

Reinforcement Learning-Based Closed-Loop Airfoil Flow Control

We systematically investigated a reinforcement learning (RL)-based closed-loop active flow control strategy to enhance the lift-to-drag ratio of a wing section with an NLF(1)-0115 airfoil at an angle of attack 5 degree. The effects of key control parameters, including actuation location, observed state, reward function, and control update interval, are evaluated at a chord-based Reynolds number of Re=20,000. Results show that all parameters significantly influence control performance, with the update interval playing a particularly critical role. Properly chosen update intervals introduce a broader spectrum of actuation frequencies, enabling more effective interactions with a wider range of flow structures and contributing to improved control effectiveness. The optimally trained RL controller is further evaluated in a three-dimensional numerical setup at the same Reynolds number. Actuation is applied using both spanwise-uniform and spanwise-varying control profiles. The results demonstrate that the pretrained controller, combined with a physics-informed spanwise distribution, achieves substantial performance gains. These findings extend the feasibility and scalability of a pretrained RL-based control strategy to more complex airfoil flows.

physics.flu-dyn

Fed-NDIF: A Noise-Embedded Federated Diffusion Model For Low-Count Whole-Body PET Denoising

Low-count positron emission tomography (LCPET) imaging can reduce patients' exposure to radiation but often suffers from increased image noise and reduced lesion detectability, necessitating effective denoising techniques. Diffusion models have shown promise in LCPET denoising for recovering degraded image quality. However, training such models requires large and diverse datasets, which are challenging to obtain in the medical domain. To address data scarcity and privacy concerns, we combine diffusion models with federated learning -- a decentralized training approach where models are trained individually at different sites, and their parameters are aggregated on a central server over multiple iterations. The variation in scanner types and image noise levels within and across institutions poses additional challenges for federated learning in LCPET denoising. In this study, we propose a novel noise-embedded federated learning diffusion model (Fed-NDIF) to address these challenges, leveraging a multicenter dataset and varying count levels. Our approach incorporates liver normalized standard deviation (NSTD) noise embedding into a 2.5D diffusion model and utilizes the Federated Averaging (FedAvg) algorithm to aggregate locally trained models into a global model, which is subsequently fine-tuned on local datasets to optimize performance and obtain personalized models. Extensive validation on datasets from the University of Bern, Ruijin Hospital in Shanghai, and Yale-New Haven Hospital demonstrates the superior performance of our method in enhancing image quality and improving lesion quantification. The Fed-NDIF model shows significant improvements in PSNR, SSIM, and NMSE of the entire 3D volume, as well as enhanced lesion detectability and quantification, compared to local diffusion models and federated UNet-based models.

eess.IV