SearcharxivSearch

arXiv subjects

Dong Zhao

Publications and source records attributed to Dong Zhao.

At least 19 recordsLinked to original sources

Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation

Test-time adaptation (TTA) aims to mitigate distribution shifts by adapting models with unlabeled target data at inference time. While TTA with vision-language models (VLMs) has shown promising results in classification, extending it to medical image segmentation remains challenging. In this setting, the adaptation gains from optimizing on VLM-generated predictions are often outweighed by the degradation to the VLM's strong pretrained features caused by noisy, update-driven learning, resulting in limited and unstable improvements. We therefore propose Memory-Supported Synergistic Adaptation (MSSA), a novel training-free TTA framework for medical image segmentation. Without updating model parameters, MSSA dynamically selects reliable image-text predictions to construct an online memory, uses them as text-guided semantic priors, and couples them with cross-image structural alignment for robust adaptation. Specifically, MSSA consists of (i) a noise-aware memory construction module that filters and stabilizes cross-modal predictions, and (ii) a relevance-driven prototype alignment module that aligns the target sample with structurally consistent memory samples and their reliable predictions to improve adaptation. Extensive experiments on multiple medical segmentation benchmarks demonstrate that MSSA consistently improves VLM-based segmentation models and outperforms existing fine-tuning-based TTA methods by a clear margin, with gains of up to 12.2% DSC and 11.7% mIoU. Project page: https://lingrayy.github.io/MSSA/ .

cs.CV

Revisiting cosmic anisotropy with the Pantheon+ compilation

We investigate cosmic anisotropy within the updated Pantheon+ sample using both the dipole fitting (DF) and hemisphere comparison (HC) methods. With the DF method, the dipole signal within the full sample is statistically weak. However, the low-$z$ subsample yields a dipole signal of $A_{\mathrm{D}} = 0.952^{+0.454}_{-0.403} \times 10^{-3}$ at $\sim 2\sigma$ significance, pointing towards $(l,b) = (149.77^\circ, -12.20^\circ)$. This signal is predominantly driven by a combined subset of surveys 5, 56, 63, and 150, which is characterized by an amplitude of $A_{\mathrm{D}} = 1.730_{-0.715}^{+0.554} \times 10^{-3}$ towards $(l,b) = (153.05^\circ, -1.25^\circ)$. For the HC method, the full sample yields a maximum anisotropy level of $\mathrm{AL}_{\mathrm{max}} = 0.289 \pm 0.052$ oriented towards $(l,b) = (127.97^\circ, 17.90^\circ)$ with a $1.56\sigma$ significance. This preferred direction is primarily determined by the highly inhomogeneous SNLS subsample, whereas the low-$z$ and high-$z$ subsamples act to suppress the anisotropy level along this axis. These subsample-dependent results suggest that the apparent anisotropy arises from local structures or the inhomogeneous distribution of the datasets rather than an intrinsic cosmic anisotropy.

astro-ph.CO

Testing cosmic anisotropy with the Combo correlation of gamma-ray bursts

We employ the sample of 244 gamma-ray bursts (GRBs; i.e., C244) with the Combo correlation to test cosmic anisotropy. Meanwhile, the Pantheon sample is introduced to verify whether the GRB sample can suppress the fake anisotropic signals induced by inhomogeneous spatial distributions. In the dipole fitting (DF) method, under the dipole-modulated $\Lambda$CDM model, the C244 sample shifts the best-fitting longitude $l$ derived from the Pantheon sample by $54.09^\circ$ and reduces the uncertainty in $l$ by approximately $40\%$. Compared to the 118 GRBs (i.e., A118) with the $E_\mathrm{p}$-$E_\mathrm{iso}$ correlation, the shift in longitude $l$ increases by additional $21.35^\circ$. In the hemisphere comparison (HC) method, the preferred direction derived from the C244+Pantheon sample deviates from that of the Pantheon-only sample by more than $1\sigma$. In contrast, the preferred direction from the A118+Pantheon sample is consistent with the Pantheon-only result within the $1\sigma$ uncertainty. The preferred direction changes significantly as the number of GRBs increases from 118 to 244. Our results show that a larger GRB sample can reduce the fake anisotropic signals caused by inhomogeneous spatial distributions. Accordingly, we suggest that GRBs have the potential to provide a reliable probe of cosmic anisotropy.

astro-ph.CO

The Deep Learning-Based Dual-Branch Multimodal Fusion Model for Solar Flare Prediction

Solar flares are intense eruptive events caused by the rapid release of magnetic energy, often impacting Earth's space environment through electromagnetic radiation and high-energy particles. Accurate flare prediction is critical for space weather forecasting. However, many existing deep learning approaches often rely on single-modal inputs or shallow feature fusion, limiting their ability to capture complementary information. In this study, we propose a dual-branch multimodal fusion deep learning model for predicting 24-hour solar flares. The model integrates magnetograms and magnetic parameters through cross-attention mechanisms, followed by cross-scale interactions at the feature level to enhance multi-scale representation. It is designed to perform both binary prediction of $\geqslant$ C-class flares and multi-class classification of C, M, and X-class flares. To ensure rigorous evaluation, we employ a stratified group five-fold cross-validation scheme to preserve class representativeness and adopt a splitting-before-sampling strategy based on NOAA active region numbers to prevent data leakage. Experimental results show that the model achieves a TSS of 0.661 and an HSS of 0.658 for binary $\geqslant$ C-class prediction, while notably attaining a TSS of 0.780 and an HSS of 0.775 for X-class flares in the multi-class task. Compared with existing approaches, the model demonstrates superior performance in predicting intense X-class flares, effectively suppresses the false alarm rate, and exhibits strong generalization capability.

astro-ph.SR

Revisiting the Independence Assumption in LEO Satellite-to-Ground Optical Links: A State-Coupled Joint Fading Model

Performance analysis of low Earth orbit (LEO) satellite-to-ground optical links relies on composite fading models that typically evaluate scintillation and angular loss under the assumption of statistical independence. While ensuring analytical tractability, this assumption decouples fading mechanisms driven by the same atmospheric turbulence and fails to capture the distinct effects of free atmosphere (FA) and boundary layer (BL) perturbations. To model this coupling while preserving tractability, this paper develops a state-coupled joint fading model. In the proposed framework, aperture-averaged scintillation and effective angular loss are jointly characterized by a discrete slow atmospheric state, parameterized by separate FA and BL scaling factors. By replacing unconditional independence with state-conditioned independence, the model enables a closed-form derivation of the outage probability, preserving the computational simplicity of the independent baseline. Numerical results show that the independent baseline can misestimate outage under non-nominal layered turbulence states. This outage prediction bias varies with elevation because the relative roles of scintillation and angular loss change with the link geometry, resulting in different residual angular correction requirements for a given outage target.

eess.SP

Spatiotemporal flat optics for terabit-per-second single-channel data transmission

Exponential growth in global data traffic demands ever-increasing transmission rates--a pursuit fundamentally constrained by the physical limitations of digital-to-analog converters (DACs). Existing strategies to overcome this bottleneck, such as multi-DAC arrays and optical time-division multiplexing, inevitably introduce system complexity and coordination overhead. Here we demonstrate an all-optical spatiotemporal transmitter that generates controllable high-repetition information-carrying femtosecond pulses at the focus of a phase-modulated planar diffractive lens (PDL) through optical-path-induced spatial-to-temporal conversion. Each pulse serves as an information bit, encoding binary data via on-axis focal intensity states corresponding to '0' and '1', achieved by switching between topological and constant phase modulations. High experimental orthogonality between arbitrary bits enables nearly error-free transmission of 15X15-pixel grayscale (8-bit coding) and colour (9-bit coding) images at a record-high single-channel rate of approximately 3 terabits per second (Tbit/s). Free from electronic and coordination bottlenecks, this all-optical transmitter establishes a scalable high-speed single-channel pathway toward ultrahigh-capacity optical communication.

physics.optics

Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

Knowledge distillation (KD) has been widely applied in semantic segmentation to compress large models, but conventional approaches primarily preserve in-domain accuracy while neglecting out-of-domain generalization, which is essential under distribution shifts. This limitation becomes more severe with the emergence of vision foundation models (VFMs): although VFMs exhibit strong robustness on unseen data, distilling them with conventional KD often compromises this ability. We propose Generalizable Knowledge Distillation (GKD), a multi-stage framework that explicitly enhances generalization. GKD decouples representation learning from task learning. In the first stage, the student acquires domain-agnostic representations through selective feature distillation, and in the second stage, these representations are frozen for task adaptation, thereby mitigating overfitting to visible domains. To further support transfer, we introduce a query-based soft distillation mechanism, where student features act as queries to teacher representations to selectively retrieve transferable spatial knowledge from VFMs. Extensive experiments on five domain generalization benchmarks demonstrate that GKD consistently outperforms existing KD methods, achieving average gains of +1.9% in foundation-to-foundation (F2F) and +10.6% in foundation-to-local (F2L) distillation. The code will be available at https://github.com/Younger-hua/GKD.

cs.CV

Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

Domain Generalization in Semantic Segmentation (DG-SS) aims to enable segmentation models to perform robustly in unseen environments. However, conventional DG-SS methods are restricted to a fixed set of known categories, limiting their applicability in open-world scenarios. Recent progress in Vision-Language Models (VLMs) has advanced Open-Vocabulary Semantic Segmentation (OV-SS) by enabling models to recognize a broader range of concepts. Yet, these models remain sensitive to domain shifts and struggle to maintain robustness when deployed in unseen environments, a challenge that is particularly severe in urban-driving scenarios. To bridge this gap, we introduce Open-Vocabulary Domain Generalization in Semantic Segmentation (OVDG-SS), a new setting that jointly addresses unseen domains and unseen categories. We introduce the first benchmark for OVDG-SS in autonomous driving, addressing a previously unexplored problem and covering both synthetic-to-real and real-to-real generalization across diverse unseen domains and unseen categories. In OVDG-SS, we observe that domain shifts often distort text-image correlations in pre-trained VLMs, which hinders the performance of OV-SS models. To tackle this challenge, we propose S2-Corr, a state-space-driven text-image correlation refinement mechanism that mitigates domain-induced distortions and produces more consistent text-image correlations under distribution changes. Extensive experiments on our constructed benchmark demonstrate that the proposed method achieves superior cross-domain performance and efficiency compared to existing OV-SS approaches.

cs.CV

Catalyzing Informed Residential Energy Retrofit Decisions via Domain-Specific LLM

Residential energy retrofit initiation is often stalled by an expertise gap, where homeowners lack the technical literacy required for structured building energy assessments and are thereby trapped in low-information environments with fragmented sources. To bridge this gap, this study reports a domain-specific large language model (LLM) designed to catalyze informed decision-making based solely on homeowner-accessible, natural-language descriptions, e.g., building age, size, and location. The model is created using the parameter-efficient low-rank adaption (LoRA) fine-tuning approach on a massive corpus grounded in physics-based energy simulations and techno-economic calculations from 536,416 U.S. residential building prototypes. Nine major retrofit categories are evaluated, including envelope upgrades, HVAC systems, and renewable energy installations. Validations against physics-grounded benchmarks show that the LLM consistently identifies high-quality retrofit options, achieving top-3 hit rates of 98.9% for maximum CO2 reduction and 93.3% for the shortest discounted payback year. Moreover, the model exhibits strong robustness under incomplete input conditions, maintaining stable performance even when basic dwelling descriptions are only 60% partially specified. By significantly lowering the information activation energy for non-expert users while maintaining the scientific rigor, this physics-based AI model offers a scalable pathway for parallelized, user-centered decision making, accelerating cumulative energy savings and emission reductions across community and national scales.

cs.CY

Analysis, detection and control of secure and safe cyber-physical control systems in a unified framework

This paper deals with analysis, simultaneous detection of faults and attacks, fault-tolerant control and attack-resilient of cyber-physical control systems. In our recent work, it has been observed that an attack detector driven by an input residual signal is capable of reliably detecting attacks. In particular, observing system dynamics from the perspective of the system input-output signal space reveals that attacks and system uncertainties act on different system subspaces. These results motivate our exploration of secure and safe cyber-physical control systems in the unified framework of control and detection. The unified framework is proposed to handle control and detection issues uniformly and in subspaces of system input-output data. Its mathematical and control-theoretic basis is system coprime factorizations with Bezout identity at its core. We firstly explore those methods and schemes of the unified framework, which serve as the major control-theoretic tool in our work. It is followed by re-visiting and examining established attack detection and resilient control schemes. The major part of our work is the endeavours to develop a control-theoretic paradigm, in which analysis, simultaneous detection of faults and attacks, fault-tolerant and attack-resilient control of cyber-physical control systems are addressed in a unified manner.

eess.SY

A dispersion-driven 3D color near-eye meta-display

Chromatic dispersion, an inherent wavelength-dependent phenomenon in optical systems, has traditionally been regarded as a detrimental effect to be minimized in imaging and display. Here, we present a paradigm shift by deliberately engineering and harnessing metalens dispersion as a functional mechanism for three-dimensional (3D) near-eye displays. Specifically, we exploit lateral dispersion to transform transverse offset between green and red objects into image-space angular separations that make their images intersected virtually, thereby creating color-merged 3D virtual-image perception. This meta-display architecture preserves compactness of conventional planar display while exhibiting less data requirements and lower hardware complexity than other near-eye 3D displays. Experimentally, we demonstrate a multi-color near-eye 3D system achieving an 11{\deg} field of view, 22 pixels-per-degree angular resolution, 0.9 m depth of field, and 19 distinct image planes. This work establishes a new pathway for metasurfaces toward visual displays and highlights great potential for future virtual/augmented reality.

physics.optics

An all-optical convolutional neural network for image identification

In modern artificial intelligence, convolutional neural networks (CNNs) have become a cornerstone for visual and perceptual tasks. However, their implementation on conventional electronic hardware faces fundamental bottlenecks in speed and energy efficiency due to resistive and capacitive losses. Photonic alternatives offer a promising route, yet the difficulty of realizing optical nonlinearities has prevented the realization of all-optical CNNs capable of end-to-end image classification. Here, we demonstrate an all-optical CNN that bypasses the need for explicit optical nonlinear activations. Our architecture comprises a single spatial-differentiation convolutional stage--using 24 directional kernels spanning 360{\deg}, along with a mean-filtering kernel--followed by a diffractive fully-connected layer. The directional convolution enhances feature selectivity, suppresses noise and crosstalk, and simplifies the classification task, allowing the weak nonlinearity inherent in optical diffraction to achieve high accuracy. We report experimentally classification accuracies of 86.8% on handwritten digits (MNIST) and 94.8% on a ten-class gesture dataset. The system delivers a computational throughput of 1.13X10^5 tera-operations per second (TOPS) and an energy efficiency of 1.51X10^3 TOPS/W--the highest reported among CNN hardware--with the potential to improve by a further 5-6 orders of magnitude using nanosecond-scale detectors. This work establishes a scalable pathway toward ultralow-latency, ultralow-energy vision processing for real-time intelligent systems.

physics.optics

A 160 {\deg} x 160 {\deg} Dynamic Holographic Meta-Projector

Holography can reconstruct immersive light fields for virtual and augmented reality by modulating optical wavefront. Due to huge pixel sizes, current spatial light modulators (SLMs) have small field-of-view (FOV) for holographic displays. Despite various methods for etendue expansion, the largest full-screen FOV for dynamic holography is only 70 {\deg} X 70 {\deg}, which remains insufficient for large-scale, high-resolution, three-dimensional displays. Here, we report a pixel-interpolation-assisted holographic meta-projector that substantially expands the FOV by integrating multiple subwavelength metasurface pixels within each microscale pixel of a traditional SLM. Leveraging large-angle diffraction of the metasurface and implementing k-space distortion correction for ultra-wide angles, we experimentally demonstrate dynamic holographic image reconstruction with a FOV of 160 {\deg} X 160 {\deg} -equivalent to a system numerical aperture of 0.985-at a high framerate of 60 Hz, surpassing the temporal resolution threshold of human vision. This system represents the state-of-the-art near-full-screen holographic dynamic display, thereby opening the door to high-dynamic-range and large-FOV holographic displays.

physics.optics

Adaptive Lighting Control in Visible Light Systems: An Integrated Sensing, Communication, and Illumination Framework

Indoor visible light communication (VLC) is a promising sixth-generation (6G) technology, as its directional and sensitive optical signals are naturally suited for integrated sensing and communication (ISAC). However, current research mainly focuses on maximizing data rates and sensing accuracy, creating a conflict between high performance, high energy consumption, and user visual comfort. This paper proposes an adaptive integrated sensing, communication, and illumination (ISCI) framework that resolves this conflict by treating energy savings as a primary objective. The framework's mechanism first partitions the receiving plane using a geometric methodology, defining an activity area and a surrounding non-activity area to match distinct user requirements. User location, determined using non-line-of-sight (NLOS) sensing, then acts as a dynamic switch for the system's optimization objective. The system adaptively shifts between minimizing total transmit power while guaranteeing communication and illumination performance in the activity area and maximizing signal-to-noise ratio (SNR) uniformity in the non-activity area. Numerical results confirm that this adaptive ISCI approach achieves 53.59% energy savings over a non-adaptive system and improves SNR uniformity by 57.79%, while satisfying all illumination constraints and maintaining a mean localization error of 0.071 m.

eess.SY

MG-HGNN: A Heterogeneous GNN Framework for Indoor Wi-Fi Fingerprint-Based Localization

Received signal strength indicator (RSSI) is the primary representation of Wi-Fi fingerprints and serves as a crucial tool for indoor localization. However, existing RSSI-based positioning methods often suffer from reduced accuracy due to environmental complexity and challenges in processing multi-source information. To address these issues, we propose a novel multi-graph heterogeneous GNN framework (MG-HGNN) to enhance spatial awareness and improve positioning performance. In this framework, two graph construction branches perform node and edge embedding, respectively, to generate informative graphs. Subsequently, a heterogeneous graph neural network is employed for graph representation learning, enabling accurate positioning. The MG-HGNN framework introduces the following key innovations: 1) multi-type task-directed graph construction that combines label estimation and feature encoding for richer graph information; 2) a heterogeneous GNN structure that enhances the performance of conventional GNN models. Evaluations on the UJIIndoorLoc and UTSIndoorLoc public datasets demonstrate that MG-HGNN not only achieves superior performance compared to several state-of-the-art methods, but also provides a novel perspective for enhancing GNN-based localization methods. Ablation studies further confirm the rationality and effectiveness of the proposed framework.

cs.LG

Dual Detection Framework for Faults and Integrity Attacks in Cyber-Physical Control Systems

Anomaly detection plays a vital role in the security and safety of cyber-physical control systems, and accurately distinguishing between different anomaly types is crucial for system recovery and mitigation. This study proposes a dual detection framework for anomaly detection and discrimination. By leveraging the dynamic characteristics of control loops and the stealthiness features of integrity attacks, the closed-loop stealthiness condition is first derived, and two dedicated detectors are designed and deployed on the controller side and the plant side, respectively, enabling joint plant fault and cyber attack detection. Moreover, by jointly analyzing the residual response of the two detectors corresponding to different anomalies, it is proved that the proposed method can distinguish between faults and integrity attacks due to the detectors' individual residual spaces. According to the detector's residual space, the fault and attack detection performance is further improved by a two-stage optimization scheme. Simulation results validate the effectiveness of the proposed approach.

eess.SY

Co-TAP: Three-Layer Agent Interaction Protocol Technical Report

This paper proposes Co-TAP (T: Triple, A: Agent, P: Protocol), a three-layer agent interaction protocol designed to address the challenges faced by multi-agent systems across the three core dimensions of Interoperability, Interaction and Collaboration, and Knowledge Sharing. We have designed and proposed a layered solution composed of three core protocols: the Human-Agent Interaction Protocol (HAI), the Unified Agent Protocol (UAP), and the Memory-Extraction-Knowledge Protocol (MEK). HAI focuses on the interaction layer, standardizing the flow of information between users, interfaces, and agents by defining a standardized, event-driven communication paradigm. This ensures the real-time performance, reliability, and synergy of interactions. As the core of the infrastructure layer, UAP is designed to break down communication barriers among heterogeneous agents through unified service discovery and protocol conversion mechanisms, thereby enabling seamless interconnection and interoperability of the underlying network. MEK, in turn, operates at the cognitive layer. By establishing a standardized ''Memory (M) - Extraction (E) - Knowledge (K)'' cognitive chain, it empowers agents with the ability to learn from individual experiences and form shareable knowledge, thereby laying the foundation for the realization of true collective intelligence. We believe this protocol framework will provide a solid engineering foundation and theoretical guidance for building the next generation of efficient, scalable, and intelligent multi-agent applications.

cs.AI

Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models

Conventional approaches to building energy retrofit decision making suffer from limited generalizability and low interpretability, hindering adoption in diverse residential contexts. With the growth of Smart and Connected Communities, generative AI, especially large language models (LLMs), may help by processing contextual information and producing practitioner readable recommendations. We evaluate seven LLMs (ChatGPT, DeepSeek, Gemini, Grok, Llama, and Claude) on residential retrofit decisions under two objectives: maximizing CO2 reduction (technical) and minimizing payback period (sociotechnical). Performance is assessed on four dimensions: accuracy, consistency, sensitivity, and reasoning, using a dataset of 400 homes across 49 US states. LLMs generate effective recommendations in many cases, reaching up to 54.5 percent top 1 match and 92.8 percent within top 5 without fine tuning. Performance is stronger for the technical objective, while sociotechnical decisions are limited by economic trade offs and local context. Agreement across models is low, and higher performing models tend to diverge from others. LLMs are sensitive to location and building geometry but less sensitive to technology and occupant behavior. Most models show step by step, engineering style reasoning, but it is often simplified and lacks deeper contextual awareness. Overall, LLMs are promising assistants for energy retrofit decision making, but improvements in accuracy, consistency, and context handling are needed for reliable practice.

cs.AI