Searcharxiv⌕ Search

arXiv subjects

Da Zhang

Publications and source records attributed to Da Zhang.

At least 37 records · Page 2Linked to original sources

UniDiff: A Unified Diffusion Framework for Multimodal Time Series Forecasting

As multimodal data proliferates across diverse real-world applications, leveraging heterogeneous information such as texts and timestamps for accurate time series forecasting (TSF) has become a critical challenge. While diffusion models demonstrate exceptional performance in generation tasks, their application to TSF remains largely confined to modeling single-modality numerical sequences, overlooking the abundant cross-modal signals inherent in complex heterogeneous data. To address this gap, we propose UniDiff, a unified diffusion framework for multimodal time series forecasting. To process the numerical sequence, our framework first tokenizes the time series into patches, preserving local temporal dynamics by mapping each patch to an embedding space via a lightweight MLP. At its core lies a unified and parallel fusion module, where a single cross-attention mechanism adaptively weighs and integrates structural information from timestamps and semantic context from texts in one step, enabling a flexible and efficient interplay between modalities. Furthermore, we introduce a novel classifier-free guidance mechanism designed for multi-source conditioning, allowing for decoupled control over the guidance strength of textual and temporal information during inference, which significantly enhances model robustness. Extensive experiments on real-world benchmark datasets across eight domains demonstrate that the proposed UniDiff model achieves state-of-the-art performance.

cs.LG↗

FAIM: Frequency-Aware Interactive Mamba for Time Series Classification

Time series classification (TSC) is crucial in numerous real-world applications, such as environmental monitoring, medical diagnosis, and posture recognition. TSC tasks require models to effectively capture discriminative information for accurate class identification. Although deep learning architectures excel at capturing temporal dependencies, they often suffer from high computational cost, sensitivity to noise perturbations, and susceptibility to overfitting on small-scale datasets. To address these challenges, we propose FAIM, a lightweight Frequency-Aware Interactive Mamba model. Specifically, we introduce an Adaptive Filtering Block (AFB) that leverages Fourier Transform to extract frequency-domain features from time series data. The AFB incorporates learnable adaptive thresholds to dynamically suppress noise and employs element-wise coupling of global and local semantic adaptive filtering, enabling in-depth modeling of the synergy among different frequency components. Furthermore, we design an Interactive Mamba Block (IMB) to facilitate efficient multi-granularity information interaction, balancing the extraction of fine-grained discriminative features and comprehensive global contextual information, thereby endowing FAIM with powerful and expressive representations for TSC tasks. Additionally, we incorporate a self-supervised pre-training mechanism to enhance FAIM's understanding of complex temporal patterns and improve its robustness across various domains and high-noise scenarios. Extensive experiments on multiple benchmarks demonstrate that FAIM consistently outperforms existing state-of-the-art (SOTA) methods, achieving a superior trade-off between accuracy and efficiency and exhibits outstanding performance.

cs.LG↗

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

Large vision-language models (VLMs) have achieved remarkable success in natural scene understanding, yet their application to underwater environments remains largely unexplored. Underwater imagery presents unique challenges including severe light attenuation, color distortion, and suspended particle scattering, while requiring specialized knowledge of marine ecosystems and organism taxonomy. To bridge this gap, we introduce UWBench, a comprehensive benchmark specifically designed for underwater vision-language understanding. UWBench comprises 15,003 high-resolution underwater images captured across diverse aquatic environments, encompassing oceans, coral reefs, and deep-sea habitats. Each image is enriched with human-verified annotations including 15,281 object referring expressions that precisely describe marine organisms and underwater structures, and 124,983 question-answer pairs covering diverse reasoning capabilities from object recognition to ecological relationship understanding. The dataset captures rich variations in visibility, lighting conditions, and water turbidity, providing a realistic testbed for model evaluation. Based on UWBench, we establish three comprehensive benchmarks: detailed image captioning for generating ecologically informed scene descriptions, visual grounding for precise localization of marine organisms, and visual question answering for multimodal reasoning about underwater environments. Extensive experiments on state-of-the-art VLMs demonstrate that underwater understanding remains challenging, with substantial room for improvement. Our benchmark provides essential resources for advancing vision-language research in underwater contexts and supporting applications in marine science, ecological monitoring, and autonomous underwater exploration. Our code and benchmark will be available.

cs.CV↗

Temporal-order-driven asymmetric quantum interference and temporal coherence enhancement in spontaneous six-wave mixing

Narrow-band multiphoton entanglement sources serve as a core enabling resource for advanced quantum information technologies. Recently, researchers have directly generated energy-time entangled triphoton W states in a hot atomic medium via spontaneous six-wave mixing for the first time. However, a rigorous theoretical framework for this process remains lacking to date, confining our understanding to a mere extension of the biphoton model. Here, we analytically investigate the generation mechanism of energy-time entangled triphotons and their classically controllable optical properties in an electromagnetically induced transparency-assisted five-level cold atomic system. Notably, triphoton generation follows strict temporal ordering, resulting in asymmetric quantum interference in triple coincidence counts--unreplicable and unexplainable by the inherently symmetric biphoton model. These results establish a rigorous physical framework for spontaneous six-wave mixing-generated triphotons, clarify their distinctions from states produced via cascaded nonlinear models, and substantially advance their utility in quantum information protocols.

quant-ph↗

Unsupervised Detection of Topological Phase Transitions with a Quantum Reservoir

In quantum many-body systems, characterizing topological phase transitions typically requires complex many-body topological invariants, which are costly to compute and measure. Inspired by quantum reservoir computing, we propose an unsupervised quantum phase detection method based on a many-body localized evolution, enabling efficient identification of phase transitions in the extended SSH model. The evolved quantum states produce feature distributions under local measurements, which, after simple post-processing and dimensionality reduction, naturally cluster according to different Hamiltonian parameters. Numerical simulations show that the evolution combined with local measurements can significantly amplify distinctions between quantum states, providing an efficient means to detect topological phase transitions. Our approach requires neither complex measurements nor full density matrix reconstruction, making it practical and feasible for noisy intermediate-scale quantum devices.

quant-ph↗

Robust and Efficient Quantum Reservoir Computing with Discrete Time Crystal

The rapid development of machine learning and quantum computing has placed quantum machine learning at the forefront of research. However, existing quantum machine learning algorithms based on quantum variational algorithms face challenges in trainability and noise robustness. In order to address these challenges, we introduce a gradient-free, noise-robust quantum reservoir computing algorithm that harnesses discrete time crystal dynamics as a reservoir. We first calibrate the memory, nonlinear, and information scrambling capacities of the quantum reservoir, revealing their correlation with dynamical phases and non-equilibrium phase transitions. We then apply the algorithm to the binary classification task and establish a comparative quantum kernel advantage. For ten-class classification, both noisy simulations and experimental results on superconducting quantum processors match ideal simulations, demonstrating the enhanced accuracy with increasing system size and confirming the topological noise robustness. Our work presents the first experimental demonstration of quantum reservoir computing for image classification based on digital quantum simulation. It establishes the correlation between quantum many-body non-equilibrium phase transitions and quantum machine learning performance, providing new design principles for quantum reservoir computing and broader quantum machine learning algorithms in the NISQ era.

quant-ph↗

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby restricting input flexibility and increasing the number of training parameters. To address these challenges, we propose StitchFusion, a straightforward yet effective modal fusion framework that integrates large-scale pre-trained models directly as encoders and feature fusers. This approach facilitates comprehensive multi-modal and multi-scale feature fusion, accommodating any visual modal inputs. Specifically, Our framework achieves modal integration during encoding by sharing multi-modal visual information. To enhance information exchange across modalities, we introduce a multi-directional adapter module (MultiAdapter) to enable cross-modal information transfer during encoding. By leveraging MultiAdapter to propagate multi-scale information across pre-trained encoders during the encoding process, StitchFusion achieves multi-modal visual information integration during encoding. Extensive comparative experiments demonstrate that our model achieves state-of-the-art performance on four multi-modal segmentation datasets with minimal additional parameters. Furthermore, the experimental integration of MultiAdapter with existing Feature Fusion Modules (FFMs) highlights their complementary nature. Our code is available at StitchFusion_repo.

cs.CV↗

SVGen: Interpretable Vector Graphics Generation with Large Language Models

Scalable Vector Graphics (SVG) is widely used in front-end development and UI/UX design due to its scalability, editability, and rendering efficiency. However, turning creative ideas into precise vector graphics remains a time-consuming challenge. To address this, we introduce SVG-1M, a large-scale dataset of high-quality SVGs paired with natural language descriptions. Through advanced data augmentation and annotation, we create well-aligned Text to SVG training pairs, including a subset with Chain of Thought annotations for enhanced semantic guidance. Based on this dataset, we propose SVGen, an end-to-end model that generates SVG code from natural language inputs. Our approach ensures semantic accuracy and structural completeness, supported by curriculum learning and reinforcement learning optimization. Experiments show that SVGen outperforms general large models and traditional rendering methods in both effectiveness and efficiency. Code, model, and dataset are available on GitHub.

cs.LG↗

Efficient quantum state tomography with auxiliary systems

Quantum state tomography is a technique in quantum information science used to reconstruct the density matrix of an unknown quantum state, providing complete information about the quantum state. It is of significant importance in fields such as quantum computation, quantum communication, and quantum simulation. However, as the size of the quantum system increases, the number of measurement settings and sampling requirements for quantum state tomography grow exponentially with the number of qubits. This not only makes experimental design and implementation more complex, but also exacerbates the consumption of experimental resources. These limitations severely hinder the application of state tomography in large-scale quantum systems. To reduce measurement settings and improve sampling efficiency, this study proposes a state tomography method based on auxiliary systems. This method can be implemented through either entanglement between the quantum system to be measured and a quantum auxiliary system or through correlation between the quantum system and a probabilistic classical auxiliary system. Measurements on the entire joint system enable more efficient extraction of information about the quantum state to be measured. This method relies on standard quantum gate operations and requires only two measurement settings, with a total sampling complexity of $O(d^2)$, significantly simplifying experimental operations and measurement processes. Additionally, this study provides two schemes for measuring purity based on the proposed circuit, one of which achieves measurement precision at the Heisenberg limit. This study validates the effectiveness of the proposed method through a detailed theoretical analysis, a series of numerical simulations, and experiments.

quant-ph↗

One-way network nonlocality of continuous variable entangled networks

Nonlocality is a key feature of quantum networks and is being studied for its potential applications in quantum communication and computing. Understanding and harnessing nonlocality in quantum networks could lead to the development of faster and more secure communication systems. All the nonclassicalities are limited to discrete variable quantum networks. We propose the first method to verify the network nonlocality of all optical quantum network consisting of two entangled states where one-way classical communication is allowed. This provides the first device-independent method to verify the quantum correlations generated from all optical continuous-variable quantum networks.

quant-ph↗

Dynamic Proxy Domain Generalizes the Crowd Localization by Better Binary Segmentation

Crowd localization targets on predicting each instance precise location within an image. Current advanced methods propose the pixel-wise binary classification to tackle the congested prediction, in which the pixel-level thresholds binarize the prediction confidence of being the pedestrian head. Since the crowd scenes suffer from extremely varying contents, counts and scales, the confidence-threshold learner is fragile and under-generalized encountering domain knowledge shift. Moreover, at the most time, the target domain is agnostic in training. Hence, it is imperative to exploit how to enhance the generalization of confidence-threshold locator to the latent target domain. In this paper, we propose a Dynamic Proxy Domain (DPD) method to generalize the learner under domain shift. Concretely, based on the theoretical analysis to the generalization error risk upper bound on the latent target domain to a binary classifier, we propose to introduce a generated proxy domain to facilitate generalization. Then, based on the theory, we design a DPD algorithm which is composed by a training paradigm and proxy domain generator to enhance the domain generalization of the confidence-threshold learner. Besides, we conduct our method on five kinds of domain shift scenarios, demonstrating the effectiveness on generalizing the crowd localization. Our code will be available at https://github.com/zhangda1018/DPD.

cs.CV↗

Performance Characterizations and Usage Guidelines of Samsung CXL Memory Module Hybrid Prototype

The growing prevalence of data-intensive workloads, such as artificial intelligence (AI), machine learning (ML), high-performance computing (HPC), in-memory databases, and real-time analytics, has exposed limitations in conventional memory technologies like DRAM. While DRAM offers low latency and high throughput, it is constrained by high costs, scalability challenges, and volatility, making it less viable for capacity-bound and persistent applications in modern datacenters. Recently, Compute Express Link (CXL) has emerged as a promising alternative, enabling high-speed, cacheline-granular communication between CPUs and external devices. By leveraging CXL technology, NAND flash can now be used as memory expansion, offering three-fold benefits: byte-addressability, scalable capacity, and persistence at a low cost. Samsung's CXL Memory Module Hybrid (CMM-H) is the first product to deliver these benefits through a hardware-only solution, i.e., it does not incur any OS and IO overheads like conventional block devices. In particular, CMM-H integrates a DRAM cache with NAND flash in a single device to deliver near-DRAM latency. This paper presents the first publicly available study for comprehensive characterizations of an FPGA-based CMM-H prototype. Through this study, we address users' concerns about whether a wide variety of applications can successfully run on a memory device backed by NAND flash medium. Additionally, based on these characterizations, we provide key insights into how to best take advantage of the CMM-H device.

cs.AR↗

FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation

Open-vocabulary segmentation aims to identify and segment specific regions and objects based on text-based descriptions. A common solution is to leverage powerful vision-language models (VLMs), such as CLIP, to bridge the gap between vision and text information. However, VLMs are typically pretrained for image-level vision-text alignment, focusing on global semantic features. In contrast, segmentation tasks require fine-grained pixel-level alignment and detailed category boundary information, which VLMs alone cannot provide. As a result, information extracted directly from VLMs can't meet the requirements of segmentation tasks. To address this limitation, we propose FGAseg, a model designed for fine-grained pixel-text alignment and category boundary supplementation. The core of FGAseg is a Pixel-Level Alignment module that employs a cross-modal attention mechanism and a text-pixel alignment loss to refine the coarse-grained alignment from CLIP, achieving finer-grained pixel-text semantic alignment. Additionally, to enrich category boundary information, we introduce the alignment matrices as optimizable pseudo-masks during forward propagation and propose Category Information Supplementation module. These pseudo-masks, derived from cosine and convolutional similarity, provide essential global and local boundary information between different categories. By combining these two strategies, FGAseg effectively enhances pixel-level alignment and category boundary information, addressing key challenges in open-vocabulary segmentation. Extensive experiments demonstrate that FGAseg outperforms existing methods on open-vocabulary semantic segmentation benchmarks.

cs.CV↗

Quantum-inspired Interpretable Deep Learning Architecture for Text Sentiment Analysis

Text has become the predominant form of communication on social media, embedding a wealth of emotional nuances. Consequently, the extraction of emotional information from text is of paramount importance. Despite previous research making some progress, existing text sentiment analysis models still face challenges in integrating diverse semantic information and lack interpretability. To address these issues, we propose a quantum-inspired deep learning architecture that combines fundamental principles of quantum mechanics (QM principles) with deep learning models for text sentiment analysis. Specifically, we analyze the commonalities between text representation and QM principles to design a quantum-inspired text representation method and further develop a quantum-inspired text embedding layer. Additionally, we design a feature extraction layer based on long short-term memory (LSTM) networks and self-attention mechanisms (SAMs). Finally, we calculate the text density matrix using the quantum complex numbers principle and apply 2D-convolution neural networks (CNNs) for feature condensation and dimensionality reduction. Through a series of visualization, comparative, and ablation experiments, we demonstrate that our model not only shows significant advantages in accuracy and efficiency compared to previous related models but also achieves a certain level of interpretability by integrating QM principles. Our code is available at QISA.

cs.CV↗

When Box Meets Graph Neural Network in Tag-aware Recommendation

Last year has witnessed the re-flourishment of tag-aware recommender systems supported by the LLM-enriched tags. Unfortunately, though large efforts have been made, current solutions may fail to describe the diversity and uncertainty inherent in user preferences with only tag-driven profiles. Recently, with the development of geometry-based techniques, e.g., box embedding, diversity of user preferences now could be fully modeled as the range within a box in high dimension space. However, defect still exists as these approaches are incapable of capturing high-order neighbor signals, i.e., semantic-rich multi-hop relations within the user-tag-item tripartite graph, which severely limits the effectiveness of user modeling. To deal with this challenge, in this paper, we propose a novel algorithm, called BoxGNN, to perform the message aggregation via combination of logical operations, thereby incorporating high-order signals. Specifically, we first embed users, items, and tags as hyper-boxes rather than simple points in the representation space, and define two logical operations to facilitate the subsequent process. Next, we perform the message aggregation mechanism via the combination of logical operations, to obtain the corresponding high-order box representations. Finally, we adopt a volume-based learning objective with Gumbel smoothing techniques to refine the representation of boxes. Extensive experiments on two publicly available datasets and one LLM-enhanced e-commerce dataset have validated the superiority of BoxGNN compared with various state-of-the-art baselines. The code is released online

cs.IR↗

U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation

Multimodal semantic segmentation is a pivotal component of computer vision and typically surpasses unimodal methods by utilizing rich information set from various sources.Current models frequently adopt modality-specific frameworks that inherently biases toward certain modalities. Although these biases might be advantageous in specific situations, they generally limit the adaptability of the models across different multimodal contexts, thereby potentially impairing performance. To address this issue, we leverage the inherent capabilities of the model itself to discover the optimal equilibrium in multimodal fusion and introduce U3M: An Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation. Specifically, this method involves an unbiased integration of multimodal visual data. Additionally, we employ feature fusion at multiple scales to ensure the effective extraction and integration of both global and local features. Experimental results demonstrate that our approach achieves superior performance across multiple datasets, verifing its efficacy in enhancing the robustness and versatility of semantic segmentation in diverse settings. Our code is available at U3M-multimodal-semantic-segmentation.

cs.CV↗

Diverse Entanglement Mechanisms in Multimode Nonlinear Continuous Variables

Non-Gaussian entangled states play a crucial role in harnessing quantum advantage in continuous-variable quantum information. However, how to fully characterize N-partite (N > 3) non-Gaussian entanglement without quantum state tomography remains elusive, leading to a very limited understanding of the underlying entanglement mechanism. Here, we propose several necessary and sufficient conditions for the positive-partial-transposition separability of multimode nonlinear quantum states resulting from high-order Hamiltonians and successive beam splitting operations. When applied to the initial state, the beam-splitter operations induce the emergence of different types of entanglement mechanisms, including pairwise high-order entanglement, collective high-order entanglement and the crossover between the two. We show numerically that for the four-mode scenario, the threshold for the existence of entanglement for any bipartition does not exceed the entanglement of the original state at fixed high-order moments. These results provide a new perspective for understanding multipartite nonlinear entanglement and will promote their application in quantum information processing.

quant-ph↗

Certification of non-Gaussian Einstein-Podolsky-Rosen Steering

Non-Gaussian quantum states are a known necessary resource for reaching a quantum advantage and for violating Bell inequalities in continuous variable systems. As one kind of manifestation of quantum correlations, Einstein-Podolsky-Rosen (EPR) steering enables verification of shared entanglement even when one of the subsystems is not characterized. However, how to detect and classify such an effect for non-Gaussian states is far from being well understood. Here, we present an efficient non-Gaussian steering criterion based on the high-order observables and conduct a systematic investigation into the hierarchy of non-Gaussian steering criteria. Moreover, we apply our criterion to three experimentally-relevant non-Gaussian states under realistic conditions and, in particular, propose a feasible scheme to create multi-component cat states with tunable size by performing a suitable high-order quadrature measurement on the steering party. Our work reveals the fundamental characteristics of non-Gaussianity and quantum correlations, and offers new insights to explore their applications in quantum information processing.

quant-ph↗