SearcharxivSearch

arXiv subjects

Peng-Fei Zhang

Publications and source records attributed to Peng-Fei Zhang.

At least 19 recordsLinked to original sources

Hausdorff dimension of images and graphs of some random complex series

Let $\{X_n= e^{2\pi i \theta_n}\}$ be a sequence of Steinhaus random variables, where $\theta_n$ are independent and uniformly distributed on $[0,1]$. We compute the almost sure Hausdorff dimension of the images and graphs of the random complex series $S(x)=\sum_{n=1}^{\infty}a_n X_n\phi_n(\lambda_nx)$, where $\lambda_n$ is an increasing sequence with $\sup_n\lambda_{n+1}/\lambda_n<\infty$ and $\phi_n$ satisfies some uniform Lipschitz and boundedness conditions. This class of series includes the famous Weierstrass and Riemann functions as well as others appeared in literature. These results help predict the exact values of the deterministic cases.

math.CA

Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models

Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical Refinement Attack (HRA), a multimodal universal attack framework for VLP models. For the image modality, we refine the optimization path by leveraging a temporal hierarchy of historical and estimated future gradients to avoid local minima and stabilize universal perturbation learning. For the text modality, it hierarchically models textual importance by considering both intra- and inter-sentence contributions to identify globally influential words, which are then used as universal text perturbations. Extensive experiments across various downstream tasks, VLP models, and datasets, demonstrate the superior transferability of the proposed universal multimodal attacks.

cs.CV

A Step Toward World Models: A Survey on Robotic Manipulation

Autonomous agents are increasingly expected to operate in complex, dynamic, and uncertain environments, performing tasks such as manipulation, navigation, and decision-making. Achieving these capabilities requires agents to understand the underlying mechanisms and dynamics of the world, moving beyond reactive control or simple replication of observed states. This motivates the development of world models as internal representations that encode environmental states, capture dynamics, and support prediction, planning, and reasoning. Despite growing interest, the definition, scope, architectures, and essential capabilities of world models remain ambiguous. In this survey, we go beyond prescribing a fixed definition and limiting our scope to methods explicitly labeled as world models. Instead, we examine approaches that exhibit the core capabilities of world models through a review of methods in robotic manipulation. We analyze their roles across perception, prediction, and control, identify key challenges and solutions, and distill the core components, capabilities, and functions that a fully realized world model should possess. Building on this analysis, we aim to motivate further development toward generalizable and practical world models for robotics.

cs.RO

MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models

Current adversarial attacks for evaluating the robustness of vision-language pre-trained (VLP) models in multi-modal tasks suffer from limited transferability, where attacks crafted for a specific model often struggle to generalize effectively across different models, limiting their utility in assessing robustness more broadly. This is mainly attributed to the over-reliance on model-specific features and regions, particularly in the image modality. In this paper, we propose an elegant yet highly effective method termed Meticulous Adversarial Attack (MAA) to fully exploit model-independent characteristics and vulnerabilities of individual samples, achieving enhanced generalizability and reduced model dependence. MAA emphasizes fine-grained optimization of adversarial images by developing a novel resizing and sliding crop (RScrop) technique, incorporating a multi-granularity similarity disruption (MGSD) strategy. Extensive experiments across diverse VLP models, multiple benchmark datasets, and a variety of downstream tasks demonstrate that MAA significantly enhances the effectiveness and transferability of adversarial attacks. A large cohort of performance studies is conducted to generate insights into the effectiveness of various model configurations, guiding future advancements in this domain.

cs.CV

Proposal of quantum repeater architecture based on Rydberg atom quantum processors

Realizing large-scale quantum networks requires the generation of high-fidelity quantum entanglement states between remote quantum nodes, a key resource for quantum communication, distributed computation and sensing applications. However, entanglement distribution between quantum network nodes is hindered by optical transmission loss and local operation errors. Here, we propose a novel quantum repeater architecture that synergistically integrates Rydberg atom quantum processors with optical cavities to overcome these challenges. Our scheme leverages cavity-mediated interactions for efficient remote entanglement generation, followed by Rydberg interaction-based entanglement purification and swapping. Numerical simulations, incorporating realistic experimental parameters, demonstrate the generation of Bell states with 99\% fidelity at rates of 1.1\,kHz between two nodes in local-area network (distance $0.1\,\mathrm{km}$), and can be extend to metropolitan-area ($25\,\mathrm{km}$) or intercity ($\mathrm{250\,\mathrm{km}}$, with the assitance of frequency converters) network with a rate of 0.1\,kHz. This scalable approach opens up near-term opportunities for exploring quantum network applications and investigating the advantages of distributed quantum information processing.

quant-ph

Universal Adversarial Perturbations for Vision-Language Pre-trained Models

Vision-language pre-trained (VLP) models have been the foundation of numerous vision-language tasks. Given their prevalence, it becomes imperative to assess their adversarial robustness, especially when deploying them in security-crucial real-world applications. Traditionally, adversarial perturbations generated for this assessment target specific VLP models, datasets, and/or downstream tasks. This practice suffers from low transferability and additional computation costs when transitioning to new scenarios. In this work, we thoroughly investigate whether VLP models are commonly sensitive to imperceptible perturbations of a specific pattern for the image modality. To this end, we propose a novel black-box method to generate Universal Adversarial Perturbations (UAPs), which is so called the Effective and T ransferable Universal Adversarial Attack (ETU), aiming to mislead a variety of existing VLP models in a range of downstream tasks. The ETU comprehensively takes into account the characteristics of UAPs and the intrinsic cross-modal interactions to generate effective UAPs. Under this regime, the ETU encourages both global and local utilities of UAPs. This benefits the overall utility while reducing interactions between UAP units, improving the transferability. To further enhance the effectiveness and transferability of UAPs, we also design a novel data augmentation method named ScMix. ScMix consists of self-mix and cross-mix data transformations, which can effectively increase the multi-modal data diversity while preserving the semantics of the original data. Through comprehensive experiments on various downstream tasks, VLP models, and datasets, we demonstrate that the proposed method is able to achieve effective and transferrable universal adversarial attacks.

cs.CV

Effective and Robust Adversarial Training against Data and Label Corruptions

Corruptions due to data perturbations and label noise are prevalent in the datasets from unreliable sources, which poses significant threats to model training. Despite existing efforts in developing robust models, current learning methods commonly overlook the possible co-existence of both corruptions, limiting the effectiveness and practicability of the model. In this paper, we develop an Effective and Robust Adversarial Training (ERAT) framework to simultaneously handle two types of corruption (i.e., data and label) without prior knowledge of their specifics. We propose a hybrid adversarial training surrounding multiple potential adversarial perturbations, alongside a semi-supervised learning based on class-rebalancing sample selection to enhance the resilience of the model for dual corruption. On the one hand, in the proposed adversarial training, the perturbation generation module learns multiple surrogate malicious data perturbations by taking a DNN model as the victim, while the model is trained to maintain semantic consistency between the original data and the hybrid perturbed data. It is expected to enable the model to cope with unpredictable perturbations in real-world data corruption. On the other hand, a class-rebalancing data selection strategy is designed to fairly differentiate clean labels from noisy labels. Semi-supervised learning is performed accordingly by discarding noisy labels. Extensive experiments demonstrate the superiority of the proposed ERAT framework.

cs.LG

Balanced and Explainable Social Media Analysis for Public Health with Large Language Models

As social media becomes increasingly popular, more and more public health activities emerge, which is worth noting for pandemic monitoring and government decision-making. Current techniques for public health analysis involve popular models such as BERT and large language models (LLMs). Although recent progress in LLMs has shown a strong ability to comprehend knowledge by being fine-tuned on specific domain datasets, the costs of training an in-domain LLM for every specific public health task are especially expensive. Furthermore, such kinds of in-domain datasets from social media are generally highly imbalanced, which will hinder the efficiency of LLMs tuning. To tackle these challenges, the data imbalance issue can be overcome by sophisticated data augmentation methods for social media datasets. In addition, the ability of the LLMs can be effectively utilised by prompting the model properly. In light of the above discussion, in this paper, a novel ALEX framework is proposed for social media analysis on public health. Specifically, an augmentation pipeline is developed to resolve the data imbalance issue. Furthermore, an LLMs explanation mechanism is proposed by prompting an LLM with the predicted results from BERT models. Extensive experiments conducted on three tasks at the Social Media Mining for Health 2023 (SMM4H) competition with the first ranking in two tasks demonstrate the superior performance of the proposed ALEX method. Our code has been released in https://github.com/YanJiangJerry/ALEX.

cs.CL

First look at data from the 13-antenna setup of GRANDProto300 in northwest China

The Giant Radio Array for Neutrino Detection (GRAND) is an envisioned observatory of ultra-high-energy neutrinos, cosmic rays, and gamma rays, with energies above 100 PeV. GRAND targets the radio signals emitted by extensive air showers induced by the interaction of ultra-high-energy particles in the atmosphere, using an array of 200,000 radio antennas split into sub-arrays deployed worldwide. GRANDProto13 (GP13) is a 13-antenna demonstrator array deployed in February 2023 in the Gansu province of China, as a precursor for GRANDProto300, which will validate the detection principle of the GRAND experiment. Its goal is to measure the radio background present at the site, validate the design of the detection units and develop an autonomous radio trigger for air showers. We will describe GP13 and its operation, and show preliminary results on noise monitoring.

astro-ph.IM

Determination of geopotential difference by hydrogen masers based on precise point positioning time-frequency transfer

According to the general relativity theory, the geopotential difference can be determined by gravity frequency shift between two clocks. Here we report on the experiments to determine the geopotential difference between two remote sites by hydrogen masers based on precise point positioning time-frequency transfer technique. The experiments include the remote clock comparison and the local clock comparison using two CH1-95 active hydrogen masers linked with global navigation satellite system time-frequency receivers. The frequency difference between two hydrogen masers at two sites is derived from the time difference series resolved by the above-mentioned technique. Considering the local clock comparison as calibration, the determined geopotential difference by our experiments is 12,142.3 (112.4) m^2/s^2, quite close to the value 12,153.3 (2.3) m^2/s^2 computed by the EIGEN-6C4 model. Results show that the proposed approach here for determining geopotential difference is feasible, operable, and promising for applications in various fields.

physics.geo-ph

FedVMR: A New Federated Learning method for Video Moment Retrieval

Despite the great success achieved, existing video moment retrieval (VMR) methods are developed under the assumption that data are centralizedly stored. However, in real-world applications, due to the inherent nature of data generation and privacy concerns, data are often distributed on different silos, bringing huge challenges to effective large-scale training. In this work, we try to overcome above limitation by leveraging the recent success of federated learning. As the first that is explored in VMR field, the new task is defined as video moment retrieval with distributed data. Then, a novel federated learning method named FedVMR is proposed to facilitate large-scale and secure training of VMR models in decentralized environment. Experiments on benchmark datasets demonstrate its effectiveness. This work is the very first attempt to enable safe and efficient VMR training in decentralized scene, which is hoped to pave the way for further study in the related research field.

cs.CV

Self-supervised Graph-based Point-of-interest Recommendation

The exponential growth of Location-based Social Networks (LBSNs) has greatly stimulated the demand for precise location-based recommendation services. Next Point-of-Interest (POI) recommendation, which aims to provide personalised POI suggestions for users based on their visiting histories, has become a prominent component in location-based e-commerce. Recent POI recommenders mainly employ self-attention mechanism or graph neural networks to model complex high-order POI-wise interactions. However, most of them are merely trained on the historical check-in data in a standard supervised learning manner, which fail to fully explore each user's multi-faceted preferences, and suffer from data scarcity and long-tailed POI distribution, resulting in sub-optimal performance. To this end, we propose a Self-s}upervised Graph-enhanced POI Recommender (S2GRec) for next POI recommendation. In particular, we devise a novel Graph-enhanced Self-attentive layer to incorporate the collaborative signals from both global transition graph and local trajectory graphs to uncover the transitional dependencies among POIs and capture a user's temporal interests. In order to counteract the scarcity and incompleteness of POI check-ins, we propose a novel self-supervised learning paradigm in \ssgrec, where the trajectory representations are contrastively learned from two augmented views on geolocations and temporal transitions. Extensive experiments are conducted on three real-world LBSN datasets, demonstrating the effectiveness of our model against state-of-the-art methods.

cs.LG

Self-induced optical non-reciprocity

Non-reciprocal optical components are indispensable in optical applications, and their realization without any magnetic field arose increasing research interests in photonics. Exciting experimental progress has been achieved by either introducing spatial-temporal modulation of the optical medium or combining Kerr-type optical nonlinearity with spatial asymmetry in photonic structures. However, extra driving fields are required for the first approach, while the isolation of noise and the transmission of the signal cannot be simultaneously achieved for the other approach. Here, we experimentally demonstrate a new concept of nonlinear non-reciprocal susceptibility for optical media and realize the completely passive isolation of optical signals without any external bias field. The self-induced isolation by the input signal is demonstrated with an extremely high isolation ratio of 63.4 dB, a bandwidth of 2.1 GHz for 60 dB isolation, and a low insertion loss of around 1 dB. Furthermore, novel functional optical devices are realized, including polarization purification and non-reciprocal leverage. The demonstrated nonlinear non-reciprocity provides a versatile tool to control light and deepen our understanding of light-matter interactions, and enables applications ranging from topological photonics to unidirectional quantum information transfer in a network.

physics.optics

Discovering Domain Disentanglement for Generalized Multi-source Domain Adaptation

A typical multi-source domain adaptation (MSDA) approach aims to transfer knowledge learned from a set of labeled source domains, to an unlabeled target domain. Nevertheless, prior works strictly assume that each source domain shares the identical group of classes with the target domain, which could hardly be guaranteed as the target label space is not observable. In this paper, we consider a more versatile setting of MSDA, namely Generalized Multi-source Domain Adaptation, wherein the source domains are partially overlapped, and the target domain is allowed to contain novel categories that are not presented in any source domains. This new setting is more elusive than any existing domain adaptation protocols due to the coexistence of the domain and category shifts across the source and target domains. To address this issue, we propose a variational domain disentanglement (VDD) framework, which decomposes the domain representations and semantic features for each instance by encouraging dimension-wise independence. To identify the target samples of unknown classes, we leverage online pseudo labeling, which assigns the pseudo-labels to unlabeled target data based on the confidence scores. Quantitative and qualitative experiments conducted on two benchmark datasets demonstrate the validity of the proposed framework.

cs.LG

Lightweight Self-Attentive Sequential Recommendation

Modern deep neural networks (DNNs) have greatly facilitated the development of sequential recommender systems by achieving state-of-the-art recommendation performance on various sequential recommendation tasks. Given a sequence of interacted items, existing DNN-based sequential recommenders commonly embed each item into a unique vector to support subsequent computations of the user interest. However, due to the potentially large number of items, the over-parameterised item embedding matrix of a sequential recommender has become a memory bottleneck for efficient deployment in resource-constrained environments, e.g., smartphones and other edge devices. Furthermore, we observe that the widely-used multi-head self-attention, though being effective in modelling sequential dependencies among items, heavily relies on redundant attention units to fully capture both global and local item-item transition patterns within a sequence. In this paper, we introduce a novel lightweight self-attentive network (LSAN) for sequential recommendation. To aggressively compress the original embedding matrix, LSAN leverages the notion of compositional embeddings, where each item embedding is composed by merging a group of selected base embedding vectors derived from substantially smaller embedding matrices. Meanwhile, to account for the intrinsic dynamics of each item, we further propose a temporal context-aware embedding composition scheme. Besides, we develop an innovative twin-attention network that alleviates the redundancy of the traditional multi-head self-attention while retaining full capacity for capturing long- and short-term (i.e., global and local) item dependencies. Comprehensive experiments demonstrate that LSAN significantly advances the accuracy and memory efficiency of existing sequential recommenders.

cs.IR

Spectral Deconvolution Analysis on Olivine-Orthopyroxene Mixtures with Simulated Space Weathering Modifications

Olivine and pyroxene are important mineral end-members for studying the sur-face material compositions of mafic bodies. The profiles of visible and near-infraredspectra of olivine-orthopyroxene mixtures systematically varied with their compositionratios. In our experiments, we combine the RELAB spectral database with a new spec-tral data obtained from some assembled olivine-orthopyroxene mixtures. We found thatthe commonly-used band area ratio (BAR, Cloutis et al. 1986) does not work well onour newly obtained spectral data. To investigate this issue, an empirical procedure basedon fitted results by modified Gaussian model is proposed to analyze the spectral curves.Following the new empirical procedure, the end-member abundances can be estimatedwith a 15% accuracy with some prior mineral absorption features. In addition, the mix-ture samples configured in our experiments are also irradiated by pulsed lasers to simulateand investigate the space weathering effects. Spectral deconvolution results confirm thatlow-content olivine on celestial bodies are difficult to measure and estimate. Therefore,the olivine abundance of space weathered materials may be underestimated from remotesensing data. This study may be used to quantify the spectral relationship of olivine-orthopyroxene mixtures and further reveal their correlation between the spectra of ordi-nary chondrites and silicate asteroids.

astro-ph.EP

A gamma-ray periodic modulation in Globular Cluster 47 Tucanae

The Globular Cluster 47 Tucanae was firstly detected in gamma-rays by the Large Area Telescope (LAT) onboard the \emph{Fermi} Gamma-ray Space Telescope, and the gamma-ray emission has been widely attributed to the millisecond pulsars. In this work, we analyze the Fermi-LAT pass 8 data ranging from 2008 August to 2017 May and report the detection of a modulation with a period of $18.416\pm0.008$ hours at a significance level of $\sim 4.8σ$. This is the first time to detect a significant modulation with a period much longer than that of millisecond pulsars in gamma-rays from Globular Clusters. The periodic modulation signal appears in the {\it Swift}-BAT data as well. The phase-folded Chandra X-ray light curve of a point source may have provided an additional clue.

astro-ph.HE

Study on Estimating Quantum Discord by Neural Network with Prior Knowledge

Machine learning has achieved success in many areas because of its powerful fitting ability, so we hope it can help us to solve some significant physical quantitative problems, such as quantum correlation. In this research we will use neural networks to predict the value of quantum discord. Quantum discord is a measure of quantum correlation which is defined as the difference between quantum mutual information and classical correlation for a bipartite system. Since the definition contains an optimization term, it makes analytically solving hard. For some special cases and small systems, such as two-qubit systems and some X-states, the explicit solutions have been calculated. However, for general cases, we still know very little. Therefore, we study the feasibility of estimating quantum discord by machine learning method on two-qubit systems. In order to get an interpretable and high performance model, we modify the ordinary neural network by introducing some prior knowledge which come from the analysis about quantum discord. Our results show that prior knowledge actually improve the performance of neural network.

quant-ph