Searcharxiv⌕ Search

arXiv subjects

Feng He

Publications and source records attributed to Feng He.

At least 37 records · Page 2Linked to original sources

Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains

In the big data era, the computer vision field benefits from large-scale datasets such as LAION-2B, LAION-400M, and ImageNet-21K, Kinetics, on which popular models like the ViT and ConvNeXt series have been pre-trained, acquiring substantial knowledge. However, numerous downstream tasks in specialized and data-limited scientific domains continue to pose significant challenges. In this paper, we propose a novel Cluster Attention Adapter (CLAdapter), which refines and adapts the rich representations learned from large-scale data to various data-limited downstream tasks. Specifically, CLAdapter introduces attention mechanisms and cluster centers to personalize the enhancement of transformed features through distribution correlation and transformation matrices. This enables models fine-tuned with CLAdapter to learn distinct representations tailored to different feature sets, facilitating the models' adaptation from rich pre-trained features to various downstream scenarios effectively. In addition, CLAdapter's unified interface design allows for seamless integration with multiple model architectures, including CNNs and Transformers, in both 2D and 3D contexts. Through extensive experiments on 10 datasets spanning domains such as generic, multimedia, biological, medical, industrial, agricultural, environmental, geographical, materials science, out-of-distribution (OOD), and 3D analysis, CLAdapter achieves state-of-the-art performance across diverse data-limited scientific domains, demonstrating its effectiveness in unleashing the potential of foundation vision models via adaptive transfer. Code is available at https://github.com/qklee-lz/CLAdapter.

cs.CV↗

Benchmarking Atomic Ionization Driven by Strong Quantum Light

The recently available high-intensity quantum light pulses provide novel tools for controlling light-matter interactions. However, the rigor of the theoretical frameworks currently used to describe the interaction of strong quantum light with atoms and molecules remains unverified. Here, we establish a rigorous benchmark by solving the fully quantized time-dependent Schrödinger equation for an atom exposed to bright squeezed vacuum light. Our \textit{ab initio} simulations reveal a critical limitation of the widely used $Q$-representation: although it accurately reproduces the total photoelectron spectrum after tracing over photon states, it completely fails to capture the electron-photon joint energy spectrum. To overcome this limitation, we develop a general theoretical framework based on the Feynman path integral that properly incorporates the electron-photon quantum entanglement. Our results provide both quantitative benchmarks and fundamental theoretical insights for the emerging field of strong-field quantum optics.

quant-ph↗

Generalizing Graph Transformers Across Diverse Graphs and Tasks via Pre-training

Graph pre-training has been concentrated on graph-level tasks involving small graphs (e.g., molecular graphs) or learning node representations on a fixed graph. Extending graph pre-trained models to web-scale graphs with billions of nodes in industrial scenarios, while avoiding negative transfer across graphs or tasks, remains a challenge. We aim to develop a general graph pre-trained model with inductive ability that can make predictions for unseen new nodes and even new graphs. In this work, we introduce a scalable transformer-based graph pre-training framework called PGT (Pre-trained Graph Transformer). Based on the masked autoencoder architecture, we design two pre-training tasks: one for reconstructing node features and the other for reconstructing local structures. Unlike the original autoencoder architecture where the pre-trained decoder is discarded, we propose a novel strategy that utilizes the decoder for feature augmentation. Our framework, tested on the publicly available ogbn-papers100M dataset with 111 million nodes and 1.6 billion edges, achieves state-of-the-art performance, showcasing scalability and efficiency. We have deployed our framework on Tencent's online game data, confirming its capability to pre-train on real-world graphs with over 540 million nodes and 12 billion edges and to generalize effectively across diverse static and dynamic downstream tasks.

cs.LG↗

The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies

Inspired by the rapid development of Large Language Models (LLMs), LLM agents have evolved to perform complex tasks. LLM agents are now extensively applied across various domains, handling vast amounts of data to interact with humans and execute tasks. The widespread applications of LLM agents demonstrate their significant commercial value; however, they also expose security and privacy vulnerabilities. At the current stage, comprehensive research on the security and privacy of LLM agents is highly needed. This survey aims to provide a comprehensive overview of the newly emerged privacy and security issues faced by LLM agents. We begin by introducing the fundamental knowledge of LLM agents, followed by a categorization and analysis of the threats. We then discuss the impacts of these threats on humans, environment, and other agents. Subsequently, we review existing defensive strategies, and finally explore future trends. Additionally, the survey incorporates diverse case studies to facilitate a more accessible understanding. By highlighting these critical security and privacy issues, the survey seeks to stimulate future research towards enhancing the security and privacy of LLM agents, thereby increasing their reliability and trustworthiness in future applications.

cs.CR↗

From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy

Light field microscopy (LFM) has become an emerging tool in neuroscience for large-scale neural imaging in vivo, notable for its single-exposure volumetric imaging, broad field of view, and high temporal resolution. However, learning-based 3D reconstruction in XLFM remains underdeveloped due to two core challenges: the absence of standardized datasets and the lack of methods that can efficiently model its angular-spatial structure while remaining physically grounded. We address these challenges by introducing three key contributions. First, we construct the XLFM-Zebrafish benchmark, a large-scale dataset and evaluation suite for XLFM reconstruction. Second, we propose Masked View Modeling for Light Fields (MVN-LF), a self-supervised task that learns angular priors by predicting occluded views, improving data efficiency. Third, we formulate the Optical Rendering Consistency Loss (ORC Loss), a differentiable rendering constraint that enforces alignment between predicted volumes and their PSF-based forward projections. On the XLFM-Zebrafish benchmark, our method improves PSNR by 7.7% over state-of-the-art baselines.

cs.CV↗

Statistical Signatures of Integrable and Non-Integrable Quantum Hamiltonians

Integrability is a cornerstone of classical mechanics, where it has a precise meaning. Extending this notion to quantum systems, however, remains subtle and unresolved. In particular, deciding whether a quantum Hamiltonian - viewed simply as a matrix - defines an integrable system is far from obvious, yet crucial for understanding non-equilibrium dynamics, spectral correlations, and correlation functions in many-body physics. We develop a statistical framework that approaches quantum integrability from a probabilistic standpoint. A key observation is that integrability requires a finite probability of vanishing energy gaps. Building on this, we propose a two-step protocol to distinguish integrable from non-integrable Hamiltonians. First, we apply a systematic Monte Carlo decimation of the spectrum, which exponentially compresses the Hilbert space and reveals whether level spacings approach Poisson statistics or remain mixed. The termination point of this decimation indicates the statistical character of the spectrum. Second, we analyze $k$-step gap distributions, which sharpen the distinction between Poisson and mixed statistics. Our procedure applies to Hamiltonians of any finite size, independent of whether their structure involves a few blocks or an exponentially fragmented Hilbert space. As a benchmark, we implement the protocol on quantum Hamiltonians built from the permutation group $\mathcal{S}_N$, demonstrating both its effectiveness and generality.

cond-mat.stat-mech↗

LPS-GNN : Deploying Graph Neural Networks on Graphs with 100-Billion Edges

Graph Neural Networks (GNNs) have emerged as powerful tools for various graph mining tasks, yet existing scalable solutions often struggle to balance execution efficiency with prediction accuracy. These difficulties stem from iterative message-passing techniques, which place significant computational demands and require extensive GPU memory, particularly when dealing with the neighbor explosion issue inherent in large-scale graphs. This paper introduces a scalable, low-cost, flexible, and efficient GNN framework called LPS-GNN, which can perform representation learning on 100 billion graphs with a single GPU in 10 hours and shows a 13.8% improvement in User Acquisition scenarios. We examine existing graph partitioning methods and design a superior graph partition algorithm named LPMetis. In particular, LPMetis outperforms current state-of-the-art (SOTA) approaches on various evaluation metrics. In addition, our paper proposes a subgraph augmentation strategy to enhance the model's predictive performance. It exhibits excellent compatibility, allowing the entire framework to accommodate various GNN algorithms. Successfully deployed on the Tencent platform, LPS-GNN has been tested on public and real-world datasets, achieving performance lifts of 8. 24% to 13. 89% over SOTA models in online applications.

cs.LG↗

Probing valence electron and hydrogen dynamics using charge-pair imaging with ultrafast electron diffraction

A key challenge in ultrafast science has been to directly track the coupled motions of electrons and nuclei in real-space and real-time. This study presents a significant step towards this goal by demonstrating the feasibility of time-resolved real-space tracking of valence electron and hydrogen dynamics during the photodissociation of ammonia (NH3) using MeV ultrafast electron diffraction. It is demonstrated that the enhanced temporal resolution, in conjunction with the analysis of the charge-pair distribution function, enables the disentanglement of the correlated motion of valence electrons and hydrogens in photoexcited ammonia molecule. The methodology employed in this study, which utilizes the charge-pair distribution function from ultrafast electron scattering to retrieve intertwined electron and nucleus dynamics, may open up new opportunities in the study of quantum dynamics for a wide range of molecules.

physics.chem-ph↗

How Robust is Model Editing after Fine-Tuning? An Empirical Study on Text-to-Image Diffusion Models

Model editing offers a low-cost technique to inject or correct a particular behavior in a pre-trained model without extensive retraining, supporting applications such as factual correction and bias mitigation. Despite this common practice, it remains unknown whether edits persist after fine-tuning or whether they are inadvertently reversed. This question has fundamental practical implications. For example, if fine-tuning removes prior edits, it could serve as a defence mechanism against hidden malicious edits. Vice versa, the unintended removal of edits related to bias mitigation could pose serious safety concerns. We systematically investigate the interaction between model editing and fine-tuning in the context of T2I diffusion models, which are known to exhibit biases and generate inappropriate content. Our study spans two T2I model families (Stable Diffusion and FLUX), two sota editing techniques, and three fine-tuning methods (DreamBooth, LoRA, and DoRA). Through an extensive empirical analysis across diverse editing tasks and evaluation metrics, our findings reveal a trend: edits generally fail to persist through fine-tuning, even when fine-tuning is tangential or unrelated to the edits. Notably, we observe that DoRA exhibits the strongest edit reversal effect. At the same time, among editing methods, UCE demonstrates greater robustness, retaining significantly higher efficacy post-fine-tuning compared to ReFACT. These findings highlight a crucial limitation in current editing methodologies, emphasizing the need for more robust techniques to ensure reliable long-term control and alignment of deployed AI systems. These findings have dual implications for AI safety: they suggest that fine-tuning could serve as a remediation mechanism for malicious edits while simultaneously highlighting the need for re-editing after fine-tuning to maintain beneficial safety and alignment properties.

cs.AI↗

ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs

Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought (Long CoT) reasoning have demonstrated remarkable cross-domain generalization capabilities. However, the underlying mechanisms supporting such transfer remain poorly understood. We hypothesize that cross-domain generalization arises from shared abstract reasoning prototypes -- fundamental reasoning patterns that capture the essence of problems across domains. These prototypes minimize the nuances of the representation, revealing that seemingly diverse tasks are grounded in shared reasoning structures.Based on this hypothesis, we propose ProtoReasoning, a framework that enhances the reasoning ability of LLMs by leveraging scalable and verifiable prototypical representations (Prolog for logical reasoning, PDDL for planning).ProtoReasoning features: (1) an automated prototype construction pipeline that transforms problems into corresponding prototype representations; (2) a comprehensive verification system providing reliable feedback through Prolog/PDDL interpreters; (3) the scalability to synthesize problems arbitrarily within prototype space while ensuring correctness. Extensive experiments show that ProtoReasoning achieves 4.7% improvement over baseline models on logical reasoning (Enigmata-Eval), 6.3% improvement on planning tasks, 4.0% improvement on general reasoning (MMLU) and 1.0% on mathematics (AIME24). Significantly, our ablation studies confirm that learning in prototype space also demonstrates enhanced generalization to structurally similar problems compared to training solely on natural language representations, validating our hypothesis that reasoning prototypes serve as the foundation for generalizable reasoning in large language models.

cs.CL↗

Coherent Control of Ion-Photoelectron Dynamics through Rabi Oscillations: An ab initio study

We present first-principles numerical simulations of photoionization in neon induced by bichromatic extreme ultraviolet pulses with frequencies $ω$ and $2ω$, specially chosen to make $ω$ equal to the energy difference between the $2s$ and $2p$ subshells. This allows for the production of photoelectrons from the $2s$ shell by $2ω$ pulse and from the $2p$ shell by $ω$ pulse with the same energy. Using the multi-configurational time-dependent Hartree-Fock method, we explore how Rabi coupling between subshells generates coherence between the corresponding photoelectron wave packets. Our \textit{ab initio} calculations confirm the analytical results derived from the essential-states approach in [K. L. Ishikawa, K. C. Prince, and K. Ueda, J. Phys. Chem. A 127, 10638 (2023)], validating the theoretical predictions. Although we focus on the Ne $2p$ and $2s$ subshells, our approach is applicable to a broad range of systems exhibiting photoionization from multiple subshells. The laser parameters employed in our simulations are available in modern Free Electron Lasers (FELs), and we anticipate that this work could stimulate experimental investigations using FELs to study ion-photoelectron coherence and entanglement.

physics.atom-ph↗

Time-dependent Hole States in Multiconfigurational Time-Dependent Hartree-Fock Approaches: Applications in Photoionization of Water Molecule

By simulating the real-time multielectron wavefunction with the multi-configurational time-dependent Hartree-Fock (MCTDHF) approach, we conduct an \textit{ab initio} study of the single-photon ionization process of a body-fixed water molecule ($\mathrm{H_2O}$) driven by attosecond pulses. To this end, we present a full-dimensional implementation of the MCTDHF method based on one-center expansions, allowing for the simulation of arbitrarily polarized lasers and multi-center polyatomic potentials. With a rigorous definition of the time-dependent hole state (TDHS) using the time-domain generalization of extended Koopmans' theorem (TD-EKT), we derive the reduced ion density matrix within the MCTDHF framework, which inherently encodes the total and channel-resolved photoionization cross sections of $\mathrm{H_2O}$. The cross sections obtained are benchmarked against existing experimental and theoretical results, validating the TDHS formalism. Furthermore, by adjusting the phase delay and intensity ratio of a pair of orthogonally polarized attosecond pulses, we explore the ultrafast control of attosecond coherence between electronic states of $\mathrm{H_2O^+}$.

physics.atom-ph↗

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

cs.CL↗

Efficient Adaptive Bandwidth Allocation for Deadline-Aware Online Admission Control in Time-Sensitive Networking

With the growing demand for dynamic real-time applications, online admission control for time-critical event-triggered (ET) traffic in Time-Sensitive Networking (TSN) has become a critical challenge. The main issue lies in dynamically allocating bandwidth with real-time guarantees in response to traffic changes while also meeting the requirements for rapid response, scalability, and high resource utilization in online scenarios. To address this challenge, we propose an online admission control method for ET traffic based on the TSN/ATS+CBS (asynchronous traffic shaper and credit-based shaper) architecture. This method provides a flexible framework for real-time guaranteed online admission control, supporting dynamic bandwidth allocation and reclamation at runtime without requiring global reconfiguration, thus improving scalability. Within this framework, we further integrate a novel strategy based on network calculus (NC) theory for efficient and high-utilization bandwidth reallocation. On the one hand, the strategy focuses on adaptively balancing residual bandwidth with deadline awareness to prevent bottleneck egress ports, thereby improving admission capacity. On the other hand, it employs a non-trivial analytical result to reduce the search space, accelerating the solving process. Experimental results from both large-scale synthetic and realistic test cases show that, compared to the state-of-the-art, our method achieves an average 56% increase in admitted flows and an average 92% reduction in admission time. Additionally, it postpones the occurrence of bottleneck egress ports and the first rejection of admission requests, thereby enhancing adaptability.

cs.NI↗

Resolving Rydberg-Electron Recapture Dynamics via Laser-driven Frustrated Tunneling Ionization

By employing two-color counter-rotating circularly polarized laser fields, we investigate the dynamics of electron recapture into Rydberg states under strong, ultrashort laser pulses, probed via coherent extreme-ultraviolet free-induction decay (XFID). Our study reveals significant distinctions between XFID and above-threshold high-order harmonic generation in terms of their ellipticity dependence on the driving-laser waveforms, yield variations with the laser-intensity ratios, and sensitivity to the driving-laser ellipticity. All these differences arise from the fundamentally distinct electron trajectories underlying the two processes. More importantly, our findings provide compelling evidence that Rydberg-electron recapture predominantly occurs at the end of the driving laser field, offering the first direct experimental confirmation of this long-proposed mechanism.

physics.optics↗

Imaging the photochemical dynamics of cyclobutanone with MeV ultrafast electron diffraction

We study the photoinduced chemical dynamics of cyclobutanone upon excitation at 200 nm to the 3s Rydberg state using MeV ultrafast electron diffraction (UED). We observe both the elastic scattering signal, which contains information about the structural dynamics, and the inelastic scattering signal, which encodes information about the electronic state. Our results suggest a sub-picosecond timescale for the photodissociation dynamics, and an excited state lifetime of about 230 femtoseconds. The dissociation is found to be dominated by the C3 channel where cyclopropane and CO are produced. The branching ratio of the C3 channel to the C2 channel where ethene and ketene are produced, is estimated to be approximately 5:3. Our data suggest that the C3 and C2 channels account for approximately 80% of the photoproducts, with the remaining 20% exhibiting ring-opened structures. It is found that the timescale associated with the dissociation process in the C2 channel is shorter compared to that in the C3 channel. Leveraging the enhanced temporal resolution of MeV UED, our results provide a real-time mapping of the nuclear wavepacket dynamics, capturing the complete photochemical dynamics from S2 minimum through the S1/S0 conical intersection, and finally to the dissociation. Our experimental results provide new insights into the Norrish Type I reaction and can be used to benchmark non-adiabatic dynamics simulations.

physics.chem-ph↗

Implicit Priors Editing in Stable Diffusion via Targeted Token Adjustment

Implicit assumptions and priors are often necessary in text-to-image generation tasks, especially when textual prompts lack sufficient context. However, these assumptions can sometimes reflect outdated concepts, inaccuracies, or societal bias embedded in the training data. We present Embedding-only Editing (Embedit), a method designed to efficiently adjust implict assumptions and priors in the model without affecting its interpretation of unrelated objects or overall performance. Given a "source" prompt (e.g., "rose") that elicits an implicit assumption (e.g., rose is red) and a "destination" prompt that specifies the desired attribute (e.g., "blue rose"), Embedit fine-tunes only the word token embedding (WTE) of the target object ("rose") to optimize the last hidden state of text encoder in Stable Diffusion, a SOTA text-to-image model. This targeted adjustment prevents unintended effects on other objects in the model's knowledge base, as the WTEs for unrelated objects and the model weights remain unchanged. Consequently, when a prompt does not contain the edited object, all representations, and the model outputs are identical to those of the original, unedited model. Our method is highly efficient, modifying only 768 parameters for Stable Diffusion 1.4 and 2048 for XL in a single edit, matching the WTE dimension of each respective model. This minimal scope, combined with rapid execution, makes Embedit highly practical for real-world applications. Additionally, changes are easily reversible by restoring the original WTE layers. Our experimental results demonstrate that Embedit consistently outperforms previous methods across various models, tasks, and editing scenarios (both single and sequential multiple edits), achieving at least a 6.01% improvement (from 87.17% to 93.18%).

cs.CV↗

Evaluation of Vortex Criteria by Virtue of the Quadruple Decomposition of Velocity Gradient Tensor

Based on the analysis of the velocity gradient tensor, we investigate in this paper the physical interpretation and limitations of four vortex criteria: $ω$, $Q$, $\varDelta$ and $λ_{ci}$, and reveal the actual physical meaning of vortex patterns which are usually illustrated by level sets of various vortex criteria. A quadruple decomposition based on the normality of the velocity gradient tensor is proposed for the first time, which resolves the motion of a fluid element into dilation, axial stretch along the normal frame, in-plane distortion, and simple shear, in order to clarify the kinematical interpretation of various vortex criteria. The mean rotation characterized by the vorticity $ω$ always consists of simple shear; the $Q$-criterion can reflect the strength of net rotation within the invariant plane relative to the axial stretch of a fluid element, and it is a sufficient but unnecessary condition for the existence of net rotation; the $\varDelta$-criterion can exactly identify the existence of net rotation, but it is not the strength of net rotation; in the case that net rotation exists, $λ_{ci}$ characterize its absolute strength. net rotation is the total effect of the normal rotation within the invariant plane and the simple shear, where the former is the most basic rotation. The newly introduced quadruple decomposition can improve our understanding of vortices in fluids and their motions.

physics.flu-dyn↗