SearcharxivSearch

arXiv subjects

Xiaohui Gao

Publications and source records attributed to Xiaohui Gao.

At least 19 recordsLinked to original sources

Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning

Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to degraded generalization. While many long-tailed semi-supervised learning (LTSSL) methods have been proposed, the mechanisms by which they implicitly debias logits remain poorly understood. In this work, we revisit LTSSL through the lens of learning dynamics and provide a theoretical characterization of logits debiasing. Specifically, we derive a step-wise decomposition of the logits updates, showing that predictions are dominated by class-imbalance bias that reliably reflects label priors. To expose this effect, we use the logits of a task-irrelevant baseline image as an indicator of accumulated bias and prove that they converge to the class prior. This provides a unified view where LTSSL remedies such as logit adjustment, reweighting, and resampling correspond to reshaping gradient dynamics. Based on this insight, we propose DyTrim, a principle-based dynamic pruning framework that reallocates gradient budget through class-aware pruning on labeled data and confidence-based soft pruning on unlabeled data. We provide theoretical guarantees that DyTrim reduces class bias and improves generalization. Extensive experiments on standard LTSSL benchmarks show consistent gains across architectures and methods. Code available at: https://jiajun0425.github.io/DyTrim

cs.LG

Noise-driven pseudovorticity multipoles in self-focusing beams with quintic saturation

We investigate pseudovorticity generation in Gaussian beams undergoing self-focusing under amplitude and phase noise, using the cubic-quintic nonlinear Schrödinger equation. Pseudovorticity, defined as the curl of the optical momentum flux, characterizes local rotational flow in the absence of phase singularities. Our numerical simulations show that thermal amplitude and phase noise induce a multipolar pseudovorticity pattern. Unlike the pure cubic case, where noise asymmetries are radiated away during collapse, the quintic saturation arrests collapse and traps the noise in the resulting soliton. Hence, pseudovorticity multipoles persist, oscillating at the focusing-refocusing period. These results suggest a potential pathway for controlling local optical torque through noise engineering.

physics.optics

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics and programming. However, the mechanistic role of Sample Difficulty in RLVR remains poorly understood. In this paper, we investigate RLVR through the lens of difficulty-wise and one-sample analysis. We find that sample difficulty has a non-monotonic effect on RLVR: easy and medium-difficulty problems yield the strongest and most stable reasoning improvements, whereas overly hard problems often provide weak learning signals, induce degenerate behaviors such as answer repetition or skipping necessary computation, and can ultimately degrade the model's pre-existing capabilities. Beyond the obverse of response, we further analyze the model's internal feature dynamics using Temporal Sparse Autoencoders (T-SAE). Easy problems mainly reinforce direct-answer and basic-computation features while suppressing deliberative-reasoning features; hard problems activate reasoning-related features but become useful only when successful trajectories are sampled; medium-difficulty problems provide a more balanced signal, strengthening both computation and multi-step reasoning features. Motivated by these findings, we propose difficulty-adaptive strategies for hard-sample utilization, using backward-reasoning reformulation and T-SAE-guided training signals to improve reward density and credit assignment during RLVR. Overall, our results identify sample difficulty as a key factor governing both the optimization dynamics and representation evolution of RLVR.

cs.AI

Orientation-Dependent Ion Acceleration from Laser-Irradiated Rectangular Nanorings

Laser-driven ion acceleration from nanostructured targets offers a promising route to compact, high-energy ion sources. In this work, we demonstrate through particle-in-cell simulations that rectangular nanoring targets significantly enhance energy absorption and increase the cutoff energy of laser-accelerated ions. The nanoring geometry enables strong field confinement within its hollow core when optimally oriented relative to the laser polarization, leading to hotter electron populations and more robust sheath acceleration. These results demonstrate that rectangular nanorings offer a versatile platform for controlling laser-plasma interactions at solid densities and advancing compact, high-repetition-rate particle sources.

physics.plasm-ph

FNF: Functional Network Fingerprint for Large Language Models

The development of large language models (LLMs) is costly and has significant commercial value. Consequently, preventing unauthorized appropriation of open-source LLMs and protecting developers' intellectual property rights have become critical challenges. In this work, we propose the Functional Network Fingerprint (FNF), a training-free, sample-efficient method for detecting whether a suspect LLM is derived from a victim model, based on the consistency between their functional network activity. We demonstrate that models that share a common origin, even with differences in scale or architecture, exhibit highly consistent patterns of neuronal activity within their functional networks across diverse input samples. In contrast, models trained independently on distinct data or with different objectives fail to preserve such activity alignment. Unlike conventional approaches, our method requires only a few samples for verification, preserves model utility, and remains robust to common model modifications (such as fine-tuning, pruning, and parameter permutation), as well as to comparisons across diverse architectures and dimensionalities. FNF thus provides model owners and third parties with a simple, non-invasive, and effective tool for protecting LLM intellectual property. The code is available at https://github.com/WhatAboutMyStar/LLM_ACTIVATION.

cs.CL

Brain-Inspired Exploration of Functional Networks and Key Neurons in Large Language Models

In recent years, the rapid advancement of large language models (LLMs) in natural language processing has sparked significant interest among researchers to understand their mechanisms and functional characteristics. Although prior studies have attempted to explain LLM functionalities by identifying and interpreting specific neurons, these efforts mostly focus on individual neuron contributions, neglecting the fact that human brain functions are realized through intricate interaction networks. Inspired by research on functional brain networks (FBNs) in the field of neuroscience, we utilize similar methodologies estabilished in FBN analysis to explore the "functional networks" within LLMs in this study. Experimental results highlight that, much like the human brain, LLMs exhibit certain functional networks that recur frequently during their operation. Further investigation reveals that these functional networks are indispensable for LLM performance. Inhibiting key functional networks severely impairs the model's capabilities. Conversely, amplifying the activity of neurons within these networks can enhance either the model's overall performance or its performance on specific tasks. This suggests that these functional networks are strongly associated with either specific tasks or the overall performance of the LLM. Code is available at https://github.com/WhatAboutMyStar/LLM_ACTIVATION.

q-bio.NC

Field enhancement of intense laser pulses in a subwavelength plasma aperture

The interaction of intense, ultra-short laser pulses with nanostructures offers promising avenues for spatiotemporal light control. While enhanced optical transmission through subwavelength apertures has been extensively studied in the linear regime, its extension to ultrashort, high-intensity pulses remains largely unexplored. Here we demonstrate, through three-dimensional particle-in-cell simulations, significant field enhancement of intense laser pulses in subwavelength plasma apertures. The enhancement exhibits a non-resonant character, remaining robust across a wide range of plasma densities and saturating above approximately $20n_c$, while showing minimal dependence on wall thickness. Analysis of the Poynting vector reveals that energy concentration arises from interference between the incident field and back-scattered longitudinal field components. This size-dependent enhanced transmission in plasma apertures enables potential applications such as plasma-based dichroic filters operating at extreme intensities.

physics.optics

High-order Mie resonance and transient field enhancement in laser-driven plasma nanoshells

We demonstrate substantial field enhancement in plasma nanoshells through high-order Mie resonances using combined Mie theory and particle-in-cell simulations. Optimal shell geometries yield approximately threefold electric field enhancement for 800 nm irradiation, with transient buildup times of tens of femtoseconds before plasma expansion disrupts resonance. Few-cycle pulses produce reduced enhancement due to insufficient resonance establishment. These findings enable optimized laser-plasma interactions for applications including diagnostics of laser-cluster interaction and energetic ion production from engineered core-shell targets, highlighting the critical role of temporal dynamics in nanoplasma resonances.

physics.plasm-ph

Jacobian-Based Interpretation of Nonlinear Neural Encoding Model

In recent years, the alignment between artificial neural network (ANN) embeddings and blood oxygenation level dependent (BOLD) responses in functional magnetic resonance imaging (fMRI) via neural encoding models has significantly advanced research on neural representation mechanisms and interpretability in the brain. However, these approaches remain limited in characterizing the brain's inherently nonlinear response properties. To address this, we propose the Jacobian-based Nonlinearity Evaluation (JNE), an interpretability metric for nonlinear neural encoding models. JNE quantifies nonlinearity by statistically measuring the dispersion of local linear mappings (Jacobians) from model representations to predicted BOLD responses, thereby approximating the nonlinearity of BOLD signals. Centered on proposing JNE as a novel interpretability metric, we validated its effectiveness through controlled simulation experiments on various activation functions and network architectures, and further verified it on real fMRI data, demonstrating a hierarchical progression of nonlinear characteristics from primary to higher-order visual cortices, consistent with established cortical organization. We further extended JNE with Sample-Specificity (JNE-SS), revealing stimulus-selective nonlinear response patterns in functionally specialized brain regions. As the first interpretability metric for quantifying nonlinear responses, JNE provides new insights into brain information processing. Code available at https://github.com/Gaitxh/JNE.

q-bio.NC

Pruning Large Language Models by Identifying and Preserving Functional Networks

Structured pruning is one of the representative techniques for compressing large language models (LLMs) to reduce GPU memory consumption and accelerate inference speed. It offers significant practical value in improving the efficiency of LLMs in real-world applications. Current structured pruning methods typically rely on assessment of the importance of the structure units and pruning the units with less importance. Most of them overlooks the interaction and collaboration among artificial neurons that are crucial for the functionalities of LLMs, leading to a disruption in the macro functional architecture of LLMs and consequently a pruning performance degradation. Inspired by the inherent similarities between artificial neural networks and functional neural networks in the human brain, we alleviate this challenge and propose to prune LLMs by identifying and preserving functional networks within LLMs in this study. To achieve this, we treat an LLM as a digital brain and decompose the LLM into functional networks, analogous to identifying functional brain networks in neuroimaging data. Afterwards, an LLM is pruned by preserving the key neurons within these functional networks. Experimental results demonstrate that the proposed method can successfully identify and locate functional networks and key neurons in LLMs, enabling efficient model pruning. Our code is available at https://github.com/WhatAboutMyStar/LLM_ACTIVATION.

cs.CL

LCGC: Learning from Consistency Gradient Conflicting for Class-Imbalanced Semi-Supervised Debiasing

Classifiers often learn to be biased corresponding to the class-imbalanced dataset, especially under the semi-supervised learning (SSL) set. While previous work tries to appropriately re-balance the classifiers by subtracting a class-irrelevant image's logit, but lacks a firm theoretical basis. We theoretically analyze why exploiting a baseline image can refine pseudo-labels and prove that the black image is the best choice. We also indicated that as the training process deepens, the pseudo-labels before and after refinement become closer. Based on this observation, we propose a debiasing scheme dubbed LCGC, which Learning from Consistency Gradient Conflicting, by encouraging biased class predictions during training. We intentionally update the pseudo-labels whose gradient conflicts with the debiased logits, representing the optimization direction offered by the over-imbalanced classifier predictions. Then, we debiased the predictions by subtracting the baseline image logits during testing. Extensive experiments demonstrate that LCGC can significantly improve the prediction accuracy of existing CISSL models on public benchmarks.

cs.CV

Efficient Trajectory Generation in 3D Environments with Multi-Level Map Construction

We propose a robust and efficient framework to generate global trajectories for ground robots in complex 3D environments. The proposed method takes point cloud as input and efficiently constructs a multi-level map using triangular patches as the basic elements. A kinematic path search is adopted on the patches, where motion primitives on different patches combine to form the global min-time cost initial trajectory. We use a same-level expansion method to locate the nearest obstacle for each trajectory waypoint and construct an objective function with curvature, smoothness and obstacle terms for optimization. We evaluate the method on several complex 3D point cloud maps. Compared to existing methods, our method demonstrates higher robustness to point cloud noise, enabling the generation of high quality trajectory while maintaining high computational efficiency. Our code will be publicly available at https://github.com/ck-tian/MLMC-planner.

cs.RO

Brain-like Functional Organization within Large Language Models

The human brain has long inspired the pursuit of artificial intelligence (AI). Recently, neuroimaging studies provide compelling evidence of alignment between the computational representation of artificial neural networks (ANNs) and the neural responses of the human brain to stimuli, suggesting that ANNs may employ brain-like information processing strategies. While such alignment has been observed across sensory modalities--visual, auditory, and linguistic--much of the focus has been on the behaviors of artificial neurons (ANs) at the population level, leaving the functional organization of individual ANs that facilitates such brain-like processes largely unexplored. In this study, we bridge this gap by directly coupling sub-groups of artificial neurons with functional brain networks (FBNs), the foundational organizational structure of the human brain. Specifically, we extract representative patterns from temporal responses of ANs in large language models (LLMs), and use them as fixed regressors to construct voxel-wise encoding models to predict brain activity recorded by functional magnetic resonance imaging (fMRI). This framework links the AN sub-groups to FBNs, enabling the delineation of brain-like functional organization within LLMs. Our findings reveal that LLMs (BERT and Llama 1-3) exhibit brain-like functional architecture, with sub-groups of artificial neurons mirroring the organizational patterns of well-established FBNs. Notably, the brain-like functional organization of LLMs evolves with the increased sophistication and capability, achieving an improved balance between the diversity of computational behaviors and the consistency of functional specializations. This research represents the first exploration of brain-like functional organization within LLMs, offering novel insights to inform the development of artificial general intelligence (AGI) with human brain principles.

q-bio.NC

LinBridge: A Learnable Framework for Interpreting Nonlinear Neural Encoding Models

Neural encoding of artificial neural networks (ANNs) links their computational representations to brain responses, offering insights into how the brain processes information. Current studies mostly use linear encoding models for clarity, even though brain responses are often nonlinear. This has sparked interest in developing nonlinear encoding models that are still interpretable. To address this problem, we propose LinBridge, a learnable and flexible framework based on Jacobian analysis for interpreting nonlinear encoding models. LinBridge posits that the nonlinear mapping between ANN representations and neural responses can be factorized into a linear inherent component that approximates the complex nonlinear relationship, and a mapping bias that captures sample-selective nonlinearity. The Jacobian matrix, which reflects output change rates relative to input, enables the analysis of sample-selective mapping in nonlinear models. LinBridge employs a self-supervised learning strategy to extract both the linear inherent component and nonlinear mapping biases from the Jacobian matrices of the test set, allowing it to adapt effectively to various nonlinear encoding models. We validate the LinBridge framework in the scenario of neural visual encoding, using computational visual representations from CLIP-ViT to predict brain activity recorded via functional magnetic resonance imaging (fMRI). Our experimental results demonstrate that: 1) the linear inherent component extracted by LinBridge accurately reflects the complex mappings of nonlinear neural encoding models; 2) the sample-selective mapping bias elucidates the variability of nonlinearity across different levels of the visual processing hierarchy. This study presents a novel tool for interpreting nonlinear neural encoding models and offers fresh evidence about hierarchical nonlinearity distribution in the visual cortex.

q-bio.NC

Understanding LLMs: A Comprehensive Overview from Training to Inference

The introduction of ChatGPT has led to a significant increase in the utilization of Large Language Models (LLMs) for addressing downstream tasks. There's an increasing focus on cost-efficient training and deployment within this context. Low-cost training and deployment of LLMs represent the future development trend. This paper reviews the evolution of large language model training techniques and inference deployment technologies aligned with this emerging trend. The discussion on training includes various aspects, including data preprocessing, training architecture, pre-training tasks, parallel training, and relevant content related to model fine-tuning. On the inference side, the paper covers topics such as model compression, parallel computation, memory scheduling, and structural optimization. It also explores LLMs' utilization and provides insights into their future development.

cs.CL

Control of spatiotemporal localization of infrared pulses in gas-filled capillaries using weak ultraviolet pulses

Manipulation of intense pulse propagation in gas-filled capillaries is desirable for various high-field applications. Tuning the parameters of the driving laser pulse and the working gas is the conventional approach, and it provides limited capability of control. Here we demonstrate through numerical simulations a practical scheme to control the propagation of intense pulses. A weak ultraviolet pulse is launched into a capillary with a negative delay with respect to a main infrared pulse. The pulses begin to temporally overlap due to dispersion. As the main pulse self-compresses, the control pulse is strongly red-shifted due to cross-phase modulation. The frequency shifts of the two pulses mitigate pulse walk-off and allow an efficient coupling, substantially extending the effective interaction length. This interesting phenomenon may benefit applications such as high-order harmonic generation.

physics.optics

Anisotropic field ionization in nano-clusters mediated by Brunel-electron driven plasma waves

Ionization is one of the most fundamental processes in intense laser-matter interaction. It is extremely efficient for clusters in laser fields and often leads to surprisingly high charge states at moderate laser intensities. Here we reveal a novel ionization mechanism in laser-cluster interaction through particle-in-cell simulations. As the laser field ionizes a cluster, Brunel electrons pushed back into the clustered plasma form an attosecond bunches, impulsively exciting plasma oscillation. The resulting localized wake field further ionizes the cluster, causing a highly ionized rod-like core along the polarization axis. This anisotropic ionization is prominent using few-cycle pulses and may be washed out using longer pulses due to collisional ionization. This newly identified ionization channel can potentially provide complicated site-specific control of ionization in nanometer-scale targets.

physics.plasm-ph

Ionization dynamics of submicron-sized clusters in intense ultrafast laser pulses

Submicron-sized targets are found in intense laser-cluster interaction experiments and laser-based material processing. Here we investigate the internal field localization due to Mie scattering and its effect on ionization dynamics in submicron-sized clusters using Mie calculation and particle-in-cell simulations. As a result of intertwined processes of pulse propagation and ionization, sub-micron nanofocusing dominates at lower intensity and gives rise to an ionization hotspot at the rear of the targets while plasma shielding wins over at a higher intensity, stopping further rear side ionization. As ionization is often the precursor of other processes, understanding the ionization dynamics of ultrafast laser pulses with wavelength-sized nanostructure can be relevant for intense laser-cluster experiments and femtosecond laser micro/nanomachining.

physics.plasm-ph