SearcharxivSearch

arXiv subjects

Jianlong Wang

Publications and source records attributed to Jianlong Wang.

5 recordsLinked to original sources

Exact Asymptotics for the Exit Time Probabilities of Scalar Ornstein-Uhlenbeck Bridges

This paper aims to derive accurate asymptotic estimates for the exit time probabilities of scalar Ornstein-Uhlenbeck (OU) bridges. The exit time probabilities are expressed as an asymptotic series in powers of a small parameter that characterizes the intensity of the noise inputs. It is shown that the series is valid in certain regions where all its terms are smooth functions. The results enable an accurate evaluation of the probability for a corresponding OU process to escape from a domain before a specified time, provided its initial and terminal states are known.

math.PR

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in applied AI. A common strategy in LLM fine-tuning is curriculum learning, which pre-orders training samples based on metrics like difficulty to improve learning efficiency compared to a random sampling strategy. However, most existing methods for LLM fine-tuning rely on a static curriculum, designed prior to training, which lacks adaptability to the model's evolving needs during fine-tuning. To address this, we propose EDCO, a novel framework based on two key concepts: inference entropy and dynamic curriculum orchestration. Inspired by recent findings that maintaining high answer entropy benefits long-term reasoning gains, EDCO prioritizes samples with high inference entropy in a continuously adapted curriculum. EDCO integrates three core components: an efficient entropy estimator that uses prefix tokens to approximate full-sequence entropy, an entropy-based curriculum generator that selects data points with the highest inference entropy, and an LLM trainer that optimizes the model on the selected curriculum. Comprehensive experiments in communication, medicine and law domains, EDCO outperforms traditional curriculum strategies for fine-tuning Qwen3-4B and Llama3.2-3B models under supervised and reinforcement learning settings. Furthermore, the proposed efficient entropy estimation reduces computational time by 83.5% while maintaining high accuracy.

cs.LG

EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture

In this paper, we propose EDIT (Encoder-Decoder Image Transformer), a novel architecture designed to mitigate the attention sink phenomenon observed in Vision Transformer models. Attention sink occurs when an excessive amount of attention is allocated to the [CLS] token, distorting the model's ability to effectively process image patches. To address this, we introduce a layer-aligned encoder-decoder architecture, where the encoder utilizes self-attention to process image patches, while the decoder uses cross-attention to focus on the [CLS] token. Unlike traditional encoder-decoder framework, where the decoder depends solely on high-level encoder representations, EDIT allows the decoder to extract information starting from low-level features, progressively refining the representation layer by layer. EDIT is naturally interpretable demonstrated through sequential attention maps, illustrating the refined, layer-by-layer focus on key image features. Experiments on ImageNet-1k and ImageNet-21k, along with transfer learning tasks, show that EDIT achieves consistent performance improvements over DeiT3 models. These results highlight the effectiveness of EDIT's design in addressing attention sink and improving visual feature extraction.

cs.CV

Investigation on the properties of Sine-Wiener noise and its induced escape in the particular limit case $D \to \infty$

Sine-Wiener noise is increasingly adopted in realistic stochastic modeling for its bounded nature. However, many features of the SW noise are still unexplored. In this paper, firstly, the properties of the SW noise and its integral process are explored as the parameter $D$ in the SW noise tends to infinite. It is found that although the distribution of the SW noise is quite different from Gaussian white noise, the integral process of the SW noise shows many similarities with the Wiener process. Inspired by the Wiener process, which uses the diffusion coefficient to denote the intensity of the Gaussian noise, a quantity is put forward to characterize the SW noise's intensity. Then we apply the SW noise to a one-dimensional double-well potential system and the Maier-Stein system to investigate the escape behaviors. A more interesting result is observed that the mean first exit time also follows the well-known Arrhenius law as in the case of the Gaussian noise, and the quasi-potential and the exit location distributions are very close to the results of the Gaussian noise.

cond-mat.stat-mech

A comparing study on optoelectronic properties of phototransistors based on MEH-PPV and PbS QD hybrids with bulk- and layer-heterojunction

As the responsivity (R) of a thin film photo detector is proportional to the product of the photo-induced carrier density (n) and mobility (u) (Z. Sun, Z. Liu, J. Li, G.-a. Tai, S. Lau and F. Yan, Adv. Mater., 24, 5878, 2012), which of the types is conducive to photo detection, choosing between layer-heterojunction (LH) and bulk-heterojunction (BH) field effect phototransistors (FEpTs), is still unknown. A comparison study is performed based on an MEH-PPV and PbS QDs hybrid.

cond-mat.mes-hall