SearcharxivSearch

arXiv subjects

Jiayu Zheng

Publications and source records attributed to Jiayu Zheng.

15 recordsLinked to original sources

From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification

Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an open-instance challenge, where intra-class variations are vast and complex, beyond existing benchmarks. While traditional video encoder models struggle to fit these diverse distributions, vision-language models (VLMs) offer superior generalization but have not fully leveraged their reasoning capabilities (intuition) for such tasks. In this paper, we bridge this gap with an intrinsic reasoning framework that evolves open-instance video classification from imitation to intuition. Our approach, namely DeepIntuit, begins with a cold-start supervised alignment to initialize reasoning capability, followed by refinement using Group Relative Policy Optimization (GRPO) to enhance reasoning coherence through reinforcement learning. Crucially, to translate this reasoning into accurate classification, DeepIntuit then introduces an intuitive calibration stage. In this stage, a classifier is trained on this intrinsic reasoning traces generated by the refined VLM, ensuring stable knowledge transfer without distribution mismatch. Extensive experiments demonstrate that for open-instance video classification, DeepIntuit benefits significantly from transcending simple feature imitation and evolving toward intrinsic reasoning. Our project is available at https://bwgzk-keke.github.io/DeepIntuit/.

cs.CV

Discrete Feynman-Kac approximation for parabolic Anderson model using random walks

In this paper, we introduce a natively positive approximation method based on the Feynman-Kac representation using random walks, to approximate the solution to the one-dimensional parabolic Anderson model of Skorokhod type, with either a flat or a Dirac delta initial condition. Assuming the driving noise is a fractional Brownian sheet with Hurst parameters $H \geq \frac{1}{2}$ and $H_* \geq \frac{1}{2}$ in time and space, respectively, we also provide an error analysis of the proposed method. The error in $L^p (Ω)$ norm is of order \[ O \big(h^{\frac{1}{2}[(2H + H_* - 1) \wedge 1] - ε}\big), \] where $h > 0$ is the step size in time (resp. $\sqrt{h}$ in space), and $ε> 0$ can be chosen arbitrarily small. This error order matches the Hölder continuity of the solution in time with a correction order $ε$, making it `almost' optimal. Furthermore, these results provide a quantitative framework for convergence of the partition function of directed polymers in Gaussian environments to the parabolic Anderson model.

math.PR

RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This stems from a critical \textit{rationale gap}, where sparse domain data is insufficient to bridge the semantic distance between complex spatio-temporal content and abstract classification labels. We propose a two-stage self-improvement paradigm to bridge this gap without new annotations. First, we prompt the VLMs to generate detailed textual rationales for each video, compelling them to articulate the domain-specific logic. The VLM is then fine-tuned on these self-generated rationales, utilizing this intermediate supervision to align its representations with the nuances of the target domain. Second, conventional supervised fine-tuning (SFT) is performed on the task labels, achieving markedly higher effectiveness as a result of the model's pre-acquired domain reasoning. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms direct SFT, validating self-generated rationale as an effective, annotation-efficient paradigm for adapting VLMs to domain-specific video analysis.

cs.CV

Do Students Rely on AI? Analysis of Student-ChatGPT Conversations from a Field Study

This study explores how college students interact with generative AI (ChatGPT-4) during educational quizzes, focusing on reliance and predictors of AI adoption. Conducted at the early stages of ChatGPT implementation, when students had limited familiarity with the tool, this field study analyzed 315 student-AI conversations during a brief, quiz-based scenario across various STEM courses. A novel four-stage reliance taxonomy was introduced to capture students' reliance patterns, distinguishing AI competence, relevance, adoption, and students' final answer correctness. Three findings emerged. First, students exhibited overall low reliance on AI and many of them could not effectively use AI for learning. Second, negative reliance patterns often persisted across interactions, highlighting students' difficulty in effectively shifting strategies after unsuccessful initial experiences. Third, certain behavioral metrics strongly predicted AI reliance, highlighting potential behavioral mechanisms to explain AI adoption. The study's findings underline critical implications for ethical AI integration in education and the broader field. It emphasizes the need for enhanced onboarding processes to improve student's familiarity and effective use of AI tools. Furthermore, AI interfaces should be designed with reliance-calibration mechanisms to enhance appropriate reliance. Ultimately, this research advances understanding of AI reliance dynamics, providing foundational insights for ethically sound and cognitively enriching AI practices.

cs.AI

Quantitative diffusion approximation for the Neutral $r$-Alleles Wright-Fisher Model with Mutations

We apply a Lindeberg principle under the Markov process setting to approximate the Wright-Fisher model with neutral $r$-alleles using a diffusion process, deriving an error rate based on a function class distance involving fourth-order bounded differentiable functions. This error rate consists of a linear combination of the maximum mutation rate and the reciprocal of the population size. Our result improves the error bound in the seminal work [PNAS,1977], where only the special case $r=2$ was studied.

math.PR

Quantum Noise of Kramers-Kronig Receiver

The Kramers-Kronig (KK) receiver provides an efficient method to reconstruct the complex-valued optical field by means of intensity detection given a minimum-phase signal. In this paper, we analytically show that for detecting coherent states through measuring the minimum-phase signal, while keeping the radial quantum fluctuation the same as the balanced heterodyne detection does, the KK receiver can indirectly recover the tangential component with fluctuation equivalently reduced to 1/3 times the radial one at the decision time, by adopting the KK relations to utilize the information of the physically measured radial component of other time of the symbol period. In consequence, the KK receiver achieves 3/2 times the signal-to-noise ratio of balanced heterodyne detection, while presenting an asymmetric quantum fluctuation distribution depending on the time-varying phase. Therefore, the KK receiver provides a feasible scheme to reduce the quantum fluctuation for obtaining the selected component to 2/ 3 times that of physically measuring the same component of the coherent state. This work provides a physical insight of the KK receiver and should enrich the knowledge of electromagnetic noise in quantum optical measurement.

quant-ph

Co-GRU Enhanced End-to-End Design for Long-haul Coherent Transmission Systems

In recent years, the end-to-end (E2E) scheme based on deep learning (DL) has been proposed as a potential scheme to jointly optimize the encoder and the decoder parameters of the optical communication system. Compared with conventional deep neural network (DNN) adopted in E2E design, center-oriented Gated Recurrent Unit (Co-GRU) network has the ability to learn and compensate for inter-symbol interference (ISI) with low computation cost while satisfying the gradient backpropagation (BP) condition. In this paper, the Co-GRU structure is adopted for both channel modeling and decoder implementation in E2E design for long-haul coherent wavelength division multiplexing (WDM) transmission systems, which can enhance the performance of general mutual information (GMI) and Q2-factor. For the E2E system with Co-GRU based decoder, the gain of GMI and Q2-factor are respectively improved 0.2 bits/sym and 0.48dB, compared to that of the conventional QAM system, for a 5-channel dual-polarization coherent system transmitting over 960km standard single mode fiber (SSMF). This work paves the way for further study of the application of the Co-GRU structure for both the data-driven channel modeling and the decoder performance improvement in E2E design.

eess.SP

Moment asymptotics for super-Brownian motions

In this paper, long time and high order moment asymptotics for super-Brownian motions (sBm's) are studied. By using a moment formula for sBm's (e.g. Theorem 3.1, Hu et al. Ann. Appl. Probab. 2023+), precise upper and lower bounds for all positive integer moments and for all time of sBm's for certain initial conditions are achieved. Then, the moment asymptotics as time goes to infinity or as the moment order goes to infinity follow immediately. Additionally, as an application of the two-sided moment bounds, the tail probability estimates of sBm's are obtained.

math.PR

On mean-field super-Brownian motions

The mean-field stochastic partial differential equation (SPDE) corresponding to a mean-field super-Brownian motion (sBm) is obtained and studied. In this mean-field sBm, the branching-particle lifetime is allowed to depend upon the probability distribution of the sBm itself, producing an SPDE whose space-time white noise coefficient has, in addition to the typical sBm square root, an extra factor that is a function of the probability law of the density of the mean-field sBm. This novel mean-field SPDE is thus motivated by population models where things like overcrowding and isolation can affect growth. A two step approximation method is employed to show existence for this SPDE under general conditions. Then, mild moment conditions are imposed to get uniqueness. Finally, smoothness of the SPDE solution is established under a further simplifying condition.

math.PR

Nonlinear McKean-Vlasov diffusions under the weak Hormander condition with quantile-dependent coefficients

In this paper, the strong existence and uniqueness for a degenerate finite system of quantile-dependent McKean-Vlasov stochastic differential equations are obtained under a weak Hörmander condition. The approach relies on the apriori bounds for the density of the solution to time inhomogeneous diffusions. The time inhomogeneous Feynman-Fac formula is used to construct a contraction map for this degenerate system.

math.PR

Mean-variance portfolio selection under partial information with drift uncertainty

In this paper, we study the mean-variance portfolio selection problem under partial information with drift uncertainty. First we show that the market model is complete even in this case while the information is not complete and the drift is uncertain. Then, the optimal strategy based on partial information is derived, which reduces to solving a related backward stochastic differential equation (BSDE). Finally, we propose an efficient numerical scheme to approximate the optimal portfolio that is the solution of the BSDE mentioned above. Malliavin calculus and the particle representation play important roles in this scheme.

q-fin.PM

Multi-Drone based Single Object Tracking with Agent Sharing Network

Drone equipped with cameras can dynamically track the target in the air from a broader view compared with static cameras or moving sensors over the ground. However, it is still challenging to accurately track the target using a single drone due to several factors such as appearance variations and severe occlusions. In this paper, we collect a new Multi-Drone single Object Tracking (MDOT) dataset that consists of 92 groups of video clips with 113,918 high resolution frames taken by two drones and 63 groups of video clips with 145,875 high resolution frames taken by three drones. Besides, two evaluation metrics are specially designed for multi-drone single object tracking, i.e. automatic fusion score (AFS) and ideal fusion score (IFS). Moreover, an agent sharing network (ASNet) is proposed by self-supervised template sharing and view-aware fusion of the target from multiple drones, which can improve the tracking accuracy significantly compared with single drone tracking. Extensive experiments on MDOT show that our ASNet significantly outperforms recent state-of-the-art trackers.

cs.CV

Stochastic maximum principle for generalized mean-field delay control problem

In this paper, we first give the existence and uniqueness theorems for generalized mean-filed delay stochastic differential equations (GMFDSDEs) and mean-field anticipated backward stochastic differential equations (MFABSDEs). Then we study the stochastic maximum principle for generalized mean-filed delay control problem. Since the state is distribution-depending, we define the adjoint equation as a MFABSDE, in which, all the derivatives of coefficients are in Fréchet sense. We deduce the stochastic maximum principle, and also obtain, under some additional assumptions, a sufficient condition for the optimality of the control.

math.OC

Pathwise uniqueness for stochastic differential equations driven by pure jump processes

Based on the weak existence and weak uniqueness, we study the pathwise uniqueness of the solutions for a class of one-dimensional stochastic differential equations driven by pure jump processes. By using Tanaka's formula and the local time technique, we show that there is no gap between the strong uniqueness and weak uniqueness when the coefficients of the Poisson random measures satisfy a suitable condition

math.PR