SearcharxivSearch

arXiv subjects

Beibei Zhang

Publications and source records attributed to Beibei Zhang.

At least 19 recordsLinked to original sources

Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision

Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise in complex and lengthy contexts can obscure localized details, misleading MLLMs to produce incorrect answers. Recent works mitigate these issues by incentivizing deep reasoning to include relevant evidence. However, these methods have two main problems: First, the reinforcement fine-tuning framework (RFT) they leveraged incurs substantial training overheads, including high annotation costs and complicated reward designs. Second, the self-reflective and iterative-perception mechanism in some methods causes lengthy outputs and high inference latency. To alleviate these problems, we propose a novel Segment-to-Video Supervision} method (S2V) to efficiently enhance fine-grained reasoning in LVU. Specifically, we generate question answer pairs (VQA) based on localized segments, and then transfer these segment-based VQA back to the whole video for training. Due to focusing on short segments, segment-based VQA can naturally notice details which tend to be overlooked from a whole-video perspective. Training on such data can enforce MLLMs to correctly associate fine-grained details with QA while avoiding distracting noise in the whole video. The S2V training involves just reinforcement learning (RL) with a simple accuracy reward based on only 10K VQA samples and the resulting S2V model predicts answer using a single forward pass with limited output tokens. Experimental results demonstrate that S2V can consistently improve LVU performance across multiple LVU benchmarks, outperforming both general MLLMs and reasoning-based methods not only in LVU accuracy but also in training and inference efficiency.

cs.CV

Temporal properties of the stochastic fractional heat equation with rough dependence in space

This paper investigates the nonlinear stochastic fractional heat equation driven by a Gaussian noise that is white in time and fractional in space with a Hurst parameter $H \in \big(\frac{3-\alpha}{4}, \frac{1}{2}\big)$. Specifically, the driving operator is the fractional Laplacian of order $\alpha/2 \in (1/2, 1)$. We characterize the asymptotic behavior of the temporal increment $u(t+\varepsilon,x)-u(t,x)$ for fixed $t\ge 0$ and $x\in\mathbb{R}$ as $\varepsilon\downarrow 0$. Utilizing these precise asymptotic estimates, we establish Khinchin's and Chung's laws of the iterated logarithm for the temporal process $t \mapsto u(t,x)$.

math.PR

Large deviation principles for SPDEs with locally Lipschitz coefficients

Consider the stochastic partial differential equation, \begin{align*} \partial_t u^{\varepsilon}(t\,,x) = \frac{1}{2} \partial^2_x u^{\varepsilon}(t\,,x) + b(t\,,u^{\varepsilon}(t\,,x)) + \sqrt{\varepsilon}\sigma(t\,,u^{\varepsilon}(t\,,x)) \dot{W}(t\,,x), \end{align*} where $(t\,,x)\in(0\,,\infty)\times\mathbb{R}$, and $\dot{W}$ denotes space-time white noise. Foondun, Khoshnevisan, and Nualart \cite{FKN24} showed that this stochastic partial differential equation is well-posed under the assumptions that the initial condition $u(0)$ is bounded and measurable, while $b$ and $\sigma$ are locally Lipschitz continuous functions with at most linear growth. A Freidlin-Wentzell large deviation principle for the stochastic partial differential equation is established by a weak convergence approach in this paper.

math.PR

Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement

Personalized product search (PPS) aims to retrieve products relevant to the given query considering user preferences within their purchase histories. Since large language models (LLM) exhibit impressive potential in content understanding and reasoning, current methods explore to leverage LLM to comprehend the complicated relationships among user, query and product to improve the search performance of PPS. Despite the progress, LLM-based PPS solutions merely take textual contents into consideration, neglecting multimodal contents which play a critical role for product search. Motivated by this, we propose a novel framework, HMPPS, for \textbf{H}arnessing \textbf{M}ultimodal large language models (MLLM) to deal with \textbf{P}ersonalized \textbf{P}roduct \textbf{S}earch based on multimodal contents. Nevertheless, the redundancy and noise in PPS input stand for a great challenge to apply MLLM for PPS, which not only misleads MLLM to generate inaccurate search results but also increases the computation expense of MLLM. To deal with this problem, we additionally design two query-aware refinement modules for HMPPS: 1) a perspective-guided summarization module that generates refined product descriptions around core perspectives relevant to search query, reducing noise and redundancy within textual contents; and 2) a two-stage training paradigm that introduces search query for user history filtering based on multimodal representations, capturing precise user preferences and decreasing the inference cost. Extensive experiments are conducted on four public datasets to demonstrate the effectiveness of HMPPS. Furthermore, HMPPS is deployed on an online search system with billion-level daily active users and achieves an evident gain in A/B testing.

cs.MM

Spectral Hardening Reveals Afterglow Emergence in Long-Duration Fast X-ray Transients: A Case Study of GRB 250404A/EP250404a

The prompt emission and afterglow phases of gamma-ray bursts (GRBs) have been extensively studied, yet the transition between these two phases remains inadequately characterized due to limited multiwavelength observational coverage. Among the recent growing samples of fast X-ray transients observed by Einstein Probe (EP), a subgroup of GRBs are captured with long-duration X-ray emission, potentially containing featured evolution from prompt emission to the afterglow phase. In this Letter, we present a detailed analysis of GRB 250404A/EP250404a, a bright fast X-ray transient detected simultaneously by EP and the Fermi Gamma-ray Burst Monitor in X-rays and gamma rays. Its continuous X-ray emission reveals a long-duration tail, accompanied by distinct spectral evolution manifested by the spectral index $\alpha_{\rm X}$ with an initial softening, followed by an evident hardening, eventually reaching a plateau at the value of $\sim$ -2. Early optical and near-infrared observations enable broadband modeling with forward- and reverse-shock components, confirming that the X-ray hardening signals the emergence of the external-shock afterglow. From this spectral hardening we infer that the prompt phase in soft X-rays lasted $\sim300\;\mathrm{s}$, which is more than 3 times longer than the gamma-ray $T_{90}$. This well-tracked soft-hard-flat spectral pattern provides a clear indication of afterglow emergence from the fading prompt emission and offers a practical criterion for identifying a distinct population of GRBs among fast X-ray transients, even when the detection of the gamma-ray counterpart or obvious temporal break is absent.

astro-ph.HE

A polychromatic continuous-variable quantum communication network enabled by optical frequency combs

In classical communication, the introduction of polychromatic resources has rapidly boosted classical networks' rate and scale. Quantum communication is now at a similar critical stage in its development, and therefore, it is essential to investigate polychromatic quantum communication networks. In this letter, we report a polychromatic continuous-variable quantum communication network enabled by optical frequency combs. The multi-mode density matrices constituted by polychromatic quantum networks are studied. Considering the limited mode isolation, the maximum amount of information that eavesdroppers can obtain is recalculated, therefore, the total secret key rate is provided. We have also demonstrated that, compared to other multiplexing techniques, polychromatic quantum networks can theoretically achieve a secret key rate without decreasing with the increase in users. In the experiment, direct-transmission type and round-trip type quantum communication networks were built using optical frequency combs and dual-comb interference detection technology. The Gaussian-modulated continuous-variable quantum key distribution (CV-QKD) protocol has been validated, with a network capacity of 19 and a total secret key rate of 8.75 Gbps at a uniform distance of 5 km (asymptotic case), 0.82 Mbps at 120 km (finite-size effect), 89.10 Mbps at 40 km (compsable security), 13.66 Mbps at 40 km (compsable finite-size security). This implementation not only provides technical support for a high-speed multi-node quantum network, but also provides a solution for the future quantum Internet with continuous variables.

quant-ph

Gradient bounds and Liouville property for a class of hypoelliptic diffusion via coupling

In this paper, we obtain the reverse Bakry-\'Emery type estimates for a class of hypoelliptic diffusion operator by coupling method. The (right and reverse) Poincar\'e inequalities and the (right and reverse) logarithmic Sobolev inequalities are presented as consequences of such estimates. Wang-Harnack inequality, Hamilton's gradient estimate and Liouville property are also presented by reverse logarithmic Sobolev inequality.

math.PR

Air-to-Ground Cooperative OAM Communications

For users in hotspot region, orbital angular momentum (OAM) can realize multifold increase of spectrum efficiency (SE), and the flying base station (FBS) can rapidly support the real-time communication demand. However, the hollow divergence and alignment requirement impose crucial challenges for users to achieve air-to-ground OAM communications, where there exists the line-of-sight path. Therefore, we propose the air-to-ground cooperative OAM communication (ACOC) scheme, which can realize OAM communications for users with size-limited devices. The waist radius is adjusted to guarantee the maximum intensity at the cooperative users (CUs). We derive the closed-form expression of the optimal FBS position, which satisfies the antenna alignment for two cooperative user groups (CUGs). Furthermore, the selection constraint is given to choose two CUGs composed of four CUs. Simulation results are provided to validate the optimal FBS position and the SE superiority of the proposed ACOC scheme.

cs.IT

Precoding Based Downlink OAM-MIMO Communications with Rate Splitting

Orbital angular momentum (OAM) and rate splitting (RS) are the potential key techniques for the future wireless communications. As a new orthogonal resource, OAM can achieve the multifold increase of spectrum efficiency to relieve the scarcity of the spectrum resource, but how to enhance the privacy performance imposes crucial challenge for OAM communications. RS technique divides the information into private and common parts, which can guarantee the privacies for all users. In this paper, we integrate the RS technique into downlink OAM-MIMO communications, and study the precoding optimization to maximize the sum capacity. First, the concentric uniform circular arrays (UCAs) are utilized to construct the downlink transmission framework of OAM-MIMO communications with RS. Particularly, users in the same user pair utilize RS technique to obtain the information and different user pairs use different OAM modes. Then, we derive the OAM-MIMO channel model, and formulate the sum capacity maximization problem. Finally, based on the fractional programming, the optimal precoding matrix is obtained to maximize the sum capacity by using quadratic transformation. Extensive simulation results show that by using the proposed precoding optimization algorithm, OAM-MIMO communications with RS can achieve higher sum capacity than the traditional communication schemes.

cs.IT

Do As I Do: Pose Guided Human Motion Copy

Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the motion of a source person. The problem is inherently challenging due to the subtle human-body texture details to be generated and the temporal consistency to be considered. Existing approaches typically adopt a conventional GAN with an L1 or L2 loss to produce the target fake video, which intrinsically necessitates a large number of training samples that are challenging to acquire. Meanwhile, current methods still have difficulties in attaining realistic image details and temporal consistency, which unfortunately can be easily perceived by human observers. Motivated by this, we try to tackle the issues from three aspects: (1) We constrain pose-to-appearance generation with a perceptual loss and a theoretically motivated Gromov-Wasserstein loss to bridge the gap between pose and appearance. (2) We present an episodic memory module in the pose-to-appearance generation to propel continuous learning that helps the model learn from its past poor generations. We also utilize geometrical cues of the face to optimize facial details and refine each key body part with a dedicated local GAN. (3) We advocate generating the foreground in a sequence-to-sequence manner rather than a single-frame manner, explicitly enforcing temporal inconsistency. Empirical results on five datasets, iPER, ComplexMotion, SoloDance, Fish, and Mouse datasets, demonstrate that our method is capable of generating realistic target videos while precisely copying motion from a source video. Our method significantly outperforms state-of-the-art approaches and gains 7.2% and 12.4% improvements in PSNR and FID respectively.

cs.CV

Multi-Objective Trajectory Planning with Dual-Encoder

Time-jerk optimal trajectory planning is crucial in advancing robotic arms' performance in dynamic tasks. Traditional methods rely on solving complex nonlinear programming problems, bringing significant delays in generating optimized trajectories. In this paper, we propose a two-stage approach to accelerate time-jerk optimal trajectory planning. Firstly, we introduce a dual-encoder based transformer model to establish a good preliminary trajectory. This trajectory is subsequently refined through sequential quadratic programming to improve its optimality and robustness. Our approach outperforms the state-of-the-art by up to 79.72\% in reducing trajectory planning time. Compared with existing methods, our method shrinks the optimality gap with the objective function value decreasing by up to 29.9\%.

cs.RO

Moirai: Towards Optimal Placement for Distributed Inference on Heterogeneous Devices

The escalating size of Deep Neural Networks (DNNs) has spurred a growing research interest in hosting and serving DNN models across multiple devices. A number of studies have been reported to partition a DNN model across devices, providing device placement solutions. The methods appeared in the literature, however, either suffer from poor placement performance due to the exponential search space or miss an optimal placement as a consequence of the reduced search space with limited heuristics. Moreover, these methods have ignored the runtime inter-operator optimization of a computation graph when coarsening the graph, which degrades the end-to-end inference performance. This paper presents Moirai that better exploits runtime inter-operator fusion in a model to render a coarsened computation graph, reducing the search space while maintaining the inter-operator optimization provided by inference backends. Moirai also generalizes the device placement algorithm from multiple perspectives by considering inference constraints and device heterogeneity.Extensive experimental evaluation with 11 large DNNs demonstrates that Moirai outperforms the state-of-the-art counterparts, i.e., Placeto, m-SCT, and GETF, up to 4.28$\times$ in reduction of the end-to-end inference latency. Moirai code is anonymously released at \url{https://github.com/moirai-placement/moirai}.

cs.DC

Semantics-Driven Cloud-Edge Collaborative Inference

With the proliferation of video data in smart city applications like intelligent transportation, efficient video analytics has become crucial but also challenging. This paper proposes a semantics-driven cloud-edge collaborative approach for accelerating video inference, using license plate recognition as a case study. The method separates semantics extraction and recognition, allowing edge servers to only extract visual semantics (license plate patches) from video frames and offload computation-intensive recognition to the cloud or neighboring edges based on load. This segmented processing coupled with a load-aware work distribution strategy aims to reduce end-to-end latency and improve throughput. Experiments demonstrate significant improvements in end-to-end inference speed (up to 5x faster), throughput (up to 9 FPS), and reduced traffic volumes (50% less) compared to cloud-only or edge-only processing, validating the efficiency of the proposed approach. The cloud-edge collaborative framework with semantics-driven work partitioning provides a promising solution for scaling video analytics in smart cities.

cs.CV

Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization

Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current state-of-the-art approaches typically learn from articulated motion sequences in the straightforward 3D Euclidean space. However, the vanilla Euclidean space is not efficient for modeling important motion characteristics such as the joint-wise angular acceleration, which reveals the driving force behind the motion. Moreover, current methods typically attend to each channel equally and lack theoretical constrains on extracting task-relevant features from the input. In this paper, we seek to tackle these challenges from three aspects: (1) We propose to incorporate an acceleration representation, explicitly modeling the higher-order variations in motion. (2) We introduce a novel Stream-GCN network equipped with multi-stream components and channel attention, where different representations (i.e., streams) supplement each other towards a more precise action recognition while attention capitalizes on those important channels. (3) We explore feature-level supervision for maximizing the extraction of task-relevant information and formulate this into a mutual information loss. Empirically, our approach sets the new state-of-the-art performance on three benchmark datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA. Our code is anonymously released at https://github.com/ActionR-Group/Stream-GCN, hoping to inspire the community.

cs.CV

Large deviation principle for reflected SPDE on infinite spatial domain

We study a large deviation principle for a reflected stochastic partial differential equation on infinite spatial domain. A new sufficient condition for the weak convergence criterion proposed by Matoussi, Sabbagh and Zhang ({\it Appl. Math. Optim.} 83: 849-879, 2021) plays an important role in the proof.

math.PR

A large deviation principle for the stochastic heat equation with general rough noise

We study Freidlin-Wentzell's large deviation principle for one dimensional nonlinear stochastic heat equation driven by a Gaussian noise: $$\frac{\partial u^\varepsilon(t,x)}{\partial t} = \frac{\partial^2 u^\varepsilon(t,x)}{\partial x^2}+\sqrt{\varepsilon} \sigma(t, x, u^\varepsilon(t,x))\dot{W}(t,x),\quad t> 0,\, x\in\mathbb{R},$$ where $\dot W$ is white in time and fractional in space with Hurst parameter $H\in(\frac 14,\frac 12)$. Recently, Hu and Wang ({\it Ann. Inst. Henri Poincar\'e Probab. Stat.} {\bf 58} (2022) 379-423) studied the well-posedness of this equation without the technical condition of $\sigma(0)=0$ which was previously assumed in Hu et al. ({\it Ann. Probab}. {\bf 45} (2017) 4561-4616). We adopt a new sufficient condition proposed by Matoussi et al. ({\it Appl. Math. Optim.} \textbf{83} (2021) 849-879) for the weak convergence criterion of the large deviation principle.

math.PR