SearcharxivSearch

arXiv subjects

Chunhui Liu

Publications and source records attributed to Chunhui Liu.

At least 19 recordsLinked to original sources

Geodesic Focusing Conditions in $f(Q)$ Gravity

We study the geodesic deviation equation in symmetric teleparallel geometry (STG), where the relative acceleration is defined with respect to the STG connection. We analyze the modified Raychaudhuri equation along a geodesic congruence in $f(Q)$ gravity under the Weyl-type ansatz, together with an additional assumption under which the metric variation term along the congruence is converted into a disformation-induced acceleration term. In contrast to the purely geometrical Raychaudhuri equation obtained in general metric-affine settings, the equation derived here contains matter-source contributions through the trace equation of $f(Q)$ gravity. Different from general relativity, focusing in $f(Q)$ gravity is not automatic, and one must impose an appropriate focusing condition. We collect the model-dependent terms in the modified Raychaudhuri equation into an effective energy-momentum trace $T_{\text{eff}}$, so that the focusing condition can be written as the inequality $T\leq T_{\text{eff}}$, where $T$ is the trace of the matter energy-momentum tensor. We also apply this condition to the flat Friedmann--Lema\^{i}tre--Robertson--Walker (FLRW) background. The homogeneous and isotropic STG connection admits three branches, each characterized by a single connection function $\gamma_i$, with $i=1,2,3$. Only the first branch with the coincident gauge is compatible with the Weyl-type ansatz. We obtain the resulting effective trace $T_{\text{eff}}=T$ for any form of $f(Q)$ satisfying $f_Q>0$ in the flat FLRW universe. The focusing inequality is saturated and imposes no additional constraint on the matter content.

gr-qc

CoRe: A Continuously Reward-Finetuned LLM Query Rewriter for Multi-Stage Context-Aware Relevance in Web-Scale Video Search

LLM-based query rewriters in production face a tension: the training reward must reflect how the rewrite is consumed by the production ranker, yet the training procedure must be cheap enough to support continuous redeployment as data drifts. We present CoRe (Context Relevance), such a system, redeployed weekly for over five months in a major short-video search engine. Our reward uses the deployed multimodal relevance model as its source and a multiplicative ratio form mirroring the production fusion algebra, closing the simulation-production gap that offline reward proxies leave open. A semi-online Mixed Preference Optimization loop makes this reward affordable at multi-million-instance weekly scale: a DPO-style pairwise objective restricts the gradient pass to a small top-k/bottom-k subset of sampled trajectories, and a phase structure reduces trainer/inference-server parameter syncs from per-step to per-phase. An automated promotion gate over reward-like and stability metrics detected and recovered from a real reward-hacking incident in production. Rewriter output is consumed as parallel relevance signals at recall, rawrank, and finerank without displacing the original signals, bounding rewriter-failure blast radius. Online A/B from two sequential production launches, first deploying the rewriter at finerank, then extending consumption to recall and rawrank, delivers statistically significant reductions in change-query rate on rewrite-impacted queries, with all headline relevance and engagement metrics moving in the expected direction.

cs.IR

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-based desmoking approaches rely on scarce paired supervision and deterministic restoration pipelines, making it difficult to perform exploration or reinforcement-driven refinement under real surgical conditions. We propose PhySe-RPO, a diffusion restoration framework optimized through Physics- and Semantics-Guided Relative Policy Optimization. The core idea is to transform deterministic restoration into a stochastic policy, enabling trajectory-level exploration and critic-free updates via group-relative optimization. A physics-guided reward imposes illumination and color consistency, while a visual-concept semantic reward learned from CLIP-based surgical concepts promotes smoke-free and anatomically coherent restoration. Together with a reference-free perceptual constraint, PhySe-RPO produces results that are physically consistent, semantically faithful, and clinically interpretable across synthetic and real robotic surgical datasets, providing a principled route to robust diffusion-based restoration under limited paired supervision.

cs.AI

On the rational points in conics of a cubic surfac

In this paper, we give a uniform upper bound on the rational points of bounded height provided by conics in a cubic surface. For this target, we give a generalized version of the global determinant method of Salberger by Arakelov geometry.

math.AG

Regularization of Gauss-Bonnet Gravity in Riemann-Cartan Geometry

We extend the conformal dimensional-derivative regularization of four-dimensional Gauss- Bonnet gravity to Riemann-Cartan geometry, obtaining a regularized action whose torsionless limit equals the well-known regularized four-dimensional Einstein-Gauss-Bonnet model. Varying independently with respect to the scalar, tetrad, and spin connection yields field equations that remain strictly second order in covariant derivatives, thereby avoiding Ostrogradsky-type instabil- ities. Within this framework we obtain static, spherically symmetric black holes carrying torsion hair, showing that the regularized Gauss-Bonnet interaction can support long-range torsion hair without invoking extra dimensions.

gr-qc

Towards AI-Native RAN: An Operator's Perspective of 6G Day 1 Standardization

Artificial Intelligence/Machine Learning (AI/ML) has become the most certain and prominent feature of 6G mobile networks. Unlike 5G, where AI/ML was not natively integrated but rather an add-on feature over existing architecture, 6G shall incorporate AI from the onset to address its complexity and support ubiquitous AI applications. Based on our extensive mobile network operation and standardization experience from 2G to 5G, this paper explores the design and standardization principles of AI-Native radio access networks (RAN) for 6G, with a particular focus on its critical Day 1 architecture, functionalities and capabilities. We investigate the framework of AI-Native RAN and present its three essential capabilities to shed some light on the standardization direction; namely, AI-driven RAN processing/optimization/automation, reliable AI lifecycle management (LCM), and AI-as-a-Service (AIaaS) provisioning. The standardization of AI-Native RAN, in particular the Day 1 features, including an AI-Native 6G RAN architecture, were proposed. For validation, a large-scale field trial with over 5000 5G-A base stations have been built and delivered significant improvements in average air interface latency, root cause identification, and network energy consumption with the proposed architecture and the supporting AI functions. This paper aims to provide a Day 1 framework for 6G AI-Native RAN standardization design, balancing technical innovation with practical deployment.

cs.NI

Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion

Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing labeled training data quality and cross-modal fusion significantly improves model performance, influencing key metrics such as quality view rates and ad revenue. High-quality annotations are crucial for advancing content modeling, yet traditional statistical-based active learning (AL) methods face limitations: they struggle to detect overconfident misclassifications and are less effective in distinguishing semantically similar items in deep neural networks. Additionally, audio information plays an increasing role, especially in short-video platforms, yet most pre-trained multimodal architectures primarily focus on text and images. While training from scratch across all three modalities is possible, it sacrifices the benefits of leveraging existing pre-trained visual-language (VL) and audio models. To address these challenges, we propose kNN-based Latent Space Broadening (LSB) to enhance AL efficiency and Vision-Language Modeling with Audio Enhancement (VLMAE), a mid-fusion approach integrating audio into VL models. This system deployed in production systems, leading to significant business gains.

cs.MM

Thermodynamic of the $f(Q)$ universe

We investigate thermodynamics of apparent horizon in the $f(Q)$ universe with trivial and nontrivial connections. We first explore the perspectives of the first law, generalized second law and $P-V$ phase transition with trivial connection. We show that the lowest-order correction of entropy has the same form as that in loop quantum gravity, and the critical exponents of the phase transition caused by the lowest-order correction are consistent with those in mean field theory. We then examine the thermodynamic implication of nontrivial connections. We find that nontrivial connections in the $f(Q)$ universe imply non-equilibrium states from the perspective of thermodynamics.

gr-qc

Gravitational wave in symmetric teleparallel gravity with different connections

We investigate the cosmological perturbations around all three branches of spatially flat universe with different connections in symmetric teleparallel gravity. The model we consider can cover both the case of f(Q) model and that of the non-minimal coupling between a scalar field and the non-metricity scalar. We focus on analyzing and comparing the propagation behavior and stability of the tensorial and non-tensorial gravitational waves on spatially flat universe with different connections.

gr-qc

LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval

Video-text retrieval is a class of cross-modal representation learning problems, where the goal is to select the video which corresponds to the text query between a given text query and a pool of candidate videos. The contrastive paradigm of vision-language pretraining has shown promising success with large-scale datasets and unified transformer architecture, and demonstrated the power of a joint latent space. Despite this, the intrinsic divergence between the visual domain and textual domain is still far from being eliminated, and projecting different modalities into a joint latent space might result in the distorting of the information inside the single modality. To overcome the above issue, we present a novel mechanism for learning the translation relationship from a source modality space $\mathcal{S}$ to a target modality space $\mathcal{T}$ without the need for a joint latent space, which bridges the gap between visual and textual domains. Furthermore, to keep cycle consistency between translations, we adopt a cycle loss involving both forward translations from $\mathcal{S}$ to the predicted target space $\mathcal{T'}$, and backward translations from $\mathcal{T'}$ back to $\mathcal{S}$. Extensive experiments conducted on MSR-VTT, MSVD, and DiDeMo datasets demonstrate the superiority and effectiveness of our LaT approach compared with vanilla state-of-the-art methods.

cs.CV

On the global determinant method

In this paper, we build the global determinant method of Salberger by Arakelov geometry explicitly. As an application, we study the dependence on the degree of the number of rational points of bounded height in plane curves. We will also explain why some constants will be more explicit if we admit the Generalized Riemann Hypothesis.

math.NT

TubeR: Tubelet Transformer for Video Action Detection

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21.

cs.CV

SSCAP: Self-supervised Co-occurrence Action Parsing for Unsupervised Temporal Action Segmentation

Temporal action segmentation is a task to classify each frame in the video with an action label. However, it is quite expensive to annotate every frame in a large corpus of videos to construct a comprehensive supervised training dataset. Thus in this work we propose an unsupervised method, namely SSCAP, that operates on a corpus of unlabeled videos and predicts a likely set of temporal segments across the videos. SSCAP leverages Self-Supervised learning to extract distinguishable features and then applies a novel Co-occurrence Action Parsing algorithm to not only capture the correlation among sub-actions underlying the structure of activities, but also estimate the temporal path of the sub-actions in an accurate and general way. We evaluate on both classic datasets (Breakfast, 50Salads) and the emerging fine-grained action dataset (FineGym) with more complex activity structures and similar sub-actions. Results show that SSCAP achieves state-of-the-art performance on all datasets and can even outperform some weakly-supervised approaches, demonstrating its effectiveness and generalizability.

cs.CV

VidTr: Video Transformer Without Convolutions

We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatio-temporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels, but with heavy memory usage. We then present VidTr which reduces the memory cost by 3.3$\times$ while keeping the same performance. To further optimize the model, we propose the standard deviation based topK pooling for attention ($pool_{topK\_std}$), which reduces the computation by dropping non-informative features along temporal dimension. VidTr achieves state-of-the-art performance on five commonly used datasets with lower computational requirement, showing both the efficiency and effectiveness of our design. Finally, error analysis and visualization show that VidTr is especially good at predicting actions that require long-term temporal reasoning.

cs.CV

Selective Feature Compression for Efficient Activity Recognition Inference

Most action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching temporal region is expensive for a real-world application. In this work, we focus on improving the inference efficiency of current action recognition backbones on trimmed videos, and illustrate that one action model can also cover then informative region by dropping non-informative features. We present Selective Feature Compression (SFC), an action recognition inference strategy that greatly increase model inference efficiency without any accuracy compromise. Differently from previous works that compress kernel sizes and decrease the channel dimension, we propose to compress feature flow at spatio-temporal dimension without changing any backbone parameters. Our experiments on Kinetics-400, UCF101 and ActivityNet show that SFC is able to reduce inference speed by 6-7x and memory usage by 5-6x compared with the commonly used 30 crops dense sampling procedure, while also slightly improving Top1 Accuracy. We thoroughly quantitatively and qualitatively evaluate SFC and all its components and show how does SFC learn to attend to important video regions and to drop temporal features that are uninformative for the task of action recognition.

cs.CV