SearcharxivSearch

arXiv subjects

Qingsong Zhao

Publications and source records attributed to Qingsong Zhao.

15 recordsLinked to original sources

Density Collapse and Gradient Catastrophe in a One-Dimensional Euler--Poisson--Cattaneo System

We study finite-time singularity formation for a one-dimensional Euler--Poisson system in Lagrangian coordinates with Cattaneo heat conduction and quadratic heat-flux corrections in the pressure and internal energy. Under the structural choice $κ=κ_0v$, two different breakdown mechanisms are obtained. First, an explicit affine solution reaches $v=0$ in finite time, so the Eulerian density $ρ=1/v$ diverges while the remaining state variables stay finite at each fixed spatial point. Second, for $a'(1)<0$, the Poisson field is localized near equilibrium by $F=ϕ_x/v$, producing a strictly hyperbolic balance law subject to the propagated constraint $F_x=1-v$. We compute the characteristic speeds and an explicit genuine-nonlinearity coefficient and show that genuine nonlinearity fails for all propagating families only when two explicit algebraic degeneracy conditions hold simultaneously. For neutral compactly supported data satisfying the Poisson constraint exactly, a narrow compressive wave generates only an $O(\varepsilon^2η)$ zero-speed Poisson mode. A constraint-compatible John--Hörmander--Bärlin bootstrap and a Riccati comparison then yield finite-time gradient catastrophe while the solution remains uniformly close to equilibrium and $F_x$ stays bounded. The two results exhibit density collapse and small-amplitude gradient blow-up as distinct breakdown mechanisms within the same Euler--Poisson--Cattaneo model.

math.AP

Local and Global Existence of Strong Solutions to the One-Dimensional Compressible Navier--Stokes Equations with General Pressure Law

We consider the Cauchy problem for the one-dimensional full compressible Navier--Stokes equations with general constitutive laws $p=p(v,θ)$ and $e=e(v,θ)$. The pressure and internal energy are assumed to be sufficiently smooth, thermodynamically compatible in the sense that $e_v=θp_θ-p$, and to satisfy $e_θ>0$. We first establish the local-in-time existence and uniqueness of strong solutions for large initial data without imposing monotonicity conditions on the pressure. For perturbations around a constant equilibrium $(\bar v,0,\barθ)$, we further assume the mechanical stability condition $p_v(\bar v,\barθ)<0$. By introducing a relative thermodynamic energy through the Gibbs relation, we derive a basic energy identity and prove its local quadratic coercivity near the equilibrium. Combining this estimate with higher-order a priori bounds, we obtain the global existence and uniqueness of strong solutions for sufficiently small $H^1(\mathbb R)$ perturbations. Moreover, the specific volume and temperature remain uniformly bounded away from zero and infinity for all time.

math.AP

Unleashing the Potential of Mamba: Boosting a LiDAR 3D Sparse Detector by Using Cross-Model Knowledge Distillation

The LiDAR 3D object detector that balances accuracy and speed is crucial for achieving real-time perception in autonomous driving. However, many existing LiDAR detection models rely on complex feature transformations, leading to poor real-time performance and high resource consumption, which limits their practical effectiveness. In this work, we propose a Faster LiDAR 3D object detection framework that Adaptively aligns Sparse voxels to enable efficient heterogeneous knowledge Distillation, called FASD. We aim to distill the Transformer's sequence modeling capability into Mamba models, significantly boosting accuracy through knowledge transfer. Specifically, we first design a cross-model knowledge distillation architecture to convey the global contextual understanding capabilities of the Transformer to Mamba. The Transformer-based teacher model employs a scale-adaptive attention mechanism to enhance multi-scale fusion. In contrast, the Mamba-based student model leverages feature alignment through spatial alignment adapters, supervised with latent-space features and span-head logit distributions, leading to improved performance and efficiency. We evaluated FASD on the Waymo and nuScenes datasets, achieving up to a 2x reduction in FLOPs and a 4x reduction in memory consumption, while improving baseline performance by 1-2 percentage points and maintaining high deployment efficiency.

cs.CV

Gradient Catastrophe for Solutions to the Conservation Laws with Source Term

This paper studies singularity formation for conservation laws with a source term. Motivated by John (1974) and Barlin (2023), we prove finite-time blow-up under initial data conditions weaker than those in Barlin. Moreover, we show that a sufficiently small compact support length of the initial data promotes blow-up. Hence, global existence can only be achieved when the initial data have a large compact support length.

math.AP

Gradient Catastrophe for Solutions to the Hyperbolic Navier-Stokes Equations

This paper studies local existence and the singularity formation of the solutions of the one-dimensional hyperbolic Navier-Stokes equations, in particular proving the gradient blow-up of the derivatives of the solutions. The underlying model introduces a relaxation mechanism that leads to hyperbolization, achieved both through a nonlinear Cattaneo law for heat conduction and through Maxwell-type constitutive relations for the stress tensor. Our main approach is to prove that the hyperbolic Navier-Stokes equations are indeed hyperbolic, and to prove that they possess two genuinely nonlinear eigenvalues, thereby establishing the blow-up of the gradient of the solution. In addition, we provide a derivation of the equation of state for the hyperbolic Navier-Stokes equations in the appendix.

math.AP

RIVER: A Real-Time Interaction Benchmark for Video LLMs

The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video intERaction Bench (RIVER Bench), designed for evaluating online video comprehension. RIVER Bench introduces a novel framework comprising Retrospective Memory, Live-Perception, and Proactive Anticipation tasks, closely mimicking interactive dialogues rather than responding to entire videos at once. We conducted detailed annotations using videos from diverse sources and varying lengths, and precisely defined the real-time interactive format. Evaluations across various model categories reveal that while offline models perform well in single question-answering tasks, they struggle with real-time processing. Addressing the limitations of existing models in online video interaction, especially their deficiencies in long-term memory and future perception, we proposed a general improvement method that enables models to interact with users more flexibly in real time. We believe this work will significantly advance the development of real-time interactive video understanding models and inspire future research in this emerging field. Datasets and code are publicly available at https://github.com/OpenGVLab/RIVER.

cs.CV

Beyond Static Artifacts: A Forensic Benchmark for Video Deepfake Reasoning in Vision Language Models

Current Vision-Language Models (VLMs) for deepfake detection excel at identifying spatial artifacts but overlook a critical dimension: temporal inconsistencies in video forgeries. Adapting VLMs to reason about these dynamic cues remains a distinct challenge. To bridge this gap, we propose Forensic Answer-Questioning (FAQ), a large-scale benchmark that formulates temporal deepfake analysis as a multiple-choice task. FAQ introduces a three-level hierarchy to progressively evaluate and equip VLMs with forensic capabilities: (1) Facial Perception, testing the ability to identify static visual artifacts; (2) Temporal Deepfake Grounding, requiring the localization of dynamic forgery artifacts across frames; and (3) Forensic Reasoning, challenging models to synthesize evidence for final authenticity verdicts. We evaluate a range of VLMs on FAQ and generate a corresponding instruction-tuning set, FAQ-IT. Extensive experiments show that models fine-tuned on FAQ-IT achieve advanced performance on both in-domain and cross-dataset detection benchmarks. Ablation studies further validate the impact of our key design choices, confirming that FAQ is the driving force behind the temporal reasoning capabilities of these VLMs.

cs.CV

Boosting SAM for Cross-Domain Few-Shot Segmentation via Conditional Point Sparsification

Motivated by the success of the Segment Anything Model (SAM) in promptable segmentation, recent studies leverage SAM to develop training-free solutions for few-shot segmentation, which aims to predict object masks in the target image based on a few reference exemplars. These SAM-based methods typically rely on point matching between reference and target images and use the matched dense points as prompts for mask prediction. However, we observe that dense points perform poorly in Cross-Domain Few-Shot Segmentation (CD-FSS), where target images are from medical or satellite domains. We attribute this issue to large domain shifts that disrupt the point-image interactions learned by SAM, and find that point density plays a crucial role under such conditions. To address this challenge, we propose Conditional Point Sparsification (CPS), a training-free approach that adaptively guides SAM interactions for cross-domain images based on reference exemplars. Leveraging ground-truth masks, the reference images provide reliable guidance for adaptively sparsifying dense matched points, enabling more accurate segmentation results. Extensive experiments demonstrate that CPS outperforms existing training-free SAM-based methods across diverse CD-FSS datasets.

cs.CV

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks neglect the fine-grained and long-horizon nature of authentic laboratory work, especially in wet-lab settings. To bridge this gap, we introduce ExpVid, the first benchmark designed to systematically evaluate MLLMs on scientific experiment videos. Curated from peer-reviewed video publications, ExpVid features a new three-level task hierarchy that mirrors the scientific process: (1) Fine-grained Perception of tools, materials, and actions; (2) Procedural Understanding of step order and completeness; and (3) Scientific Reasoning that connects the full experiment to its published conclusions. Our vision-centric annotation pipeline, combining automated generation with multi-disciplinary expert validation, ensures that tasks require visual grounding. We evaluate 19 leading MLLMs on ExpVid and find that while they excel at coarse-grained recognition, they struggle with disambiguating fine details, tracking state changes over time, and linking experimental procedures to scientific outcomes. Our results reveal a notable performance gap between proprietary and open-source models, particularly in high-order reasoning. ExpVid not only provides a diagnostic tool but also charts a roadmap for developing MLLMs capable of becoming trustworthy partners in scientific experimentation.

cs.CV

Rethinking the Zigzag Flattening for Image Reading

Sequence ordering of word vector matters a lot to text reading, which has been proven in natural language processing (NLP). However, the rule of different sequence ordering in computer vision (CV) was not well explored, e.g., why the ``zigzag" flattening (ZF) is commonly utilized as a default option to get the image patches ordering in vision networks. Notably, when decomposing multi-scale images, the ZF could not maintain the invariance of feature point positions. To this end, we investigate the Hilbert fractal flattening (HF) as another method for sequence ordering in CV and contrast it against ZF. The HF has proven to be superior to other curves in maintaining spatial locality, when performing multi-scale transformations of dimensional space. And it can be easily plugged into most deep neural networks (DNNs). Extensive experiments demonstrate that it can yield consistent and significant performance boosts for a variety of architectures. Finally, we hope that our studies spark further research about the flattening strategy of image reading.

cs.CV

DROP: Decouple Re-Identification and Human Parsing with Task-specific Features for Occluded Person Re-identification

The paper introduces the Decouple Re-identificatiOn and human Parsing (DROP) method for occluded person re-identification (ReID). Unlike mainstream approaches using global features for simultaneous multi-task learning of ReID and human parsing, or relying on semantic information for attention guidance, DROP argues that the inferior performance of the former is due to distinct granularity requirements for ReID and human parsing features. ReID focuses on instance part-level differences between pedestrian parts, while human parsing centers on semantic spatial context, reflecting the internal structure of the human body. To address this, DROP decouples features for ReID and human parsing, proposing detail-preserving upsampling to combine varying resolution feature maps. Parsing-specific features for human parsing are decoupled, and human position information is exclusively added to the human parsing branch. In the ReID branch, a part-aware compactness loss is introduced to enhance instance-level part differences. Experimental results highlight the efficacy of DROP, especially achieving a Rank-1 accuracy of 76.8% on Occluded-Duke, surpassing two mainstream methods. The codebase is accessible at https://github.com/shuguang-52/DROP.

cs.CV

Adaptive Discriminative Regularization for Visual Classification

How to improve discriminative feature learning is central in classification. Existing works address this problem by explicitly increasing inter-class separability and intra-class similarity, whether by constructing positive and negative pairs for contrastive learning or posing tighter class separating margins. These methods do not exploit the similarity between different classes as they adhere to i.i.d. assumption in data. In this paper, we embrace the real-world data distribution setting that some classes share semantic overlaps due to their similar appearances or concepts. Regarding this hypothesis, we propose a novel regularization to improve discriminative learning. We first calibrate the estimated highest likelihood of one sample based on its semantically neighboring classes, then encourage the overall likelihood predictions to be deterministic by imposing an adaptive exponential penalty. As the gradient of the proposed method is roughly proportional to the uncertainty of the predicted likelihoods, we name it adaptive discriminative regularization (ADR), trained along with a standard cross entropy loss in classification. Extensive experiments demonstrate that it can yield consistent and non-trivial performance improvements in a variety of visual classification tasks (over 10 benchmarks). Furthermore, we find it is robust to long-tailed and noisy label data distribution. Its flexible design enables its compatibility with mainstream classification architectures and losses.

cs.LG

Towards Privacy-Preserving Person Re-identification via Person Identify Shift

Recently privacy concerns of person re-identification (ReID) raise more and more attention and preserving the privacy of the pedestrian images used by ReID methods become essential. De-identification (DeID) methods alleviate privacy issues by removing the identity-related of the ReID data. However, most of the existing DeID methods tend to remove all personal identity-related information and compromise the usability of de-identified data on the ReID task. In this paper, we aim to develop a technique that can achieve a good trade-off between privacy protection and data usability for person ReID. To achieve this, we propose a novel de-identification method designed explicitly for person ReID, named Person Identify Shift (PIS). PIS removes the absolute identity in a pedestrian image while preserving the identity relationship between image pairs. By exploiting the interpolation property of variational auto-encoder, PIS shifts each pedestrian image from the current identity to another with a new identity, resulting in images still preserving the relative identities. Experimental results show that our method has a better trade-off between privacy-preserving and model performance than existing de-identification methods and can defend against human and model attacks for data privacy.

cs.CV

Radially Symmetric Stationary Wave for Two-dimensional Burgers Equation

We are concerned with the radially symmetric stationary wave for the exterior problem of two-dimensional Burgers equation. A sufficient and necessary condition to guarantee the existence of such a stationary wave is given and it is also shown that such a stationary wave satisfies nice decay estimates and is time-asymptotically nonlinear stable under radially symmetric perturbation.

math.AP

Asymptotics of Radially Symmetric Solutions for the Exterior Problem of Multidimensional Burgers Equation

We are concerned with the large-time behavior of the radially symmetric solution for multidimensional Burgers equation on the exterior of a ball $\mathbb{B}_{r_0}(0)\subset \mathbb{R}^n$ for $n\geq 3$ and some positive constant $r_0>0$, where the boundary data $v_-$ and the far field state $v_+$ of the initial data are prescribed and correspond to a stationary wave. It is shown in \cite{Hashimoto-Matsumura-JDE-2019} that a sufficient condition to guarantee the existence of such a stationary wave is $v_+<0, v_-\leq |v_+|+μ(n-1)/r_0$. Since the stationary wave is no longer monotonic, its nonlinear stability is justified only recently in \cite{Hashimoto-Matsumura-JDE-2019} for the case when $v_\pm<0, v_-\leq v_++μ(n-1)/r_0$. The main purpose of this paper is to verify the time asymptotically nonlinear stability of such a stationary wave for the whole range of $v_\pm$ satisfying $v_+<0, v_-\leq |v_+|+μ(n-1)/r_0$. Furthermore, we also derive the temporal convergence rate, both algebraically and exponentially. Our stability analysis is based on a space weighted energy method with a suitable chosen weight function, while for the temporal decay rates, in addition to such a space weighted energy method, we also use the space-time weighted energy method employed in \cite{Kawashima-Matsumura-CMP-1985} and \cite{Yin-Zhao-KRM-2009}.

math.AP