Searcharxiv⌕ Search

arXiv subjects

Sibei Yang

Publications and source records attributed to Sibei Yang.

At least 73 records · Page 4Linked to original sources

Predual Spaces of Hardy Spaces Related to Fractional Schrödinger Operators

Let $n\in\mathbb{N}$ and $α\in(0,\min\{2,n\})$. For any $a\in[a^\ast,\infty)$, the fractional Schrödinger operator $L_α$ is defined by \begin{equation*} L_α:=(-Δ)^{α/2}+a{|x|}^{-α}, \end{equation*} where $a^*:=-{\frac{2^αΓ((d+α)/4)^2}{Γ((d-α)/4)^2}}$. Let $γ\in[0,\fracα{n})$. In this paper, we introduce the VMO-type spaces $\mathrm{VMO}_{L_α}^γ(\mathbb{R}^{n})$ associated with $L_α$, and characterize these spaces via some tent spaces. We also prove that, for any given $p\in(\frac{n}{n+α},1]$, the space $\mathrm{VMO}_{L_α}^{\frac{1}{p}-1} (\mathbb{R}^{n})$ is the predual space of the Hardy space $H_{L_α}^p\left(\mathbb{R}^{n}\right)$ related to $L_α$.

math.FA↗

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the transfer of these advancements to multimodal contexts introduces heightened challenges, including but not limited to the impractical need for labor-intensive annotation and the limitations in terms of flexibility, generalizability, and explainability. To evoke CoT reasoning in multimodality, this work first conducts an in-depth analysis of these challenges posed by multimodality and presents two key insights: "keeping critical thinking" and "letting everyone do their jobs" in multimodal CoT reasoning. Furthermore, this study proposes a novel DDCoT prompting that maintains a critical attitude through negative-space prompting and incorporates multimodality into reasoning by first dividing the reasoning responsibility of LLMs into reasoning and recognition and then integrating the visual recognition capability of visual models into the joint reasoning process. The rationales generated by DDCoT not only improve the reasoning abilities of both large and small language models in zero-shot prompting and fine-tuning learning, significantly outperforming state-of-the-art methods but also exhibit impressive generalizability and explainability.

cs.CV↗

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a semantic-coherent video, exhibiting a rich portrayal of temporal semantics such as the whole process of flower blooming rather than a set of "moving images", we propose a novel Free-Bloom pipeline that harnesses large language models (LLMs) as the director to generate a semantic-coherence prompt sequence, while pre-trained latent diffusion models (LDMs) as the animator to generate the high fidelity frames. Furthermore, to ensure temporal and identical coherence while maintaining semantic coherence, we propose a series of annotative modifications to adapting LDMs in the reverse process, including joint noise sampling, step-aware attention shift, and dual-path interpolation. Without any video data and training requirements, Free-Bloom generates vivid and high-quality videos, awe-inspiring in generating complex scenes with semantic meaningful frame sequences. In addition, Free-Bloom is naturally compatible with LDMs-based extensions.

cs.CV↗

LoGoPrompt: Synthetic Text Images Can Be Good Visual Prompts for Vision-Language Models

Prompt engineering is a powerful tool used to enhance the performance of pre-trained models on downstream tasks. For example, providing the prompt "Let's think step by step" improved GPT-3's reasoning accuracy to 63% on MutiArith while prompting "a photo of" filled with a class name enables CLIP to achieve $80$\% zero-shot accuracy on ImageNet. While previous research has explored prompt learning for the visual modality, analyzing what constitutes a good visual prompt specifically for image recognition is limited. In addition, existing visual prompt tuning methods' generalization ability is worse than text-only prompting tuning. This paper explores our key insight: synthetic text images are good visual prompts for vision-language models! To achieve that, we propose our LoGoPrompt, which reformulates the classification objective to the visual prompt selection and addresses the chicken-and-egg challenge of first adding synthetic text images as class-wise visual prompts or predicting the class first. Without any trainable visual prompt parameters, experimental results on 16 datasets demonstrate that our method consistently outperforms state-of-the-art methods in few-shot learning, base-to-new generalization, and domain generalization.

cs.CV↗

Temporal Collection and Distribution for Referring Video Object Segmentation

Referring video object segmentation aims to segment a referent throughout a video sequence according to a natural language expression. It requires aligning the natural language expression with the objects' motions and their dynamic associations at the global video level but segmenting objects at the frame level. To achieve this goal, we propose to simultaneously maintain a global referent token and a sequence of object queries, where the former is responsible for capturing video-level referent according to the language expression, while the latter serves to better locate and segment objects with each frame. Furthermore, to explicitly capture object motions and spatial-temporal cross-modal reasoning over objects, we propose a novel temporal collection-distribution mechanism for interacting between the global referent token and object queries. Specifically, the temporal collection mechanism collects global information for the referent token from object queries to the temporal motions to the language expression. In turn, the temporal distribution first distributes the referent token to the referent sequence across all frames and then performs efficient cross-frame reasoning between the referent sequence and object queries in every frame. Experimental results show that our method outperforms state-of-the-art methods on all benchmarks consistently and significantly.

cs.CV↗

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

Embodied Reference Understanding studies the reference understanding in an embodied fashion, where a receiver is required to locate a target object referred to by both language and gesture of the sender in a shared physical environment. Its main challenge lies in how to make the receiver with the egocentric view access spatial and visual information relative to the sender to judge how objects are oriented around and seen from the sender, i.e., spatial and visual perspective-taking. In this paper, we propose a REasoning from your Perspective (REP) method to tackle the challenge by modeling relations between the receiver and the sender and the sender and the objects via the proposed novel view rotation and relation reasoning. Specifically, view rotation first rotates the receiver to the position of the sender by constructing an embodied 3D coordinate system with the position of the sender as the origin. Then, it changes the orientation of the receiver to the orientation of the sender by encoding the body orientation and gesture of the sender. Relation reasoning models the nonverbal and verbal relations between the sender and the objects by multi-modal cooperative reasoning in gesture, language, visual content, and spatial position. Experiment results demonstrate the effectiveness of REP, which consistently surpasses all existing state-of-the-art algorithms by a large margin, i.e., +5.22% absolute accuracy in terms of Prec0.5 on YouRefIt.

cs.CV↗

CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection

Task driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too diverse to be limited to a closed set of object vocabulary for traditional object detection. Simply mapping categories and visual features of common objects to the task cannot address the challenge. In this paper, we propose to explore fundamental affordances rather than object categories, i.e., common attributes that enable different objects to accomplish the same task. Moreover, we propose a novel multi-level chain-of-thought prompting (MLCoT) to extract the affordance knowledge from large language models, which contains multi-level reasoning steps from task to object examples to essential visual attributes with rationales. Furthermore, to fully exploit knowledge to benefit object recognition and localization, we propose a knowledge-conditional detection framework, namely CoTDet. It conditions the detector from the knowledge to generate object queries and regress boxes. Experimental results demonstrate that our CoTDet outperforms state-of-the-art methods consistently and significantly (+15.6 box AP and +14.8 mask AP) and can generate rationales for why objects are detected to afford the task.

cs.CV↗

EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment

Vision-language models such as CLIP have boosted the performance of open-vocabulary object detection, where the detector is trained on base categories but required to detect novel categories. Existing methods leverage CLIP's strong zero-shot recognition ability to align object-level embeddings with textual embeddings of categories. However, we observe that using CLIP for object-level alignment results in overfitting to base categories, i.e., novel categories most similar to base categories have particularly poor performance as they are recognized as similar base categories. In this paper, we first identify that the loss of critical fine-grained local image semantics hinders existing methods from attaining strong base-to-novel generalization. Then, we propose Early Dense Alignment (EDA) to bridge the gap between generalizable local semantics and object-level prediction. In EDA, we use object-level supervision to learn the dense-level rather than object-level alignment to maintain the local fine-grained semantics. Extensive experiments demonstrate our superior performance to competing approaches under the same strict setting and without using external training resources, i.e., improving the +8.4% novel box AP50 on COCO and +3.9% rare mask AP on LVIS.

cs.CV↗

Contrastive Grouping with Transformer for Referring Image Segmentation

Referring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classification frameworks, which attempt straightforwardly to align vision and language at the pixel level, thus failing to capture critical object-level information. In this paper, we propose a mask classification framework, Contrastive Grouping with Transformer network (CGFormer), which explicitly captures object-level information via token-based querying and grouping strategy. Specifically, CGFormer first introduces learnable query tokens to represent objects and then alternately queries linguistic features and groups visual features into the query tokens for object-aware cross-modal reasoning. In addition, CGFormer achieves cross-level interaction by jointly updating the query tokens and decoding masks in every two consecutive layers. Finally, CGFormer cooperates contrastive learning to the grouping strategy to identify the token and its mask corresponding to the referent. Experimental results demonstrate that CGFormer outperforms state-of-the-art methods in both segmentation and generalization settings consistently and significantly.

cs.CV↗

Grounded Image Text Matching with Mismatched Relation Reasoning

This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating pre-trained models on this task, with a focus on the challenging settings of limited data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained models lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. RCRN can be interpreted as a modular program and delivers strong performance in both length generalization and data efficiency.

cs.CV↗

Hardy Spaces Associated with Non-Negative Self-Adjoint Operators and Ball Quasi-Banach Function Spaces on Doubling Metric Measure Spaces and Their Applications

Let $(\mathcal{X},d,μ)$ be a doubling metric measure space in the sense of R. R. Coifman and G. Weiss, $L$ a non-negative self-adjoint operator on $L^2(\mathcal{X})$ satisfying the Davies--Gaffney estimate, and $X(\mathcal{X})$ a ball quasi-Banach function space on $\mathcal{X}$ satisfying some mild assumptions. In this article, the authors introduce the Hardy type space $H_{X,\,L}(\mathcal{X})$ by the Lusin area function associated with $L$ and establish the atomic and the molecular characterizations of $H_{X,\,L}(\mathcal{X}).$ As an application of these characterizations of $H_{X,\,L}(\mathcal{X})$, the authors obtain the boundedness of spectral multiplies on $H_{X,\,L}(\mathcal{X})$. Moreover, when $L$ satisfies the Gaussian upper bound estimate, the authors further characterize $H_{X,\,L}(\mathcal{X})$ in terms of the Littlewood--Paley functions $g_L$ and $g_{λ,\,L}^\ast$ and establish the boundedness estimate of Schrödinger groups on $H_{X,\,L}(\mathcal{X})$. Specific spaces $X(\mathcal{X})$ to which these results can be applied include Lebesgue spaces, Orlicz spaces, weighted Lebesgue spaces, and variable Lebesgue spaces. This shows that the results obtained in the article have extensive generality.

math.FA↗

DreamFace: Progressive Generation of Animatable 3D Faces under Text Guidance

Emerging Metaverse applications demand accessible, accurate, and easy-to-use tools for 3D digital human creations in order to depict different cultures and societies as if in the physical world. Recent large-scale vision-language advances pave the way to for novices to conveniently customize 3D content. However, the generated CG-friendly assets still cannot represent the desired facial traits for human characteristics. In this paper, we present DreamFace, a progressive scheme to generate personalized 3D faces under text guidance. It enables layman users to naturally customize 3D facial assets that are compatible with CG pipelines, with desired shapes, textures, and fine-grained animation capabilities. From a text input to describe the facial traits, we first introduce a coarse-to-fine scheme to generate the neutral facial geometry with a unified topology. We employ a selection strategy in the CLIP embedding space, and subsequently optimize both the details displacements and normals using Score Distillation Sampling from generic Latent Diffusion Model. Then, for neutral appearance generation, we introduce a dual-path mechanism, which combines the generic LDM with a novel texture LDM to ensure both the diversity and textural specification in the UV space. We also employ a two-stage optimization to perform SDS in both the latent and image spaces to significantly provides compact priors for fine-grained synthesis. Our generated neutral assets naturally support blendshapes-based facial animations. We further improve the animation ability with personalized deformation characteristics by learning the universal expression prior using the cross-identity hypernetwork. Notably, DreamFace can generate of realistic 3D facial assets with physically-based rendering quality and rich animation ability from video footage, even for fashion icons or exotic characters in cartoons and fiction movies.

cs.GR↗

PCRLv2: A Unified Visual Information Preservation Framework for Self-supervised Pre-training in Medical Image Analysis

Recent advances in self-supervised learning (SSL) in computer vision are primarily comparative, whose goal is to preserve invariant and discriminative semantics in latent representations by comparing siamese image views. However, the preserved high-level semantics do not contain enough local information, which is vital in medical image analysis (e.g., image-based diagnosis and tumor segmentation). To mitigate the locality problem of comparative SSL, we propose to incorporate the task of pixel restoration for explicitly encoding more pixel-level information into high-level semantics. We also address the preservation of scale information, a powerful tool in aiding image understanding but has not drawn much attention in SSL. The resulting framework can be formulated as a multi-task optimization problem on the feature pyramid. Specifically, we conduct multi-scale pixel restoration and siamese feature comparison in the pyramid. In addition, we propose non-skip U-Net to build the feature pyramid and develop sub-crop to replace multi-crop in 3D medical imaging. The proposed unified SSL framework (PCRLv2) surpasses its self-supervised counterparts on various tasks, including brain tumor segmentation (BraTS 2018), chest pathology identification (ChestX-ray, CheXpert), pulmonary nodule detection (LUNA), and abdominal organ segmentation (LiTS), sometimes outperforming them by large margins with limited annotations.

cs.CV↗

Maximal Function and Riesz Transform Characterizations of Hardy Spaces Associated with Homogeneous Higher Order Elliptic Operators and Ball Quasi-Banach Function Spaces

Let $L$ be a homogeneous divergence form higher order elliptic operator with complex bounded measurable coefficients on $\mathbb{R}^n$ and $X$ a ball quasi-Banach function space on $\mathbb{R}^n$ satisfying some mild assumptions. Denote by $H_{X,\, L}(\mathbb{R}^n)$ the Hardy space, associated with both $L$ and $X$, which is defined via the Lusin area function related to the semigroup generated by $L$. In this article, the authors establish both the maximal function and the Riesz transform characterizations of $H_{X,\, L}(\mathbb{R}^n)$. The results obtained in this article have a wide range of generality and can be applied to the weighted Hardy space, the variable Hardy space, the mixed-norm Hardy space, the Orlicz--Hardy space, the Orlicz-slice Hardy space, and the Morrey--Hardy space, associated with $L$. In particular, even when $L$ is a second order divergence form elliptic operator, both the maximal function and the Riesz transform characterizations of the mixed-norm Hardy space, the Orlicz-slice Hardy space, and the Morrey--Hardy space, associated with $L$, obtained in this article, are totally new.

math.FA↗

Preservational Learning Improves Self-supervised Medical Image Models by Reconstructing Diverse Contexts

Preserving maximal information is one of principles of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preservation. Moreover, it is necessary and complemental to introduce an explicit solution to preserve more information. From this perspective, we introduce Preservational Learning to reconstruct diverse image contexts in order to preserve more information in learned representations. Together with the contrastive loss, we present Preservational Contrastive Representation Learning (PCRL) for learning self-supervised medical representations. PCRL provides very competitive results under the pretraining-finetuning protocol, outperforming both self-supervised and supervised counterparts in 5 classification/segmentation tasks substantially.

cs.CV↗

Heat Kernels and Hardy Spaces on Non-Tangentially Accessible Domains with Applications to Global Regularity of Inhomogeneous Dirichlet Problems

Let $n\ge2$ and $Ω$ be a bounded non-tangentially accessible domain (for short, NTA domain) of $\mathbb{R}^n$. Assume that $L_D$ is a second-order divergence form elliptic operator having real-valued, bounded, measurable coefficients on $L^2(Ω)$ with the Dirichlet boundary condition. The main aim of this article is threefold. First, the authors prove that the heat kernels $\{K_t^{L_D}\}_{t>0}$ generated by $L_D$ are Hölder continuous. Second, for any $p\in(0,1]$, the authors introduce the `geometrical' Hardy space $H^p_r(Ω)$ by restricting any element of the Hardy space $H^p(\mathbb{R}^n)$ to $Ω$, and show that, when $p\in(\frac{n}{n+δ_0},1]$, $H^p_r(Ω)=H^p(Ω)=H^p_{L_D}(Ω)$ with equivalent quasi-norms, where $H^p(Ω)$ and $H^p_{L_D}(Ω)$ respectively denote the Hardy space on $Ω$ and the Hardy space associated with $L_D$, and $δ_0\in(0,1]$ is the critical index of the Hölder continuity for the kernels $\{K_t^{L_D}\}_{t>0}$. Third, as applications, the authors obtain the global gradient estimates in both $L^p(Ω)$, with $p\in(1,p_0)$, and $H^p_z(Ω)$, with $p\in(\frac{n}{n+1},1]$, for the inhomogeneous Dirichlet problem of second-order divergence form elliptic equations on bounded NTA domains, where $p_0\in(2,\infty)$ is a constant depending only on $n$, $Ω$, and the coefficient matrix of $L_D$. It is worth pointing out that the range $p\in(1,p_0)$ for the global gradient estimate in the scale of Lebesgue spaces $L^p(Ω)$ is sharp and the above results are established without any additional assumptions on both the coefficient matrix of $L_D$, and the domain $Ω$.

math.AP↗

Global Gradient Estimates for Dirichlet Problems of Elliptic Operators with a BMO Anti-Symmetric Part

Let $n\ge2$ and $Ω\subset\mathbb{R}^n$ be a bounded NTA domain. In this article, the authors investigate (weighted) global gradient estimates for Dirichlet boundary value problems of second order elliptic equations of divergence form with an elliptic symmetric part and a BMO anti-symmetric part in $Ω$. More precisely, for any given $p\in(2,\infty)$, the authors prove that a weak reverse Hölder inequality with exponent $p$ implies the global $W^{1,p}$ estimate and the global weighted $W^{1,q}$ estimate, with $q\in[2,p]$ and some Muckenhoupt weights, of solutions to Dirichlet boundary value problems. As applications, the authors establish some global gradient estimates for solutions to Dirichlet boundary value problems of second order elliptic equations of divergence form with small $\mathrm{BMO}$ symmetric part and small $\mathrm{BMO}$ anti-symmetric part, respectively, on bounded Lipschitz domains, quasi-convex domains, Reifenberg flat domains, $C^1$ domains, or (semi-)convex domains, in weighted Lebesgue spaces. Furthermore, as further applications, the authors obtain the global gradient estimate, respectively, in (weighted) Lorentz spaces, (Lorentz--)Morrey spaces, (Musielak--)Orlicz spaces, and variable Lebesgue spaces. Even on global gradient estimates in Lebesgue spaces, the results obtained in this article improve the known results via weakening the assumption on the coefficient matrix.

math.AP↗

A Two-Weight Boundedness Criterion and Its Applications

In this article, the authors establish a general (two-weight) boundedness criterion for a pair of functions, $(F,f)$, on $\mathbb{R}^n$ in the scale of weighted Lebesgue spaces, weighted Lorentz spaces, (Lorentz--)Morrey spaces, and variable Lebesgue spaces. As applications, the authors give a unified approach to prove the (two-weight) boundedness of Calderón--Zygmund operators, Littlewood--Paley $g$-functions, Lusin area functions, Littlewood--Paley $g^\ast_λ$-functions, and fractional integral operators, in the aforementioned function spaces. Moreover, via applying the above (two-weight) boundedness criterion, the authors further obtain the (two-weight) boundedness of Riesz transforms, Littlewood--Paley $g$-functions, and fractional integral operators associated with second-order divergence elliptic operators with complex bounded measurable coefficients on $\mathbb{R}^n$ in the aforementioned function spaces.

math.AP↗