SearcharxivSearch

arXiv subjects

Zipeng Wang

Publications and source records attributed to Zipeng Wang.

At least 19 recordsLinked to original sources

Compact Toeplitz operators via the Berezin transform on radial weighted Bergman spaces

Let $\omega$ be a radial $\widehat{\mathcal D}$-weight and $u$ be a bounded function on the unit disk $\mathbb D$. We prove that the Toeplitz operator \(T_{\omega,u}\) is compact on \(A_\omega^2\) if and only if its Berezin transform vanishes at the boundary. Our approach is based on a polynomial frame for $A_\omega^2$ and a detailed localization analysis of the resulting infinite matrix representation of $T_{\omega,u}$. Even in the unweighted Bergman space \(A^2\), our argument is new and does not rely on the classical translation operators. We further show that this Axler--Zheng compactness characterization does not extend, in general, to products of Toeplitz operators on \(\widehat{\mathcal D}\)-weighted Bergman spaces, and hence to the corresponding Toeplitz algebra generated by bounded symbols. More precisely, we construct a radial log-subharmonic \(\widehat{\mathcal D}\)-weight \(\omega\) and bounded symbols \(u,v\) such that the product \(T_{\omega,v}T_{\omega,u}\) is noncompact, whereas its Berezin transform vanishes at the boundary.

math.CV

GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization

Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinates, however, has lagged behind. Predicting a place name for a geocoding API is discrete and lossy: it ignores image evidence and collapses multi-granular semantics into a coarse lookup. We argue that the bottleneck has shifted from what a model reasons to how that reasoning is represented for a continuous, geometry-aware decoder. We present GeoBridge, a role-decoupled conditioning mechanism that connects a frozen semantic MLLM to a frozen Riemannian flow-matching head that generates coordinates on the sphere. The central obstacle is arole conflict: supervising the condition with discrete semantic labels biases its representation toward class-discriminative geometry, at odds with the smooth manifold the generative head requires. GeoBridge keeps the semantic supervision decoupled from the condition interface: a separate projection forms the continuous condition the frozen head expects, injecting geographic priors without disturbing the spherical decoder. On IM2GPS3K, GeoBridge reaches 38.67/52.89/70.37 at the 25/200/750 km thresholds, improving over a place-name-to-API pipeline and reasoning-augmented direct prediction at these precision-relevant scales. GeoBridge is a decode-side algorithmic contribution, orthogonal and complementary to chain-of-thought reasoning. Code will be made publicly available.

cs.CV

Strong and weak-type estimates for radial weighted Bergman projections

We completely characterize the $L^p$-boundedness and the weak-type (1,1) estimate of radial weighted Bergman projections on the unit disk. Our result, in particular, confirms a conjecture proposed by Pel\'{a}ez and R\"{a}tty\"{a} in 2021 and thereby settles a longstanding problem in the area that was formally posed by Dostani\'{c} in 2004. Consequently, we establish the dichotomy that a radial weighted Bergman projection is bounded either only for $p=2$, or for all $p\in(1,\infty)$.

math.CV

A new proof of maximal theorem on Heisenberg groups

Given $0\leq\alpha<1$, we define \[\begin{array}{lr} \mathbf{M}_\alpha f(u,v,t) = \sup_{ \mathbf{R} \ni (0,0,0)} {\rm vol} \{\mathbf{R}\}^{\alpha-1} \iiint_\mathbf{R}\left|f [(u,v,t)\odot(\xi,\eta,\tau)^{-1}]\right|d\xi d\eta d\tau \end{array}\] where $\mathbf{R}\subset\mathbb{R}^{2n+1}$ is a rectangle parallel to the coordinates. Moreover, $\odot$ denotes the multiplication law on a real Heisenberg group. The $\mathbf{L}^p$-boundedness of $\mathbf{M}_0$ has been previously proved by M. Christ. We show $\mathbf{M}_\alpha\colon\mathbf{L}^p(\mathbb{R}^{2n+1}) \to \mathbf{L}^q(\mathbb{R}^{2n+1})$ for $\alpha={1\over p}-{1\over q},~ 1<p\leq q<\infty$ by applying a geometric covering lemma due to C\'{o}rdoba and Fefferman.

math.CA

Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D

Current 2D-to-3D conversion methods achieve geometric accuracy but are artistically deficient, failing to replicate the immersive and emotionally resonant experience of professional 3D cinema. This is because geometric reconstruction paradigms mistake deliberate artistic intent, such as strategic zero-plane shifts for pop-out effects and local depth sculpting, for data noise or ambiguity. This paper argues for a new paradigm: Artistic Disparity Synthesis, shifting the goal from physically accurate disparity estimation to artistically coherent disparity synthesis. We propose Art3D, a preliminary framework exploring this paradigm. Art3D uses a dual-path architecture to decouple global depth parameters (macro-intent) from local artistic effects (visual brushstrokes) and learns from professional 3D film data via indirect supervision. We also introduce a preliminary evaluation method to quantify cinematic alignment. Experiments show our approach demonstrates potential in replicating key local out-of-screen effects and aligning with the global depth styles of cinematic 3D content, laying the groundwork for a new class of artistically-driven conversion tools.

cs.CV

HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its success heavily relies on dense camera coverage; under sparse-view conditions, insufficient supervision leads to irregular Gaussian distributions, characterized by globally sparse coverage, blurred background, and distorted high-frequency areas. To address this, we propose HeroGS, Hierarchical Guidance for Robust 3D Gaussian Splatting, a unified framework that establishes hierarchical guidance across the image, feature, and parameter levels. At the image level, sparse supervision is converted into pseudo-dense guidance, globally regularizing the Gaussian distributions and forming a consistent foundation for subsequent optimization. Building upon this, Feature-Adaptive Densification and Pruning (FADP) at the feature level leverages low-level features to refine high-frequency details and adaptively densifies Gaussians in background regions. The optimized distributions then support Co-Pruned Geometry Consistency (CPG) at parameter level, which guides geometric consistency through parameter freezing and co-pruning, effectively removing inconsistent splats. The hierarchical guidance strategy effectively constrains and optimizes the overall Gaussian distributions, thereby enhancing both structural fidelity and rendering quality. Extensive experiments demonstrate that HeroGS achieves high-fidelity reconstructions and consistently surpasses state-of-the-art baselines under sparse-view conditions.

cs.CV

The optimal hypercontractive constants for $\mathbb{Z}_3$ and biased Bernoulli random variables

We resolve a folklore problem of determining the optimal hypercontractive constants $r_{p,q}(\mathbb{Z}_3)$ for the cyclic group $\mathbb{Z}_3$ for all $1 < p < q < \infty$. More precisely, we have \[ r_{p,q}(\mathbb{Z}_3) = \frac{(1 + 2x)(1 - y)}{(1 + 2y)(1 - x)}, \] where $(x,y)$ is the unique solution in the open unit square $(0,1)\times (0,1)$ to the system of equations \begin{align*} \left\{ \begin{aligned} &\frac{1}{1+2x}\Big(\frac{1+2x^p}{3}\Big)^{\frac{1}{p}}=\frac{1}{1+2y}\Big(\frac{1+2y^q}{3}\Big)^{\frac{1}{q}},\\ &\frac{(1-x)(1-x^{p-1})}{1+2x^p}=\frac{(1-y)(1-y^{q-1})}{1+2y^q}. \end{aligned} \right. \end{align*} Consequently, for rational $p, q\in \mathbb{Q}$, the constants $r_{p,q}(\mathbb{Z}_3)$ are algebraic numbers which generally admit no radical expressions, since their often rather complicated minimal polynomials may have non-solvable Galois groups. Our formalism relies on a key observation: the existence of nontrivial critical extremizers. This approach can also be adapted to resolve a long-standing open problem -- determining all optimal $(p,q)$-hypercontractive constants for biased Bernoulli random variables, which are closely related to noise operators. Several noteworthy phenomena emerge from numerical simulations: the monotonicity of the hypercontractive constants in the parameters, and the appearance of intriguing limit shapes. These phenomena merit further investigation.

math.FA

$L^p$--$L^q$ estimates for Shimorin-type integral operators

Let $\nu$ be a positive measure on $[0,1]$. A Shimorin-type operator $T_\nu$ is an integral operator on the unit disk given by \[ T_\nu f(z) = \int_{\mathbb{D}} \frac{1}{1 - z\overline{\lambda}} \left( \int_0^1 \frac{d\nu(r)}{1 - r z \overline{\lambda}} \right) f(\lambda) \, dA(\lambda), \] which originates from Shimorin's work on Bergman-type kernel representations for logarithmically subharmonic weighted Bergman spaces. In this paper, we study $L^p$--$L^q$ estimates for $T_\nu$. Unlike classical Bergman-type operators, the critical line on the $(1/p,1/q)$-plane that separates the boundedness and unboundedness regions of $T_\nu$ is not immediately evident. Moreover, even along this line, new phenomena arise. In the present work, by introducing a quantity $c_\nu$, \begin{itemize} \item we first determine the critical boundary in the $(1/p,1/q)$-plane for bounded $T_\nu$; \item furthermore, on this critical line, we establish necessary and sufficient conditions for $T_\nu$ which have standard Bergman-type $L^p$--$L^q$ estimates, meaning that it is bounded in the interior of the region and admits weak-type and BMO-type estimates at endpoints. \end{itemize}

math.CV

FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention

3D reconstruction from multi-view images is a core challenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to traditional per-scene optimization techniques. Among them, state-of-the-art models like the Visual Geometry Grounding Transformer (VGGT) leverage full self-attention over all image tokens to capture global relationships. However, this approach suffers from poor scalability due to the quadratic complexity of self-attention and the large number of tokens generated in long image sequences. In this work, we introduce FlashVGGT, an efficient alternative that addresses this bottleneck through a descriptor-based attention mechanism. Instead of applying dense global attention across all tokens, FlashVGGT compresses spatial information from each frame into a compact set of descriptor tokens. Global attention is then computed as cross-attention between the full set of image tokens and this smaller descriptor set, significantly reducing computational overhead. Moreover, the compactness of the descriptors enables online inference over long sequences via a chunk-recursive mechanism that reuses cached descriptors from previous chunks. Experimental results show that FlashVGGT achieves reconstruction accuracy competitive with VGGT while reducing inference time to just 9.3% of VGGT for 1,000 images, and scaling efficiently to sequences exceeding 3,000 images. Our project page is available at https://wzpscott.github.io/flashvggt_page/.

cs.CV

Stein-Weiss inequality revisit on Heisenberg group

We study a family of fractional integral operators defined on Heisenberg group whose kernels satisfy Zygmund dilation. We give a characterization between a two-weight norm inequality and the necessary constraints by considering the weights to be suitable powers. As a result, we obtain a Stein-Weiss inequality on Heisenberg group.

math.CA

PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes

Recent advancements in personalized image generation have significantly improved facial identity preservation, particularly in fields such as entertainment and social media. However, existing methods still struggle to achieve precise control over facial attributes in a per-subject-tuning-free (PSTF) way. Tuning-based techniques like PreciseControl have shown promise by providing fine-grained control over facial features, but they often require extensive technical expertise and additional training data, limiting their accessibility. In contrast, PSTF approaches simplify the process by enabling image generation from a single facial input, but they lack precise control over facial attributes. In this paper, we introduce a novel, PSTF method that enables both precise control over facial attributes and high-fidelity preservation of facial identity. Our approach utilizes a face recognition model to extract facial identity features, which are then mapped into the $W^+$ latent space of StyleGAN2 using the e4e encoder. We further enhance the model with a Triplet-Decoupled Cross-Attention module, which integrates facial identity, attribute features, and text embeddings into the UNet architecture, ensuring clean separation of identity and attribute information. Trained on the FFHQ dataset, our method allows for the generation of personalized images with fine-grained control over facial attributes, while without requiring additional fine-tuning or training data for individual identities. We demonstrate that our approach successfully balances personalization with precise facial attribute control, offering a more efficient and user-friendly solution for high-quality, adaptable facial image synthesis. The code is publicly available at https://github.com/UnicomAI/PSTF-AttControl.

cs.CV

Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarity

With the rapid advancement of large language models (LLMs), natural language processing (NLP) has achieved remarkable progress. Nonetheless, significant challenges remain in handling texts with ambiguity, polysemy, or uncertainty. We introduce the Fuzzy Reasoning Chain (FRC) framework, which integrates LLM semantic priors with continuous fuzzy membership degrees, creating an explicit interaction between probability-based reasoning and fuzzy membership reasoning. This transition allows ambiguous inputs to be gradually transformed into clear and interpretable decisions while capturing conflicting or uncertain signals that traditional probability-based methods cannot. We validate FRC on sentiment analysis tasks, where both theoretical analysis and empirical results show that it ensures stable reasoning and facilitates knowledge transfer across different model scales. These findings indicate that FRC provides a general mechanism for managing subtle and ambiguous expressions with improved interpretability and robustness.

cs.CL

HyRF: Hybrid Radiance Fields for Memory-efficient and High-quality Novel View Synthesis

Recently, 3D Gaussian Splatting (3DGS) has emerged as a powerful alternative to NeRF-based approaches, enabling real-time, high-quality novel view synthesis through explicit, optimizable 3D Gaussians. However, 3DGS suffers from significant memory overhead due to its reliance on per-Gaussian parameters to model view-dependent effects and anisotropic shapes. While recent works propose compressing 3DGS with neural fields, these methods struggle to capture high-frequency spatial variations in Gaussian properties, leading to degraded reconstruction of fine details. We present Hybrid Radiance Fields (HyRF), a novel scene representation that combines the strengths of explicit Gaussians and neural fields. HyRF decomposes the scene into (1) a compact set of explicit Gaussians storing only critical high-frequency parameters and (2) grid-based neural fields that predict remaining properties. To enhance representational capacity, we introduce a decoupled neural field architecture, separately modeling geometry (scale, opacity, rotation) and view-dependent color. Additionally, we propose a hybrid rendering scheme that composites Gaussian splatting with a neural field-predicted background, addressing limitations in distant scene representation. Experiments demonstrate that HyRF achieves state-of-the-art rendering quality while reducing model size by over 20 times compared to 3DGS and maintaining real-time performance. Our project page is available at https://wzpscott.github.io/hyrf/.

cs.CV