SearcharxivSearch

arXiv subjects

Yanhong Yang

Publications and source records attributed to Yanhong Yang.

9 recordsLinked to original sources

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.

cs.CV

MovePose: A High-performance Human Pose Estimation Algorithm on Mobile and Edge Devices

We present MovePose, an optimized lightweight convolutional neural network designed specifically for real-time body pose estimation on CPU-based mobile devices. The current solutions do not provide satisfactory accuracy and speed for human posture estimation, and MovePose addresses this gap. It aims to maintain real-time performance while improving the accuracy of human posture estimation for mobile devices. Our MovePose algorithm has attained an Mean Average Precision (mAP) score of 68.0 on the COCO \cite{cocodata} validation dataset. The MovePose algorithm displayed efficiency with a performance of 69+ frames per second (fps) when run on an Intel i9-10920x CPU. Additionally, it showcased an increased performance of 452+ fps on an NVIDIA RTX3090 GPU. On an Android phone equipped with a Snapdragon 8 + 4G processor, the fps reached above 11. To enhance accuracy, we incorporated three techniques: deconvolution, large kernel convolution, and coordinate classification methods. Compared to basic upsampling, deconvolution is trainable, improves model capacity, and enhances the receptive field. Large kernel convolution strengthens these properties at a decreased computational cost. In summary, MovePose provides high accuracy and real-time performance, marking it a potential tool for a variety of applications, including those focused on mobile-side human posture estimation. The code and models for this algorithm will be made publicly accessible.

cs.CV

Asymmetric Dual-Decoder U-Net for Joint Rain and Haze Removal

This work studies the joint rain and haze removal problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading to a performance drop in the visual applications, such as autonomous driving. However, jointly removing the rain and haze in scene images is ill-posed and challenging, where the existence of haze and rain and the change of atmosphere light, can both degrade the scene information. Current methods focus on the contamination removal part, thus ignoring the restoration of the scene information affected by the change of atmospheric light. We propose a novel deep neural network, named Asymmetric Dual-decoder U-Net (ADU-Net), to address the aforementioned challenge. The ADU-Net produces both the contamination residual and the scene residual to efficiently remove the rain and haze while preserving the fidelity of the scene information. Extensive experiments show our work outperforms the existing state-of-the-art methods by a considerable margin in both synthetic data and real-world data benchmarks, including RainCityscapes, BID Rain, and SPA-Data. For instance, we improve the state-of-the-art PSNR value by 2.26/4.57 on the RainCityscapes/SPA-Data, respectively. Codes will be made available freely to the research community.

cs.CV

Uniformization of $p$-adic curves via Higgs-de Rham flows

Let $k$ be an algebraic closure of a finite field of odd characteristic. We prove that for any rank two graded Higgs bundle with maximal Higgs field over a generic hyperbolic curve $X_1$ defined over $k$, there exists a lifting $X$ of the curve to the ring $W(k)$ of Witt vectors as well as a lifting of the Higgs bundle to a periodic Higgs bundle over $X/W$. As a consequence, it gives rise to a two-dimensional absolutely irreducible representation of the arithmetic fundamental group $π_1(X_K)$ of the generic fiber of $X$. This curve $X$ and its associated representation is in close relation with the canonical curve and its associated canonical crystalline representation in the $p$-adic Teichmüller theory for curves due to S. Mochizuki. Our result may be viewed as an analogue of the Hitchin-Simpson's uniformization theory of hyperbolic Riemann surfaces via Higgs bundles.

math.AG

Rationality of Euler-Chow series and finite generation of Cox rings

In this paper we work with a series whose coefficients are the Euler characteristic of Chow varieties of a given projective variety. For varieties where the Cox ring is defined, it is easy to see that in this case the ring associated to the series is the Cox ring. If this ring is noetherian then the series is rational. It is an open question whether the converse holds. In this paper we give an example showing the converse fails. However we conjecture that it holds when the variety is rationally connected. As an evidence of this conjecture, It is proved that the series is not rational, and in a sense defined, not algebraic, in the case of the blowup of the projective plane at nine or more points in general position. Furthermore, we also treat some other examples of varieties with infinitely generated Cox ring, studied by Mukai and Hassett-Tschinkel. These are the first examples known where the series is not rational. We also compute the series for Del Pezzo surfaces.

math.AG

Semistable Higgs bundles of small ranks are strongly Higgs semistable

We give a proof of the conjecture that a semistable Higgs bundle is strongly Higgs semistable in the case of small ranks, based upon the fact that there exists a gr-semistable Griffiths-transverse filtration on a $\nabla$-invariant semistable vector bundle. The latter is a generalization of a result of Carlos Simpson.

math.AG

The Fixed Point Locus of the Verschiebung on M_x(2,0) for Genus-2 Curves X in Charateristic 2

In this note, we prove that for every ordinary genus-2 curve $X$ over a finite field $κ$ of characteristic 2 with $\text{Aut}(X/κ)=\db{Z}/2\db{Z} \times S_3$, there exist $\text{SL}(2,κ\sembrack{s})$-representations of $π_1(X)$ such that the image of $π_1(\bar{X})$ is infinite. This result gives a geometric interpretation of Laszlo's counterexample [12] to a question regarding the finiteness of the geometric monodromy of representations of the fundamental group [4].

math.AG

An Improvement of de Jong-Oort's Purity Theorem

Consider an $F$-crystal over a noetherian scheme $S$. De Jong--Oort's purity theorem states that the associated Newton polygons over all points of $S$ are constant if this is true outside a subset of codimension bigger than 1. In this paper we show an improvement of the theorem, which says that the Newton polygons over all points of $S$ have a common break point if this is true outside a subset of codimension bigger than 1.

math.AG

On special geometry of the moduli space of string vacua with fluxes

In this paper we construct a special geometry over the moduli space of type II string vacua with both NS and RR fluxes turning on. Depending on what fluxes are turning on we divide into three cases of moduli space of generalized structures. They are respectively generalized Calabi-Yau structures, generalized Calabi-Yau metric structures and ${\cal N} =1$ generalized string vacua. It is found that the $d d^{\cal J}$ lemma can be established for all three cases. With the help of the $d d^{\cal J}$ lemma we identify the moduli space locally as a subspace of $d_{H}$ cohomologies. This leads naturally to the special geometry of the moduli space. It has a flat symplectic structure and a K$\ddot{\rm a}$hler metric with the Hitchin functional (modified if RR fluxes are included) the K$\ddot{\rm a}$hler potential. Our work is based on previous works of Hitchin and recent works of Gra$\tilde{\rm n}$a-Louis-Waldram, Goto, Gualtieri, Yi Li and Tomasiello. The special geometry is useful in flux compactifications of type II string theories.

hep-th