SearcharxivSearch

arXiv subjects

Yufei Cai

Publications and source records attributed to Yufei Cai.

9 recordsLinked to original sources

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.

cs.CV

r2py: AI-Assisted Conversion of R Statistical Packages to Python

Thousands of R packages hold statistical methods with no native Python equivalent. Runtime bridges require an R installation; hand-written ports do not scale. Translation fails silently where the languages diverge, as in transform normalization, integer width, and argument evaluation. We present r2py, a framework that converts an R package into a native Python library using orchestrated language-model agents under human supervision, with correctness established by numerical comparison against the original at declared tolerances. The compiled code is retained unmodified, so any divergence lies in the translation. Seven phases decompose the work for independent invocations: structural analysis fixes conversion order, every base-R construct's rendering is settled in reviewable guides before code generation, and four verification methods each expose defects their predecessors miss. Packages reaching compiled code through .Call() add a five-phase prologue reconstructing the R C API they use. Conversions of KernSmooth and rpart reproduce R across 518 and 846 tests.

stat.CO

Missing pairs in open cluster catalogs

Open clusters (OCs) in our Galaxy can be found in pairs, possibly forming physical binaries, or in groups. These objects offer unique insights into the process of star formation and testify to the dynamical interactions at local and galactic scales. Therefore, building as complete a census as possible is a valuable endeavor. This work is aimed at identifying and characterizing new OC pair candidates that had been overlooked in previous studies. Two recent comprehensive catalogs were cross-matched to identify OCs in the first catalog that had been missing from the second one. From this list, counterparts in the second catalog were searched within a 3D distance of 50 pc. Candidate pairs were then selected by applying constraints on the tangential velocity (TV) difference. An orbital integration was performed to assess gravitational binding. The similarity in terms of the radial velocity (RV) and age was evaluated. We identified seven isolated binary cluster candidates, comprising two likely bound systems with stable orbits over 100 Myr; two pairs with a possible common origin but lacking RV confirmation; and three pairs with significant velocity discrepancies, suggesting they are unbound or in transitional states. We also identified six cluster group candidates, while refining the membership of known complexes such as UBC\_672 and NGC\_1977, and discovering a new group around FSR\_0198. Notably, the UBC\_392 group exhibits coherent proper motions but inconsistent RVs and large age spreads, indicating that it is not gravitationally bound. Additionally, we reconciled 15 clusters with discrepant nomenclature between the two catalogs. Multi-catalog integration combined with kinematic and dynamical validation is essential for establishing a complete census of Galactic cluster pairs.

astro-ph.GA

The morphological stability of open clusters: a new 2D perspective

Open clusters (OCs) usually evolve gradually as the number of their members changes, which can be manifested in their morphological characteristics. We aim to investigate the morphological stability of 1,490 OCs and further explore the potential change of morphological stability of the OCs at different spatial positions, using the OC catalog from the literature. We define for the first time a new morphological stability parameter Ncore/Nouter, a ratio of member numbers between cluster core and outer areas within tidal radii, which has a significant positive correlation against N, with a slope of 1.140$\pm$0.039, significantly steeper than the 0.720$\pm$0.026 measured for Score/Souter. This demonstrates that the stellar density in the core is a more sensitive tracer for morphological stability than geometry. Spatially, the radial sample OCs have larger slopes of Ncore/Nouter and Score/Souter against N, with 1.083$\pm$0.116 and 0.733$\pm$0.080, respectively, whereas those in the tangential direction 1.013$\pm$0.110 and 0.529$\pm$0.075, respectively, which means that the impact on sample OCs from tidal forces directed toward the Galactic center is possibly stronger than that from the shear force caused by the differential rotation of the Galactic disk. Besides, the sample OCs younger than 30 Myr display a shallow slope of 0.751$\pm$0.166, with those older than 800 Myr (1.442$\pm$0.128), reflecting that young OCs likely endure both internal disruptions, such as early dynamical heating weakening core binding and more severe external disturbances, compared to older OCs.

astro-ph.GA

EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models

The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing motion transfer methods explored the motion representations of reference videos to guide generation. Nevertheless, these methods typically rely on sample-specific optimization strategy, resulting in high computational burdens. In this paper, we propose EfficientMT, a novel and efficient end-to-end framework for video motion transfer. By leveraging a small set of synthetic paired motion transfer samples, EfficientMT effectively adapts a pretrained T2V model into a general motion transfer framework that can accurately capture and reproduce diverse motion patterns. Specifically, we repurpose the backbone of the T2V model to extract temporal information from reference videos, and further propose a scaler module to distill motion-related information. Subsequently, we introduce a temporal integration mechanism that seamlessly incorporates reference motion features into the video generation process. After training on our self-collected synthetic paired samples, EfficientMT enables general video motion transfer without requiring test-time optimization. Extensive experiments demonstrate that our EfficientMT outperforms existing methods in efficiency while maintaining flexible motion controllability. Our code will be available https://github.com/PrototypeNx/EfficientMT.

cs.CV

Decoupled Textual Embeddings for Customized Image Generation

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned concept, limiting the potential to compose concept into new scenes. To address these issues, we propose the DETEX, a novel approach that learns the disentangled concept embedding for flexible customized text-to-image generation. Unlike conventional methods that learn a single concept embedding from the given images, our DETEX represents each image using multiple word embeddings during training, i.e., a learnable image-shared subject embedding and several image-specific subject-unrelated embeddings. To decouple irrelevant attributes (i.e., background and pose) from the subject embedding, we further present several attribute mappers that encode each image as several image-specific subject-unrelated embeddings. To encourage these unrelated embeddings to capture the irrelevant information, we incorporate them with corresponding attribute words and propose a joint training strategy to facilitate the disentanglement. During inference, we only use the subject embedding for image generation, while selectively using image-specific embeddings to retain image-specified attributes. Extensive experiments demonstrate that the subject embedding obtained by our method can faithfully represent the target concept, while showing superior editability compared to the state-of-the-art methods. Our code will be made published available.

cs.CV

Maximum Likelihood Estimates of Parameters in Generalized Gamma Distribution with SeLF Algorithm

This undergraduate thesis focuses on calculating maximum likelihood estimates of parameters in the generalized Gamma distribution using the SeLF algorithm. As an extension of the Gamma distribution, the generalized Gamma distribution can better fit real data and has been widely applied. The research begins by exploring the definition of the generalized Gamma distribution and its similarities and differences from the traditional Gamma distribution. Then, the SeLF and US algorithms are discussed in detail. The SeLF algorithm is a new algorithm based on the Minorization-Maximization algorithm, which can obtain the local optimal solution with few iterations, with the advantages of fast computation, high accuracy, and good convergence. The US algorithm is a method for finding the zeros of a function, which stands at a higher level than the SeLF algorithm and can improve the convergence speed and stability. This thesis proposes a method for calculating maximum likelihood estimates of the parameters in the generalized Gamma distribution using the SeLF and US algorithms, and presents the practical implementation of the algorithms, as well as simulations and data analysis to evaluate the performance of the proposed methods. The results demonstrate that the SeLF algorithm can achieve more stable and accurate estimates of the parameters in the generalized Gamma distribution more quickly, compared to traditional Newton's method, which can be useful in various applications. This thesis provides a comprehensive and in-depth exploration of the generalized Gamma distribution and the SeLF algorithm, and proposes a new method for calculating maximum likelihood estimates of parameters, contributing to the development of statistical methods for parameter estimation in complex models. The proposed method in this thesis has important practical significance and application value for solving practical problems.

stat.ME

Denotational validation of higher-order Bayesian inference

We present a modular semantic account of Bayesian inference algorithms for probabilistic programming languages, as used in data science and machine learning. Sophisticated inference algorithms are often explained in terms of composition of smaller parts. However, neither their theoretical justification nor their implementation reflects this modularity. We show how to conceptualise and analyse such inference algorithms as manipulating intermediate representations of probabilistic programs using higher-order functions and inductive types, and their denotational semantics. Semantic accounts of continuous distributions use measurable spaces. However, our use of higher-order functions presents a substantial technical difficulty: it is impossible to define a measurable space structure over the collection of measurable functions between arbitrary measurable spaces that is compatible with standard operations on those functions, such as function application. We overcome this difficulty using quasi-Borel spaces, a recently proposed mathematical structure that supports both function spaces and continuous distributions. We define a class of semantic structures for representing probabilistic programs, and semantic validity criteria for transformations of these representations in terms of distribution preservation. We develop a collection of building blocks for composing representations. We use these building blocks to validate common inference algorithms such as Sequential Monte Carlo and Markov Chain Monte Carlo. To emphasize the connection between the semantic manipulation and its traditional measure theoretic origins, we use Kock's synthetic measure theory. We demonstrate its usefulness by proving a quasi-Borel counterpart to the Metropolis-Hastings-Green theorem.

cs.PL

A Theory of Changes for Higher-Order Languages - Incrementalizing λ-Calculi by Static Differentiation

If the result of an expensive computation is invalidated by a small change to the input, the old result should be updated incrementally instead of reexecuting the whole computation. We incrementalize programs through their derivative. A derivative maps changes in the program's input directly to changes in the program's output, without reexecuting the original program. We present a program transformation taking programs to their derivatives, which is fully static and automatic, supports first-class functions, and produces derivatives amenable to standard optimization. We prove the program transformation correct in Agda for a family of simply-typed λ-calculi, parameterized by base types and primitives. A precise interface specifies what is required to incrementalize the chosen primitives. We investigate performance by a case study: We implement in Scala the program transformation, a plugin and improve performance of a nontrivial program by orders of magnitude.

cs.PL