SearcharxivSearch

arXiv subjects

Yifan Cheng

Publications and source records attributed to Yifan Cheng.

At least 19 recordsLinked to original sources

Compression of virtual spaces in transcorrelated methods via singular value decomposition: application to the G2 set

We introduce a new singular-value-decomposition-based scheme for constructing small virtual spaces out of large basis sets for transcorrelated (TC) calculations, termed SVD-TC. This work builds on the recent finding that the residual basis error in the TC reference energy converges more slowly than that of the correlation energy. Within the new workflow, the post Hartree-Fock TC calculation is performed in a compressed virtual orbital subspace, obtained by projecting the canonical virtual orbitals from a large basis set onto a smaller basis set through singular value decomposition (SVD). This allows us to achieve the high accuracy allowed by the large basis, whilst the bottleneck steps - TC integral calculation and post-HF correlation method such as CCSD(T) - incur the cost of only a small virtual space calculation. The method therefore is highly efficient, whilst avoiding the composite nature of the reference correction method. Using the new scheme, we widen the scope of benchmark-quality TC results into more complex molecules than previously considered: using the G2-1 set of 55 molecules with first- and second-row atoms, we apply SVD-xTC-CCSD(T) to compute atomization energies. We compare our results against the near-exact semistochastic heat-bath configuration interaction (SHCI) reference values and experiment. We find that SVD-xTC-CCSD(T) delivers chemical accuracy already with triple-$\zeta$ basis sets. Finally, we use the quadruple-$\zeta$ results to analyze the accuracy of pseudopotentials within the TC method, and show that pseudopotential TC workflow provides faster basis-set convergence than all-electron TC. We also present timings for computing the atomization energies on G2-1 set, demonstrating the efficiency of our TC workflows.

physics.chem-ph

Interpolative Separable Density-Fitting for Transcorrelated Hamiltonians

The transcorrelated (TC) method dramatically accelerates the convergence of correlated calculations toward the complete-basis-set (CBS) limit by folding a Jastrow correlator into the Hamiltonian via a similarity transformation, incorporating the electron--electron cusp into the effective interaction. We make the TC framework practical for large systems and flexible, multi-center correlators by compressing the grid-evaluated TC integrals with the interpolative separable density-fitting (ISDF) approximation, combined with the effective two-body (xTC) treatment of the three-body operator. This low-rank representation reduces storage and integration costs by orders of magnitude, and a multi-GPU implementation with automatic differentiation of the correlator makes the construction routine for large basis sets. We demonstrate the resulting ISDF-xTC-CCSD method on the linear hydrogen chain, reaching the joint thermodynamic and CBS limits with basis sets up to cc-pV5Z in agreement with state-of-the-art many-body references to within about 1~mHa/atom, and on the benzene ground-state energy with up to 1200 orbitals (cc-pCV5Z), where the method attains state-of-the-art accuracy at the coupled cluster singles and doubles level and its CBS extrapolation is markedly more robust than that of conventional coupled-cluster methods.

physics.chem-ph

An Additive Reference Correction Scheme for the Transcorrelated Method

We introduce an additive reference correction for the transcorrelated (TC) method and its three-body mean-field approximation (xTC), to improve energy differences computed in small orbital basis sets. The correction is motivated by the observation that, for xTC atomization energies, the dominant error in double-{\zeta} bases originates from the reference contribution rather than from the correlation energy. In the proposed reference-corrected scheme (RC-xTC), the small-basis correlation energy is retained, while the corresponding TC reference energy is replaced by its value from a larger basis. Benchmark calculations for the non-relativistic HEAT set with the Dunning basis-set family show that RC-xTC substantially improves both total and atomization energies relative to standard xTC in double-{\zeta} bases. At the CCSD(T) level, RC-xTC yields better atomization energies than CCSD(T)-F12a in the double-{\zeta} regime, while preserving the favorable total-energy accuracy of xTC. At the CCSD level, RC-xTC improves atomization energies relative to F12a throughout the full basis-set sequence. As the basis set is enlarged, xTC and RC-xTC become progressively identical, as expected from the construction of the correction.

physics.chem-ph

Fish Audio S2 Technical Report

We introduce Fish Audio S2, an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and, most importantly, instruction-following control via natural-language descriptions. To scale training, we develop a multi-stage training recipe together with a staged data pipeline covering video captioning and speech captioning, voice-quality assessment, and reward modeling. To push the frontier of open-source TTS, we release our model weights, fine-tuning code, and an SGLang-based inference engine. The inference engine is production-ready for streaming, achieving an RTF of 0.195 and a time-to-first-audio below 100 ms.Our code and weights are available on GitHub (https://github.com/fishaudio/fish-speech) and Hugging Face (https://huggingface.co/fishaudio/s2-pro). We highly encourage readers to visit https://fish.audio to try custom voices.

cs.SD

Modular Construction of Jastrow Factors for the Transcorrelated Method

In this work, we explore the reuse of terms in the Jastrow factor between systems for use in the transcorrelated method, to reduce the number of optimisable parameters for a given system. In particular, we propose a workflow in which atom-specific parts of Jastrow factors, optimised in atoms, may be reused in the molecule, with only a few parameters in the electron-electron part of the Jastrow left to optimise, while maintaining performance. We find that the modified workflow not only reduces the number of terms needing to be optimised, but also improves the accuracy of xTC-CCSD(T) energies.

physics.chem-ph

Micro-Expression Recognition via Fine-Grained Dynamic Perception

Facial micro-expression recognition (MER) is a challenging task, due to the transience, subtlety, and dynamics of micro-expressions (MEs). Most existing methods resort to hand-crafted features or deep networks, in which the former often additionally requires key frames, and the latter suffers from small-scale and low-diversity training data. In this paper, we develop a novel fine-grained dynamic perception (FDP) framework for MER. We propose to rank frame-level features of a sequence of raw frames in chronological order, in which the rank process encodes the dynamic information of both ME appearances and motions. Specifically, a novel local-global feature-aware transformer is proposed for frame representation learning. A rank scorer is further adopted to calculate rank scores of each frame-level feature. Afterwards, the rank features from rank scorer are pooled in temporal dimension to capture dynamic representation. Finally, the dynamic representation is shared by a MER module and a dynamic image construction module, in which the former predicts the ME category, and the latter uses an encoder-decoder structure to construct the dynamic image. The design of dynamic image construction task is beneficial for capturing facial subtle actions associated with MEs and alleviating the data scarcity issue. Extensive experiments show that our method (i) significantly outperforms the state-of-the-art MER methods, and (ii) works well for dynamic image construction. Particularly, our FDP improves by 4.05%, 2.50%, 7.71%, and 2.11% over the previous best results in terms of F1-score on the CASME II, SAMM, CAS(ME)^2, and CAS(ME)^3 datasets, respectively. The code is available at https://github.com/CYF-cuber/FDP.

cs.CV

Effective Bayesian Modeling of Large Spatiotemporal Count Data Using Autoregressive Gamma Processes

We put forward a new Bayesian modeling strategy for spatiotemporal count data that enables efficient posterior sampling. Most previous models for such data decompose logarithms of the response Poisson rates into fixed effects and spatial random effects, where the latter is typically assumed to follow a latent Gaussian process, the conditional autoregressive model, or the intrinsic conditional autoregressive model. Since log-Gaussian is not conjugate to Poisson, such implementations must resort to either approximation methods like INLA or Metropolis moves on latent states in MCMC algorithms for model fitting and exhibit several approximation and posterior sampling challenges. Instead of modeling logarithms of spatiotemporal frailties jointly as a Gaussian process, we construct a spatiotemporal autoregressive gamma process guaranteed stationary across the time dimension. We decompose latent Poisson variables to permit fully conjugate Gibbs sampling of spatiotemporal frailties and design a sparse spatial dependence structure to get a linear computational complexity that facilitates efficient posterior computation. Our model permits convenient Bayesian predictive machinery based on posterior samples that delivers satisfactory performance in predicting at new spatial locations and time intervals. We have performed extensive simulation experiments and real data analyses, which corroborated our model's accurate parameter estimation, model fitting, and out-of-sample prediction capabilities.

stat.ME

MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style Convolution

Facial micro-expression recognition (MER) is a challenging problem, due to transient and subtle micro-expression (ME) actions. Most existing methods depend on hand-crafted features, key frames like onset, apex, and offset frames, or deep networks limited by small-scale and low-diversity datasets. In this paper, we propose an end-to-end micro-action-aware deep learning framework with advantages from transformer, graph convolution, and vanilla convolution. In particular, we propose a novel F5C block composed of fully-connected convolution and channel correspondence convolution to directly extract local-global features from a sequence of raw frames, without the prior knowledge of key frames. The transformer-style fully-connected convolution is proposed to extract local features while maintaining global receptive fields, and the graph-style channel correspondence convolution is introduced to model the correlations among feature patterns. Moreover, MER, optical flow estimation, and facial landmark detection are jointly trained by sharing the local-global features. The two latter tasks contribute to capturing facial subtle action information for MER, which can alleviate the impact of insufficient training data. Extensive experiments demonstrate that our framework (i) outperforms the state-of-the-art MER methods on CASME II, SAMM, and SMIC benchmarks, (ii) works well for optical flow estimation and facial landmark detection, and (iii) can capture facial subtle muscle actions in local regions associated with MEs. The code is available at https://github.com/CYF-cuber/MOL.

cs.CV

ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation

Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and, noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding.

cs.SD

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning.

cs.SD

Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Text-to-Speech (TTS) systems face ongoing challenges in processing complex linguistic features, handling polyphonic expressions, and producing natural-sounding multilingual speech - capabilities that are crucial for future AI applications. In this paper, we present Fish-Speech, a novel framework that implements a serial fast-slow Dual Autoregressive (Dual-AR) architecture to enhance the stability of Grouped Finite Scalar Vector Quantization (GFSQ) in sequence generation tasks. This architecture improves codebook processing efficiency while maintaining high-fidelity outputs, making it particularly effective for AI interactions and voice cloning. Fish-Speech leverages Large Language Models (LLMs) for linguistic feature extraction, eliminating the need for traditional grapheme-to-phoneme (G2P) conversion and thereby streamlining the synthesis pipeline and enhancing multilingual support. Additionally, we developed FF-GAN through GFSQ to achieve superior compression ratios and near 100\% codebook utilization. Our approach addresses key limitations of current TTS systems while providing a foundation for more sophisticated, context-aware speech synthesis. Experimental results show that Fish-Speech significantly outperforms baseline models in handling complex linguistic scenarios and voice cloning tasks, demonstrating its potential to advance TTS technology in AI applications. The implementation is open source at \href{https://github.com/fishaudio/fish-speech}{https://github.com/fishaudio/fish-speech}.

cs.SD

V2I-Calib++: A Multi-terminal Spatial Calibration Approach in Urban Intersections for Collaborative Perception

Urban intersections, dense with pedestrian and vehicular traffic and compounded by GPS signal obstructions from high-rise buildings, are among the most challenging areas in urban traffic systems. Traditional single-vehicle intelligence systems often perform poorly in such environments due to a lack of global traffic flow information and the ability to respond to unexpected events. Vehicle-to-Everything (V2X) technology, through real-time communication between vehicles (V2V) and vehicles to infrastructure (V2I), offers a robust solution. However, practical applications still face numerous challenges. Calibration among heterogeneous vehicle and infrastructure endpoints in multi-end LiDAR systems is crucial for ensuring the accuracy and consistency of perception system data. Most existing multi-end calibration methods rely on initial calibration values provided by positioning systems, but the instability of GPS signals due to high buildings in urban canyons poses severe challenges to these methods. To address this issue, this paper proposes a novel multi-end LiDAR system calibration method that does not require positioning priors to determine initial external parameters and meets real-time requirements. Our method introduces an innovative multi-end perception object association technique, utilizing a new Overall Distance metric (oDist) to measure the spatial association between perception objects, and effectively combines global consistency search algorithms with optimal transport theory. By this means, we can extract co-observed targets from object association results for further external parameter computation and optimization. Extensive comparative and ablation experiments conducted on the simulated dataset V2X-Sim and the real dataset DAIR-V2X confirm the effectiveness and efficiency of our method. The code for this method can be accessed at: https://github.com/MassimoQu/v2i-calib.

cs.RO

Observation of Transient Trion Induced by Ultrafast Charge Transfer in Graphene/MoS2 Heterostructure

Van der Waals (Vdw) heterostructures constructed from TMDCs provide an ideal platform for exploring various quasiparticle behaviors, with trion-composed of neutral exciton and charged carrier-being a notable example. There are typically three methods to generate trion: electrical doping, chemical doping, and direct optical doping. The first two methods generate static trion, while the last gives rise to transient trion. Here, we present an indirect optical doping approach to generate transient trion via ultrafast charge transfer (CT) and achieve control over the trion-to-exciton ratio by adjusting CT in Gr/MoS2 heterostructure. Furthermore, we demonstrated that dynamics of the transient trion generated with this method, which shows slightly longer lifetime than that of exciton accounted for the Coulomb interactions between trion and charged defect. This study provides fresh perspectives on the construction of new quasiparticles, dynamical characterization and the control of the many-body interaction in two-dimensional structure.

cond-mat.mes-hall

Efficient simulation of inhomogeneously correlated systems using block interaction product states

The strength of DMRG lies in its treatment of identical sites that are energetically degenerate and spatially similar. However, this becomes a drawback when applied to quantum chemistry calculations for large systems, as entangled orbitals often span broad ranges in energy and space, with notably inhomogeneous interactions. In this study, we propose addressing strong intra-fragment and weak inter-fragment correlations separately using a multi-configurational block interaction product state (BIPS) framework. The strong correlation is captured in electronic states on fragments, considering entanglement between fragments and their environments. This method has been tested in various chemical systems and shows high accuracy and efficiency in addressing inhomogeneous effects in quantum chemistry.

quant-ph

MM Algorithms for Statistical Estimation in Quantile Regression

Quantile regression \parencite{Koenker1978} is a robust and practically useful way to efficiently model quantile varying correlation and predict varied response quantiles of interest. This article constructs and tests MM algorithms, which are simple to code and have been suggested superior to some other prominent quantile regression methods in nonregularized problems \parencite{Pietrosanu2017}, in an array of linear quantile regression settings. Simulation studies comparing MM to existing tested methods and applications to various real data sets have corroborated our algorithms' effectiveness.

stat.ME

Enhancing Scalability in Bayesian Nonparametric Factor Analysis of Spatiotemporal Data

This article introduces novel and practicable Bayesian factor analysis frameworks that are computationally feasible for moderate to large spatiotemporal data. Previous Bayesian analysis of spatiotemporal data has utilized a Bayesian factor model with separable temporal latent factors and spatial factor loadings, along with stick-breaking process priors on the loadings to enable clustering of spatial locations. Such a flexible Bayesian model, however, faces a prohibitively high computational cost in posterior sampling when the spatial and temporal dimensions increase to a couple hundred. We address this computational challenge with several speed-up proposals. We integrate a new slice sampling algorithm that permits varying numbers of spatial mixture components across all latent factors and guarantees them to be non-increasing through the posterior sampling iterations, thus effectively reducing the number of mixture parameters. Additionally, we introduce a spatial latent nearest-neighbor Gaussian process prior and new sequential updating algorithms for the spatially varying latent variables in the stick-breaking process prior. Our new models and sampling algorithms exhibit significantly enhanced computational scalability and storage efficiency and possess powerful inferential capabilities for both spatiotemporal prediction and clustering of spatial locations with similar temporal trajectories. The improvement in computational efficiency and inferential performance is substantiated by extensive simulation experiments.

stat.ME

Spectroscopic Evidence for Interfacial Charge Separation and Recombination in Graphene-MoS2 Vertical Heterostructures

Vertical van der Waals (vdW) heterostructures consisting of graphene (Gr) and transition metal dichalcogenides (TMDs) have created a fascinating platform for exploring optical and electronic properties in the two-dimensional limit. Previous study has revealed the ultrafast formation of interfacial excitons and the exciton dynamics in the Gr/MoS2 heterostructure. However, a fully understanding of interfacial charge separation and the subsequent dynamics in graphene-based heterostructures remains elusive. Here, we investigate the carrier dynamics of Gr-MoS2 (including Gr/MoS2 and MoS2/Gr stacking sequences) heterostructures under different photoexcitation energies and stacking sequences by comprehensive ultrafast means, including time-resolved terahertz spectroscopy (TRTS), terahertz emission spectroscopy (TES) and transient absorption spectroscopy (TAS). We demonstrate that the Gr/MoS2 heterostructure generates hot electron injection from graphene into the MoS2 layer with photoexcitation of sub-A-exciton of MoS2, while the interfacial charge separation in the MoS2/Gr could be partially blocked by the electric field of substrate. Charge transfer (CT) occurs in same directions for the Gr-MoS2 heterostructures with opposite stacking order, resulting in the opposite orientations of the interfacial photocurrent, as directly demonstrated by the terahertz (THz) emission. Moreover, we demonstrate that the recombination time of interfacial charges after CT is on a timescale of 18 ps to 1 ns, depending on the density of defect states in MoS2 layer. This work provides a comprehensive and unambiguous picture of the interfacial charge dynamics of graphene-based heterostructures, which is essential for developing Gr/TMDs based optoelectronic devices.

physics.optics

High-Power Near-Concentric Fabry-Perot Cavity for Phase Contrast Electron Microscopy

Transmission electron microscopy (TEM) of vitrified biological macromolecules (cryo-EM) is limited by the weak phase contrast signal that is available from such samples. Using a phase plate would thus substantially improve the signal-to-noise ratio. We have previously demonstrated the use of a high-power Fabry-Perot cavity as a phase plate for TEM. We now report improvements to our laser cavity that allow us to achieve record continuous-wave intensities of over 450 GW/cm$^{2}$, sufficient to produce the optimal 90° phase shift for 300 keV electrons. In addition, we have performed the first cryo-EM reconstruction using a laser phase plate, demonstrating that the stability of this laser phase plate is sufficient for use during standard cryo-EM data collection.

physics.ins-det