SearcharxivSearch

arXiv subjects

Zhifeng Deng

Publications and source records attributed to Zhifeng Deng.

5 recordsLinked to original sources

Diffeomorphic Logarithm of Special Orthogonal Matrices

The special orthogonal group $\mathbb{SO}_n$ is a Lie group whose geometry and local structure are encoded by the exponential map in its Lie algebra $\mathbf{Skew}_n$, the set of skew-symmetric matrices. The associated multi-valued inverse problem -- the matrix logarithm -- in $\mathbb{SO}_n$ exhibits a highly nontrivial local diffeomorphism structure, which differs from the matrix logarithm for invertible matrices. This work characterizes the local diffeomorphism structure of the exponential in the set of skew-symmetric matrices where its derivative is invertible. We show that this set with an invertible derivative can be organized into diffeomorphic regions, using a canonical alignment of Schur decompositions. In particular, the region that contains the principal logarithm has a special multiplicity structure: each matrix in $\mathbb{SO}_n$ admits at most two skew-symmetric preimages in this region. Based on this geometric framework, we introduce the diffeomorphic logarithm of special orthogonal matrices together with an efficient and stable algorithm. Moreover, it is applied to the Karcher mean problem in $\mathbb{SO}_n$, demonstrating continuous behavior of the mean under perturbations of the data, which is not captured by the principal logarithm.

math.DG

TTS-1 Technical Report

We introduce Inworld TTS-1, a set of two Transformer-based autoregressive text-to-speech (TTS) models. Our largest model, TTS-1-Max, has 8.8B parameters and is designed for utmost quality and expressiveness in demanding applications. TTS-1 is our most efficient model, with 1.6B parameters, built for real-time speech synthesis and on-device use cases. By scaling train-time compute and applying a sequential process of pre-training, fine-tuning, and RL-alignment of the speech-language model (SpeechLM) component, both models achieve state-of-the-art performance on a variety of benchmarks, demonstrating exceptional quality relying purely on in-context learning of the speaker's voice. Inworld TTS-1 and TTS-1-Max can generate high-resolution 48 kHz speech with low latency, and support 11 languages with fine-grained emotional control and non-verbal vocalizations through audio markups. We additionally open-source our training and modeling code under an MIT license.

cs.CL

The Exponential of Skew-Symmetric Matrices: A Nearby Inverse and Efficient Computation of Derivatives

The matrix exponential restricted to skew-symmetric matrices has numerous applications, notably in view of its interpretation as the Lie group exponential and Riemannian exponential for the special orthogonal group. We characterize the invertibility of the derivative of the skew-restricted exponential, thereby providing a simple expression of the tangent conjugate locus of the orthogonal group. In view of the skew restriction, this characterization differs from the classic result on the invertibility of the derivative of the exponential of real matrices. Based on this characterization, for every skew-symmetric matrix $A$ outside the (zero-measure) tangent conjugate locus, we explicitly construct the domain and image of a smooth inverse -- which we term \emph{nearby logarithm} -- of the skew-restricted exponential around $A$. This nearby logarithm reduces to the classic principal logarithm of special orthogonal matrices when $A=\mathbf{0}$. The symbolic formulae for the differentiation and its inverse are derived and implemented efficiently. The extensive numerical experiments show that the proposed formulae are up to $3.9$-times and $3.6$-times faster than the current state-of-the-art robust formulae for the differentiation and its inversion, respectively.

math.DG

Moment switching in nanotube magnetic force probes

A recent advance in improving the spatial resolution of magnetic force microscopy (MFM) uses as sensor tips carbon nanotubes grown at the apex of conventional silicon cantilever pyramids and coated with a thin ferromagnetic layer. Magnetic images of high density vertically recorded media using these tips exhibit a doubling of the spatial frequency under some conditions. Here we demonstrate that this spatial frequency doubling is due to the switching of the moment direction of the nanotube tip. This results in a signal which is proportional to the absolute value of the signal normally observed in MFM. Our modeling indicates that a significant fraction of the tip volume is involved in the observed switching, and that it should be possible to image very high bit densities with nanotube magnetic force sensors.

cond-mat.mes-hall