SearcharxivSearch

arXiv subjects

Shidong Li

Publications and source records attributed to Shidong Li.

16 recordsLinked to original sources

Fish Audio S2 Technical Report

We introduce Fish Audio S2, an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and, most importantly, instruction-following control via natural-language descriptions. To scale training, we develop a multi-stage training recipe together with a staged data pipeline covering video captioning and speech captioning, voice-quality assessment, and reward modeling. To push the frontier of open-source TTS, we release our model weights, fine-tuning code, and an SGLang-based inference engine. The inference engine is production-ready for streaming, achieving an RTF of 0.195 and a time-to-first-audio below 100 ms.Our code and weights are available on GitHub (https://github.com/fishaudio/fish-speech) and Hugging Face (https://huggingface.co/fishaudio/s2-pro). We highly encourage readers to visit https://fish.audio to try custom voices.

cs.SD

Performance analysis of tail-minimization and the linear rate of convergence of a proximal algorithm for sparse signal recovery

Recovery error bounds of tail-minimization and the rate of convergence of an efficient proximal alternating algorithm for sparse signal recovery are considered in this article. Tail-minimization focuses on minimizing the energy in the complement $T^c$ of an estimated support $T$. Under the restricted isometry property (RIP) condition, we prove that tail-$\ell_1$ minimization can exactly recover sparse signals in the noiseless case for a given $T$. In the noisy case, two recovery results for the tail-$\ell_1$ minimization and the tail-lasso models are established. Error bounds are improved over existing results. Additionally, we show that the RIP condition becomes surprisingly relaxed, allowing the RIP constant to approach $1$ as the estimation $T$ closely approximates the true support $S$. Finally, an efficient proximal alternating minimization algorithm is introduced for solving the tail-lasso problem using Hadamard product parametrization. The linear rate of convergence is established using the Kurdyka-{\L}ojasiewicz inequality. Numerical results demonstrate that the proposed algorithm significantly improves signal recovery performance compared to state-of-the-art techniques.

cs.IT

Dosimetric Evaluation of a New Rotating Gamma System for Stereotactic Radiosurgery

Purpose: A novel rotating gamma stereotactic radiosurgery (SRS) system (Galaxy RTi) with real-time image guidance technology has been developed for high-precision SRS and frameless fractionated stereotactic radiotherapy (SRT). This work investigated the dosimetric quality of Galaxy by comparing both the machine treatment parameters and plan dosimetry parameters with those of the widely used Leksell Gamma Knife (LGK) systems for SRS. Methods: The Galaxy RTi system uses 30 cobalt-60 sources on a rotating gantry to deliver non-coplanar, non-overlapping arcs simultaneously while the LGK 4C uses 201 static cobalt-60 sources to deliver noncoplanar beams. Ten brain cancer patients were unarchived from our clinical database, which were previously treated on the LGK 4C. The lesion volume for these cases varied from 0.1 cm3 to 15.4 cm3. Galaxy plans were generated using the Prowess TPS (Prowess, Concord, CA) with the same dose constraints and optimization parameters. Treatment quality metrics such as target coverage (%volume receiving the prescription dose), conformity index (CI), cone size, shots number, beam-on time were compared together with DVH curves and dose distributions. Results: Superior treatment plans were generated for the Galaxy system that met our clinical acceptance criteria. For the 10 patients investigated, the mean CI and dose coverage for Galaxy was 1.77 and 99.24 compared to 1.94 and 99.19 for LGK, respectively. The beam-on time for Galaxy was 17.42 minutes compared to 21.34 minutes for LGK (both assuming dose rates at the initial installation). The dose fall-off is much faster for Galaxy, compared with LGK. Conclusion: The Galaxy RTi system can provide dose distributions with similar quality to that of LGK with less beam-on time and faster dose fall-off. The system is also capable of real-time image guidance at treatment position to ensure accurate dose delivery for SRS.

physics.med-ph

Visualizing and Quantifying Wettability Alteration by Silica Nanofluids

An aqueous suspension of silica nanoparticles or nanofluid can alter the wettability of surfaces, specifically by making them hydrophilic and oil-repellent under water. Wettability alteration by nanofluids have important technological applications, including for enhanced oil recovery and heat transfer processes. A common way to characterize the wettability alteration is by measuring the contact angles of an oil droplet with and without nanoparticles. While easy to perform, contact angle measurements do not fully capture the wettability changes to the surface. Here, we employed several complementary techniques, such as cryo-scanning electron microscopy, confocal fluorescence and reflection interference contrast microscopy and droplet probe atomic force Microscopy (AFM), to visualize and quantify the wettability alterations by fumed silica nanoparticles. We found that nanoparticles adsorbed onto glass surfaces to form a porous layer with hierarchical micro- and nano-structures. The porous layer is able to trap a thin water film, which reduces contact between the oil droplet and the solid substrate. As a result, even a small addition of nanoparticles (0.1 wt%) lowers the adhesion force for a 20-$μ$m-sized oil droplet by more than 400 times from 210$\pm$10 nN to 0.5$\pm$0.3 nN as measured using droplet probe AFM. Finally, we show that silica nanofluids can improve oil recovery rates by 8% in a micromodel with glass channels that resemble a physical rock network.

physics.app-ph

Orthogonal subspace based fast iterative thresholding algorithms for joint sparsity recovery

Sparse signal recoveries from multiple measurement vectors (MMV) with joint sparsity property have many applications in signal, image, and video processing. The problem becomes much more involved when snapshots of the signal matrix are temporally correlated. With signal's temporal correlation in mind, we provide a framework of iterative MMV algorithms based on thresholding, functional feedback and null space tuning. Convergence analysis for exact recovery is established. Unlike most of iterative greedy algorithms that select indices in a measurement/solution space, we determine indices based on an orthogonal subspace spanned by the iterative sequence. In addition, a functional feedback that controls the amount of energy relocation from the "tails" is implemented and analyzed. It is seen that the principle of functional feedback is capable to lower the number of iteration and speed up the convergence of the algorithm. Numerical experiments demonstrate that the proposed algorithm has a clearly advantageous balance of efficiency, adaptivity and accuracy compared with other state-of-the-art algorithms.

cs.IT

Efficient iterative thresholding algorithms with functional feedbacks and convergence analysis

An accelerated class of adaptive scheme of iterative thresholding algorithms is studied analytically and empirically. They are based on the feedback mechanism of the null space tuning techniques (NST+HT+FB). The main contribution of this article is the accelerated convergence analysis and proofs with a variable/adaptive index selection and different feedback principles at each iteration. These convergence analysis require no longer a priori sparsity information $s$ of a signal. %key theory in this paper is the concept that the number of indices selected at each iteration should be considered in order to speed up the convergence. It is shown that uniform recovery of all $s$-sparse signals from given linear measurements can be achieved under reasonable (preconditioned) restricted isometry conditions. Accelerated convergence rate and improved convergence conditions are obtained by selecting an appropriate size of the index support per iteration. The theoretical findings are sufficiently demonstrated and confirmed by extensive numerical experiments. It is also observed that the proposed algorithms have a clearly advantageous balance of efficiency, adaptivity and accuracy compared with all other state-of-the-art greedy iterative algorithms.

cs.IT

Local sparsity and recovery of fusion frames structured signals

The problem of recovering signals of high complexity from low quality sensing devices is analyzed via a combination of tools from signal processing and harmonic analysis. By using the rich structure offered by the recent development in fusion frames, we introduce a compressed sensing framework in which we split the dense information into sub-channel or local pieces and then fuse the local estimations. Each piece of information is measured by potentially low quality sensors, modeled by linear matrices and recovered via compressed sensing -- when necessary. Finally, by a fusion process within the fusion frames, we are able to recover accurately the original signal. Using our new method, we show, and illustrate on simple numerical examples, that it is possible, and sometimes necessary, to split a signal via local projections and / or filtering for accurate, stable, and robust estimation. In particular, we show that by increasing the size of the fusion frame, a certain robustness to noise can also be achieved. While the computational complexity remains relatively low, we achieve stronger recovery performance compared to usual single-device compressed sensing systems.

cs.IT

The finite steps of convergence of the fast thresholding algorithms with feedbacks

Iterative algorithms based on thresholding, feedback and null space tuning (NST+HT+FB) for sparse signal recovery are exceedingly effective and fast, particularly for large scale problems. The core algorithm is shown to converge in finitely many steps under a (preconditioned) restricted isometry condition. In this paper, we present a new perspective to analyze the algorithm, which turns out that the efficiency of the algorithm can be further elaborated by an estimate of the number of iterations for the guaranteed convergence. The convergence condition of NST+HT+FB is also improved. Moreover, an adaptive scheme (AdptNST+HT+FB) without the knowledge of the sparsity level is proposed with its convergence guarantee. The number of iterations for the finite step of convergence of the AdptNST+HT+FB scheme is also derived. It is further shown that the number of iterations can be significantly reduced by exploiting the structure of the specific sparse signal or the random measurement matrix.

math.NA

The convergence guarantee of the iterative thresholding algorithm with suboptimal feedbacks for large systems

Thresholding based iterative algorithms have the trade-off between effectiveness and optimality. Some are effective but involving sub-matrix inversions in every step of iterations. For systems of large sizes, such algorithms can be computationally expensive and/or prohibitive. The null space tuning algorithm with hard thresholding and feedbacks (NST+HT+FB) has a mean to expedite its procedure by a suboptimal feedback, in which sub-matrix inversion is replaced by an eigenvalue-based approximation. The resulting suboptimal feedback scheme becomes exceedingly effective for large system recovery problems. An adaptive algorithm based on thresholding, suboptimal feedback and null space tuning (AdptNST+HT+subOptFB) without a prior knowledge of the sparsity level is also proposed and analyzed. Convergence analysis is the focus of this article. Numerical simulations are also carried out to demonstrate the superior efficiency of the algorithm compared with state-of-the-art iterative thresholding algorithms at the same level of recovery accuracy, particularly for large systems.

cs.IT

Spark Level Sparsity and the $\ell_1$ Tail Minimization

Solving compressed sensing problems relies on the properties of sparse signals. It is commonly assumed that the sparsity s needs to be less than one half of the spark of the sensing matrix A, and then the unique sparsest solution exists, and recoverable by $\ell_1$-minimization or related procedures. We discover, however, a measure theoretical uniqueness exists for nearly spark-level sparsity from compressed measurements Ax = b. Specifically, suppose A is of full spark with m rows, and suppose $\frac{m}{2}$ < s < m. Then the solution to Ax = b is unique for x with $\|x\|_0 \leq s$ up to a set of measure 0 in every s-sparse plane. This phenomenon is observed and confirmed by an $\ell_1$-tail minimization procedure, which recovers sparse signals uniquely with s > $\frac{m}{2}$ in thousands and thousands of random tests. We further show instead that the mere $\ell_1$-minimization would actually fail if s > $\frac{m}{2}$ even from the same measure theoretical point of view.

cs.IT

Tight and random nonorthogonal fusion frames

First we show that tight nonorthogonal fusion frames a relatively easy to com by. In order to do this we need to establish a classification of how to to wire a self adjoint operator as a product of (nonorthogonal) projection operators. We also discuss the link between nonorthogonal fusion frames and positive operator valued measures, we define and study a nonorthogonal fusion frame potential, and we introduce the idea of random nonorthogonal fusion frames.

math.FA

Fast thresholding algorithms with feedbacks for sparse signal recovery

We provide another framework of iterative algorithms based on thresholding, feedback and null space tuning for sparse signal recovery arising in sparse representations and compressed sensing. Several thresholding algorithms with various feedbacks are derived, which are seen as exceedingly effective and fast. Convergence results are also provided. The core algorithm is shown to converge in finite many steps under a (preconditioned) restricted isometry condition. The algorithms are seen as particularly effective for large scale problems. Numerical studies about the effectiveness and the speed of the algorithms are also presented.

cs.IT

Compressed Sensing with General Frames via Optimal-dual-based $\ell_1$-analysis

Compressed sensing with sparse frame representations is seen to have much greater range of practical applications than that with orthonormal bases. In such settings, one approach to recover the signal is known as $\ell_1$-analysis. We expand in this article the performance analysis of this approach by providing a weaker recovery condition than existing results in the literature. Our analysis is also broadly based on general frames and alternative dual frames (as analysis operators). As one application to such a general-dual-based approach and performance analysis, an optimal-dual-based technique is proposed to demonstrate the effectiveness of using alternative dual frames as analysis operators. An iterative algorithm is outlined for solving the optimal-dual-based $\ell_1$-analysis problem. The effectiveness of the proposed method and algorithm is demonstrated through several experiments.

cs.IT

Performance Analysis of $\ell_1$-synthesis with Coherent Frames

Signals with sparse frame representations comprise a much more realistic model of nature than that with orthonomal bases. Studies about the signal recovery associated with such sparsity models have been one of major focuses in compressed sensing. In such settings, one important and widely used signal recovery approach is known as $\ell_1$-synthesis (or Basis Pursuit). We present in this article a more effective performance analysis (than what are available) of this approach in which the dictionary $\Dbf$ may be highly, and even perfectly correlated. Under suitable conditions on the sensing matrix $\Phibf$, an error bound of the recovered signal $\hat{\fbf}$ (by the $\ell_1$-synthesis method) is established. Such an error bound is governed by the decaying property of $\tilde{\Dbf}_{\text{o}}^*\fbf$, where $\fbf$ is the true signal and $\tilde{\Dbf}_{\text{o}}$ denotes the optimal dual frame of $\Dbf$ in the sense that $\|\tilde{\Dbf}_{\text{o}}^*\hat{\fbf}\|_1$ produces the smallest $\|\tilde{\Dbf}^*\tilde{\fbf}\|_1$ in value among all dual frames $\tilde{\Dbf}$ of $\Dbf$ and all feasible signals $\tilde{\fbf}$. This new performance analysis departs from the usual description of the combo $\Phibf\Dbf$, and places the description on $\Phibf$. Examples are demonstrated to show that when the usual analysis fails to explain the working performance of the synthesis approach, the newly established results do.

cs.IT

Non-orthogonal fusion frames and the sparsity of fusion frame operators

Fusion frames have become a major tool in the implementation of distributed systems. The effectiveness of fusion frame applications in distributed systems is reflected in the efficiency of the end fusion process. This in turn is reflected in the efficiency of the inversion of the fusion frame operator $S_{\cW}$, which in turn is heavily dependent on the sparsity of $S_{\cW}$. We will show that sparsity of the fusion frame operator naturally exists by introducing a notion of {\it non-orthogonal fusion frames}. We show that for a fusion frame $\{W_i,v_i\}_{i\in I}$, if $\text{dim}(W_i)=k_i$, then the matrix of the non-orthogonal fusion frame operator $\cSw$ has in its corresponding location at most a $k_i\times k_i$ block matrix. We provide necessary and sufficient conditions for which the new fusion frame operator $\cSw$ is diagonal and/or a multiple of an identity. A set of other critical questions are also addressed. A scheme of {\it multiple fusion frames} whose corresponding fusion frame operator becomes an diagonal operator is also examined.

math.FA

Fusion Frames and Distributed Processing

Let $\{W_i\}_{i\in I}$ be a (redundant) sequence of subspaces each being endowed with a weight $v_i$, and let $\mathcal{H}$ be the closed linear span of the $W_i$'s, a composite Hilbert space. Provided that $\{(W_i,v_i)\}_{i \in I}$ satisfies a certain property which controls the weighted overlaps of the subspaces, it is called a {\em fusion frame}. These systems contain conventional frames as a special case, however they go far ``beyond frame theory''. In case each subspace $W_i$ is equipped with a frame system $\{f_{ij}\}_{j \in J_i}$ by which it is spanned, we refer to $\{(W_i,v_i,\{f_{ij}\}_{j \in J_i})\}_{i \in I}$ as a {\em fusion frame system}. In this paper, we describe a weighted and distributed processing procedure that fuse together information in all subspaces $W_i$ of a fusion frame system to obtain the global information in $\mathcal{H}$. The weighted and distributed processing technique described in fusion frames is not only a natural fit in distributed processing systems such as sensor networks, but also an efficient scheme for parallel processing of very large frame systems. We further provide an extensive study of the robustness of fusion frame systems.

math.FA