SearcharxivSearch

arXiv subjects

Farzan Haddadi

Publications and source records attributed to Farzan Haddadi.

At least 19 recordsLinked to original sources

Denoising Diffusion Generative Models Secretly Calculate Attentions

Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.

cs.AI

Matrix Completion via Nonsmooth Regularization of Fully Connected Neural Networks

Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values. It has been shown that enhanced performance could be attained by using nonlinear estimators such as deep neural networks. Deep fully connected neural networks (FCNNs), one of the most suitable architectures for matrix completion, suffer from over-fitting due to their high capacity, which leads to low generalizability. In this paper, we control over-fitting by regularizing the FCNN model in terms of the $\ell_{1}$ norm of intermediate representations and nuclear norm of weight matrices. As such, the resulting regularized objective function becomes nonsmooth and nonconvex, i.e., existing gradient-based methods cannot be applied to our model. We propose a variant of the proximal gradient method and investigate its convergence to a critical point. In the initial epochs of FCNN training, the regularization terms are ignored, and through epochs, the effect of that increases. The gradual addition of nonsmooth regularization terms is the main reason for the better performance of the deep neural network with nonsmooth regularization terms (DNN-NSR) algorithm. Our simulations indicate the superiority of the proposed algorithm in comparison with existing linear and nonlinear algorithms.

cs.IT

L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework

State-space models (SSMs) have recently emerged as efficient alternatives to computationally intensive architectures such as Transformers for sequence modeling. However, their training typically relies on static loss functions, which may be suboptimal at different stages of learning. In this work, we introduce a hybrid model that integrates the Hyena architecture with a Dynamic Loss Network (DLN) under a Learning-to-Teach (L2T) paradigm, referred to as L2T-DLN. In this framework, the Hyena model serves as a student whose loss function is adapted online, while a teacher model, equipped with a memory of the student's past performance, guides the DLN to dynamically trade off the primary cross-entropy objective and a regularization term. We evaluate the proposed L2T-Hyena model on the Penn Treebank (PTB) and WikiText-103 language modeling benchmarks and compare it against both a vanilla Hyena SSM and a Transformer baseline. On PTB, our model achieves a validation perplexity of 102.6, representing a substantial improvement over the 110.5 obtained by the vanilla Hyena trained with a static loss function and 121.28 achieved by the Transformer baseline. Similar gains are observed on WikiText-103, where L2T-Hyena reaches a validation perplexity of 68.3, outperforming vanilla Hyena (73.7) and Transformer (89.8). These results indicate that coupling SSMs with adaptive loss functions can significantly enhance both the quality and efficiency of deep learning models for sequential data and hold strong promise for applications in natural language processing, time-series analysis, and biological signal processing.

cs.IT

Hybrid Quantum-Classical Selective State Space Artificial Intelligence

Hybrid Quantum Classical (HQC) algorithms constitute one of the most effective paradigms for exploiting the computational advantages of quantum systems in large-scale numerical tasks. By operating in high-dimensional Hilbert spaces, quantum circuits enable exponential speed-ups and provide access to richer representations of cost landscapes compared to purely classical methods. These capabilities are particularly relevant for machine learning, where state-of-the-art models especially in Natural Language Processing (NLP) suffer from prohibitive time complexity due to massive matrix multiplications and high-dimensional optimization. In this manuscript, we propose a Hybrid Quantum Classical selection mechanism for the Mamba architecture, designed specifically for temporal sequence classification problems. Our approach leverages Variational Quantum Circuits (VQCs) as quantum gating modules that both enhance feature extraction and improve suppression of irrelevant information. This integration directly addresses the computational bottlenecks of deep learning architectures by exploiting quantum resources for more efficient representation learning. We analyze how introducing quantum subroutines into large language models (LLMs) impacts their generalization capability, expressivity, and parameter efficiency. The results highlight the potential of quantum-enhanced gating mechanisms as a path toward scalable, resource-efficient NLP models, in a limited simulation step. Within the first four epochs on a reshaped MNIST dataset with input format (batch, 784, d_model), our hybrid model achieved 24.6% accuracy while using one quantum layer and achieve higher expressivity, compared to 21.6% obtained by a purely classical selection mechanism. we state No founding

quant-ph

Multi-weight Nuclear Norm Minimization for Low-rank Matrix Recovery in Presence of Subspace Prior Information

Weighted nuclear norm minimization has been recently recognized as a technique for reconstruction of a low-rank matrix from compressively sampled measurements when some prior information about the column and row subspaces of the matrix is available. In this work, we study the recovery conditions and the associated recovery guarantees of weighted nuclear norm minimization when multiple weights are allowed. This setup might be used when one has access to prior subspaces forming multiple angles with the column and row subspaces of the ground-truth matrix. While existing works in this field use a single weight to penalize all the angles, we propose a multi-weight problem which is designed to penalize each angle independently using a distinct weight. Specifically, we prove that our proposed multi-weight problem is stable and robust under weaker conditions for the measurement operator than the analogous conditions for single-weight scenario and standard nuclear norm minimization. Moreover, it provides better reconstruction error than the state of the art methods. We illustrate our results with extensive numerical experiments that demonstrate the advantages of allowing multiple weights in the recovery procedure.

cs.IT

Off-the-grid Recovery of Time and Frequency Shifts with Multiple Measurement Vectors

We address the problem of estimating time and frequency shifts of a known waveform in the presence of multiple measurement vectors (MMVs). This problem naturally arises in radar imaging and wireless communications. Specifically, a signal ensemble is observed, where each signal of the ensemble is formed by a superposition of a small number of scaled, time-delayed, and frequency shifted versions of a known waveform sharing the same continuous-valued time and frequency components. The goal is to recover the continuous-valued time-frequency pairs from a small number of observations. In this work, we propose a semidefinite programming which exactly recovers $s$ pairs of time-frequency shifts from $L$ regularly spaced samples per measurement vector under a minimum separation condition between the time-frequency shifts. Moreover, we prove that the number $s$ of time-frequency shifts scales linearly with the number $L$ of samples up to a log-factor. Extensive numerical results are also provided to validate the effectiveness of the proposed method over the single measurement vectors (SMVs) problem. In particular, we find that our approach leads to a relaxed minimum separation condition and reduced number of required samples.

cs.IT

Blind Two-Dimensional Super Resolution in Multiple Input Single Output Linear Systems

In this paper, we consider a multiple-input single-output (MISO) linear time-varying system whose output is a superposition of scaled and time-frequency shifted versions of inputs. The goal of this paper is to determine system characteristics and input signals from the single output signal. More precisely, we want to recover the continuous time-frequency shift pairs, the corresponding (complex-valued) amplitudes and the input signals from only one output vector. This problem arises in a variety of applications such as radar imaging, microscopy, channel estimation and localization problems. While this problem is naturally ill-posed, by constraining the unknown input waveforms to lie in separate known low-dimensional subspaces, it becomes tractable. More explicitly, we propose a semidefinite program which exactly recovers time-frequency shift pairs and input signals. We prove uniqueness and optimality of the solution to this program. Moreover, we provide a grid-based approach which can significantly reduce computational complexity in exchange for adding a small gridding error. Numerical results confirm the ability of our proposed method to exactly recover the unknowns.

cs.IT

Adaptive Recovery of Dictionary-sparse Signals using Binary Measurements

One-bit compressive sensing (CS) is an advanced version of sparse recovery in which the sparse signal of interest can be recovered from extremely quantized measurements. Namely, only the sign of each measurement is available to us. In many applications, the ground-truth signal is not sparse itself, but can be represented in a redundant dictionary. A strong line of research has addressed conventional CS in this signal model including its extension to one-bit measurements. However, one-bit CS suffers from the extremely large number of required measurements to achieve a predefined reconstruction error level. A common alternative to resolve this issue is to exploit adaptive schemes. Adaptive sampling acts on the acquired samples to trace the signal in an efficient way. In this work, we utilize an adaptive sampling strategy to recover dictionary-sparse signals from binary measurements. For this task, a multi-dimensional threshold is proposed to incorporate the previous signal estimates into the current sampling procedure. This strategy substantially reduces the required number of measurements for exact recovery. Our proof approach is based on the recent tools in high dimensional geometry in particular random hyperplane tessellation and Gaussian width. We show through rigorous and numerical analysis that the proposed algorithm considerably outperforms state of the art approaches. Further, our algorithm reaches an exponential error decay in terms of the number of quantized measurements.

cs.IT

Living near the edge: A lower-bound on the phase transition of total variation minimization

This work is about the total variation (TV) minimization which is used for recovering gradient-sparse signals from compressed measurements. Recent studies indicate that TV minimization exhibits a phase transition behavior from failure to success as the number of measurements increases. In fact, in large dimensions, TV minimization succeeds in recovering the gradient-sparse signal with high probability when the number of measurements exceeds a certain threshold; otherwise, it fails almost certainly. Obtaining a closed-form expression that approximates this threshold is a major challenge in this field and has not been appropriately addressed yet. In this work, we derive a tight lower-bound on this threshold in case of any random measurement matrix whose null space is distributed uniformly with respect to the Haar measure. In contrast to the conventional TV phase transition results that depend on the simple gradient-sparsity level, our bound is highly affected by generalized notions of gradient-sparsity. Our proposed bound is very close to the true phase transition of TV minimization confirmed by simulation results.

cs.IT

A Greedy Algorithm for Matrix Recovery with Subspace Prior Information

Matrix recovery is the problem of recovering a low-rank matrix from a few linear measurements. Recently, this problem has gained a lot of attention as it is employed in many applications such as Netflix prize problem, seismic data interpolation and collaborative filtering. In these applications, one might access to additional prior information about the column and row spaces of the matrix. These extra information can potentially enhance the matrix recovery performance. In this paper, we propose an efficient greedy algorithm that exploits prior information in the recovery procedure. The performance of the proposed algorithm is measured in terms of the rank restricted isometry property (R-RIP). Our proposed algorithm with prior subspace information converges under a more milder condition on the R-RIP in compared with the case that we do not use prior information. Additionally, our algorithm performs much better than nuclear norm minimization in terms of both computational complexity and success rate.

cs.IT

Exploiting Prior Information in Block Sparse Signals

We study the problem of recovering a block-sparse signal from under-sampled observations. The non-zero values of such signals appear in few blocks, and their recovery is often accomplished using a $\ell_{1,2}$ optimization problem. In applications such as DNA micro-arrays, some prior information about the block support, i.e., blocks containing non-zero elements, is available. A typical way to consider the extra information in recovery procedures is to solve a weighted $\ell_{1,2}$ problem. In this paper, we consider a block sparse model, where the block support has intersection with some given subsets of blocks with known probabilities. Our goal in this work is to minimize the number of required linear Gaussian measurements for perfect recovery of the signal by tuning the weights of a weighted $\ell_{1,2}$ problem. For this goal, we apply tools from conic integral geometry and derive closed-form expressions for the optimal weights. We show through precise analysis and simulations that the weighted $\ell_{1,2}$ problem with optimal weights significantly outperforms the regular $\ell_{1,2}$ problem. We further examine the sensitivity of the optimal weights to the mismatch of block probabilities, and conclude stability under small probability deviations.

cs.IT

On the Error in Phase Transition Computations for Compressed Sensing

Evaluating the statistical dimension is a common tool to determine the asymptotic phase transition in compressed sensing problems with Gaussian ensemble. Unfortunately, the exact evaluation of the statistical dimension is very difficult and it has become standard to replace it with an upper-bound. To ensure that this technique is suitable, [1] has introduced an upper-bound on the gap between the statistical dimension and its approximation. In this work, we first show that the error bound in [1] in some low-dimensional models such as total variation and $\ell_1$ analysis minimization becomes poorly large. Next, we develop a new error bound which significantly improves the estimation gap compared to [1]. In particular, unlike the bound in [1] that is not applicable to settings with overcomplete dictionaries, our bound exhibits a decaying behavior in such cases.

cs.IT

Optimal Weighted Low-rank Matrix Recovery with Subspace Prior Information

Matrix sensing is the problem of reconstructing a low-rank matrix from a few linear measurements. In many applications such as collaborative filtering, the famous Netflix prize problem, and seismic data interpolation, there exists some prior information about the column and row spaces of the ground-truth low-rank matrix. In this paper, we exploit this prior information by proposing a weighted optimization problem where its objective function promotes both rank and prior subspace information. Using the recent results in conic integral geometry, we obtain the unique optimal weights that minimize the required number of measurements. As simulation results confirm, the proposed convex program with optimal weights requires substantially fewer measurements than the regular nuclear norm minimization.

cs.IT

One Bit Spectrum Sensing in Cognitive Radio Sensor Networks

This paper proposes a spectrum sensing algorithm from one bit measurements in a cognitive radio sensor network. A likelihood ratio test (LRT) for the one bit spectrum sensing problem is derived. Different from the one bit spectrum sensing research work in the literature, the signal is assumed to be a discrete random correlated Gaussian process, where the correlation is only available within immediate successive samples of the received signal. The employed model facilitates the design of a powerful detection criteria with measurable analytical performance. One bit spectrum sensing criterion is derived for one sensor which is then generalized to multiple sensors. Performance of the detector is analyzed by obtaining closed-form formulas for the probability of false alarm and the probability of detection. Simulation results corroborate the theoretical findings and confirm the efficacy of the proposed detector in the context of highly correlated signals and large number of sensors.

eess.SP

Distribution-aware Block-sparse Recovery via Convex Optimization

We study the problem of reconstructing a block-sparse signal from compressively sampled measurements. In certain applications, in addition to the inherent block-sparse structure of the signal, some prior information about the block support, i.e. blocks containing non-zero elements, might be available. Although many block-sparse recovery algorithms have been investigated in Bayesian framework, it is still unclear how to incorporate the information about the probability of occurrence into regularization-based block-sparse recovery in an optimal sense. In this work, we bridge between these fields by the aid of a new concept in conic integral geometry. Specifically, we solve a weighted optimization problem when the prior distribution about the block support is available. Moreover, we obtain the unique weights that minimize the expected required number of measurements. Our simulations on both synthetic and real data confirm that these weights considerably decrease the required sample complexity.

cs.IT

Eigenvectors of Deformed Wigner Random Matrices

We investigate eigenvectors of rank-one deformations of random matrices $\boldsymbol B = \boldsymbol A + θ\boldsymbol {uu}^*$ in which $\boldsymbol A \in \mathbb R^{N \times N}$ is a Wigner real symmetric random matrix, $θ\in \mathbb R^+$, and $\boldsymbol u$ is uniformly distributed on the unit sphere. It is well known that for $θ> 1$ the eigenvector associated with the largest eigenvalue of $\boldsymbol B$ closely estimates $\boldsymbol u$ asymptotically, while for $θ< 1$ the eigenvectors of $\boldsymbol B$ are uninformative about $\boldsymbol u$. We examine $\mathcal O(\frac{1}{N})$ correlation of eigenvectors with $\boldsymbol u$ before phase transition and show that eigenvectors with larger eigenvalue exhibit stronger alignment with deforming vector through an explicit inverse law. This distribution function will be shown to be the ordinary generating function of Chebyshev polynomials of second kind. These polynomials form an orthogonal set with respect to the semicircle weighting function. This law is an increasing function in the support of semicircle law for eigenvalues $(-2\: ,+2)$. Therefore, most of energy of the unknown deforming vector is concentrated in a $cN$-dimensional ($c<1$) known subspace of $\boldsymbol B$. We use a combinatorial approach to prove the result.

math.ST

Improved Recovery of Analysis Sparse Vectors in Presence of Prior Information

In this work, we consider the problem of recovering analysis-sparse signals from under-sampled measurements when some prior information about the support is available. We incorporate such information in the recovery stage by suitably tuning the weights in a weighted $\ell_1$ analysis optimization problem. Indeed, we try to set the weights such that the method succeeds with minimum number of measurements. For this purpose, we exploit the upper-bound on the statistical dimension of a certain cone to determine the weights. Our numerical simulations confirm that the introduced method with tuned weights outperforms the standard $\ell_1$ analysis technique.

cs.IT

Strong Interference Alignment

Interference alignment (IA) adjusts signaling scheme such that all interfering signals are squeezed in interference subspace. IA mostly achieves its performance via infinite extension of the channel, which is a major challenge for IA in practical systems. In this paper, we make part of interference very strong and achieve perfect IA within limited number of channel extensions. A single-hop $3$ user single antenna interference channel (IFC) is considered and it is shown that only one of the interfering signal streams needs to be strong so that perfect IA is feasible.

cs.IT