SearcharxivSearch

arXiv subjects

Mariano Rivera

Publications and source records attributed to Mariano Rivera.

13 recordsLinked to original sources

COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification

We introduce CoLoRA (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs). CoLoRA extends LoRA to convolutional layers by decomposing kernel updates into lightweight depthwise and pointwise components. This design reduces the number of trainable convolutional-update parameters by over 80\% compared with full convolutional fine-tuning, while allowing the learned updates to be merged into the pretrained convolutional kernels, thereby preserving the original model size and inference complexity. Experiments on MedMNIST datasets, particularly OCTMNISTv2, demonstrate that CoLoRA applied to VGG16 and ResNet50 achieves competitive classification performance while substantially reducing the number of trainable parameters. Comparisons with transfer learning, adapters, BitFit, and convolutional LoRA variants further characterize the trade-offs among predictive performance, trainable parameters, and training cost. Additional experiments on CIFAR-100 and Cats vs. Dogs provide preliminary evidence that the proposed adaptation strategy also transfers to non-medical image-classification tasks. Peak GPU-memory measurements further show that parameter efficiency does not translate directly into proportional training-memory savings, with memory consumption depending strongly on the placement of the adapted convolutional layers. Overall, CoLoRA provides a parameter-efficient and deployment-efficient alternative to full fine-tuning for convolutional models.

cs.CV

How to train your VAE

Variational Autoencoders (VAEs) have become a cornerstone in generative modeling and representation learning within machine learning. This paper explores a nuanced aspect of VAEs, focusing on interpreting the Kullback-Leibler (KL) Divergence, a critical component within the Evidence Lower Bound (ELBO) that governs the trade-off between reconstruction accuracy and regularization. Meanwhile, the KL Divergence enforces alignment between latent variable distributions and a prior imposing a structure on the overall latent space but leaves individual variable distributions unconstrained. The proposed method redefines the ELBO with a mixture of Gaussians for the posterior probability, introduces a regularization term to prevent variance collapse, and employs a PatchGAN discriminator to enhance texture realism. Implementation details involve ResNetV2 architectures for both the Encoder and Decoder. The experiments demonstrate the ability to generate realistic faces, offering a promising solution for enhancing VAE-based generative models.

cs.LG

Attentive VQ-VAE

We present a novel approach to enhance the capabilities of VQ-VAE models through the integration of a Residual Encoder and a Residual Pixel Attention layer, named Attentive Residual Encoder (AREN). The objective of our research is to improve the performance of VQ-VAE while maintaining practical parameter levels. The AREN encoder is designed to operate effectively at multiple levels, accommodating diverse architectural complexities. The key innovation is the integration of an inter-pixel auto-attention mechanism into the AREN encoder. This approach allows us to efficiently capture and utilize contextual information across latent vectors. Additionally, our models uses additional encoding levels to further enhance the model's representational power. Our attention layer employs a minimal parameter approach, ensuring that latent vectors are modified only when pertinent information from other pixels is available. Experimental results demonstrate that our proposed modifications lead to significant improvements in data representation and generation, making VQ-VAEs even more suitable for a wide range of applications as the presented.

cs.CV

EXTRACTER: Efficient Texture Matching with Attention and Gradient Enhancing for Large Scale Image Super Resolution

Recent Reference-Based image super-resolution (RefSR) has improved SOTA deep methods introducing attention mechanisms to enhance low-resolution images by transferring high-resolution textures from a reference high-resolution image. The main idea is to search for matches between patches using LR and Reference image pair in a feature space and merge them using deep architectures. However, existing methods lack the accurate search of textures. They divide images into as many patches as possible, resulting in inefficient memory usage, and cannot manage large images. Herein, we propose a deep search with a more efficient memory usage that reduces significantly the number of image patches and finds the $k$ most relevant texture match for each low-resolution patch over the high-resolution reference patches, resulting in an accurate texture match. We enhance the Super Resolution result adding gradient density information using a simple residual architecture showing competitive metrics results: PSNR and SSMI.

cs.CV

Hadamard Layer to Improve Semantic Segmentation

The Hadamard Layer, a simple and computationally efficient way to improve results in semantic segmentation tasks, is presented. This layer has no free parameters that require to be trained. Therefore it does not increase the number of model parameters, and the extra computational cost is marginal. Experimental results show that the new Hadamard layer substantially improves the performance of the investigated models (variants of the Pix2Pix model). The performance's improvement can be explained by the Hadamard layer forcing the network to produce an internal encoding of the classes so that all bins are active. Therefore, the network computation is more distributed. In a sort that the Hadamard layer requires that to change the predicted class, it is necessary to modify $2^{k-1}$ bins, assuming $k$ bins in the encoding. A specific loss function allows a stable and fast training convergence.

cs.CV

AxonNet: A self-supervised Deep Neural Network for Intravoxel Structure Estimation from DW-MRI

We present a method for estimating intravoxel parameters from a DW-MRI based on deep learning techniques. We show that neural networks (DNNs) have the potential to extract information from diffusion-weighted signals to reconstruct cerebral tracts. We present two DNN models: one that estimates the axonal structure in the form of a voxel and the other to calculate the structure of the central voxel using the voxel neighborhood. Our methods are based on a proposed parameter representation suitable for the problem. Since it is practically impossible to have real tagged data for any acquisition protocol, we used a self-supervised strategy. Experiments with synthetic data and real data show that our approach is competitive, and the computational times show that our approach is faster than the SOTA methods, even if training times are considered. This computational advantage increases if we consider the prediction of multiple images with the same acquisition protocol.

eess.IV

Deep neural network for fringe pattern filtering and normalisation

We propose a new framework for processing Fringe Patterns (FP). Our novel approach builds upon the hypothesis that the denoising and normalisation of FPs can be learned by a deep neural network if enough pairs of corrupted and ideal FPs are provided. The main contributions of this paper are the following: (1) We propose the use of the U-net neural network architecture for FP normalisation tasks; (2) we propose a modification for the distribution of weights in the U-net, called here the V-net model, which is more convenient for reconstruction tasks, and we conduct extensive experimental evidence in which the V-net produces high-quality results for FP filtering and normalisation. (3) We also propose two modifications of the V-net scheme, namely, a residual version called ResV-net and a fast operating version of the V-net, to evaluate the potential improvements when modify our proposal. We evaluate the performance of our methods in various scenarios: FPs corrupted with different degrees of noise, and corrupted with different noise distributions. We compare our methodology versus other state-of-the-art methods. The experimental results (on both synthetic and real data) demonstrate the capabilities and potential of this new paradigm for processing interferograms.

eess.IV

Robust Two-Step phase estimation using the Simplified Lissajous Ellipse Fitting method with Gabor Filter Banks preprocessing

We present the Simplified Lissajous Ellipse Fitting (SLEF) method for the calculation of the random phase step and the phase distribution from two phase-shifted interferograms. We consider interferograms with spatial and temporal dependency of background intensities, amplitude modulations and noise. Given these problems, the use of the Gabor Filters Bank (GFB) allows us to filter--out the noise, normalize the amplitude and eliminate the background. The normalized patterns permit to implement the SLEF algorithm, which is based on reducing the number of estimated coefficients of the ellipse equation, from five terms to only two. Our method consists of three stages. First, we preprocess the interferograms with GFB methodology in order to normalize the fringe patterns. Second, we calculate the phase step by using the proposed SLEF technique and third, we estimate the phase distribution using a two--steps formula. For the calculation of the phase step, we present two alternatives: the use of the Least Squares (LS) method to approximate the values of the coefficients and, in order to improve the LS estimation, a robust estimation based on the Leclerc's potential. The SLEF method's performance is evaluated through synthetic and experimental data to demonstrate its feasibility.

eess.IV

Two-Step Phase Shifting Algorithms: Where Are We?

Two steps phase shifting interferometry has been a hot topic in the recent years. We present a comparison study of 12 representative self--tunning algorithms based on two-steps phase shifting interferometry. We evaluate the performance of such algorithms by estimating the phase step of synthetic and experimental fringe patterns using 3 different normalizing processes: Gabor Filters Bank (GFB), Deep Neural Networks (DNNs) and Hilbert Huang Transform (HHT); in order to retrieve the background, the amplitude modulation and noise. We present the variants of state-of-the-art phase step estimation algorithms by using the GFB and DNNs as normalization preprocesses, as well as the use of a robust estimator such as the median to estimate the phase step. We present experimental results comparing the combinations of the normalization processes and the two steps phase shifting algorithms. Our study demonstrates that the quality of the retrieved phase from of two-step interferograms is more dependent of the normalizing process than the phase step estimation method.

eess.IV

Computation of the phase step between two-step fringe patterns based on Gram--Schmidt algorithm

We present the evaluation of a closed form formula for the calculation of the original step between two randomly shifted fringe patterns. Our proposal extends the Gram--Schmidt orthonormalization algorithm for fringe pattern. Experimentally, the phase shift is introduced by a electro--mechanical devices (such as piezoelectric or moving mounts).The estimation of the actual phase step allows us to improve the phase shifting device calibration. The evaluation consists of three cases that represent different pre-normalization processes: First, we evaluate the accuracy of the method in the orthonormalization process by estimating the test step using synthetic normalized fringe patterns with no background, constant amplitude and different noise levels. Second, we evaluate the formula with a variable amplitude function on the fringe patterns but with no background. Third, we evaluate non-normalized noisy fringe patterns including the comparison of pre-filtering processes such as the Gabor filter banks and the isotropic normalization process, in order to emphasize how they affect in the calculation of the phase step.

eess.IV

Half-quadratic transportation problems

We present a primal--dual memory efficient algorithm for solving a relaxed version of the general transportation problem. Our approach approximates the original cost function with a differentiable one that is solved as a sequence of weighted quadratic transportation problems. The new formulation allows us to solve differentiable, non-- convex transportation problems.

math.OC

Diffusion MRI microstructure models with in vivo human brain Connectom data: results from a multi-group comparison

A large number of mathematical models have been proposed to describe the measured signal in diffusion-weighted (DW) magnetic resonance imaging (MRI) and infer properties about the white matter microstructure. However, a head-to-head comparison of DW-MRI models is critically missing in the field. To address this deficiency, we organized the "White Matter Modeling Challenge" during the International Symposium on Biomedical Imaging (ISBI) 2015 conference. This competition aimed at identifying the DW-MRI models that best predict unseen DW data. in vivo DW-MRI data was acquired on the Connectom scanner at the A.A.Martinos Center (Massachusetts General Hospital) using gradients strength of up to 300 mT/m and a broad set of diffusion times. We focused on assessing the DW signal prediction in two regions: the genu in the corpus callosum, where the fibres are relatively straight and parallel, and the fornix, where the configuration of fibres is more complex. The challenge participants had access to three-quarters of the whole dataset, and their models were ranked on their ability to predict the remaining unseen quarter of data. In this paper we provide both an overview and a more in-depth description of each evaluated model, report the challenge results, and infer trends about the model characteristics that were associated with high model ranking. This work provides a much needed benchmark for DW-MRI models. The acquired data and model details for signal prediction evaluation are provided online to encourage a larger scale assessment of diffusion models in the future.

physics.med-ph

Two step robust fringe analysis method with random shift

We propose a two steps fringe analysis method assuming random phase step and changes in the illumination conditions. Our method constructs on a Gabor Filter--Bank (GFB) that independently estimates the phase from the fringe patterns and filters noise. As result of the GFB we obtain the two phase maps except by a random sign map. We show that such a random sign map is common to the independently computed phases and can be estimated from the residual between the phases. We estimate the final phase with a robust unwrapping procedure that interpolates unreliable phase regions. We present numerical experiments with synthetic and real date that demonstrate our method performance.

physics.optics