Searcharxiv⌕ Search

arXiv subjects

Jürgen Frikel

Publications and source records attributed to Jürgen Frikel.

18 recordsLinked to original sources

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis

Document type classification in visually rich documents remains challenging, as relevant information is distributed across textual, visual, and layout modalities. To capture this complexity, current approaches rely on diverse multimodal modeling strategies, resulting in heterogeneous architectures that complicate systematic comparison. This variability is also reflected in existing comparative studies, which often rely on heterogeneous evaluation setups, further complicating systematic comparison and making it difficult to assess progress. To address these limitations, this work provides a structured analysis of multimodal design strategies across transformer- and LLM-based architectures, combined with a controlled empirical comparison within a unified experimental framework. Specifically, four representative models (LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B) are evaluated on the RVL-CDIP benchmark to systematically analyze the contributions of text, image, and layout information for document type classification, with a particular focus on contrasting OCR-dependent and OCR-free approaches. The results show that specialized multimodal Transformers outperform LLM-based approaches on visually rich and layout-intensive documents. Image information contributes most strongly to reliable classification, while OCR-derived text provides useful but secondary support. These findings highlight that multimodal processing remains essential for documents with pronounced layout structure. Overall, the study provides a systematic basis for comparing multimodal architectures and offers practical guidance for selecting effective feature combinations and model designs for document type classification.

cs.CV↗

Data-proximal null-space networks for inverse problems

Inverse problems are inherently ill-posed and therefore require regularization techniques to achieve a stable solution. While traditional variational methods have well-established theoretical foundations, recent advances in machine learning based approaches have shown remarkable practical performance. However, the theoretical foundations of learning-based methods in the context of regularization are still underexplored. In this paper, we propose a general framework that addresses the current gap between learning-based methods and regularization strategies. In particular, our approach emphasizes the crucial role of data consistency in the solution of inverse problems and introduces the concept of data-proximal null-space networks as a key component for their solution. We provide a complete convergence analysis by extending the concept of regularizing null-space networks with data proximity in the visual part. We present numerical results for limited-view computed tomography to illustrate the validity of our framework.

math.NA↗

Translation invariant diagonal frame decomposition for the Radon transform

In this article, we address the challenge of solving the ill-posed reconstruction problem in computed tomography using a translation invariant diagonal frame decomposition (TI-DFD). First, we review the concept of a TI-DFD for general linear operators and the corresponding filter-based regularization. We then introduce the TI-DFD for the Radon transform on $L^2(\R^2)$ and provide an exemplary construction using the TI wavelet transform. Presented numerical results clearly demonstrate the benefits of our approach over non-translation invariant counterparts.

math.NA↗

Data-proximal complementary $\ell^1$-TV reconstruction for limited data CT

In a number of tomographic applications, data cannot be fully acquired, resulting in a severely underdetermined image reconstruction. In such cases, conventional methods lead to reconstructions with significant artifacts. To overcome these artifacts, regularization methods are applied that incorporate additional information. An important example is TV reconstruction, which is known to be efficient at compensating for missing data and reducing reconstruction artifacts. At the same time, however, tomographic data is also contaminated by noise, which poses an additional challenge. The use of a single regularizer must therefore account for both the missing data and the noise. However, a particular regularizer may not be ideal for both tasks. For example, the TV regularizer is a poor choice for noise reduction across multiple scales, in which case $\ell^1$ curvelet regularization methods are well suited. To address this issue, in this paper we introduce a novel variational regularization framework that combines the advantages of different regularizers. The basic idea of our framework is to perform reconstruction in two stages, where the first stage mainly aims at accurate reconstruction in the presence of noise, and the second stage aims at artifact reduction. Both reconstruction stages are connected by a data proximity condition. The proposed method is implemented and tested for limited-view CT using a combined curvelet-TV approach. We define and implement a curvelet transform adapted to the limited-view problem and illustrate the advantages of our approach in numerical experiments.

math.NA↗

Translation invariant diagonal frame decomposition of inverse problems and their regularization

Solving inverse problems is central to a variety of important applications, such as biomedical image reconstruction and non-destructive testing. These problems are characterized by the sensitivity of direct solution methods with respect to data perturbations. To stabilize the reconstruction process, regularization methods have to be employed. Well-known regularization methods are based on frame expansions, such as the wavelet-vaguelette (WVD) decomposition, which are well adapted to the underlying signal class and the forward model and furthermore allow efficient implementation. However, it is well known that the lack of translational invariance of wavelets and related systems leads to specific artifacts in the reconstruction. To overcome this problem, in this paper we introduce and analyze the translation invariant diagonal frame decomposition (TI-DFD) of linear operators as a novel concept generalizing the SVD. We characterize ill-posedness via the TI-DFD and prove that a TI-DFD combined with a regularizing filter leads to a convergent regularization method with optimal convergence rates. As illustrative example, we construct a wavelet-based TI-DFD for one-dimensional integration, where we also investigate our approach numerically. The results indicate that filtered TI-DFDs eliminate the typical wavelet artifacts when using standard wavelets and provide a fast, accurate, and stable solution scheme for inverse problems.

math.NA↗

Regularization of Inverse Problems by Filtered Diagonal Frame Decomposition

The characteristic feature of inverse problems is their instability with respect to data perturbations. In order to stabilize the inversion process, regularization methods have to be developed and applied. In this work we introduce and analyze the concept of filtered diagonal frame decomposition which extends the standard filtered singular value decomposition to the frame case. Frames as generalized singular system allows to better adapt to a given class of potential solutions. In this paper, we show that filtered diagonal frame decomposition yield a convergent regularization method. Moreover, we derive convergence rates under source type conditions and prove order optimality under the assumption that the considered frame is a Riesz-basis.

math.NA↗

Feature reconstruction from incomplete tomographic data without detour

In this paper, we consider the problem of feature reconstruction from incomplete x-ray CT data. Such problems occurs, e.g., as a result of dose reduction in the context medical imaging. Since image reconstruction from incomplete data is a severely ill-posed problem, the reconstructed images may suffer from characteristic artefacts or missing features, and significantly complicate subsequent image processing tasks (e.g., edge detection or segmentation). In this paper, we introduce a novel framework for the robust reconstruction of convolutional image features directly from CT data, without the need of computing a reconstruction firs. Within our framework we use non-linear (variational) regularization methods that can be adapted to a variety of feature reconstruction tasks and to several limited data situations . In our numerical experiments, we consider several instances of edge reconstructions from angularly undersampled data and show that our approach is able to reliably reconstruct feature maps in this case.

eess.IV↗

Combining reconstruction and edge detection in computed tomography

We present two methods that combine image reconstruction and edge detection in computed tomography (CT) scans. Our first method is as an extension of the prominent filtered backprojection algorithm. In our second method we employ $\ell^{1}$-regularization for stable calculation of the gradient. As opposed to the first method, we show that this approach is able to compensate for undersampled CT data.

math.NA↗

A new 3D model for magnetic particle imaging using realistic magnetic field topologies for algebraic reconstruction

We derive a new 3D model for magnetic particle imaging (MPI) that is able to incorporate realistic magnetic fields in the reconstruction process. In real MPI scanners, the generated magnetic fields have distortions that lead to deformed magnetic low-field volumes (LFV) with the shapes of ellipsoids or bananas instead of ideal field-free points (FFP) or lines (FFL), respectively. Most of the common model-based reconstruction schemes in MPI use however the idealized assumption of an ideal FFP or FFL topology and, thus, generate artifacts in the reconstruction. Our model-based approach is able to deal with these distortions and can generally be applied to dynamic magnetic fields that are approximately parallel to their velocity field. We show how this new 3D model can be discretized and inverted algebraically in order to recover the magnetic particle concentration. To model and describe the magnetic fields, we use decompositions of the fields in spherical harmonics. We complement the description of the new model with several simulations and experiments.

math.NA↗

Sparse regularization of inverse problems by operator-adapted frame thresholding

We analyze sparse frame based regularization of inverse problems by means of a diagonal frame decomposition (DFD) for the forward operator, which generalizes the SVD. The DFD allows to define a non-iterative (direct) operator-adapted frame thresholding approach which we show to provide a convergent regularization method with linear convergence rates. These results will be compared to the well-known analysis and synthesis variants of sparse $\ell^1$-regularization which are usually implemented thorough iterative schemes. If the frame is a basis (non-redundant case), the three versions of sparse regularization, namely synthesis and analysis variants of $\ell^1$ regularization as well as the DFD thresholding are equivalent. However, in the redundant case, those three approaches are pairwise different.

math.NA↗

Mathematical Analysis of the 1D Model and Reconstruction Schemes for Magnetic Particle Imaging

Magnetic particle imaging (MPI) is a promising new in-vivo medical imaging modality in which distributions of super-paramagnetic nanoparticles are tracked based on their response in an applied magnetic field. In this paper we provide a mathematical analysis of the modeled MPI operator in the univariate situation. We provide a Hilbert space setup, in which the MPI operator is decomposed into simple building blocks and in which these building blocks are analyzed with respect to their mathematical properties. In turn, we obtain an analysis of the MPI forward operator and, in particular, of its ill-posedness properties. We further get that the singular values of the MPI core operator decrease exponentially. We complement our analytic results by some numerical studies which, in particular, suggest a rapid decay of the singular values of the MPI operator.

math.NA↗

Efficient regularization with wavelet sparsity constraints in PAT

In this paper we consider the reconstruction problem of photoacoustic tomography (PAT) with a flat observation surface. We develop a direct reconstruction method that employs regularization with wavelet sparsity constraints. To that end, we derive a wavelet-vaguelette decomposition (WVD) for the PAT forward operator and a corresponding explicit reconstruction formula in the case of exact data. In the case of noisy data, we combine the WVD reconstruction formula with soft-thresholding which yields a spatially adaptive estimation method. We demonstrate that our method is statistically optimal for white random noise if the unknown function is assumed to lie in any Besov-ball. We present generalizations of this approach and, in particular, we discuss the combination of vaguelette soft-thresholding with a TV prior. We also provide an efficient implementation of the vaguelette transform that leads to fast image reconstruction algorithms supported by numerical results.

math.OC↗

Limited data problems for the generalized Radon transform in $\mathbb{R}^n$

We consider the generalized Radon transform (defined in terms of smooth weight functions) on hyperplanes in $\mathbb{R}^n$. We analyze general filtered backprojection type reconstruction methods for limited data with filters given by general pseudodifferential operators. We provide microlocal characterizations of visible and added singularities in $\mathbb{R}^n$ and define modified versions of reconstruction operators that do not generate added artifacts. We calculate the symbol of our general reconstruction operators as pseudodifferential operators, and provide conditions for the filters under which the reconstruction operators are elliptic for the visible singularities. If the filters are chosen according to those conditions, we show that almost all visible singularities can be recovered reliably. Our work generalizes the results for the classical line transforms in $\mathbb{R}^2$ and the classical reconstruction operators (that use specific filters). In our proofs, we employ a general paradigm that is based on the calculus of Fourier integral operators. Since this technique does not rely on explicit expressions of the reconstruction operators, it enables us to analyze more general imaging situations.

math.AP↗

On Artifacts in Limited Data Spherical Radon Transform: Curved Observation Surface

In this article, we consider the limited data problem for spherical mean transform. We characterize the generation and strength of the artifacts in a reconstruction formula. In contrast to the third's author work [Ngu15b], the observation surface considered in this article is not flat. Our results are comparable to those obtained in [Ngu15b] for flat observation surface. For the two dimensional problem, we show that the artifacts are $k$ orders smoother than the original singularities, where $k$ is vanishing order of the smoothing function. Moreover, if the original singularity is conormal, then the artifacts are $k+\frac{1}{2}$ order smoother than the original singularity. We provide some numerical examples and discuss how the smoothing effects the artifacts visually. For three dimensional case, although the result is similar to that [Ngu15b], the proof is significantly different. We introduce a new idea of lifting the space.

math.AP↗

Joint Image Reconstruction and Segmentation Using the Potts Model

We propose a new algorithmic approach to the non-smooth and non-convex Potts problem (also called piecewise-constant Mumford-Shah problem) for inverse imaging problems. We derive a suitable splitting into specific subproblems that can all be solved efficiently. Our method does not require a priori knowledge on the gray levels nor on the number of segments of the reconstruction. Further, it avoids anisotropic artifacts such as geometric staircasing. We demonstrate the suitability of our method for joint image reconstruction and segmentation. We focus on Radon data, where we in particular consider limited data situations. For instance, our method is able to recover all segments of the Shepp-Logan phantom from $7$ angular views only. We illustrate the practical applicability on a real PET dataset. As further applications, we consider spherical Radon data as well as blurred data.

math.OC↗

Artifacts in incomplete data tomography - with applications to photoacoustic tomography and sonar

We develop a paradigm using microlocal analysis that allows one to characterize the visible and added singularities in a broad range of incomplete data tomography problems. We give precise characterizations for photo- and thermoacoustic tomography and Sonar, and provide artifact reduction strategies. In particular, our theorems show that it is better to arrange Sonar detectors so that the boundary of the set of detectors does not have corners and is smooth. To illustrate our results, we provide reconstructions from synthetic spherical mean data as well as from experimental photoacoustic data.

math.AP↗

A paradigm for the characterization of artifacts in tomography

We present a paradigm for characterization of artifacts in limited data tomography problems. In particular, we use this paradigm to characterize artifacts that are generated in reconstructions from limited angle data with generalized Radon transforms and general filtered backprojection type operators. In order to find when visible singularities are imaged, we calculate the symbol of our reconstruction operator as a pseudodifferential operator.

math.AP↗

Sparse regularization in limited angle tomography

We investigate the reconstruction problem of limited angle tomography. Such problems arise naturally in applications like digital breast tomosynthesis, dental tomography, electron microscopy etc. Since the acquired tomographic data is highly incomplete, the reconstruction problem is severely ill-posed and the traditional reconstruction methods, such as filtered backprojection (FBP), do not perform well in such situations. To stabilize the reconstruction procedure additional prior knowledge about the unknown object has to be integrated into the reconstruction process. In this work, we propose the use of the sparse regularization technique in combination with curvelets. We argue that this technique gives rise to an edge-preserving reconstruction. Moreover, we show that the dimension of the problem can be significantly reduced in the curvelet domain. To this end, we give a characterization of the kernel of limited angle Radon transform in terms of curvelets and derive a characterization of solutions obtained through curvelet sparse regularization. In numerical experiments, we will present the practical relevance of these results.

math.NA↗