Searcharxiv⌕ Search

arXiv subjects

Shintaro Yoshizawa

Publications and source records attributed to Shintaro Yoshizawa.

8 recordsLinked to original sources

A Homotopical Geometry of the Multinomial and Negative-Multinomial Potentials

We study the one-parameter family of Hessian metrics $g_α=αh_++(1-α)h_-$, $α\in[0,1]$, formed by convexly combining the log-partition functions (not the densities) of the multinomial and negative multinomial Fisher information Hessians. Four results are proved in full. (1) $g_α\succ0$ for all $α\in[0,1]$; its Gaussian curvature ($n=2$) is $-\tfrac14$ at $α=0$ and $+\tfrac14$ at $α=1$, but non-constant for intermediate $α$, changing sign at a closed-form critical value $α^\ast(w)$, with an analytic uniqueness proof. (2) Viewing $g_α$ as the spatial metric of a Kasner-type volume equation, positive-definiteness forces one conserved quantity $β\le0$, while a second, independent quantity $C=4\det K$ governs the equation's branch (bounded oscillatory, parabolic, or unbounded growth) and is not fixed by $β$ or positive-definiteness alone; we show this discriminant dictionary is a theorem about the abstract matrix equation, to which $g_α$ supplies initial data of either sign of $C$ without the trajectories remaining on $\{g_α(θ)\}$ itself. (3) $Ψ_α$ is identified as the log-partition function of a compound distribution: a negative-binomial latent count $M$ (parameter $r=1-α$), augmented by a Bernoulli parity bit and split multinomially, with non-negative base measure for every $α$. (4) The associated Bregman divergence obeys a generalised Pythagorean theorem whose projection onto the symmetric locus is the arithmetic mean, independent of $α$; the endpoints' curvature sign matches the Gauss--Bonnet excess of genuine geodesic triangles. Numerical checks are kept separate from analytic proofs throughout.

math.DG↗

Halley's Method for Rectangular Matrix Variables and the Matrix Schwarzian Derivative

Alefeld (1981) recast Halley's cubic convergence as Newton's method on $g=f/\sqrt{f'}$, linked to the Schwarzian derivative. We generalize this to matrix gradient fields $F=\nablaϕ$ ($X\in\R^{m\times n}$) via a matrix Schwarzian derivative interpreted through information geometry ($α$-connections). Simplifying Palmore (1994), we construct this operator square-root-free derivative from the third Fréchet derivative of the Newton map. Our main theorem proves local cubic convergence with an explicit error constant, without self-adjointness, commutativity, or gradient-field assumptions. A matrix-free algorithm (Hessian-vector products only) validates the theory. We contrast this with a power-Newton family (whose naive matrix extension fails) and the matrix Laguerre family (which requires an operator square root). A case study on the Oja-type system $\dot X=AXB-XBX^TAX$ reveals that convergence depends on target eigenvalue gaps, cubic gains grow with ill-conditioning, and 70-digit tests confirm exact theoretical orders alongside a working-precision accuracy budget. Finally, we note coupled formulations excel primarily when sequential deflation is inapplicable.

math.NA↗

A Generalized Ridge Regression and Convolutional LASSO

We derive the complete duality theory underlying the Hodrick--Prescott filter, whose rank-deficient second-difference penalty admits infinitely many equivalent trend representations via generalized inverses. Constructing two canonical choices---the Moore--Penrose-based \emph{B-representation} and an alternative \emph{A-representation}---we prove that the extracted trend is invariant across representations, obtain a closed-form Bregman-type divergence quantifying their disagreement under a shared coefficient vector, and show this divergence vanishes as $λ\to\infty$. We further mollify the non-smooth $\ell_1$ trend filter with a compactly supported biweight kernel to obtain a closed-form $C^3$ \emph{Convolutional LASSO} that restores Newton-type quadratic convergence without sacrificing the kink- setecting character of $\ell_1$ regularization. Theoretical results are verified numerically and benchmarked against ADMM/IRLS on NVIDIA's 2013--2018 daily closing prices, where the sparse filter isolates genuine growth-regime breaks---chiefly the November 2016 post-earnings acceleration---cleanly separated from the smooth $L_2$ trend.

stat.ME↗

Information Geometry of Gradient Flows

Taking classical information geometry as its point of departure, this paper investigates, through gradient flows, how dually flat geometry extends beyond regular convexity, non-degeneracy, and smoothness. The regular theory is developed from the log-determinant potential on positive definite Gram matrices, establishing its Legendre dual, Fisher--Rao metric, Bregman divergence, and generalized Pythagorean theorem. We connect this framework to Craig--Sakamoto deformation, Wolfe duality, and, via Yoshizawa's embedding, Brockett--Bloch--Ratiu double-bracket flows, linking isospectral dynamics, Stiefel optimization, and component learning. The Bures--Wasserstein geometry provides a complementary gradient-flow structure. The singular theory emerges from boundary behavior: difference-of-convex deformations produce indefinite or degenerate Hessians while retaining pseudo-Hessian, dually flat, Legendre-self-dual structures. Newton flows exhibit finite-time collapse or Łojasiewicz-controlled convergence near non-Morse critical sets. Fisher-metric degeneracies on the Birkhoff polytope and elliptic-curve moduli are resolved by explicit blow-ups, yielding a birationally invariant exponential decay law. We further derive a closed-form Kirillov Jacobian and introduce cross curvature as a spectral diagnostic of local escape rates, including a new Box--Cox interpolation. Reproducible numerical experiments support the closed-form results. Rather than claiming a completed theory, the paper provides foundations for singular information geometry centered on degenerate pencils, indefinite dual flatness, blow-up geometry, and Łojasiewicz-type convergence.

math.DS↗

A Power Geometry and Hypergeometric Functions

This paper is a corrected English translation of an original Japanese work introducing "power geometry," a research programme centred on power transformations, deformations and dualities of functions in the sense of optimisation theory. Its two main motivations are the origin of probability distributions (via maximum-likelihood principles, rational-function parameters and ODEs linked to information geometry, the Pearson system and Schwarzian equations) and the construction of matrix dynamical systems that compute eigenvalues, singular values and principal or minor subspaces. Building on Brockett's double-bracket flows, the author extends these systems to rectangular matrices, flag manifolds and difference-of-convex potentials, and shows how a continuous power parameter can interpolate between largest- and smallest-eigenvalue problems. Dual potentials, topological distinctions between principal and minor subspace flows, and links via matrix Schwarzian--Riccati equations are examined, with a view toward design principles for classical and quantum information-processing algorithms.

math.DS↗

Spatial Process Mining

We propose a new framework that focuses on on-site entities in the digital twin, a pairing of the real world and digital space. Characteristics include active sensing to generate event logs, spatial and temporal partitioning of complex processes, and visualization and analysis of processes that can be scaled in space and time. As a specific example, a cell production system is composed of connected manufacturing spaces called cells in a manufacturing process. A cell is sensed by ceiling cameras to generate a Gantt chart that provides a bird's-eye view of the process according to the cycle of events that occur in the cell. This Gantt chart is easy to understand for experienced operators, but we also propose a method for finding the focus of causes of deviations from the usual process without special experience or knowledge. This method captures the characteristics of the processes occurring in a cell by using our own event node ranking algorithm, a modification of HITS (Hypertext Induced Topic Selection), which scores web pages against a complex network generated from a process model.

stat.AP↗

CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps

This paper describes a multi-modal data association method for global localization using object-based maps and camera images. In global localization, or relocalization, using object-based maps, existing methods typically resort to matching all possible combinations of detected objects and landmarks with the same object category, followed by inlier extraction using RANSAC or brute-force search. This approach becomes infeasible as the number of landmarks increases due to the exponential growth of correspondence candidates. In this paper, we propose labeling landmarks with natural language descriptions and extracting correspondences based on conceptual similarity with image observations using a Vision Language Model (VLM). By leveraging detailed text information, our approach efficiently extracts correspondences compared to methods using only object categories. Through experiments, we demonstrate that the proposed method enables more accurate global localization with fewer iterations compared to baseline methods, exhibiting its efficiency.

cs.CV↗

Dynamical systems for eigenvalue problems of axisymmetric matrices with positive eigenvalues

We consider the eigenvalues and eigenvectors of an axisymmetric matrix$A$ with some special structures. We propose S-Oja-Brockett equation $\frac{dX}{dt}=AXB-XBX^TSAX,$ where $X(t) \in {\mathbb R}^{n \times m}$ with $m \leq n$, $S$ is a positive definite symmetric solution of the Sylvester equation $A^TS = SA$ and $B$ is a real positive definite diagonal matrix whose diagonal elements are distinct each other, and show the S-Oja-Brockett equation has the global convergence to eigenvalues and its eigenvectors of $A$.

math.DS↗