Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

Improving the James-Stein estimator via finite-sum truncation of its positive-part

For estimating the mean vector of a $p$-variate normal distribution ($p \ge 3$) under quadratic loss, the positive-part James--Stein estimator dominates the original James--Stein estimator, but it possesses a non-smooth thresholding boundary. In this paper, by truncating the infinite series representation of the positive-part function to a finite sum of degree $m$, we propose a new class of smooth shrinkage estimators. We prove that the proposed estimator dominates the James--Stein estimator for any dimension $p \ge 3$, provided the truncation degree satisfies $m \ge 0.95\sqrt{p-2}$.

math.ST↗

Weak, stable, and ordinary Lusternik-Schnirelmann category of finite spaces

The weak, stable, and ordinary Lusternik--Schnirelmann categories of a finite $T_0$-space $X$ satisfy $\operatorname{cat}_w(X)\leq \operatorname{cat}_s(X)\leq \operatorname{cat}(X)$. We give a general construction answering the simultaneous-strictness question of Cárdenas, Flores, Quintero, and Villar-Liñán. If $P$ is weakly contractible but noncontractible and the deletion of one point makes $P$ contractible, then adjoining $m\geq 2$ incomparable maximal points produces a connected finite space with category triple $(1,2,m)$. Moreover, after any positive number of barycentric subdivisions its ordinary category is equal to $2$. Applying the construction to a nine-point space yields examples on $m+9$ points and, in particular, a twelve-point example for which both inequalities are strict. Using additivity under disjoint unions, we also realize every triple $(a,b,c)$ with $1\leq a<b\leq 2a$ and $c\geq b$, as well as every triple $(a,a,c)$ with $2\leq a\leq c$. We conclude with questions concerning connected realizations and the minimum cardinality of a finite space exhibiting simultaneous strictness.

math.AT↗

Deep Truncated FBSDE Method: A Robust Solver for High-Dimensional Nonlinear PDEs and Fully Coupled FBSDEs

In this paper, we introduce a deep truncated forward-backward stochastic differential equation (FBSDE) method for high-dimensional partial differential equations (PDEs). Compared with existing deep-learning solvers for fully coupled FBSDEs, where strong coupling may lead to numerical instability, our approach exhibits improved stability. The proposed method combines gradient-truncated iterative decoupling with fictitious-play averaging to separate the forward and backward processes in a coupled framework. This preserves the coupled dynamics while reducing the unstable feedback induced by parameter-dependent forward paths during optimization. Furthermore, we incorporate a pathwise consistency term to create explicit local gradient shortcuts, thereby providing a structural mechanism that may mitigate gradient vanishing. We also derive a residual-based error estimate and establish conditional convergence of the fully discrete numerical approximations under suitable conditions, in which the pathwise consistency loss is not required. Our approach is particularly effective for convection-dominated equations, where the coupled formulation provides a stable representation of nonlinear transport without introducing singular terms into the BSDE. Numerical experiments demonstrate improved accuracy and stability in both low- and high-dimensional problems and robust performance for strongly coupled problems.

math.NA↗

TO-mdiSPAs: Topology Optimization of multi-directional Soft Pneumatic Actuators

Soft pneumatic actuators (SPAs) are highly promising and have been extensively explored within the field of soft robotics. Under pneumatic pressure loads, an SPA undergoes bending deformation to perform specific tasks. A multi- or omni-directional SPA can bend and move in any direction within a 3D space, leveraging its multiple degrees of freedom to realize various mechanical functions. This work presents a systematic methodology using topology optimization (TO) to achieve an optimized design for a multi-directional SPAs. To ensure manufacturing robustness, a three-field TO formulation considering blueprint, dilated, and eroded designs is implemented. Additionally, Darcy's law, incorporating a drainage term, is used to model the design-dependent nature of the pneumatic loading. A min-max optimization problem is formulated based on the target output deformations of the SPA unit and solved using the Method of Moving Asymptotes. The optimization yields a high-performing, unconventional geometric design. Finally, numerical simulations demonstrate that the optimized SPA successfully achieves versatile multi-directional movements.

cs.RO↗

Sub-quorum colorings of some infinite families of caterpillars

A partition $π=\{V_{1},V_{2},...,V_{k}\}$ of the vertex set $V$ of a graph $G$ into $k$ color classes $V_{i},$ with $i\in\{1,...,k\}$ is called a {\it quorum coloring} if for every vertex $v\in V,$ at least half of the vertices in the closed neighborhood $N[v]$ of $v$ have the same color as $v.$ The maximum cardinality of a quorum coloring of $G$ is called the {\it quorum coloring number} of $G$ and is denoted by $ψ_{q}(G).$ A {\it sub-quorum coloring} of $G$ is an onto partial function $f:V\rightarrow\left\{1,2,\ldots,\ell\right\}$ having the property that for every vertex $v\in V,$ if $f(v)$ is defined, then at least half of the vertices in $N[v]$ having an image by $f$, have the same color as $v.$ The {\it sub-quorum coloring number} $ψ_{sq}(G)$ equals the maximum value $\ell$ in a sub-quorum coloring of $G.$ In this paper, we determine the exact value of the sub-quorum coloring number for some infinite families of caterpillars including complete $n$-tuple caterpillars and complete caterpillars with minimum spine-vertex degree three.

math.CO↗

Type I estimates for the Kahler-Ricci flow I

We establish the Type I estimate for the scalar curvature of the Kahler-Ricci flow on compact Kahler manifolds, for arbitrary smooth initial metrics, whenever the flow develops finite time singularities. This estimate is a building block for the analytic minimal model program with Ricci flow.

math.DG↗

Weakly Measured Loops for Quantum Amplitude Amplification: Oracle Savings and Adaptive Search with Unknown Target Probability

We present a family of quantum amplitude amplification algorithms that interleave ordinary Grover rotations with tunable weak measurements in loops. A successful measurement yields the target state and stops the algorithm, while after a failure, the loop resumes from the post-measurement state without restarting. We express the leading costs in expected Grover iterations as $p \to 0$. For a known target probability $p$, an exact weak measurement-conditioned loop for states near the angle $π/4$ uses $(π/8+1/4+o(1))/\sqrt{p}$ iterations, improving on the standard Grover constant $π/4$ and on optimised restart Grover, and matching the corresponding infimum in our continuous-angle analysis. When only a lower bound $0 < p_0 \leq p < 1/2$ is available, a discrete Lyapunov equation gives the exact expected oracle cost for every fixed measurement strength and a closed-form optimal strength. Using the strength selected from $p_0$, the expected number of Grover iterations is at most $\frac{1}{2\sqrt{2}}(\frac{1}{\sqrt{p}}+\frac{1}{\sqrt{p_0}})$, where the expected cost decreases as the actual $p$ increases, and gives the leading constant $1/\sqrt{2}$ at the promise boundary $p = p_0$. For completely unknown $p$, we identify measurement strength $κ(t)=Θ(1/t)$ as the critical scale within the regular schedules considered here and analyse $κ_b(t)=\min\{1/2,b/t\}$. For each fixed $b > 2$, we rigorously derive a closed-form Gamma-function expression $C(b)$ giving $(C(b)/2+o(1))/\sqrt{p}$ expected iterations. Numerical minimisation of this explicit expression gives $b \approx 5.2$ and $C(b)/2 \approx 1.01$. The results show that weak measurements preserve the $Θ(1/\sqrt{p})$ search scale while adding an explicit fixed-point control mechanism and provable oracle savings.

quant-ph↗

The sharp radius in Korenblum's maximum principle for the Fock space

Korenblum's maximum principle states that if $|f| \le |g|$ near the boundary, then $\| f \| \le \| g \|$. The optimal size of the region of domination has been studied in Bergman spaces since 1991 and in Fock spaces since 2006, but only bounds have been obtained. We determine the optimal size for the Fock space $F^2$ of entire functions that are square integrable with respect to $e^{-|z|^2}\,dA(z)$: if $f, g\in F^2$ and $|f(z) \le |g(z)|$ for all $|z|>1$, then $\| f \| \le \| g \|$. The radius $1$ is optimal and we determine all pairs for which equality holds. To our knowledge, this is the first space of Bergman or Fock type in which the optimal radius in Korenblum's principle has been determined. The proof combines a Schwarz-Pick estimate centred at infinity with a moment-duality argument, and requires no numerical computation. The same method gives a moment criterion for weighted Fock-type spaces. For every non-increasing radial weight, the optimal radius equals the upper bound given by the pairs $f\equiv c$, $g(z)=z$. In particular, the upper bounds of Wee and Le for such weighted Fock spaces are sharp in the Hilbert space case. We also show that the optimal radius is $\sqrt{(β+1)/α}$ for the weight $|z|^{2β}e^{-α|z|^2}$, whenever $-1<β\le5.44$.

math.CV↗

Semantic Watermarking for Malicious Image Manipulation Detection

The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images while preserving visual plausibility, with concrete consequences for public discourse and vulnerable populations. We propose a robust semantic watermarking framework that reframes the watermark as a recoverable semantic reference rather than an opaque identifier. Our framework combines a $β$-VAE-based binary watermark (CLIP-VAE) with explicit channel-aware training---random bit-flip noise is injected during training so that the decoder learns graceful degradation under the noisy watermarking channel. As a downstream application, a lightweight module SDA-Net uses the recovered semantic embedding to expose not only whether but in which semantic direction an image has been altered. In a 5-way comparison against representative binary hashing baselines (SimHash, ITQ, HashNet, and their robust-MLP variants), CLIP-VAE achieves the highest reconstruction cosine similarity to the original CLIP embedding under realistic InstructPix2Pix bit-error rates, and uniquely supports direction-of-drift detection---a forensic complement to existing content-moderation pipelines.

cs.CV↗

ExpandDiff: Dynamic Range Expanding Diffusion for Single-Image HDR Reconstruction

Single-image HDR reconstruction requires inferring missing detail while preserving the visible content of an LDR image. Differences in sensor dynamic range and exposure cause LDR images to lose varying amounts of information in shadows and highlights. We present ExpandDiff, a conditional diffusion pipeline that jointly reconstructs clipped shadows and highlights. To account for this variation, we introduce Dynamic Clipping Synthesis (DCS), which randomly samples shadow and highlight clipping percentiles when constructing training inputs from HDR targets. A pixel-space diffusion model guided by spatially-adaptive normalization then predicts perceptually encoded HDR through a bounded output head, reconstructing both clipping directions in one sampling trajectory. On the SI-HDR benchmark, ExpandDiff variants improve HDR reconstruction accuracy by 3.43 dB in PU21-PSNR over the strongest evaluated competing method, and by 7.34 dB under two-sided clipping. The code and supplementary material are available at https://memreandiran.github.io/expanddiff/.

cs.CV↗

D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP~2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at https://jiahaozhang-public.github.io/d-scope/.

cs.CV↗

Parameterization method of reservoir properties for ensemble-based data assimilation using intermediate latent space of StyleGAN

Ensemble smoothers are the most successful and efficient techniques currently available for history matching. However, because these methods rely on Gaussian assumptions, their performance is severely degraded when the prior geology is described in terms of complex facies distributions (non-Gaussian). In this way, for these methods, we need to apply efficient parameterization techniques. Currently, the most efficient methods for performing parameterization are deep learning models. However, given the variety of existing deep learning models, studies have not identified which is most suitable for use with ensemble-based methods, although some important models had already been evaluated. Based on a recent literature review, the most promising models selected were VAE-GAN, Latent Diffusion, and StyleGAN models. As a novel aspect of this work, data assimilation with the second generation of StyleGAN (StyleGAN2) model was performed using the latent z-space and intermediate w-space, separately. They were applied in two 2D case studies: one categorical (three facies) and the other continuous. The results demonstrated that all three models are highly efficient, with the StyleGAN2 model standing out for generating samples with geological realism and achieving excellent data matching in the cases studied. Our findings show that performing data assimilation with StyleGAN2 using the intermediate space (w-space) yielded better results than the traditional application in the latent space (z-space). This is due to the fact that ESMDA uses linear updates and the w-space is much more linear and disentangled than the highly entangled z-space, thereby ensuring that the updated vectors remain close to realistic geological patterns. These results were validated using main geostatistical and history matching metrics.

cs.LG↗

Introduction to Computer Vision

This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organized as 44 short chapters across three parts, the book builds each topic from first principles: image arithmetic and morphology; convolution, pyramids, and frequency-domain filtering; feature detection, optical flow, and stereo; projective geometry, camera calibration, and structure from motion; and the full arc of modern deep learning, from a single neuron through convolutional networks, backpropagation, classic architectures, transfer learning, object detection, and semantic and instance segmentation, concluding with engineering considerations like mixed-precision and parallel training. Every technique is implemented directly in Python and NumPy or PyTorch and checked numerically against the corresponding OpenCV or PyTorch library function, so readers see not just the mathematics but its concrete behavior on real and synthetic data. The material was distilled with AI assistance from freely available online course notes, condensing extensive working code into concise mathematical exposition while preserving verified, reproducible results throughout. It is intended as a self-contained reference for students and practitioners who want to understand computer vision algorithms and their Python implementations.

cs.CV↗

MIND: Marginal-Invariant Neural Dependency Diffusion for Mixed-Type Tabular Generation

This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.

cs.LG↗

Certification-Based Differentially Private Learning

Differential privacy (DP) in machine learning is typically achieved by adding noise to model parameters (private learning) or to model outputs (private prediction). Recent work uses formal methods, namely abstract interpretation, to provide tighter privacy guarantees, but only for private prediction in classification settings. In this work, we investigate the use of formal methods as a general tool for tighter privacy analysis. First, we generalize the abstract gradient training (AGT) framework to private prediction in continuous, unbounded regression. Second, by reducing learning in parameterized models to a regression problem over the parameter space, we introduce Abstract Gradient Sampling (AGS), an algorithm that enables reachability-based analysis to provide guarantees for private learning. In both private prediction and private learning, we provide tightened privacy accounting for the AGT framework and a theoretical analysis demonstrating when our smooth sensitivity upper-bounds yield favourable privacy-utility trade-off. In practice, we validate that our regression bounds are tighter than global-sensitivity baselines on regression benchmarks, and, notably, yield the first finite privacy guarantees in settings where global prediction sensitivity is a priori unbounded. We also find that under matched conditions, our private learning algorithm can outperform standard private learners.

cs.LG↗

PEG-Tab: Sampling-Time Record Repair and Release Control for Tabular Synthesis

Pretrained tabular generators can reproduce training records even when aggregate utility remains high. When retraining is unavailable or too costly, sampling and release are the remaining intervention points. We present PEG-Tab (Post-Training Energy Guidance for Tabular Synthesis), a post-training repair and release-control framework for frozen tabular generators. For each generated row, a generator-native operator creates two alternatives. A shared calibrated score compares the three candidates, favours lower-risk records, and applies a final release check. We instantiate this interface for GReaT, CTGAN, TVAE, and TabDDPM without updating their parameters. Across five datasets and four generator families, PEG-Tab reduces mean Near Copy from $0.078$ to $0.027$ and lowers aggregate Exact Copy to zero. Relative to a $3\times$ post hoc filter, it retains higher utility in 12 of 16 transfer settings and Pareto-dominates the filter in eight. Gains are concentrated in copy and proximity-related risks.

cs.LG↗

Kirin: Cloud-native WebAssembly Service Orchestration

Modern cloud computing infrastructure relies heavily on virtualization to provide workload isolation and resource efficiency. While containers have become the dominant deployment primitive due to their fast orchestration compared to traditional virtual machines, they still introduce non-trivial overhead. WebAssembly (Wasm) has emerged as an alternative isolation technology, offering a lightweight execution model. Because of compatibility limitations, WebAssembly has mainly been applied to Function-as-a-Service (FaaS) and Edge Computing scenarios, where it has attracted growing research interest. However, recent advances in the WebAssembly ecosystem have significantly matured the technology, raising the question of whether it can serve as a viable replacement for containers in cloud-native workloads. In this paper, we present Kirin, an orchestrator for cloud-native WebAssembly services. It supports core orchestration responsibilities, including resource scheduling, lifecycle management, scaling, and request routing. We evaluate Kirin against container-based deployments for several service workloads and demonstrate promising results, particularly in service lifecycle management. Simultaneously, we confirm known limitations in CPU-intensive workloads. However, for typical service workloads, Kirin demonstrates competitive performance, suggesting that WebAssembly is a viable candidate for broader adoption as a cloud-native deployment technology.

cs.DC↗

Towards Better Exploration in Sequential Test-Time Scaling

Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.

cs.LG↗