SearcharxivSearch

arXiv subjects

Dinesh Singh

Publications and source records attributed to Dinesh Singh.

At least 19 recordsLinked to original sources

STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition

Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements. Existing methods rely heavily on apex-onset frames, overlook fine-grained inter-frame dynamics, and separately model spatial and temporal information, limiting generalization across datasets. To address these challenges, we propose STAG, a dynamic ROI-AU-coupled spatial-temporal network that jointly models motion flow and adaptive facial connectivity. The framework extracts optical flow from discriminative frames using magnitude-based selection and temporal attention. A dual-branch architecture combines an enhanced graph attention network for structured spatial reasoning with a transformer encoder for temporal modeling. A bidirectional cross-attention module enables mutual refinement of spatial and temporal features, while AU-guided dynamic connectivity adapts facial region interactions according to muscle activation patterns. The transformer captures subtle temporal dynamics beyond apex-based approaches, improving semantic consistency and interpretability for explainable micro-expression recognition. The fused representation is optimized using focal loss and evaluated on CASME II, 4DME, DFME, NaME, SAMM, and SMIC-HS. Extensive experiments demonstrate improved robustness, generalization, interpretability, and computational efficiency, confirming the effectiveness of adaptive relational reasoning, AU-guided dynamic connectivity, and deep spatial-temporal feature fusion for accurate cross-dataset micro-expression recognition.

cs.CV

A Gleason-Kahane-\.Zelazko Theorem for $H^p_{\alpha, \beta}$ spaces

We study two entities that have proved to be of interest and importance in their own right viz. the $H^p_{\alpha, \beta}$ spaces and the classical Gleason-Kahane-\.Zelazko (GKZ) theorem. We establish a GKZ-type theorem on the $H^p_{\alpha, \beta}$ spaces by identifying a natural class of functions serving as the counterpart of invertible elements, and proving that every continuous linear functional nonvanishing on this class is a point evaluation. As an application, we characterize weighted composition operators on these spaces. It must be noted that the $H^p_{\alpha, \beta}$ spaces do not possess the various structural advantages of the classical Hardy spaces where the GKZ theorem already exists and thus we have had to modify and establish methods that rely on suitable extensions and refinements of the known techniques. Additionally, we provide examples that illustrate the natural class of functions arising in our GKZ-type theorem and demonstrate the sharpness of our characterization of weighted composition operators under the assumption of surjectivity.

math.FA

Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated remarkable performance in the face generation task. However, they fail to effectively capture cross-modality semantic information in craniofacial reconstruction when translating from the skull (x-ray) to the face (optical) domain, due to a mismatch in the alignment of structural identity across modalities. To address this issue, we propose Cranio-Diff, a diffusion-based framework for cross-domain cranio-facial reconstruction from 2D X-ray skull images. The proposed approach integrates skull-conditioned structural guidance through ControlNet with biometric text conditioning to generate a face which is more semantically and structurally aligned with the given skull. The proposed Cranio-diff method is evaluated on skull-face dataset obtained from X-ray scans of 120 subjects in lateral and frontal views. To enable controlled evaluation, each face image is synthesised across three age groups (25, 45, 65) and three BMI variations of -10%, baseline and +10%, yielding 4320 paired samples. To the best of our knowledge, this is the only X-ray-face dataset with this magnitude. Extensive experiments showed that the proposed method outperforms recent existing approaches in both generated image quality and retrieval task. Finally, to evaluate the performance of our proposed method, we have evaluated the quality of the generated image using FID, IS, SSIM, LPIPS, PSNR and ArcFace score. Additionally, retrieval performance is evaluated using recall@k, mAP@k and MRR@k. Obtained experimental results demonstrate that the proposed method can be used as an alternate tool in providing aid in forensic investigations.

cs.CV

Nearly invariant subspaces in the real Hardy space

The objective of this article is to study nearly invariant subspaces of the backward shift operator on the real Hardy space. We also investigate nearly invariant subspaces with finite defect, and as a consequence, provide a characterization for almost invariant subspaces of the backward shift operator.

math.FA

Doubly twisted near-isometries: Classification and a Wold-type decomposition

We introduce and study doubly twisted near-isometries. A doubly twisted near-isometry is a tuple of near-isometries satisfying certain relations determined by a prescribed family of unitaries, thereby generalizing the notion of doubly commuting near-isometries. We establish necessary and sufficient conditions for a tuple of near-isometries to admit a Wold-type decomposition and prove that the existence of such a decomposition automatically ensures its uniqueness by providing an explicit description of the summands. Furthermore, we show that every doubly twisted near-isometry admits a Wold-type decomposition. We also characterize unitary equivalence within the class of doubly twisted near-isometries and construct an analytic model for them. Several examples are included to highlight the distinctions between our results and the corresponding results in the setting of doubly twisted isometries.

math.FA

Perturbations of Toeplitz operators on vector-valued Hardy spaces

In this article, we completely classify invariant subspaces of finite-rank perturbations of a class of Toeplitz operators on vector-valued Hardy spaces. As a consequence, in the vector-valued setting, we characterize invariant and almost invariant subspaces of a class of Toeplitz operators, as well as nearly invariant subspaces associated with certain Blaschke-based operators. We further treat the finite defect case for these nearly invariant subspaces.

math.FA

SPOT-Face: Forensic Face Identification using Attention Guided Optimal Transport

Person identification in forensic investigations becomes very challenging when common identification means for DNA (i.e., hair strands, soft tissue) are not available. Current methods utilize deep learning methods for face recognition. However, these methods lack effective mechanisms to model cross-domain structural correspondence between two different forensic modalities. In this paper, we introduce a SPOT-Face, a superpixel graph-based framework designed for cross-domain forensic face identification of victims using their skeleton and sketch images. Our unified framework involves constructing a superpixel-based graph from an image and then using different graph neural networks(GNNs) backbones to extract the embeddings of these graphs, while cross-domain correspondence is established through attention-guided optimal transport mechanism. We have evaluated our proposed framework on two publicly available dataset: IIT\_Mandi\_S2F (S2F) and CUFS. Extensive experiments were conducted to evaluate our proposed framework. The experimental results show significant improvement in identification metrics ( i.e., Recall, mAP) over existing graph-based baselines. Furthermore, our framework demonstrates to be highly effective for matching skulls and sketches to faces in forensic investigations.

cs.CV

Cranio-ID: Graph-Based Craniofacial Identification via Automatic Landmark Annotation in 2D Multi-View X-rays

In forensic craniofacial identification and in many biomedical applications, craniometric landmarks are important. Traditional methods for locating landmarks are time-consuming and require specialized knowledge and expertise. Current methods utilize superimposition and deep learning-based methods that employ automatic annotation of landmarks. However, these methods are not reliable due to insufficient large-scale validation studies. In this paper, we proposed a novel framework Cranio-ID: First, an automatic annotation of landmarks on 2D skulls (which are X-ray scans of faces) with their respective optical images using our trained YOLO-pose models. Second, cross-modal matching by formulating these landmarks into graph representations and then finding semantic correspondence between graphs of these two modalities using cross-attention and optimal transport framework. Our proposed framework is validated on the S2F and CUHK datasets (CUHK dataset resembles with S2F dataset). Extensive experiments have been conducted to evaluate the performance of our proposed framework, which demonstrates significant improvements in both reliability and accuracy, as well as its effectiveness in cross-domain skull-to-face and sketch-to-face matching in forensic science.

cs.CV

FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction

Craniofacial reconstruction in forensics is one of the processes to identify victims of crime and natural disasters. Identifying an individual from their remains plays a crucial role when all other identification methods fail. Traditional methods for this task, such as clay-based craniofacial reconstruction, require expert domain knowledge and are a time-consuming process. At the same time, other probabilistic generative models like the statistical shape model or the Basel face model fail to capture the skull and face cross-domain attributes. Looking at these limitations, we propose a generic framework for craniofacial reconstruction from 2D X-ray images. Here, we used various generative models (i.e., CycleGANs, cGANs, etc) and fine-tune the generator and discriminator parts to generate more realistic images in two distinct domains, which are the skull and face of an individual. This is the first time where 2D X-rays are being used as a representation of the skull by generative models for craniofacial reconstruction. We have evaluated the quality of generated faces using FID, IS, and SSIM scores. Finally, we have proposed a retrieval framework where the query is the generated face image and the gallery is the database of real faces. By experimental results, we have found that these generative models can be used as an assisting tool for craniofacial identifications in forensic science.

cs.CV

Exp-Graph: How Connections Learn Facial Attributes in Graph-based Expression Recognition

Facial expression recognition is crucial for human-computer interaction applications such as face animation, video surveillance, affective computing, medical analysis, etc. Since the structure of facial attributes varies with facial expressions, incorporating structural information into facial attributes is essential for facial expression recognition. In this paper, we propose Exp-Graph, a novel framework designed to represent the structural relationships among facial attributes using graph-based modeling for facial expression recognition. For facial attributes graph representation, facial landmarks are used as the graph's vertices. At the same time, the edges are determined based on the proximity of the facial landmark and the similarity of the local appearance of the facial attributes encoded using the vision transformer. Additionally, graph convolutional networks are utilized to capture and integrate these structural dependencies into the encoding of facial attributes, thereby enhancing the accuracy of expression recognition. Thus, Exp-Graph learns from the facial attribute graphs highly expressive semantic representations. On the other hand, the vision transformer and graph convolutional blocks help the framework exploit the local and global dependencies among the facial attributes that are essential for the recognition of facial expressions. We conducted comprehensive evaluations of the proposed Exp-Graph model on three benchmark datasets: Oulu-CASIA, eNTERFACE05, and AFEW. The model achieved recognition accuracies of 98.09\%, 79.01\%, and 56.39\%, respectively. These results indicate that Exp-Graph maintains strong generalization capabilities across both controlled laboratory settings and real-world, unconstrained environments, underscoring its effectiveness for practical facial expression recognition applications.

cs.CV

Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet

Craniofacial reconstruction in forensic science is crucial for the identification of the victims of crimes and disasters. The objective is to map a given skull to its corresponding face in a corpus of faces with known identities using recent advancements in computer vision, such as deep learning. In this paper, we presented a framework for the identification of a person given the X-ray image of a skull using convolutional Siamese networks for cross-domain identity representation. Siamese networks are twin networks that share the same architecture and can be trained to discover a feature space where nearby observations that are similar are grouped and dissimilar observations are moved apart. To do this, the network is exposed to two sets of comparable and different data. The Euclidean distance is then minimized between similar pairs and maximized between dissimilar ones. Since getting pairs of skull and face images are difficult, we prepared our own dataset of 40 volunteers whose front and side skull X-ray images and optical face images were collected. Experiments were conducted on the collected cross-domain dataset to train and validate the Siamese networks. The experimental results provide satisfactory results on the identification of a person from the given skull.

cs.CV

A several variables Kowalski-S\lodkowski theorem for topological spaces

In this paper, we provide a version of the classical result of Kowalski and Słodkowski that generalizes the famous Gleason-Kahane-$\dot{\rm Z}$elazko (GKZ) theorem by characterizing multiplicative linear functionals amongst all complex-valued functions on a Banach algebra. We first characterize maps on $\mathcal{A}$-valued polynomials of several variables that satisfy some conditions, motivated by the result of Kowalski and Słodkowski, as a composition of a multiplicative linear functional on $\mathcal{A}$ and a point evaluation on the polynomials, where $\mathcal{A}$ is a complex Banach algebra with identity. We then apply it to prove an analogue of Kowalski and Słodkowski's result on topological spaces of vector-valued functions of several variables. These results extend our previous work from \cite{jaikishan2024multiplicativity}; however, the techniques used differ from those used in \cite{jaikishan2024multiplicativity}. Furthermore, we characterize weighted composition operators between Hardy spaces over the polydisc amongst the continuous functions between them. Additionally, we register a partial but noteworthy success toward a multiplicative GKZ theorem for Hardy spaces.

math.FA

Transition Metal-Driven Variations in Structure, Magnetism, and Photocatalysis of Monoclinic M3Se4 (M = Fe, Co, Ni) Nanoparticles

The transition metal selenides (MxSey) have gained attention for their unique physical and chemical properties, especially those associated with the transition metal (M). Despite advancements in synthesis, fabricating these selenides is challenging due to their complex stoichiometry and high asymmetry. One such system is monoclinic iron selenide (Fe3Se4), which can be used in permanent-magnet technologies and serve as a model system for understanding magnetism. This study focuses on fabricating monoclinic M3Se4 (M = Fe, Co, or Ni) compounds via thermal decomposition, examining how solution chemistry influences their morphology and properties. With a Curie temperature of about 322 K, Fe3Se4 is ferrimagnetic, whereas Co3Se4 and Ni3Se4 are paramagnetic between 5 and 300 K. The latter two compounds also show higher catalytic activity for hydrogen evolution in water splitting, with maximum H2-evolution rates of 1.01, 5.16, and 6.83 mmol h-1g-1 for Fe3Se4, Co3Se4, and Ni3Se4, respectively.

cond-mat.mtrl-sci

Unified Anomaly Detection methods on Edge Device using Knowledge Distillation and Quantization

With the rapid advances in deep learning and smart manufacturing in Industry 4.0, there is an imperative for high-throughput, high-performance, and fully integrated visual inspection systems. Most anomaly detection approaches using defect detection datasets, such as MVTec AD, employ one-class models that require fitting separate models for each class. On the contrary, unified models eliminate the need for fitting separate models for each class and significantly reduce cost and memory requirements. Thus, in this work, we experiment with considering a unified multi-class setup. Our experimental study shows that multi-class models perform at par with one-class models for the standard MVTec AD dataset. Hence, this indicates that there may not be a need to learn separate object/class-wise models when the object classes are significantly different from each other, as is the case of the dataset considered. Furthermore, we have deployed three different unified lightweight architectures on the CPU and an edge device (NVIDIA Jetson Xavier NX). We analyze the quantized multi-class anomaly detection models in terms of latency and memory requirements for deployment on the edge device while comparing quantization-aware training (QAT) and post-training quantization (PTQ) for performance at different precision widths. In addition, we explored two different methods of calibration required in post-training scenarios and show that one of them performs notably better, highlighting its importance for unsupervised tasks. Due to quantization, the performance drop in PTQ is further compensated by QAT, which yields at par performance with the original 32-bit Floating point in two of the models considered.

cs.CV

Generic gravito-magnetic clock effects

General relativity predicts that two counter-orbiting clocks around a spinning mass differ in the time required to complete the same orbit. The difference in these two values for the orbital period is generally referred to as the gravito-magnetic (GM) clock effect. It has been proposed to measure the GM clock effect using atomic clocks carried by satellites in prograde and retrograde orbits around the Earth. The precision and stability required for satellites to accurately perform this measurement remains a challenge for current instrumentation. One of the most accurate clocks in the Universe is a millisecond pulsar, which emits periodic radio pulses with high stability. Timing of the pulsed signals from millisecond pulsars has proven to be very successful in testing predictions of general relativity and the GM clock effect is potentially measurable in binary systems. In this work we derive the generic GM clock effect by considering a slowly-spinning binary system on an elliptical orbit, with both arbitrary mass ratio and arbitrary spin orientations. The spin-orbit interaction introduces a perturbation to the orbit, causing the orbital plane to precess and nutate. We identify several different contributions to the clock effects: the choice of spin supplementary condition and the observer-dependent definition of a full revolution and "nearly-identical" orbits. We discuss the impact of these subtle definitions on the formula for GM clock effects and show that most of the existing formulae in the literature can be recovered under appropriate assumptions.

gr-qc

Nearly invariant brangesian subspaces

This article describes Hilbert spaces contractively contained in certain reproducing kernel Hilbert spaces of analytic functions on the open unit disc which are nearly invariant under division by an inner function. We extend Hitt's theorem on nearly invariant subspaces of the backward shift operator on $H^2(\bb D)$ as well as its many generalizations to the setting of de Branges spaces.

math.FA

Multiplicativity of linear functionals on function spaces on an open unit disc

This paper presents a fairly general version of the well-known Gleason-Kahane-$\dot{\text{Z}}$elazko (GKZ) theorem in the spirit of a GKZ type theorem obtained recently by Mashreghi and Ransford for Hardy spaces. In effect, we characterize a class of linear functionals as point evaluations on the vector space of all complex polynomials $\cl P$. We do not make any topological assumptions on $\cl P$. We then apply this characterization to present a version of the GKZ theorem for a vast class of topological spaces of complex-valued functions including the Hardy, Bergman, Dirichlet, and many more well-known function spaces. We obtain this result under the assumption of continuity of the linear functional, which we show, with the help of an example, to be a necessary condition for the desired conclusion. Lastly, we use the GKZ theorem for polynomials to obtain a version of the GKZ theorem for strictly cyclic weighted Hardy spaces.

math.FA

Invariant subspaces of powers of some unicellular operators

In this paper we study subspaces which are invariant under squares and cubes (separately as well as jointly) of unicellular backward weighted shift operators on a separable Hilbert space. The finite-dimensional subspaces are characterized for all weights and the infinite-dimensional subspaces are characterized for two classes of weights.

math.FA