SearcharxivSearch

arXiv subjects

Ankit Yadav

Publications and source records attributed to Ankit Yadav.

At least 19 recordsLinked to original sources

Linear Codes over $\mathbb{F}_{q}+u\mathbb{F}_{q}$ associated with Simplicial Complexes, Their Gray Images, and Subfield Codes

In recent years, simplicial complexes have gained considerable attention as a useful tool for constructing distance-optimal codes over finite fields. In this article, we construct four infinite families of linear codes over the ring $\mathcal{R}=\mathbb{F}_{q}+u\mathbb{F}_{q}$ with $u^2=0$ using simplicial complexes with one or two maximal elements, and completely determine their Lee weight distributions via exponential-sum techniques. By employing a Gray map on $\mathcal{R}$, we obtain infinite families of distance-optimal codes over $\mathbb{F}_{q}$, including a near-Griesmer family, and establish sufficient conditions for their minimality. Furthermore, we investigate the corresponding subfield codes and derive sufficient conditions for their distance-optimality and minimality, yielding infinite families of Griesmer and near-Griesmer codes.

cs.IT

New Constructions of Additive MDS TRS Codes

Additive codes over finite fields generalize linear codes, and additive MDS codes provide a natural extension of linear MDS codes. In this article, we study additive twisted Reed--Solomon (TRS) codes and obtain new constructions of additive MDS codes. First, for additive TRS codes with twist $t=2$ and an arbitrary hook, we establish necessary and sufficient conditions for the codes to be additive MDS, thereby generalizing the results in Section 3 of [Jiayu Ma et al., New families of additive non-Reed-Solomon MDS codes]. In particular, we show that the existence of an additive MDS TRS code with $t=2$ and hook $h=0$ yields codes of larger lengths than those obtained for $t=2$ and $h=k-1$ in [Jiayu Ma et al., New families of additive non-Reed-Solomon MDS codes]. Next, we consider additive TRS codes with twist vector $\mathbf{t}=(1,2)$ and hook vector $\mathbf{h}=(0,0)$, and derive necessary and sufficient conditions for them to be additive MDS. We further establish the existence of such codes. Using the Schur square technique, we obtain mild conditions under which the constructed families are inequivalent to additive Reed--Solomon (RS) codes. Finally, we determine parity-check matrices for both families of additive MDS codes considered in this article.

cs.IT

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that overconfident visual embeddings suppress output diversity under stochastic decoding, causing SE to underestimate uncertainty in such cases. Recent methods instead probe output diversity through input perturbations, including textual paraphrasing or joint text-image perturbations, and show improved performance. We study these approaches and reveals that the resulting variability is often dominated by textual changes rather than visual evidence, causing uncertainty estimates to reflect prompt sensitivity rather than visual ambiguity. We therefore propose Visual Semantic Entropy (VSE), which perturbs only the image to probe nearby visual variations while keeping the text query fixed. VSE measures uncertainty by clustering generated answers into semantic prototypes and computing the mass-weighted dispersion among them. Extensive evaluation across five modern vision-language models and five diverse VQA benchmarks demonstrates that VSE effectively captures visual ambiguity, establishing a new state-of-the-art for VLM uncertainty estimation.

cs.CV

New Quaternary codes with small Plotkin-defects from two-generator simplicial complexes

In this article, we construct infinite families of quaternary (that is, over the ring $\mathbb{Z}_4$) $\mathcal{C}_{D}$-codes, where the defining set $D$ is derived utilizing a two-generator simplicial complex, and determine their Lee weight distributions. As a result, we find three quaternary linear code families with Plotkin-defect 1 \& 2 and report at least 32 new or improved parameters having small Plotkin-defects, including 19 projective and 7 optimal parameters. We additionally report 5 quaternary linear codes with best-known parameters that are also projective. Further, we establish necessary and sufficient conditions for their Gray image to be linear, which in turn gives two infinite families of distance-optimal, one infinite family of at least almost dimension-optimal binary linear codes and five infinite families of minimal binary linear codes.

cs.IT

STRIDE: Training-Free Diversity Guidance via PCA-Directed Feature Perturbation in Single-Step Diffusion Models

Distilled one-step (T=1) or few-step (T$\leq$4) diffusion models enable real-time image generation but often exhibit reduced sample diversity compared to their multi-step counterparts. In multi-step diffusion, diversity can be introduced through schedules, trajectories, or iterative optimization; however, these mechanisms are unavailable in the few-step or single-step setting, limiting the effectiveness of existing diversity-enhancing methods. A natural alternative is to perturb intermediate features, but naive feature perturbation is often ineffective, either yielding limited diversity gains or degrading generation quality. We argue that effective diversity injection in few-step models requires perturbations that respect the model's learned feature geometry. Based on this insight, we propose STRIDE, a training-free and optimization-free method that operates in a single forward pass. STRIDE injects spatially coherent (pink) noise into intermediate transformer features, projected onto the principal components of the model's own activations, ensuring that perturbations lie on the learned feature manifold. This design enables controlled variation along meaningful directions in the representation space. Extensive experiments on FLUX.1-schnell and SD3.5 Turbo across COCO, DrawBench, PartiPrompts, and GenEval show that STRIDE consistently improves diversity while maintaining strong text alignment. In particular, STRIDE reduces intra-batch similarity with minimal impact on CLIP score, and Pareto-dominates existing training-free baselines on the diversity-fidelity frontier. These results highlight that, in the absence of iterative refinement, improving diversity in few-step and one-step diffusion depends not on increasing perturbation strength, but on aligning perturbations with the model's internal representation structure.

cs.CV

Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR

Reinforcement learning with verifiable rewards (RLVR) has become a highly effective method for improving the reasoning abilities of Large Language Models (LLMs). Recent research shows that Negative Sample Reinforcement (NSR) -- which focuses on penalizing incorrect steps rather than simply rewarding correct ones -- can match or even exceed the performance of more complex frameworks like PPO and GRPO across the entire Pass@k spectrum. However, current NSR techniques usually apply a fixed penalty throughout the training process and treat every incorrect response with the same weight. To address these limitations, we propose two extensions to the NSR framework: Adaptive Negative Sample Reinforcement. Rather than using a fixed update rule, A-NSR uses time-dependent scheduling functions. In the initial training phases, the system focuses heavily on correcting errors to stabilize the model. As training continues, it shifts toward more subtle and controlled updates. We also introduce Confidence-Weighted Negative Reinforcement, which operates on the principle that different mistakes carry different levels of importance. CW-NSR assigns specific penalty weights based on the model's normalized sequence likelihood. If the model is highly confident in a wrong path, it receives a larger penalty and for uncertain errors -- where the model is effectively exploring -- are penalized less strictly. Our formal analysis shows how these mechanisms govern token-level updates, allowing the model to leverage prior-guided probability redistribution while providing a natural defense against overfitting. We evaluated these methods on difficult reasoning datasets, including MATH, AIME 2025, and AMC23, using the Qwen2.5-Math-1.5B architecture.

cs.LG

Ru Alloying in Ni/Al Reactive Multilayers: Experimental Observations and Molecular Dynamics Simulations

Reactive multilayer thin films, a class of energetic materials, are increasingly recognized for their potential in joining applications, utilizing the chemical energy released as heat during exothermic reactions. These materials hold also promise for additional diverse technological applications, which require precise control over heat release rates and reaction propagation velocities. The microstructural properties of reactive multilayers play a critical role in determining their chemical reaction behavior. Among these, Ni/Al reactive multilayers have been extensively studied and used due to their favorable characteristics. In this study, we explore the incorporation of ruthenium (Ru) as a co-alloying element with nickel (Ni) in the Ni/Al system to investigate its impact on the materials properties, with a particular focus on reaction velocity and temperature. Ru enhances the reaction rates, but also causes a composition dependent phase transition in the as-deposited state from fcc to hcp. Additionally, molecular dynamics simulations are employed to examine the effects of Ru co-alloying with Ni, providing deeper insights into the underlying mechanisms. This work aims to advance the understanding of Ru's role in influencing the performance of Al/Ni-based reactive multilayers for advanced applications.

cond-mat.mtrl-sci

EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance

In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free guidance (CFG) is the de facto choice in modern systems and achieves this by contrasting conditional and unconditional samples. Recent work explores contrasting negative samples at inference using a weaker model, via strong/weak model pairs, attention-based masking, stochastic block dropping, or perturbations to the self-attention energy landscape. While these strategies refine the generation quality, they still lack a reliable control over the granularity or difficulty of the negative samples, and target-layer selection is often fixed. We propose Exponential Moving Average Guidance (EMAG), a training-free mechanism that modifies attention at inference time in diffusion transformers, with a statistics-based, adaptive layer-selection rule. Unlike prior methods, EMAG produces harder, semantically faithful negatives (fine-grained degradations), surfacing difficult failure modes, enabling the denoiser to refine subtle artifacts, boosting the quality and human preference score (HPS) by +0.54 over CFG. We further demonstrate that EMAG naturally composes with advanced orthogonal guidance techniques, such as APG and CADS, further improving HPS.

cs.CV

$\mathbb{F}_q\mathbb{F}_{q^2}$-additive cyclic codes and their Gray images

We investigate additive cyclic codes over the alphabet $\mathbb{F}_{q}\mathbb{F}_{q^2}$, where $q$ is a prime power. First, its generator polynomials and minimal spanning set are determined. Then, examples of $\mathbb{F}_{q^2}$-additive cyclic codes that satisfy the well-known Singleton bound are constructed. Using a Gray map, we produce certain optimal linear codes over $\mathbb{F}_{3}$. Finally, we obtain a few optimal ternary linear complementary dual (LCD) codes from $\mathbb{F}_{3}\mathbb{F}_{9}$-additive codes.

cs.IT

Optimal binary codes from $\mathcal{C}_{D}$-codes over a non-chain ring

In \cite{shi2022few-weight}, Shi and Li studied $\mathcal{C}_D$-codes over the ring $\mathcal{R}:=\mathbb{F}_2[x,y]/\langle x^2, y^2, xy-yx\rangle$ and their binary Gray images, where $D$ is derived using certain simplicial complexes. We study the subfield codes $\mathcal{C}_{D}^{(2)}$ of $\mathcal{C}_{D}$-codes over $\mathcal{R},$ where $D$ is as in \cite{shi2022few-weight} and more. We find the Hamming weight distribution and the parameters of $\mathcal{C}_D^{(2)}$ for various $D$, and identify several infinite families of codes that are distance-optimal. Besides, we provide sufficient conditions under which these codes are minimal and self-orthogonal. Two families of strongly regular graphs are obtained as an application of the constructed two-weight codes.

cs.IT

Revisiting Vision Language Foundations for No-Reference Image Quality Assessment

Large-scale vision language pre-training has recently shown promise for no-reference image-quality assessment (NR-IQA), yet the relative merits of modern Vision Transformer foundations remain poorly understood. In this work, we present the first systematic evaluation of six prominent pretrained backbones, CLIP, SigLIP2, DINOv2, DINOv3, Perception, and ResNet, for the task of No-Reference Image Quality Assessment (NR-IQA), each finetuned using an identical lightweight MLP head. Our study uncovers two previously overlooked factors: (1) SigLIP2 consistently achieves strong performance; and (2) the choice of activation function plays a surprisingly crucial role, particularly for enhancing the generalization ability of image quality assessment models. Notably, we find that simple sigmoid activations outperform commonly used ReLU and GELU on several benchmarks. Motivated by this finding, we introduce a learnable activation selection mechanism that adaptively determines the nonlinearity for each channel, eliminating the need for manual activation design, and achieving new state-of-the-art SRCC on CLIVE, KADID10K, and AGIQA3K. Extensive ablations confirm the benefits across architectures and regimes, establishing strong, resource-efficient NR-IQA baselines.

cs.CV

An Open Quantum System of Coupled Rotors

A quantum mechanical system of two coupled rotors (particles constrained to move on a circle) is studied from an open quantum systems point of view. One of the rotors is integrated out and the reduced density operator of the other rotor is studied. It's eigenvalues are worked out explicitly using the properties of Mathieu functions, and the von Neumann entropy, which is a standard measure of entanglement, is computed in terms of the Fourier coefficients defining the Mathieu functions. Furthermore, upon introducing a time-periodic delta kick and making one of the rotors much heavier than the other, the two-rotor system can be interpreted as a system-bath model, allowing us to introduce a series of approximations to derive a master equation of the Lindblad type describing the time-evolution of the reduced density operator.

quant-ph

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shapes using a benchmark focused on controlled 2D shape configurations with variations in spatial positioning, occlusion, rotation, size, and shape attributes such as type, quadrant, center-coordinates, rotation, occlusion status, and color as shown in Figure 1 and supplementary Figures S3-S81. We fine-tune state-of-the-art VLMs (2B-8B parameters) using Low-Rank Adaptation (LoRA) and validate them on multiple out-of-domain (OD) scenarios from our proposed benchmark. Our findings reveal that coherent sentence-based outputs outperform tuple formats, particularly in OD scenarios with large domain gaps. Additionally, we demonstrate that scaling numeric tokens during loss computation enhances numerical approximation capabilities, further improving performance on spatial and measurement tasks. These results highlight the importance of output format design, loss scaling strategies, and robust generalization techniques in enhancing the training and fine-tuning of VLMs, particularly for tasks requiring precise spatial approximations and strong OD generalization.

cs.CV

Poisson Geometric Formulation of Quantum Mechanics

We study the Poisson geometrical formulation of quantum mechanics for finite dimensional mixed and pure states. Equivalently, we show that quantum mechanics can be understood in the language of classical mechanics. We review the symplectic structure of the Hilbert space and identify its canonical coordinates. We extend the geometric picture to the space of density matrices $D_N^+$. We find it is not symplectic but admits a linear $\mathfrak{su}(N)$ Poisson structure. We identify Casimir surfaces of $D_N^+$ and show that the space of pure states $P_N \equiv \mathbb{C}P^{N-1}$ is one of its symplectic submanifolds which is an intersection of primitive Casimirs. We identify generic symplectic submanifolds of $D_N^+$ and calculate their dimensions. We find that $D_N^+$ is singularly foliated by the symplectic leaves of varying dimensions, also known as coadjoint orbits. We also find an ascending chain of Poisson submanifolds $D_N^M \subset D_N^{M+1}$ for $ 1 \leq M \leq N-1$. Each such Poisson submanifold $D_N^M$ is obtained by tracing out the $\mathbb{C}^M$ states from the bipartite system $\mathbb{C}^N \times \mathbb{C}^M$ and is an intersection of $N-M$ primitive Casimirs of $D_N^+$. Their Poisson structure is induced from the symplectic structure of the bipartite system. We also show their foliations. Finally, we study the positive semi-definite geometry of the symplectic submanifold $E_N^M$ consisting of the mixed states with maximum entropy in $D_N^M$.

quant-ph

Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models

Large Language Models (LLMs) are increasingly ubiquitous, yet their ability to retain and reason about temporal information remains limited, hindering their application in real-world scenarios where understanding the sequential nature of events is crucial. Our study experiments with 12 state-of-the-art models (ranging from 2B to 70B+ parameters) on a novel numerical-temporal dataset, \textbf{TempUN}, spanning from 10,000 BCE to 2100 CE, to uncover significant temporal retention and comprehension limitations. We propose six metrics to assess three learning paradigms to enhance temporal knowledge acquisition. Our findings reveal that open-source models exhibit knowledge gaps more frequently, suggesting a trade-off between limited knowledge and incorrect responses. Additionally, various fine-tuning approaches significantly improved performance, reducing incorrect outputs and impacting the identification of 'information not available' in the generations. The associated dataset and code are available at (https://github.com/lingoiitgn/TempUN).

cs.CL

A Visually Attentive Splice Localization Network with Multi-Domain Feature Extractor and Multi-Receptive Field Upsampler

Image splice manipulation presents a severe challenge in today's society. With easy access to image manipulation tools, it is easier than ever to modify images that can mislead individuals, organizations or society. In this work, a novel, "Visually Attentive Splice Localization Network with Multi-Domain Feature Extractor and Multi-Receptive Field Upsampler" has been proposed. It contains a unique "visually attentive multi-domain feature extractor" (VA-MDFE) that extracts attentional features from the RGB, edge and depth domains. Next, a "visually attentive downsampler" (VA-DS) is responsible for fusing and downsampling the multi-domain features. Finally, a novel "visually attentive multi-receptive field upsampler" (VA-MRFU) module employs multiple receptive field-based convolutions to upsample attentional features by focussing on different information scales. Experimental results conducted on the public benchmark dataset CASIA v2.0 prove the potency of the proposed model. It comfortably beats the existing state-of-the-arts by achieving an IoU score of 0.851, pixel F1 score of 0.9195 and pixel AUC score of 0.8989.

cs.CV

Towards Effective Image Forensics via A Novel Computationally Efficient Framework and A New Image Splice Dataset

Splice detection models are the need of the hour since splice manipulations can be used to mislead, spread rumors and create disharmony in society. However, there is a severe lack of image splicing datasets, which restricts the capabilities of deep learning models to extract discriminative features without overfitting. This manuscript presents two-fold contributions toward splice detection. Firstly, a novel splice detection dataset is proposed having two variants. The two variants include spliced samples generated from code and through manual editing. Spliced images in both variants have corresponding binary masks to aid localization approaches. Secondly, a novel Spatio-Compression Lightweight Splice Detection Framework is proposed for accurate splice detection with minimum computational cost. The proposed dual-branch framework extracts discriminative spatial features from a lightweight spatial branch. It uses original resolution compression data to extract double compression artifacts from the second branch, thereby making it 'information preserving.' Several CNNs are tested in combination with the proposed framework on a composite dataset of images from the proposed dataset and the CASIA v2.0 dataset. The best model accuracy of 0.9382 is achieved and compared with similar state-of-the-art methods, demonstrating the superiority of the proposed framework.

cs.CV

Datasets, Clues and State-of-the-Arts for Multimedia Forensics: An Extensive Review

With the large chunks of social media data being created daily and the parallel rise of realistic multimedia tampering methods, detecting and localising tampering in images and videos has become essential. This survey focusses on approaches for tampering detection in multimedia data using deep learning models. Specifically, it presents a detailed analysis of benchmark datasets for malicious manipulation detection that are publicly available. It also offers a comprehensive list of tampering clues and commonly used deep learning architectures. Next, it discusses the current state-of-the-art tampering detection methods, categorizing them into meaningful types such as deepfake detection methods, splice tampering detection methods, copy-move tampering detection methods, etc. and discussing their strengths and weaknesses. Top results achieved on benchmark datasets, comparison of deep learning approaches against traditional methods and critical insights from the recent tampering detection methods are also discussed. Lastly, the research gaps, future direction and conclusion are discussed to provide an in-depth understanding of the tampering detection research arena.

cs.CV