SearcharxivSearch

arXiv subjects

Deepak Kumar

Publications and source records attributed to Deepak Kumar.

At least 37 records · Page 2Linked to original sources

Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning

Assessing chronic wound infection from photographs is challenging because visual appearance varies across wound etiologies, anatomical locations, and imaging conditions. Prior image-based deep learning methods have mainly focused on classification with limited interpretability, despite the need for evidence-grounded explanations to support point-of-care decision making. We present Infection-Reasoner, a compact 4B-parameter reasoning vision-language model for chronic wound infection classification and rationale generation. To address the scarcity of expert-labeled wound images with reasoning annotations, Infection-Reasoner is trained using a two-stage pipeline: (1) reasoning distillation, in which GPT-5.1 generates chain-of-thought rationales for unlabeled wound images to initialize wound-specific reasoning in a smaller student model (Qwen3-VL-4B-Thinking), and (2) reinforcement learning post-training with Group Relative Policy Optimization on a small labeled infection dataset to refine classification reasoning. On a held-out heterogeneous wound dataset, Infection-Reasoner achieved 86.8\% accuracy, 86.4\% sensitivity, and 87.1\% specificity, outperforming several strong baselines, including GPT-5.1. Rationale quality was further evaluated using both multimodal large language model (MLLM) judges and wound expert review. Across four MLLM judges, visual-support agreement scores ranged from 0.722 to 0.903, while expert review rated 61.8\% of rationales as Correct and 32.4\% as Partially Correct.

cs.CV

Impact of CSIR, SIC, and Hardware Impairments on the Ergodic Rate of Downlink RSMA

This work investigates the ergodic rate performance analysis of rate-splitting multiple access (RSMA) in a downlink communication system under practical impairments. Closed-form expressions are derived for key performance metrics such as ergodic rate, energy efficiency, sum-rate, and Jains fairness index, capturing the joint effects of imperfect channel state information at the receiver (CSIR), imperfect successive interference cancellation (SIC), and hardware impairments. Numerical simulations validate the accuracy of the analytical expressions and reveal several insightful trends. At low transmit powers, imperfect CSIR is the dominant performance-limiting factor, followed by hardware impairments and imperfect SIC. However, as the transmit power increases, hardware impairments become the primary bottleneck, with the impact of imperfect CSIR gradually diminishing, and imperfect SIC becoming a more prominent bottleneck. Moreover, RSMA consistently outperforms non-orthogonal multiple access (NOMA) in terms of ergodic rate, fairness, and sum-rate, even under severe non-idealities. These findings underscore the importance of incorporating fairness as a core design objective alongside rate and energy efficiency, positioning RSMA as a robust and strong multiple access candidate for next-generation wireless networks.

eess.SP

RSMA-Aided Full-Duplex Networks Under Imperfect CSI and SIC: Performance Evaluation

This work investigates a full-duplex (FD)-enhanced Rate-Splitting Multiple Access (RSMA) system under practical constraints, including imperfect channel state information (CSI) and successive interference cancellation (SIC). We derive closed-form expressions for key performance metrics, such as outage probability and throughput, for both uplink and downlink users. The analysis considers co-channel interference (CCI) from uplink to downlink users and models the self-interference (SI) channel as a random variable. Monte Carlo simulations validate the analytical results and highlight the impact of system imperfections on RSMA-FD performance. At low transmit power, imperfect CSI significantly affects the system, though this effect weakens as power increases. In contrast, imperfect SIC becomes more detrimental at high transmit power, causing severe degradation. Additionally, neglecting CCI and assuming perfect SI cancellation leads to substantial overestimation of performance. Lastly, we demonstrate that the SI cancellation factor must be carefully selected to suppress interference effectively. Otherwise, a poor choice limits the full potential of FD technology.

eess.SP

GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos

Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, contextual influences, and behavioral cues, making its quantitative modeling a challenging computational social systems problem. However, computational modeling of group affect in in-the-wild scenarios remains challenging due to limited large-scale annotated datasets and the inherent complexity of multimodal social interactions shaped by contextual and behavioral variability. The lack of comprehensive datasets annotated with multimodal and contextual information further limits advances in the field. To address this, we introduce the Group Affect from ViDeos (GAViD) dataset, comprising 5091 video clips with multimodal data (video, audio and context), annotated with ternary valence and discrete emotion labels and enriched with VideoGPT-generated contextual metadata and human-annotated action cues. We also present Context-Aware Group Affect Recognition Network (CAGNet) for multimodal context-aware group affect recognition. CAGNet achieves 63.20\% test accuracy on GAViD, comparable to state-of-the-art performance. The dataset and code are available at github.com/deepakkumar-iitr/GAViD.

cs.CV

Predictions of Modular Symmetry Fixed Points on Neutrino Masses, Mixing, and Leptogenesis

In recently proposed framework of non-holomorphic modular symmetry introduces the concept of negative and zero modular weight of Yukawa couplings. These Yukawa couplings are function of complex modulus $τ$, which is responsible for the CP asymmetry produced during leptogenesis. In this work, we restrict the $τ$ on the fixed points of modular symmetry rather than its fundamental domain in such manner Yukawa couplings are also get fixed. We have adopt this framework and propose a type III seesaw mechanism. The model is tested against neutrino oscillation data through a $χ^2$ analysis using NuFIT~6.1. To test the stability of these predictions, we also analyze regions near each fixed point by introducing a deviation $τ\rightarrow τ_{\rm fixed}(1 + εe^{iϕ})$ with $ε\in (0,0.1)$ and $ϕ\in (-π,π)$. Our results show that certain fixed points, along with their nearby regions, are capable of producing viable neutrino phenomenology while also generating the observed baryon asymmetry of the Universe.

hep-ph

SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback

We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at kappa=0.75, 8 frontier models detect only 15-31% of human-flagged issues on the diff-only configuration, demonstrating that AI code review remains far below human expert performance despite strong results on code generation benchmarks. Pull requests are drawn from active open-source repositories, filtered from 700 candidates using a Repository Quality Score, and evaluated under three frozen context configurations: diff only (config_A), diff with file content (config_B), and full context (config_C), enabling systematic ablation of context provision strategies. All 8 models degrade monotonically from config_A to config_C, even when context is provided via structured semantic layers including AST-extracted function context and import graph resolution. The dominant mechanism is a collapse of Type2_Contextual issue detection at config_B, consistent with attention dilution in long contexts: a structured 2,000-token diff-with-summary prompt outperforms a 2,500-token full-context prompt enriched with execution context, behaviour mapping, and test signatures across all 8 models. The top four models are statistically indistinguishable (mean score 0.147-0.153) while a clear tier gap separates them from the remaining four (mean score <= 0.113). Dataset, contexts, annotations, and evaluation harness are released publicly.

cs.SE

Non-Fermi liquid behavior in La$_3$Ni$_2$O$_7$ thin films under hydrostatic pressure

The discovery of superconductivity in bilayer nickel-oxides has revived an intense effort to understand the potential of high-temperature superconductivity in these materials and their relation to cuprate superconductors. In this work, we investigate the growth and properties of bilayer La$_3$Ni$_2$O$_7$ thin films as a function of substrate, oxygen treatment and applied pressure in order to study the evolution of transport properties. We report epitaxial growth of La$_3$Ni$_2$O$_7$ thin films on LaAlO$_3$ (LAO) (001) and SrLaAlO$_4$ (SLAO) (001) substrates, and the effects of ex-situ annealing in a high pressure furnace under an oxygen-rich environment. Transport measurements show that the La$_3$Ni$_2$O$_7$ thin films on LAO(001) exhibit Fermi liquid-like metallic behavior with a slight Kondo-like upturn at low temperatures, which evolves with the application of modest hydrostatic pressures toward non-Fermi liquid behavior with a temperature dependence of resistance approaching $\sim$ T$^{1.4}$ at 1.41 GPa. The ability to tune the normal state resistivity of La$_3$Ni$_2$O$_7$ films to display non-Fermi liquid behavior under such a modest hydrostatic pressure range - only 6 - 8 % of that typically applied via diamond anvil cell (DAC) in La$_3$Ni$_2$O$_7$ single crystals to achieve comparable effects - is both noteworthy and unexpected. These findings imply the strong tunability of La$_3$Ni$_2$O$_7$ in thin film form and the likely proximity of a strongly fluctuating ordered state leading to non-Fermi liquid behavior under even modest applied pressures.

cond-mat.str-el

First Positronium Lifetime Imaging using $^{52}$Mn and $^{55}$Co with a plastic-based PET scanner

Positronium Lifetime Imaging (PLI) extends positron emission tomography by using the lifetime of positronium atoms as a probe of tissue molecular architecture. In this work, we report the first PLI measurements performed with $^{52}$Mn and $^{55}$Co using the modular J-PET. Four samples were studied in each experiment: two Certified Reference Materials (polycarbonate and fused silica) and two human tissues (cardiac myxoma and adipose). The selection of PLI events was based on the registration of two 511~keV annihilation photons and one prompt gamma in triple coincidence. From the resulting lifetime spectra we extracted the mean ortho-positronium lifetime $τ_{\text{oPs}}$ and the mean positron lifetime $ΔT_{\text{mean}}$ for each sample. The measured values of $τ_{\text{oPs}}$ in polycarbonate using both isotopes matches well with the certified reference values. Furthermore, $^{55}$Co reproduced identical results for fused-silica measurements at their respective uncertainty levels. In contrast, measurements with $^{52}$Mn in fused silica show a minor deviation, which could be caused by the Parafilm spacer. In myxoma and adipose tissue, the reduced $τ_{\text{oPs}}$ values are mainly linked to the long storage history of the samples rather than to the choice of isotope. Comparing peak-to-background ratios and spectral purity, $^{55}$Co provides cleaner PLI data under the same experimental conditions. Although $^{52}$Mn offers a longer half-life and a multi gamma cascade enhancing $β^{+}$ + $γ$ coincidences, but at the expense of higher background. In this study, we demonstrate that the applied selection criteria on the data measured with the modular J-PET can be used for PLI studies even with radionuclides with complex decay patterns.

physics.med-ph

The Effect of Corneal Topography and Mucins on Tear Film Rupture

Tear film rupture on the corneal surface plays a critical role in ocular health and visual comfort. Conventional theoretical approaches often idealize the cornea as a perfectly smooth surface, ignoring the surface roughness that are characteristic of healthy as well as diseased eyes. In this study, we develop a comprehensive mathematical model to investigate tear film dynamics over the corneal surface incorporating the effects of surface roughness, slip, van der Waals forces, and lipid transport at the film-air interface. The corneal surface is represented by a small-amplitude periodic modulation. Steady-state solutions obtained using asymptotics reveal nonlinear corrections to the base profile at $O(η^2)$, which are confirmed numerically. Linear stability analysis performed using the Floquet theory demonstrates that an increase in the amplitude of roughness destabilizes the film. Specifically, both the dominant growth rate and the most unstable wavenumber increase with the roughness amplitude. Nonlinear simulations show that surface roughness significantly accelerates tear-film rupture. The slip coefficient, amplitude of roughness of the corneal surface and the initial film profile are found to significantly influence the rupture time. Moreover, the location of the rupture is sensitive to the initial disturbance. These results highlight the crucial role of surface topography and slip in determining tear film stability. The predicted rupture times are consistent with the experimental observations. The proposed model provides a realistic and accurate prediction of tear film dynamics and rupture over the corneal surface. This study offers a new perspective on tear film instability and will help address challenges such as contact lens failure which is related to tear film behavior.

physics.flu-dyn

First ex-vivo positronium imaging of tissues with modular J-PET scanner using $^{44}$Sc radionuclide

This study presents the first ex-vivo positronium imaging of human tissues using the modular J-PET scanner with the $^{44}$Sc radionuclide. The $^{44}$Sc isotope was produced via the $^{44}$Ca(p, n)$^{44}$Sc nuclear reaction and used to perform positronium imaging of phantom composed of human adipose tissue, cardiac myxoma tissue, thrombi blood clot, and also porous polymer XAD4, and a certified reference material (CRM) made from fused silica. The experiment demonstrates the suitability of $^{44}$Sc as a positron source for positronium imaging. The performance of J-PET for positronium imaging with $^{44}$Sc was validated by proper reconstruction of the mean orthopositronium lifetime for CRM material and XAD-4 polymer. The mean ortho-positronium (oPs) lifetimes determined for adipose tissue, cardiac myxoma tissues and thrombi were consistent with results of previous experiments. The study highlights the potential $^{44}$Sc radionuclide for positronium lifetime imaging (PLI).

physics.med-ph

How Far Can Pretrained LLMs Go in Symbolic Music? Controlled Comparisons of Supervised and Preference-based Adaptation

Music often shares notable parallels with language, motivating the use of pretrained large language models (LLMs) for symbolic music understanding and generation. Despite growing interest, the practical effectiveness of adapting instruction-tuned LLMs to symbolic music remains insufficiently characterized. We present a controlled comparative study of finetuning strategies for ABC-based generation and understanding, comparing an off-the-shelf instruction-tuned backbone to domain-adapted variants and a music-specialized LLM baseline. Across multiple symbolic music corpora and evaluation signals, we provide some insights into adaptation choices for symbolic music applications. We highlight the domain adaptation vs.~preserving prior information tradeoff as well as the distinct behaviour of metrics used to measure the domain adaptation for symbolic music.

cs.SD

Taming Toxic Talk: Using chatbots to intervene with users posting toxic comments

Generative AI chatbots have proven surprisingly effective at persuading people to change their beliefs and attitudes in lab settings. However, the practical implications of these findings are not yet clear. In this work, we explore the impact of rehabilitative conversations with generative AI chatbots on users who share toxic content online. Toxic behaviors -- like insults or threats of violence, are widespread in online communities. Strategies to deal with toxic behavior are typically punitive, such as removing content or banning users. Rehabilitative approaches are rarely attempted, in part due to the emotional and psychological cost of engaging with aggressive users. In collaboration with seven large Reddit communities, we conducted a large-scale field experiment (N=893) to invite people who had recently posted toxic content to participate in conversations with AI chatbots. A qualitative analysis of the conversations shows that many participants engaged in good faith and even expressed remorse or a desire to change. However, we did not observe a significant change in toxic behavior in the following month compared to a control group. We discuss possible explanations for our findings, as well as theoretical and practical implications based on our results.

cs.HC

Feasibility study of the positronium lifetime imaging with the Biograph Vision Quadra and J-PET tomographs

Background: After its first ex-vivo and in-vivo demonstration, Positronium Lifetime Imaging (PLI) has received considerable interest as a potential new diagnostic biomarker. High sensitivity Positron Emission Tomography (PET) systems are needed for PLI since it requires simultaneous registration of annihilation photons and prompt gamma. In this simulation-based study, a~feasibility of PLI with the long axial field-of-view Biograph Vision Quadra (Quadra) and the Total Body J-PET scanner was investigated. Methods: The study was performed using the GATE software. Background radiation, present within the Quadra tomograph, was added to the simulation. First, the optimal placement of the energy window for the registration of the prompt gamma was investigated. Next, the organ-wise sensitivity of Quadra was calculated for the $^{68}$Ga, $^{44}$Sc, $^{22}$Na and $^{124}$I radioisotopes. Finally, the sensitivity for the scandium isotope was compared to the sensitivities obtainable with the Total Body J-PET scanner, as well as with the modular J-PET prototype. Results: The PLI sensitivities for the Quadra with the background radiation are estimated to 9.22(3), 10.46(4), 5.91(3), and 15.39(4) cps/kBq for the $^{44}$Sc, $^{68}$Ga, $^{22}$Na and $^{124}$I radioisotopes, respectively. The highest sensitivity was obtained when the energy window for the deexcitation photon is adjacent to the energy window for the annihilation photons. The determined PLI sensitivities with Quadra and the Total Body J-PET are in the order of sensitivities of standard PET imaging with the short axial field-of-view ($\sim$20 cm) PET scanners. Conclusion: The organ-wise PLI sensitivity of Quadra has been computed for the $^{68}$Ga, $^{44}$Sc, $^{22}$Na and $^{124}$I radioisotopes. A sensitivity gain by a factor of 150 was estimated relative to the modular J-PET system previously used for the first in-vivo PLI.

physics.med-ph

What Comes After Harm? Mapping Reparative Actions in AI through Justice Frameworks

As Artificial Intelligence (AI) systems are integrated into more aspects of society, they offer new capabilities but also cause a range of harms that are drawing increasing scrutiny. A large body of work in the Responsible AI community has focused on identifying and auditing these harms. However, much less is understood about what happens after harm occurs: what constitutes reparation, who initiates it, and how effective these reparations are. In this paper, we develop a taxonomy of AI harm reparation based on a thematic analysis of real-world incidents. The taxonomy organizes reparative actions into four overarching goals: acknowledging harm, attributing responsibility, providing remedies, and enabling systemic change. We apply this framework to a dataset of 1,060 AI-related incidents, analyzing the prevalence of each action and the distribution of stakeholder involvement. Our findings show that reparation efforts are concentrated in early, symbolic stages, with limited actions toward accountability or structural reform. Drawing on theories of justice, we argue that existing responses fall short of delivering meaningful redress. This work contributes a foundation for advancing more accountable and reparative approaches to Responsible AI.

cs.HC

Surface energy-driven crumpling transition in a thin sheet under compression

In our common experience, crumpling a sheet requires external compressive force and leads to a random network of folds. However, thin sheets have been theoretically predicted to spontaneously transition from a flat to a crumpled state driven by thermal fluctuations, a phenomenon that has been elusive in experiments. We report the first observation of a similar crumpling transition driven instead by surface energy. Using a sensitive experimental protocol, when we gently compress a thin polymer sheet weakly adhered to a hydrogel substrate it transitions to a self-crumpling state at a well defined critical compression independent of system details. The transition is marked by the percolation of a fold network, and a power law increase in fold density. Most remarkably, the crumpled state shows a tunable order of folds establishing the phenomenon's potential as a simple and scalable technique to do origami with extremely thin sheets.

cond-mat.soft

FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning

Pressure ulcers (PUs) are a serious and prevalent healthcare concern. Accurate classification of PU severity (Stages I-IV) is essential for proper treatment but remains challenging due to subtle visual distinctions and subjective interpretation, leading to variability among clinicians. Prior AI-based approaches using Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) achieved promising accuracy but offered limited interpretability. We present FT-ARM (Fine-Tuned Agentic Reflection Multimodal model), a fine-tuned multimodal large language model (MLLM) with an agentic self-reflection mechanism for pressure ulcer severity classification. Inspired by clinician-style diagnostic reassessment, FT-ARM iteratively refines its predictions by reasoning over visual features and encoded clinical knowledge from text, enhancing both accuracy and consistency. On the publicly available Pressure Injury Image Dataset (PIID), FT-ARM, fine-tuned from LLaMA 3.2 90B, achieved 85% accuracy in classifying PU stages I-IV, surpassing prior CNN-based models by +4%. Unlike earlier CNN/ViT studies that relied solely on offline evaluations, FT-ARM is designed and tested for live inference, reflecting real-time deployment conditions. Furthermore, it produces clinically grounded natural-language explanations, improving interpretability and trust. By integrating fine-tuning and reflective reasoning across multimodal inputs, FT-ARM advances the reliability, transparency, and clinical applicability of automated wound assessment systems, addressing the critical need for consistent and explainable PU staging to support improved patient care.

cs.CV

Stop the Nonconsensual Use of Nude Images in Research

In order to train, test, and evaluate nudity detection models, machine learning researchers typically rely on nude images scraped from the Internet. Our research finds that this content is collected and, in some cases, subsequently distributed by researchers without consent, leading to potential misuse and exacerbating harm against the subjects depicted. This position paper argues that the distribution of nonconsensually collected nude images by researchers perpetuates image-based sexual abuse and that the machine learning community should stop the nonconsensual use of nude images in research. To characterize the scope and nature of this problem, we conducted a systematic review of papers published in computing venues that collect and use nude images. Our results paint a grim reality: norms around the usage of nude images are sparse, leading to a litany of problematic practices like distributing and publishing nude images with uncensored faces, and intentionally collecting and sharing abusive content. We conclude with a call-to-action for publishing venues and a vision for research in nudity detection that balances user agency with concrete research objectives.

cs.CY

First positronium imaging using $^{44}$Sc with the J-PET scanner: a case study on the NEMA-Image Quality phantom

Positronium Lifetime Imaging (PLI), an emerging extension of conventional positron emission tomography (PET) imaging, offers a novel window for probing the submolecular properties of biological tissues by imaging the mean lifetime of the positronium atom. Currently, the method is under rapid development in terms of reconstruction and detection systems. Recently, the first in vivo PLI of the human brain was performed using the J-PET scanner utilizing the $^{68}$Ga isotope. However, this isotope has limitations due to its comparatively low prompt gamma yields, which is crucial for positronium lifetime measurement. Among alternative radionuclides, $^{44}$Sc stands out as a promising isotope for PLI, characterized by a clinically suitable half-life (4.04 hours) emitting 1157 keV prompt gamma in 100% cases after the emission of the positron. This study reports the first experimental demonstration of PLI with $^{44}$Sc, carried out on a NEMA-Image Quality (IQ) phantom using the Modular J-PET tomograph-the first plastic scintillators-based PET scanner.

physics.med-ph