SearcharxivSearch

arXiv subjects

Vishal Jain

Publications and source records attributed to Vishal Jain.

18 recordsLinked to original sources

Frozen DINO Localizes Image Edits Without a Localizer

Localized image edits can change a photograph's meaning while leaving most of it authentic, so forensic analysis must identify where an edit occurred. We show that patch-level perturbation responses from frozen DINO encoders are themselves localization maps. Training-free Localization of AI-image Edits from patch-token Drift (TRAIL) applies one global Haar perturbation and maps cosine drift between corresponding patch tokens. On 80 source-disjoint CocoGlide test images, TRAIL reaches .903 patch AUROC versus .912 for the mask-supervised Detective SAM; fixed-threshold Dice is .619 versus .709, while an oracle threshold raises TRAIL to .790. Transferred unchanged to Poisson image interpolation, TRAIL reaches .855 AUROC versus .864, showing that the cue persists without a generator. Across sixteen DINO encoders, the best block lies at normalized depth .80-.94. Global context matters: AUROC falls from .903 globally to .857 for local-in-canvas perturbations and .735 for independently encoded crops. Frozen DINO patch tokens therefore contain a strong late-layer localization signal whose visibility depends on the perturbation and preserved context. Code: https://github.com/VishalJ99/trail-image-edit-localization.

cs.CV

Vision-Language Models as Zero-Annotation Oracles in Histopathology

Foreground segmentation is the critical first step of every computational pathology pipeline, yet existing methods rely on hand-tuned heuristics or supervised models that overfit to narrow stain and scanner distributions, failing silently on specialised stains such as Jones silver or Elastica van Gieson. We propose a coarse-to-fine approach that recasts foreground segmentation as a visual perception task and leverages general-purpose vision-language models (VLMs) as zero-annotation oracles. Our key insight is that tissue-versus-background discrimination is a natural-image recognition problem, not a histopathological one, so VLMs trained on internet-scale corpora generalise where domain-specific models cannot. We introduce Leica-75, a benchmark of 75 renal transplant whole-slide images spanning three stain families. On Leica-75, our method achieves the highest segmentation quality on out-of-distribution stains (Dice 0.858 +/- 0.027 on Jones, 0.853 +/- 0.041 on EVG) with 7x lower cross-stain variance than the best supervised baseline, while remaining competitive on in-distribution H&E. Few-shot prompting with automatically curated exemplars (Auto-context) rescues hard cases on Stress-32 (n=32), a curated stress-test subset (Dice 0.470 to 0.819 for the 2B model). VLM-based annotation review matches human expert consensus (kappa=0.989 for blur detection; mean precision/recall grading accuracy 0.708 vs. human 0.646 for segmentation mask review). The resulting pseudo-labels are used to distil lightweight student models that are as performant as the teacher model while running for a fraction of the cost. Our framework provides a principled, scalable solution to a persistent infrastructure bottleneck in digital pathology.

cs.CV

Reflections and New Directions for Human-Centered Large Language Models

Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, finance, healthcare, law, and science. With this rise in global influence comes greater urgency to build, evaluate, and deploy these systems in a manner that prioritizes not only technical capabilities but also human priorities. This work presents a framework for developing Human-Centered Large Language Models (HCLLMs), which integrates perspectives from Natural Language Processing (NLP), Human-Computer Interaction (HCI), and responsible AI. Considering the ethics, economics, and technical objectives of language modeling, we argue that model developers need to address human concerns, preferences, values, and goals, not only during a cursory post-training stage, but rather with rigor and care at every stage of the pipeline. This paper offers human-centered insights and recommendations for developers at each stage, from system design to data sourcing, model training, evaluation, and responsible deployment. Then we conclude with a case study, applying these insights to understand the future of work with HCLLMs.

cs.CL

NEMO: Neural Electro-Mechano-Optic Sensors for Multiplexed Neural Interfaces

We introduce a novel electro-optomechanic neural sensor for realizing ultra-compact neural recording probes that can detect and relay electrophysiology signals from within neural tissue. This technology addresses outstanding challenges faced by existing neural recording technologies, including the resolution trade-off with signal-to-noise-ratio (SNR) due to the high impedances of small electrodes, and lingering stimulation artifacts. The sensor employs a highly miniaturized NEMS (nano-electromechanical systems) electrostatic transducer that modulates a silicon photonic microdisk resonator to convert electrical signals to an optical signal modulation. We have been able to achieve a limit of detection down to 110 microvolts, making the sensor sensitive enough to detect neural signals. This sensitive electro-optomechanic sensor directly detects electrophysiology signals and converts them to optomechanic modulation for effective transmission to outside the brain, which provides the unique potential for massive multiplexing of neural recordings. This design eliminates the need for bulky backend headstages that limit neural recording on awake free-roaming subjects. The ability of the device to record electrophysiological signals has been demonstrated using benchtop characterization and ex-vivo recordings from live neural tissue.

physics.optics

Synchronization-dissipation dynamics in the cardiorespiratory system

Dissipative coupling is known to induce synchronization. Conversely it may be hypothesized that oscillators driven to synchronize may reduce power dissipation in their coupling. The latter scenario is realized in the human cardiorespiratory system where cardiac and respiratory rhythms are controlled by the central nervous system while interacting viscoelastically through the pulmonary vasculature. Here we examine the functional significance of this coupling which is observed in respiratory sinus arrhythmia (RSA). By modelling electrical and viscoelastic interactions within the cardiorespiratory system, we identify the conditions leading to synchronization. We demonstrate that, when present, synchronization reduces cardiac power losses by 10% in humans and up to 55% in other species. The predicted gain in cardiac output is compared to the gain observed in-vivo by pacing the heart with a device restoring RSA. It is therefore surmised that RSA may improve cardiac pumping efficiency by reducing dynamic stress and power dissipation in the pulmonary vasculature.

physics.bio-ph

Ultra-sensitive graphene-based electro-optic sensors for optically-multiplexed neural recording

Large-scale neural recording with high spatio-temporal resolution is essential for understanding information processing in brain, yet current neural interfaces fall far short of comprehensively capturing brain activity due to extremely high neuronal density and limited scalability. Although recent advances have miniaturized neural probes and increased channel density, fundamental design constraints still prevent dramatic scaling of simultaneously recorded channels. To address this limitation, we introduce a novel electro-optic sensor that directly converts ultra-low-amplitude neural electrical signals into optical signals with high signal-to-noise ratio. By leveraging the ultra-high bandwidth and intrinsic multiplexing capability of light, this approach offers a scalable path toward massively parallel neural recording beyond the limits of traditional electrical interfaces. The sensor integrates an on-chip photonic microresonator with a graphene layer, enabling direct detection of neural signals without genetically encoded optical indicators or tissue modification, making it suitable for human translation. Neural signals are locally transduced into amplified optical modulations and transmitted through on-chip waveguides, enabling interference-free recording without bulky electromagnetic shielding. Arrays of wavelength-selective sensors can be multiplexed on a single bus waveguide using wavelength-division multiplexing (WDM), greatly improving scalability while maintaining a minimal footprint to reduce tissue damage. We demonstrate detection of evoked neural signals as small as 25 $\mu$V with 3 dB SNR from mouse brain tissue and show multiplexed recording from 10 sensors on a single waveguide. These results establish a proof-of-concept for optically multiplexed neural recording and point toward scalable, high-density neural interfaces for neurological research and clinical applications.

physics.optics

Cardiorespiratory coupling improves cardiac pumping efficiency in heart failure

Recent trials of a neuronal pacemaker have shown that cardiac pumping efficiency increases when respiratory sinus arrhythmia (RSA) is artificially restored in animal models of heart failure. This novel device sheds new light on the functional role of RSA, which has long been debated, by allowing the strength of cardiorespiratory coupling to be artificially varied. Here we show that RSA minimizes the cardiac power dissipated within the cardiovascular network. The cardiorespiratory system is found to exhibit mode-locked synchronized regions within which viscoelastic dissipation is reduced relative to the scenario where cardiorespiratory coupling is absent. We determine the gain in cardiac output as the magnitude of RSA increases. We find that cardiac pumping efficiency improves up and until the cardiac frequency, within each breadth intake, is approximately 1.5 times greater than the cardiac frequency in the expiratory phase, at which point it reaches a plateau. RSA was found to be most effective at low cardiac frequencies, in good agreement with clinical evidence. Simulation of the cardiac power saved under RSA is in good agreement with the 17-20% increase in cardiac output observed in RSA-paced animal models.

q-bio.TO

Learning to Infer Unobserved Behaviors: Estimating User's Preference for a Site over Other Sites

A site's recommendation system relies on knowledge of its users' preferences to offer relevant recommendations to them. These preferences are for attributes that comprise items and content shown on the site, and are estimated from the data of users' interactions with the site. Another form of users' preferences is material too, namely, users' preferences for the site over other sites, since that shows users' base level propensities to engage with the site. Estimating users' preferences for the site, however, faces major obstacles because (a) the focal site usually has no data of its users' interactions with other sites; these interactions are users' unobserved behaviors for the focal site; and (b) the Machine Learning literature in recommendation does not offer a model of this situation. Even if (b) is resolved, the problem in (a) persists since without access to data of its users' interactions with other sites, there is no ground truth for evaluation. Moreover, it is most useful when (c) users' preferences for the site can be estimated at the individual level, since the site can then personalize recommendations to individual users. We offer a method to estimate individual user's preference for a focal site, under this premise. In particular, we compute the focal site's share of a user's online engagements without any data from other sites. We show an evaluation framework for the model using only the focal site's data, allowing the site to test the model. We rely upon a Hierarchical Bayes Method and perform estimation in two different ways - Markov Chain Monte Carlo and Stochastic Gradient with Langevin Dynamics. Our results find good support for the approach to computing personalized share of engagement and for its evaluation.

cs.IR

A Comprehensive Review on Digital Image Watermarking

The advent of the Internet led to the easy availability of digital data like images, audio, and video. Easy access to multimedia gives rise to the issues such as content authentication, security, copyright protection, and ownership identification. Here, we discuss the concept of digital image watermarking with a focus on the technique used in image watermark embedding and extraction of the watermark. The detailed classification along with the basic characteristics, namely visual imperceptibility, robustness, capacity, security of digital watermarking is also presented in this work. Further, we have also discussed the recent application areas of digital watermarking such as healthcare, remote education, electronic voting systems, and the military. The robustness is evaluated by examining the effect of image processing attacks on the signed content and the watermark recoverability. The authors believe that the comprehensive survey presented in this paper will help the new researchers to gather knowledge in this domain. Further, the comparative analysis can enkindle ideas to improve upon the already mentioned techniques.

cs.MM

Decentralized Age-of-Information Bandits

Age-of-Information (AoI) is a performance metric for scheduling systems that measures the freshness of the data available at the intended destination. AoI is formally defined as the time elapsed since the destination received the recent most update from the source. We consider the problem of scheduling to minimize the cumulative AoI in a multi-source multi-channel setting. Our focus is on the setting where channel statistics are unknown and we model the problem as a distributed multi-armed bandit problem. For an appropriately defined AoI regret metric, we provide analytical performance guarantees of an existing UCB-based policy for the distributed multi-armed bandit problem. In addition, we propose a novel policy based on Thomson Sampling and a hybrid policy that tries to balance the trade-off between the aforementioned policies. Further, we develop AoI-aware variants of these policies in which each source takes its current AoI into account while making decisions. We compare the performance of various policies via simulations.

eess.SY

Algorithmic Improvements for Deep Reinforcement Learning applied to Interactive Fiction

Text-based games are a natural challenge domain for deep reinforcement learning algorithms. Their state and action spaces are combinatorially large, their reward function is sparse, and they are partially observable: the agent is informed of the consequences of its actions through textual feedback. In this paper we emphasize this latter point and consider the design of a deep reinforcement learning agent that can play from feedback alone. Our design recognizes and takes advantage of the structural characteristics of text-based games. We first propose a contextualisation mechanism, based on accumulated reward, which simplifies the learning problem and mitigates partial observability. We then study different methods that rely on the notion that most actions are ineffectual in any given situation, following Zahavy et al.'s idea of an admissible action. We evaluate these techniques in a series of text-based games of increasing difficulty based on the TextWorld framework, as well as the iconic game Zork. Empirically, we find that these techniques improve the performance of a baseline deep reinforcement learning agent applied to text-based games.

cs.AI

Contextual Bandits Evolving Over Finite Time

Contextual bandits have the same exploration-exploitation trade-off as standard multi-armed bandits. On adding positive externalities that decay with time, this problem becomes much more difficult as wrong decisions at the start are hard to recover from. We explore existing policies in this setting and highlight their biases towards the inherent reward matrix. We propose a rejection based policy that achieves a low regret irrespective of the structure of the reward probability matrix.

cs.LG

Intersubband Quantum Disc-in-Nanowire Photodetectors with Normal-incidence Response in the Long-wavelength Infrared

Semiconductor nanowires offer great potential for realizing broadband photodetectors that are compatible with silicon technology. However, the spectral range of such detectors has so far been limited to selected regions in the ultraviolet, visible and near infrared. Here, we report on broadband nanowire heterostructure array photodetectors exhibiting a photoresponse from the visible to long-wavelength infrared. In particular, the infrared response from 3-20 um is enabled by normal incidence excitation of intersubband transitions in low-bandgap InAsP quantum discs synthesized axially within InP nanowires. The optical characteristics are explained by the excitation of the longitudinal component of optical modes in the photonic crystal formed by the nanostructured portion of the detectors, combined with a non-symmetric potential profile of the discs resulting from synthesis. Our results provide a generalizable insight into how broadband nanowire photodetectors may be designed, and how engineered nanowire heterostructures open up new fascinating opportunities for optoelectronics.

physics.optics

InP/InAsP Nanowire-based Spatially Separate Absorption and Multiplication Avalanche Photodetectors

Avalanche photodetectors (APDs) are key components in optical communication systems due to their increased photocurrent gain and short response time as compared to conventional photodetectors. A detector design where the multiplication region is implemented in a large bandgap material is desired to avoid detrimental Zener tunneling leakage currents, a concern otherwise in smaller bandgap materials required for absorption at 1.3/1.55 um. Self-assembled III-V semiconductor nanowires offer key advantages such as enhanced absorption due to optical resonance effects, strain-relaxed heterostructures and compatibility with main-stream silicon technology. Here, we present electrical and optical characteristics of single InP and InP/InAsP nanowire APD structures. Temperature-dependent breakdown characteristics of p+-n-n+ InP nanowire devices were investigated first. A clear trap-induced shift in breakdown voltage was inferred from I-V measurements. An improved contact formation to the p+-InP segment was observed upon annealing, and its effect on breakdown characteristics was investigated. The bandgap in the absorption region was subsequently varied from pure InP to InAsP to realize spatially separate absorption and multiplication APDs in heterostructure nanowires. In contrast to the homojunction APDs, no trap-induced shifts were observed for the heterostructure APDs. A gain of 12 was demonstrated for selective optical excitation of the InAsP segment. Additional electron beam-induced current measurements were carried out to investigate the effect of local excitation along the nanowire on the I-V characteristics. Our results provide important insight for optimization of avalanche photodetector devices based on III-V nanowires.

physics.app-ph

Multi Agent Driven Data Mining For Knowledge Discovery in Cloud Computing

Today, huge amount of data is available on the web. Now there is a need to convert that data in knowledge which can be useful for different purposes. This paper depicts the use of data mining process, OLAP with the combination of multi agent system to find the knowledge from data in cloud computing. For this, I am also trying to explain one case study of online shopping of one Bakery Shop. May be we can increase the sale of items by using the model, which I am trying to represent.

cs.DB

Improving Statistical Multimedia Information Retrieval Model by using Ontology

A typical IR system that delivers and stores information is affected by problem of matching between user query and available content on web. Use of Ontology represents the extracted terms in form of network graph consisting of nodes, edges, index terms etc. The above mentioned IR approaches provide relevance thus satisfying users query. The paper also emphasis on analyzing multimedia documents and performs calculation for extracted terms using different statistical formulas. The proposed model developed reduces semantic gap and satisfies user needs efficiently.

cs.IR

Ontology Based Pivoted normalization using Vector Based Approach for information Retrieval

The proposed methodology is procedural i.e. it follows finite number of steps that extracts relevant documents according to users query. It is based on principles of Data Mining for analyzing web data. Data Mining first adapts integration of data to generate warehouse. Then, it extracts useful information with the help of algorithm. The task of representing extracted documents is done by using Vector Based Statistical Approach that represents each document in set of Terms.

cs.IR

Information Retrieval (IR) through Semantic Web (SW): An Overview

A large amount of data is present on the web. It contains huge number of web pages and to find suitable information from them is very cumbersome task. There is need to organize data in formal manner so that user can easily access and use them. To retrieve information from documents, we have many Information Retrieval (IR) techniques. Current IR techniques are not so advanced that they can be able to exploit semantic knowledge within documents and give precise results. IR technology is major factor responsible for handling annotations in Semantic Web (SW) languages and in the present paper knowledgeable representation languages used for retrieving information are discussed.

cs.IR