SearcharxivSearch

arXiv subjects

Manoj Kumar

Publications and source records attributed to Manoj Kumar.

At least 37 records · Page 2Linked to original sources

PaliGemma: A versatile 3B VLM for transfer

PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly knowledgeable base model that is effective to transfer. It achieves strong performance on a wide variety of open-world tasks. We evaluate PaliGemma on almost 40 diverse tasks including standard VLM benchmarks, but also more specialized tasks such as remote-sensing and segmentation.

cs.CV

Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations

Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation techniques involve adapting a text-to-image model to reconstruct an input image conditioned on the same image. We first demonstrate that a diffusion model trained to reconstruct an input image from frozen embeddings, can reconstruct the image with minor variations. Second, inspired by how text-to-image models learn from web-scale text-image pairs, we explore a new pretraining strategy to generate image variations using a large collection of image pairs. Our diffusion model \textit{Semantica} receives a random (encoded) image from a webpage as conditional input and denoises another noisy random image from the same webpage. We carefully examine various design choices for the image encoder, given its crucial role in extracting relevant context from the input image. Once trained, \textit{Semantica} can adaptively generate new images from a dataset by simply using images from that dataset as input. Finally, we identify limitations in standard image consistency metrics for evaluating image variations and propose alternative metrics based on few-shot generation.

cs.CV

Emergent dynamics due to chemo-hydrodynamic self-interactions in active polymers

The field of synthetic active matter has, thus far, been led by efforts to create point-like, isolated (yet interacting) self-propelled objects (\emph{e.g.} colloids, droplets, microrobots) and understanding their collective dynamics. The design of flexible, freely jointed active assemblies from autonomously powered components remains a challenge. Here, we report freely-jointed active polymers created using self-propelled droplets as monomeric units. Our experiments reveal that the self-shaping chemo-hydrodynamic interactions between the monomeric droplets give rise to an emergent rigidity (the acquisition of a stereotypical asymmetric C-shape) and associated ballistic propulsion of the active polymers. The rigidity and propulsion of the chains vary systematically with their lengths. Using simulations of a minimal model, we establish that the emergent polymer dynamics are a generic consequence of quasi two-dimensional confinement and auto-repulsive trail-mediated chemical interactions between the freely jointed active droplets. Finally, we tune the interplay between the chemical and hydrodynamic fields to experimentally demonstrate oscillatory dynamics of the rigid polymer propulsion. Altogether, our work highlights the possible first steps towards synthetic self-morphic active matter.

cond-mat.soft

Rigid flocks, undulatory gaits, and chiral foldamers in a chemically active polymer

Active matter systems - such as a collection of active colloidal particles - operate far from equilibrium with complex inter-particle interactions that govern their collective dynamics. Predicting the collective dynamics of such systems may aid the design of self-shaping structures comprised of active colloidal units with a prescribed dynamical function. Here, using simulations and theory, we study the collective dynamics of a chain consisting of active Brownian particles with internal interactions via trail-mediated chemicals, connected by harmonic springs in two dimensions to obtain design principles for active colloidal molecules. We show that two-dimensional confinement and chemo-repulsive interactions between the freely-jointed particles lead to an emergent rigidity of the chain in the steady-state dynamics. In the chemo-attractive regime, the chain collapses into crystals that abruptly halt their motion. Further, in a chain consisting of a binary mixture of monomers, we show that non-reciprocal chemical affinities between distinct species give rise to novel phenomena, such as chiral molecules with tunable dynamics, sustained undulatory gaits and reversal of the direction of motion. Our results suggest a novel interpretation of the role of trail-mediated interactions, in addition to providing active self-assembly principles arising due to non-reciprocal interactions.

cond-mat.soft

Frozen Feature Augmentation for Few-Shot Image Classification

Training a linear classifier or lightweight model on top of pretrained vision model outputs, so-called 'frozen features', leads to impressive performance on a number of downstream few-shot tasks. Currently, frozen features are not modified during training. On the other hand, when networks are trained directly on images, data augmentation is a standard recipe that improves performance with no substantial overhead. In this paper, we conduct an extensive pilot study on few-shot image classification that explores applying data augmentations in the frozen feature space, dubbed 'frozen feature augmentation (FroFA)', covering twenty augmentations in total. Our study demonstrates that adopting a deceptively simple pointwise FroFA, such as brightness, can improve few-shot performance consistently across three network architectures, three large pretraining datasets, and eight transfer datasets.

cs.CV

Towards Building Autonomous Data Services on Azure

Modern cloud has turned data services into easily accessible commodities. With just a few clicks, users are now able to access a catalog of data processing systems for a wide range of tasks. However, the cloud brings in both complexity and opportunity. While cloud users can quickly start an application by using various data services, it can be difficult to configure and optimize these services to gain the most value from them. For cloud providers, managing every aspect of an ever-increasing set of data services, while meeting customer SLAs and minimizing operational cost is becoming more challenging. Cloud technology enables the collection of significant amounts of workload traces and system telemetry. With the progress in data science (DS) and machine learning (ML), it is feasible and desirable to utilize a data-driven, ML-based approach to automate various aspects of data services, resulting in the creation of autonomous data services. This paper presents our perspectives and insights on creating autonomous data services on Azure. It also covers the future endeavors we plan to undertake and unresolved issues that still need attention.

cs.DC

Narrowband THz Emission from a Plasma Oscillator Imbedded in a Plasma Density Gradient

A novel method is presented for generating radiation using the beat wave associated with a bi-frequency laser pulse, to excite plasma oscillations in a plasma slab with a density gradient. By resonantly exciting a plasma wave, it can be localised and transformed into a plasma oscillator that produces a beam of radially polarised terahertz radiation. Particle-in-cell simulations and analytic theory are used to demonstrate its main characteristics, which includes narrow bandwidth. The radiator should have useful applications such as terahertz-band particle accelerators and pump-probe experiments.

physics.plasm-ph

Image Captioners Are Scalable Vision Learners Too

Contrastive pretraining on image-text pairs from the web is one of the most popular large-scale pretraining strategies for vision backbones, especially in the context of large multimodal models. At the same time, image captioning on this type of data is commonly considered an inferior pretraining strategy. In this paper, we perform a fair comparison of these two pretraining strategies, carefully matching training data, compute, and model capacity. Using a standard encoder-decoder transformer, we find that captioning alone is surprisingly effective: on classification tasks, captioning produces vision encoders competitive with contrastively pretrained encoders, while surpassing them on vision & language tasks. We further analyze the effect of the model architecture and scale, as well as the pretraining data on the representation quality, and find that captioning exhibits the same or better scaling behavior along these axes. Overall our results show that plain image captioning is a more powerful pretraining strategy than was previously believed.

cs.CV

Dual PatchNorm

We propose Dual PatchNorm: two Layer Normalization layers (LayerNorms), before and after the patch embedding layer in Vision Transformers. We demonstrate that Dual PatchNorm outperforms the result of exhaustive search for alternative LayerNorm placement strategies in the Transformer block itself. In our experiments, incorporating this trivial modification, often leads to improved accuracy over well-tuned Vision Transformers and never hurts.

cs.CV

Domain statistics in the relaxation of the one-dimensional Ising model with strong long-range interactions

After a zero temperature quench, we study the kinetics of the one-dimensional Ising model with long-range interactions between spins at distance $r$ decaying as $r^{-α}$, with $α\le 1$. As shown in our recent study [SciPost Phys 10, 109 (2021)] that only a fraction of the non-equilibrium trajectories is characterized by the presence of coarsening domains while in the remaining ones the system is quickly driven towards a magnetised state. Restricting to realisations displaying coarsening we compute numerically the probability distribution of the size of the domains and find that it exhibits a scaling behaviour with an unusual $α$-dependent power-law decay. This peculiar behaviour is also related to the divergence of the average size of domains with system size at finite times. Such a scenario differs from the one observed when $α>1$, where the distribution decays exponentially. Finally, based on numerical results and on analytical calculations we argue that the average domain size grows asymptotically linearly in time.

cond-mat.stat-mech

Scaling Vision Transformers to 22 Billion Parameters

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.

cs.CV

Pego theorem on compact groups

The Pego theorem characterizes the precompact subsets of the square-integrable functions on $\mathbb{R}^n$ via the Fourier transform. We prove the analogue of the Pego theorem on compact groups (not necessarily abelian).

math.FA

Large language models can segment narrative events similarly to humans

Humans perceive discrete events such as "restaurant visits" and "train rides" in their continuous experience. One important prerequisite for studying human event perception is the ability of researchers to quantify when one event ends and another begins. Typically, this information is derived by aggregating behavioral annotations from several observers. Here we present an alternative computational approach where event boundaries are derived using a large language model, GPT-3, instead of using human annotations. We demonstrate that GPT-3 can segment continuous narrative text into events. GPT-3-annotated events are significantly correlated with human event annotations. Furthermore, these GPT-derived annotations achieve a good approximation of the "consensus" solution (obtained by averaging across human annotations); the boundaries identified by GPT-3 are closer to the consensus, on average, than boundaries identified by individual human annotators. This finding suggests that GPT-3 provides a feasible solution for automated event annotations, and it demonstrates a further parallel between human cognition and prediction in large language models. In the future, GPT-3 may thereby help to elucidate the principles underlying human event perception.

cs.CL

Critical behavior of the three-state random-field Potts model in three dimensions

Enormous advances have been made in the past 20 years in our understanding of the random-field Ising model, and there is now consensus on many aspects of its behavior at least in thermal equilibrium. In contrast, little is known about its generalization to the random-field Potts model which has wide-ranging applications. Here we start filling this gap with an investigation of the three-state random-field Potts model in three dimensions. Building on the success of ground-state calculations for the Ising system, we use a recently developed approximate scheme based on graph-cut methods to study the properties of the zero-temperature random fixed point of the system that determines the zero and non-zero temperature transition behavior. We find compelling evidence for a continuous phase transition. Implementing an extensive finite-size scaling (FSS) analysis, we determine the critical exponents and compare them to those of the random-field Ising model.

cond-mat.stat-mech

Do better ImageNet classifiers assess perceptual similarity better?

Perceptual distances between images, as measured in the space of pre-trained deep features, have outperformed prior low-level, pixel-based metrics on assessing perceptual similarity. While the capabilities of older and less accurate models such as AlexNet and VGG to capture perceptual similarity are well known, modern and more accurate models are less studied. In this paper, we present a large-scale empirical study to assess how well ImageNet classifiers perform on perceptual similarity. First, we observe a inverse correlation between ImageNet accuracy and Perceptual Scores of modern networks such as ResNets, EfficientNets, and Vision Transformers: that is better classifiers achieve worse Perceptual Scores. Then, we examine the ImageNet accuracy/Perceptual Score relationship on varying the depth, width, number of training steps, weight decay, label smoothing, and dropout. Higher accuracy improves Perceptual Score up to a certain point, but we uncover a Pareto frontier between accuracies and Perceptual Score in the mid-to-high accuracy regime. We explore this relationship further using a number of plausible hypotheses such as distortion invariance, spatial frequency sensitivity, and alternative perceptual functions. Interestingly we discover shallow ResNets and ResNets trained for less than 5 epochs only on ImageNet, whose emergent Perceptual Score matches the prior best networks trained directly on supervised human perceptual judgements. The checkpoints for the models in our study are available at https://console.cloud.google.com/storage/browser/gresearch/perceptual_similarity.

cs.CV

Intense multicycle THz pulses generated by laser-produced nanoplasmas

We present a novel scheme to obtain robust, narrowband, and tunable THz emission by using a nano-dimensional overdense plasma target that is irradiated by two counter-propagating detuned laser pulses. So far, no narrowband THz sources with a field strength of GV/m-level have been reported from laser-solid interaction. We report intense THz pulses at beat-frequency ($\simeq$30THz) generated due to strong plasma current produced by the beat ponderomotive force in the colliding region, with an unprecedentedly high peak field strength of 11.9GV/m and spectral width $(Δf/f$$\simeq$$5.3\%)$ from 2D PIC simulations. Such an extremely bright narrowband THz source is suitable for various ambitious applications.

physics.plasm-ph

A Unified Framework for Optimization-Based Graph Coarsening

Graph coarsening is a widely used dimensionality reduction technique for approaching large-scale graph machine learning problems. Given a large graph, graph coarsening aims to learn a smaller-tractable graph while preserving the properties of the originally given graph. Graph data consist of node features and graph matrix (e.g., adjacency and Laplacian). The existing graph coarsening methods ignore the node features and rely solely on a graph matrix to simplify graphs. In this paper, we introduce a novel optimization-based framework for graph dimensionality reduction. The proposed framework lies in the unification of graph learning and dimensionality reduction. It takes both the graph matrix and the node features as the input and learns the coarsen graph matrix and the coarsen feature matrix jointly while ensuring desired properties. The proposed optimization formulation is a multi-block non-convex optimization problem, which is solved efficiently by leveraging block majorization-minimization, $\log$ determinant, Dirichlet energy, and regularization frameworks. The proposed algorithms are provably convergent and practically amenable to numerous tasks. It is also established that the learned coarsened graph is $ε\in(0,1)$ similar to the original graph. Extensive experiments elucidate the efficacy of the proposed framework for real-world applications.

stat.ML

Functional Optimization Reinforcement Learning for Real-Time Bidding

Real-time bidding is the new paradigm of programmatic advertising. An advertiser wants to make the intelligent choice of utilizing a \textbf{Demand-Side Platform} to improve the performance of their ad campaigns. Existing approaches are struggling to provide a satisfactory solution for bidding optimization due to stochastic bidding behavior. In this paper, we proposed a multi-agent reinforcement learning architecture for RTB with functional optimization. We designed four agents bidding environment: three Lagrange-multiplier based functional optimization agents and one baseline agent (without any attribute of functional optimization) First, numerous attributes have been assigned to each agent, including biased or unbiased win probability, Lagrange multiplier, and click-through rate. In order to evaluate the proposed RTB strategy's performance, we demonstrate the results on ten sequential simulated auction campaigns. The results show that agents with functional actions and rewards had the most significant average winning rate and winning surplus, given biased and unbiased winning information respectively. The experimental evaluations show that our approach significantly improve the campaign's efficacy and profitability.

cs.AI