SearcharxivSearch

arXiv subjects

Nilesh Kulkarni

Publications and source records attributed to Nilesh Kulkarni.

At least 19 recordsLinked to original sources

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation

Fine-tuning large-scale text-to-video diffusion models to add new generative controls, such as those over physical camera parameters (e.g., shutter speed or aperture), typically requires vast, high-fidelity datasets that are difficult to acquire. In this work, we propose a data-efficient fine-tuning strategy that learns these controls from sparse, low-quality synthetic data. We show that not only does fine-tuning on such simple data enable the desired controls, it actually yields superior results to models fine-tuned on photorealistic "real" data. Beyond demonstrating these results, we provide a framework that justifies this phenomenon both intuitively and quantitatively.

cs.CV

A Rapid Thermal Chemical Vapor Deposition System for Fast Synthesis of Epitaxial Graphene Under Ambient Pressure

Graphene has emerged as a promising material for next-generation electronic and thermal devices owing to its exceptional charge transport and thermal conductivity. However, high-quality samples are predominantly obtained via mechanical exfoliation from graphite crystals, a process that inherently lacks scalability. Despite extensive efforts toward large-area synthesis, cost-effective approaches for producing high-quality, large-area, single-crystalline graphene with fast turnaround time remain limited. Here, we report the design, fabrication, and performance benchmarking of a rapid thermal chemical vapor deposition (RTCVD) system capable of synthesizing epitaxial monolayer graphene under atmospheric pressure. The entire growth process, from sample loading to unloading, is achieved within $25$ minutes with a temperature ramp rate exceeding $23^\circ\mathrm{C}/s$. Growth at atmospheric pressure eliminates the need for vacuum components, thereby reducing both system complexity and operational costs. The structural and electronic quality of epitaxial graphene is comprehensively characterized using Raman spectroscopy, selected area electron diffraction (SAED), and magnetotransport measurements, which reveal signatures of quantum Hall effect in synthesized graphene samples. Furthermore, we demonstrate van der Waals epitaxial growth of palladium (Pd) thin films on graphene transferred to Si/SiO$_{2}$ substrates, establishing its single-crystalline nature over a large area and its potential as a versatile platform for subsequent heteroepitaxial growth.

cond-mat.mtrl-sci

Imprinting electrically switchable scalar spin chirality by anisotropic strain in a Kagome antiferromagnet

Topological chiral antiferromagnets, such as Mn$_{3}$Sn, are emerging as promising materials for next-generation spintronic devices due to their intrinsic transport properties linked to exotic magnetic configurations. Here, we demonstrate that anisotropic strain in Mn$_{3}$Sn thin films offers a novel approach to manipulate the magnetic ground state, unlocking new functionalities in this material. Anisotropic strain reduces the point group symmetry of the manganese (Mn) Kagome triangles from $C_{3v}$ to $C_{1}$, significantly altering the energy landscape of the magnetic states in Mn$_{3}$Sn. This symmetry reduction enables even a tiny in-plane Dzyaloshinskii-Moriya (DM) interaction to induce canting of the Mn spins out of the Kagome plane. The modified magnetic ground state introduces a finite scalar spin chirality and results in a significant Berry phase in momentum space. Consequently, a large anomalous Hall effect emerges in the Kagome plane at room temperature - an effect that is absent in the bulk material. Moreover, this two-fold degenerate magnetic state enables the creation of multiple-stable, non-volatile anomalous Hall resistance (AHR) memory states. These states are field-stable and can be controlled by thermal assisted current-induced magnetization switching requiring modest current densities and small bias fields, thereby offering a compelling new functionality in Mn$_{3}$Sn for spintronic applications.

cond-mat.mtrl-sci

Four-fold Anisotropic Magnetoresistance in Antiferromagnetic Epitaxial Thin Films of MnPt$_{x}$Pd$_{1-x}$

Antiferromagnets are emerging as promising alternatives to ferromagnets in spintronics applications. A key feature of antiferromagnets is their anisotropic magnetoresistance (AMR), which has the potential to serve as a sensitive marker for the antiferromagnetic order parameter. However, the underlying origins of this behavior remains poorly understood, particularly, in thin film geometries. In this study, we report the observation of AMR in epitaxial thin films of the collinear L1$_{0}$ antiferromagnet MnPt$_{x}$Pd$_{1-x}$. In the thicker films, AMR is dominated by a non-crystalline two-fold component, which emerges from domain reconfiguration and spin canting under applied magnetic field. As the film thickness is reduced, however, a crystalline four-fold component emerges, accompanied by the appearance of uncompensated magnetic moment, which strongly modifies the magnetotransport properties in the thinner films. We demonstrate that interfacial interactions lead to a large density of states (DOS) at the Fermi level. This enhanced DOS, combined with disorder in the thinner films, stabilizes the uncompensated moment and results in a four-fold modulation of the DOS as the Neel vector rotates, explaining the observed AMR behavior.

cond-mat.mtrl-sci

SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion Model

The computer vision community has developed numerous techniques for digitally restoring true scene information from single-view degraded photographs, an important yet extremely ill-posed task. In this work, we tackle image restoration from a different perspective by jointly denoising multiple photographs of the same scene. Our core hypothesis is that degraded images capturing a shared scene contain complementary information that, when combined, better constrains the restoration problem. To this end, we implement a powerful multi-view diffusion model that jointly generates uncorrupted views by extracting rich information from multi-view relationships. Our experiments show that our multi-view approach outperforms existing single-view image and even video-based methods on image deblurring and super-resolution tasks. Critically, our model is trained to output 3D consistent images, making it a promising tool for applications requiring robust multi-view integration, such as 3D reconstruction or pose estimation.

cs.CV

Linear non-saturating magnetoresistance and superconductivity in epitaxial thin films of YbSb$_{2}$

Rare-earth diantimonides display intriguing ground states often associated with structural order, which can be manipulated in thin film geometries. In this study, we report epitaxial synthesis of one such compound, YbSb$_{2}$, on III-V substrates using molecular-beam epitaxy. The synthesized thin films exhibit large, non-saturating, linear magnetoresistance across a wide magnetic field range. Additionally, they demonstrate superconducting properties, with a critical temperature of $\approx$ 1.025 K and a critical field of $\approx$ 83.85 Oe, consistent with the reports in bulk single crystals. While YbSb$_{2}$ has been classified as a Type-I superconductor in its bulk form, our findings provide evidence of a mixed state in the epitaxial thin films. This work paves the way for controlling the electronic ground state in this class of materials through thin film engineering.

cond-mat.supr-con

3DFIRES: Few Image 3D REconstruction for Scenes with Hidden Surface

This paper introduces 3DFIRES, a novel system for scene-level 3D reconstruction from posed images. Designed to work with as few as one view, 3DFIRES reconstructs the complete geometry of unseen scenes, including hidden surfaces. With multiple view inputs, our method produces full reconstruction within all camera frustums. A key feature of our approach is the fusion of multi-view information at the feature level, enabling the production of coherent and comprehensive 3D reconstruction. We train our system on non-watertight scans from large-scale real scene dataset. We show it matches the efficacy of single-view reconstruction methods with only one input and surpasses existing techniques in both quantitative and qualitative measures for sparse-view 3D reconstruction.

cs.CV

FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose Estimation

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose directly using neural networks are more robust to limited overlap and can infer absolute translation scale, but at the expense of reduced precision. We show how to combine the best of both methods; our approach yields results that are both precise and robust, while also accurately inferring translation scales. At the heart of our model lies a Transformer that (1) learns to balance between solved and learned pose estimations, and (2) provides a prior to guide a solver. A comprehensive analysis supports our design choices and demonstrates that our method adapts flexibly to various feature extractors and correspondence estimators, showing state-of-the-art performance in 6DoF pose estimation on Matterport3D, InteriorNet, StreetLearn, and Map-free Relocalization.

cs.CV

NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

We address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object, which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field guides the sampling of an object-conditioned human motion diffusion model, so as to encourage plausible contacts and affordance semantics. To support interactions with scarcely available data, we propose an automated synthetic data pipeline. For this, we seed a pre-trained motion model, which has priors for the basics of human movement, with interaction-specific anchor poses extracted from limited motion capture data. Using our guided diffusion model trained on generated synthetic data, we synthesize realistic motions for sitting and lifting with several objects, outperforming alternative approaches in terms of motion quality and successful action completion. We call our framework NIFTY: Neural Interaction Fields for Trajectory sYnthesis.

cs.CV

Learning to Predict Scene-Level Implicit 3D from Posed RGBD Data

We introduce a method that can learn to predict scene-level implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often been tied to meshes, we show that we can train one using only a set of posed RGBD images. This setting may help 3D reconstruction unlock the sea of accelerometer+RGBD data that is coming with new phones. Our system, D2-DRDF, can match and sometimes outperform current methods that use mesh supervision and shows better robustness to sparse data.

cs.CV

What's Behind the Couch? Directed Ray Distance Functions (DRDF) for 3D Scene Reconstruction

We present an approach for full 3D scene reconstruction from a single unseen image. We train on dataset of realistic non-watertight scans of scenes. Our approach predicts a distance function, since these have shown promise in handling complex topologies and large spaces. We identify and analyze two key challenges for predicting such image conditioned distance functions that have prevented their success on real 3D scene data. First, we show that predicting a conventional scene distance from an image requires reasoning over a large receptive field. Second, we analytically show that the optimal output of the network trained to predict these distance functions does not obey all the distance function properties. We propose an alternate distance function, the Directed Ray Distance Function (DRDF), that tackles both challenges. We show that a deep network trained to predict DRDFs outperforms all other methods quantitatively and qualitatively on 3D reconstruction from single image on Matterport3D, 3DFront, and ScanNet.

cs.CV

Collision Replay: What Does Bumping Into Things Tell You About Scene Geometry?

What does bumping into things in a scene tell you about scene geometry? In this paper, we investigate the idea of learning from collisions. At the heart of our approach is the idea of collision replay, where we use examples of a collision to provide supervision for observations at a past frame. We use collision replay to train convolutional neural networks to predict a distribution over collision time from new images. This distribution conveys information about the navigational affordances (e.g., corridors vs open spaces) and, as we show, can be converted into the distance function for the scene geometry. We analyze this approach with an agent that has noisy actuation in a photorealistic simulator.

cs.CV

Implicit Mesh Reconstruction from Unannotated Image Collections

We present an approach to infer the 3D shape, texture, and camera pose for an object from a single RGB image, using only category-level image collections with foreground masks as supervision. We represent the shape as an image-conditioned implicit function that transforms the surface of a sphere to that of the predicted mesh, while additionally predicting the corresponding texture. To derive supervisory signal for learning, we enforce that: a) our predictions when rendered should explain the available image evidence, and b) the inferred 3D structure should be geometrically consistent with learned pixel to surface mappings. We empirically show that our approach improves over prior work that leverages similar supervision, and in fact performs competitively to methods that use stronger supervision. Finally, as our method enables learning with limited supervision, we qualitatively demonstrate its applicability over a set of about 30 object categories.

cs.CV

Articulation-aware Canonical Surface Mapping

We tackle the tasks of: 1) predicting a Canonical Surface Mapping (CSM) that indicates the mapping from 2D pixels to corresponding points on a canonical template shape, and 2) inferring the articulation and pose of the template corresponding to the input image. While previous approaches rely on keypoint supervision for learning, we present an approach that can learn without such annotations. Our key insight is that these tasks are geometrically related, and we can obtain supervisory signal via enforcing consistency among the predictions. We present results across a diverse set of animal object categories, showing that our method can learn articulation and CSM prediction from image collections using only foreground mask labels for training. We empirically show that allowing articulation helps learn more accurate CSM prediction, and that enforcing the consistency with predicted CSM is similarly critical for learning meaningful articulation.

cs.CV

3D-RelNet: Joint Object and Relational Network for 3D Prediction

We propose an approach to predict the 3D shape and pose for the objects present in a scene. Existing learning based methods that pursue this goal make independent predictions per object, and do not leverage the relationships amongst them. We argue that reasoning about these relationships is crucial, and present an approach to incorporate these in a 3D prediction framework. In addition to independent per-object predictions, we predict pairwise relations in the form of relative 3D pose, and demonstrate that these can be easily incorporated to improve object level estimates. We report performance across different datasets (SUNCG, NYUv2), and show that our approach significantly improves over independent prediction approaches while also outperforming alternate implicit reasoning methods.

cs.CV

Canonical Surface Mapping via Geometric Cycle Consistency

We explore the task of Canonical Surface Mapping (CSM). Specifically, given an image, we learn to map pixels on the object to their corresponding locations on an abstract 3D model of the category. But how do we learn such a mapping? A supervised approach would require extensive manual labeling which is not scalable beyond a few hand-picked categories. Our key insight is that the CSM task (pixel to 3D), when combined with 3D projection (3D to pixel), completes a cycle. Hence, we can exploit a geometric cycle consistency loss, thereby allowing us to forgo the dense manual supervision. Our approach allows us to train a CSM model for a diverse set of classes, without sparse or dense keypoint annotation, by leveraging only foreground mask labels for training. We show that our predictions also allow us to infer dense correspondence between two images, and compare the performance of our approach against several methods that predict correspondence by leveraging varying amount of supervision.

cs.CV

Syllable-level Neural Language Model for Agglutinative Language

Language models for agglutinative languages have always been hindered in past due to myriad of agglutinations possible to any given word through various affixes. We propose a method to diminish the problem of out-of-vocabulary words by introducing an embedding derived from syllables and morphemes which leverages the agglutinative property. Our model outperforms character-level embedding in perplexity by 16.87 with 9.50M parameters. Proposed method achieves state of the art performance over existing input prediction methods in terms of Key Stroke Saving and has been commercialized.

cs.CL

An Embedded Deep Learning based Word Prediction

Recent developments in deep learning with application to language modeling have led to success in tasks of text processing, summarizing and machine translation. However, deploying huge language models for mobile device such as on-device keyboards poses computation as a bottle-neck due to their puny computation capacities. In this work we propose an embedded deep learning based word prediction method that optimizes run-time memory and also provides a real time prediction environment. Our model size is 7.40MB and has average prediction time of 6.47 ms. We improve over the existing methods for word prediction in terms of key stroke savings and word prediction rate.

cs.CL