SearcharxivSearch

arXiv subjects

Shiv Shankar

Publications and source records attributed to Shiv Shankar.

At least 19 recordsLinked to original sources

Video2Reaction: Training Foundation Video Models to Predict Audience Reaction

We introduce Video2Reaction, a multimodal dataset that maps short movie segments to the induced emotional reactions of viewers in the wild, as expressed through social media comments. Video2Reaction captures the natural diversity of emotional responses by aggregating reactions from online comments at scale, modeling labels as distributions over categorical emotions to better reflect the subjective and ambiguous nature of emotional perception. We benchmark two vision-language models (VLMs) finetuned with LoRA, showing that VLMs learn effectively from Video2Reaction and outperform specialized baselines on dominant reaction prediction. We further demonstrate that VLMs pre-finetuned on Video2Reaction transfer effectively to VCE, another induced emotion dataset with a different taxonomy and video domain. Notably, LLaVA-NeXT-Video-7B pre-finetuned on Video2Reaction and adapted on only 1% of VCE training data achieves a top-3 accuracy of 0.682, on par with the best reported VCE performance trained on the full dataset. The dataset is available at https://huggingface.co/datasets/infofusionlab/Video2Reaction

cs.CV

Task-Oriented Candidate-Latent Feedback for Coarse-to-Fine Sensing in Distributed OFDM-ISAC Networks

Future integrated sensing and communication (ISAC) architectures separate the sensing entity (SE) that acquires measurements from the sensing function (SF) that performs inference, creating a need for compact, task-oriented feedback on the SE-SF interface. Forwarding the raw channel frequency response or full per-link delay-Doppler-azimuth-elevation (DDAE) tensor is prohibitively expensive, while peak-only reporting discards target-discriminative structure under clutter. We propose a learning-based coarse-to-fine sensing pipeline with candidate-latent feedback for single-target estimation. At the SE, a lightweight convolutional scorer produces a dense delay-Doppler proposal map from pilot-based OFDM channel estimates, and a learned encoder constructs K compact C-dimensional candidate tokens by fusing per-candidate azimuth-elevation patches, normalized position, and confidence cues. The latents are uniformly quantized post-training to b bits and transmitted under a finite budget B_fb = bKC + 18K + 16 bits to the SF, which performs cross-candidate refinement, reranking, and joint four-parameter estimation. On a ray-traced urban scene with static and dynamic clutter, three operating points in the (K, C, b) design space achieve 96.33-98.88% detection at 107-806 bytes per coherent processing interval, compression ratios of 1.2-9.2 x 10^4 over the 8-bit DDAE magnitude tensor, reducing the SE-SF interface from multi-Gbit/s to sub-Mbit/s rates. Cross-scene evaluation on an independent campus-scale environment achieves 98.79-99.50% detection and at-or-better angular accuracy without retraining, indicating that the learned representation captures target-relevant structure that transports across scenes of comparable or lower clutter density.

eess.SP

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reaction prediction and other content engagement applications, we introduce $\textbf{Video2Reaction}$, a multimodal dataset that maps short movie segments to a distribution of $\textit{induced emotions}$ of viewers in the wild, as expressed through social media. $\textbf{Video2Reaction}$ spans more than 10,000 videos and serves as a reliable benchmark as well as a training resource for audience reaction prediction. To enable cost-effective continuous annotations as reactions may change over time, we develop a two-stage multi-agent pipeline using only open-source LLMs, achieving 86% correctness under blind human verification despite the inherently noisy and subjective nature of the task. We establish the first benchmark for video-to-reaction-distribution prediction in the wild and show that pretrained foundation video models fail in zero-shot settings, while finetuning transforms them into state-of-the-art predictors capable of modeling both full reaction distributions and dominant responses from video alone. However, the task remains challenging: even the strongest methods achieve only 77% Top-3 F1 in dominant reaction prediction (LLaVA-Next), highlighting a substantial gap in modeling collective audience reaction. \modification{Dataset and code are available at our project page: https://information-fusion-lab-umass.github.io/video2reaction-bench.github.io

cs.CV

Deep Learning assisted Port-Cycling based Channel Sounding for Precoder Estimation in Massive MIMO Arrays

Future wireless systems are expected to employ a substantially larger number of transmit ports for channel state information (CSI) estimation compared to current specifications. Although scaling ports improves spectral efficiency, it also increases the resource overhead to transmit reference signals across the time-frequency grid, ultimately reducing achievable data throughput. In this work, we propose an deep learning (DL)-based CSI reconstruction framework that serves as an enabler for reliable CSI acquisition in future 6G systems. The proposed solution involves designing a port-cycling mechanism that sequentially sounds different portions of CSI ports across time, thereby lowering the overhead while preserving channel observability. The proposed CSI Adaptive Network (CsiAdaNet) model exploits the resulting sparse measurements and captures both spatial and temporal correlations to accurately reconstruct the full-port CSI. The simulation results show that our method achieves overhead reduction while maintaining high CSI reconstruction accuracy.

eess.SP

Mechanistic insights into hydrogen reduction of multicomponent oxides via in-situ high-energy X-ray diffraction

Co-reduction of multicomponent oxides with hydrogen provides a carbon-neutral approach toward sustainable alloy design. Herein, we investigate the hydrogen-based direct reduction, using in-situ high-energy X-ray diffraction of two precursor variants: mechanically mixed powders and pre-sintered oxide mixtures, targeting an equiatomic CoFeMnNi alloy. We find distinct reduction pathways and microstructure evolution depending on initial precursors. Mixed powders at 700 {\deg}C are reduced to body-centered-cubic, face-centered-cubic, and MnO phases via halite, spinel, and Mn3O4 intermediates, whereas the pre-sintered material directly transforms into a mixture of metallic and oxide phases. The post-reduction microstructures are also different: mixed oxides show loosely packed morphology, whereas pre-sintered material reveals metallic nanoparticles supported on nanoporous MnO. The formation of nanoporous metallic networks is strongly governed by the precursor state, highlighting the role of initial precursors on the final microstructure. This precursor design strategy offers a single-step route to nanoporous alloys with potential applications in catalysis and energy technologies.

cond-mat.mtrl-sci

Audio-Visual Speech Separation via Bottleneck Iterative Network

Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.

cs.SD

Sustainable Pre-reduction of Ferromanganese Oxides with Hydrogen: Heating Rate-Dependent Reduction Pathways and Microstructure Evolution

The reduction of ferromanganese ores into metallic feedstock is an energy-intensive process with substantial carbon emissions, necessitating sustainable alternatives. Hydrogen-based pre-reduction of manganese-rich ores offers a low-emission pathway to augment subsequent thermic Fe-Mn alloy production. However, reduction dynamics and microstructure evolution under varying thermal conditions remain poorly understood. This study investigates the influence of heating rate on the hydrogen-based direct reduction of natural Nchwaning ferromanganese ore and a synthetic analog. Non-isothermal thermogravimetric analysis revealed a complex multistep reduction process with overlapping kinetic regimes. Isoconversional kinetic analysis showed increased activation energy with reduction degree, indicating a transition from surface-reaction to diffusion-controlled reduction mechanisms. Interrupted X-ray diffraction experiments suggested that slow heating enables complete conversion to MnO and metallic Fe, while rapid heating promotes Fe- and Mn-oxides intermixing. Thermodynamic calculations for the Fe-Mn-O system predicted the equilibrium phase evolution, indicating Mn stabilized Fe-containing spinel and halite phases. Microstructural analysis revealed that slow heating rate yields fine and dispersed Fe particles in a porous MnO matrix, while fast heating leads to sporadic Fe-rich agglomerates. These findings suggest heating rate as a critical parameter governing reduction pathway, phase distribution, and microstructure evolution, thus offering key insights for optimizing hydrogen-based pre-reduction strategies towards more efficient and sustainable ferromanganese production.

cond-mat.mtrl-sci

Hydrogen-based direct reduction of multicomponent oxides: Insights from powder and pre-sintered precursors toward sustainable alloy design

The co-reduction of metal oxide mixtures using hydrogen as a reductant in conjunction with compaction and sintering of the evolving metallic blends offers a promising alternative toward sustainable alloy production through a single, integrated, and synergistic process. Herein, we provide fundamental insights into hydrogen-based direct reduction (HyDR) of distinct oxide precursors that differ by phase composition and morphology. Specifically, we investigate the co-reduction of multicomponent metal oxides targeting a 25Co-25Fe-25Mn-25Ni (at.%) alloy, by using either a compacted powder (mechanically mixed oxides) comprising Co3O4-Fe2O3-Mn2O3-NiO or a pre-sintered compound (chemically mixed oxides) comprising a Co,Ni-rich halite and a Fe,Mn-rich spinel. Thermogravimetric analysis (TGA) at a heating rate of 10 {\deg}C/min reveals that the reduction onset temperature for the compacted powder was ~175 {\deg}C, whereas it was significantly delayed to ~525 {\deg}C for the pre-sintered sample. Nevertheless, both sample types attained a similar reduction degree (~80%) after isothermal holding for 1 h at 700 {\deg}C. Phase analysis and microstructural characterization of reduced samples confirmed the presence of metallic Co, Fe, and Ni alongside MnO. A minor fraction of Fe remains unreduced, stabilized in the (Fe,Mn)O halite phase, in accord with thermodynamic calculations. Furthermore, ~1 wt.% of BCC phase was found only in the reduced pre-sintered sample, owing to the different reduction pathways. The kinetics and thermodynamics effects were decoupled by performing HyDR experiments on pulverized pre-sintered samples. These findings demonstrate that initial precursor states influence both the reduction behavior and the microstructural evolution, providing critical insights for the sustainable production of multicomponent alloys.

cond-mat.mtrl-sci

Unraveling the thermodynamics and mechanism behind the lowering of reduction temperatures in oxide mixtures

Hydrogen-based direct reduction offers a sustainable pathway to decarbonize the metal production industry. However, stable metal oxides, like Cr$_2$O$_3$, are notoriously difficult to reduce, requiring extremely high temperatures (above 1300 $^\circ$C). Herein, we show how reducing mixed oxides can be leveraged to lower hydrogen-based reduction temperatures of stable oxides and produce alloys in a single process. Using a newly developed thermodynamic framework, we predict the precise conditions (oxygen partial pressure, temperature, and oxide composition) needed for co-reduction. We showcase this approach by reducing Cr$_2$O$_3$ mixed with Fe$_2$O$_3$ at 1100 $^\circ$C, significantly lowering reduction temperatures (by $\geq$200 $^\circ$C). Our model and post-reduction atom probe tomography analysis elucidate that the temperature-lowering effect is driven by the lower chemical activity of Cr in the metallic phase. This strategy achieves low-temperature co-reduction of mixed oxides, dramatically reducing energy consumption and CO$_2$ emissions, while unlocking transformative pathways toward sustainable alloy design.

cond-mat.mtrl-sci

Learning Straight Flows by Learning Curved Interpolants

Flow matching models typically use linear interpolants to define the forward/noise addition process. This, together with the independent coupling between noise and target distributions, yields a vector field which is often non-straight. Such curved fields lead to a slow inference/generation process. In this work, we propose to learn flexible (potentially curved) interpolants in order to learn straight vector fields to enable faster generation. We formulate this via a multi-level optimization problem and propose an efficient approximate procedure to solve it. Our framework provides an end-to-end and simulation-free optimization procedure, which can be leveraged to learn straight line generative trajectories.

cs.LG

A/B testing under Interference with Partial Network Information

A/B tests are often required to be conducted on subjects that might have social connections. For e.g., experiments on social media, or medical and social interventions to control the spread of an epidemic. In such settings, the SUTVA assumption for randomized-controlled trials is violated due to network interference, or spill-over effects, as treatments to group A can potentially also affect the control group B. When the underlying social network is known exactly, prior works have demonstrated how to conduct A/B tests adequately to estimate the global average treatment effect (GATE). However, in practice, it is often impossible to obtain knowledge about the exact underlying network. In this paper, we present UNITE: a novel estimator that relax this assumption and can identify GATE while only relying on knowledge of the superset of neighbors for any subject in the graph. Through theoretical analysis and extensive experiments, we show that the proposed approach performs better in comparison to standard estimators.

cs.LG

Adaptive Instrument Design for Indirect Experiments

Indirect experiments provide a valuable framework for estimating treatment effects in situations where conducting randomized control trials (RCTs) is impractical or unethical. Unlike RCTs, indirect experiments estimate treatment effects by leveraging (conditional) instrumental variables, enabling estimation through encouragement and recommendation rather than strict treatment assignment. However, the sample efficiency of such estimators depends not only on the inherent variability in outcomes but also on the varying compliance levels of users with the instrumental variables and the choice of estimator being used, especially when dealing with numerous instrumental variables. While adaptive experiment design has a rich literature for direct experiments, in this paper we take the initial steps towards enhancing sample efficiency for indirect experiments by adaptively designing a data collection policy over instrumental variables. Our main contribution is a practical computational procedure that utilizes influence functions to search for an optimal data collection policy, minimizing the mean-squared error of the desired (non-linear) estimator. Through experiments conducted in various domains inspired by real-world applications, we showcase how our method can significantly improve the sample efficiency of indirect experiments.

cs.LG

Privacy Aware Experiments without Cookies

Consider two brands that want to jointly test alternate web experiences for their customers with an A/B test. Such collaborative tests are today enabled using \textit{third-party cookies}, where each brand has information on the identity of visitors to another website. With the imminent elimination of third-party cookies, such A/B tests will become untenable. We propose a two-stage experimental design, where the two brands only need to agree on high-level aggregate parameters of the experiment to test the alternate experiences. Our design respects the privacy of customers. We propose an estimater of the Average Treatment Effect (ATE), show that it is unbiased and theoretically compute its variance. Our demonstration describes how a marketer for a brand can design such an experiment and analyze the results. On real and simulated data, we show that the approach provides valid estimate of the ATE with low variance and is robust to the proportion of visitors overlapping across the brands.

stat.ME

Optimization using Parallel Gradient Evaluations on Multiple Parameters

We propose a first-order method for convex optimization, where instead of being restricted to the gradient from a single parameter, gradients from multiple parameters can be used during each step of gradient descent. This setup is particularly useful when a few processors are available that can be used in parallel for optimization. Our method uses gradients from multiple parameters in synergy to update these parameters together towards the optima. While doing so, it is ensured that the computational and memory complexity is of the same order as that of gradient descent. Empirical results demonstrate that even using gradients from as low as \textit{two} parameters, our method can often obtain significant acceleration and provide robustness to hyper-parameter settings. We remark that the primary goal of this work is less theoretical, and is instead aimed at exploring the understudied case of using multiple gradients during each step of optimization.

cs.LG

Off-Policy Evaluation for Action-Dependent Non-Stationary Environments

Methods for sequential decision-making are often built upon a foundational assumption that the underlying decision process is stationary. This limits the application of such methods because real-world problems are often subject to changes due to external factors (passive non-stationarity), changes induced by interactions with the system itself (active non-stationarity), or both (hybrid non-stationarity). In this work, we take the first steps towards the fundamental challenge of on-policy and off-policy evaluation amidst structured changes due to active, passive, or hybrid non-stationarity. Towards this goal, we make a higher-order stationarity assumption such that non-stationarity results in changes over time, but the way changes happen is fixed. We propose, OPEN, an algorithm that uses a double application of counterfactual reasoning and a novel importance-weighted instrument-variable regression to obtain both a lower bias and a lower variance estimate of the structure in the changes of a policy's past performances. Finally, we show promising results on how OPEN can be used to predict future performances for several domains inspired by real-world applications that exhibit non-stationarity.

cs.LG

Progressive Fusion for Multimodal Integration

Integration of multimodal information from various sources has been shown to boost the performance of machine learning models and thus has received increased attention in recent years. Often such models use deep modality-specific networks to obtain unimodal features which are combined to obtain "late-fusion" representations. However, these designs run the risk of information loss in the respective unimodal pipelines. On the other hand, "early-fusion" methodologies, which combine features early, suffer from the problems associated with feature heterogeneity and high sample complexity. In this work, we present an iterative representation refinement approach, called Progressive Fusion, which mitigates the issues with late fusion representations. Our model-agnostic technique introduces backward connections that make late stage fused representations available to early layers, improving the expressiveness of the representations at those stages, while retaining the advantages of late fusion designs. We test Progressive Fusion on tasks including affective sentiment detection, multimedia analysis, and time series fusion with different models, demonstrating its versatility. We show that our approach consistently improves performance, for instance attaining a 5% reduction in MSE and 40% improvement in robustness on multimodal time series prediction.

cs.LG

Implicit Training of Energy Model for Structure Prediction

Most deep learning research has focused on developing new model and training procedures. On the other hand the training objective has usually been restricted to combinations of standard losses. When the objective aligns well with the evaluation metric, this is not a major issue. However when dealing with complex structured outputs, the ideal objective can be hard to optimize and the efficacy of usual objectives as a proxy for the true objective can be questionable. In this work, we argue that the existing inference network based structure prediction methods ( Tu and Gimpel 2018; Tu, Pang, and Gimpel 2020) are indirectly learning to optimize a dynamic loss objective parameterized by the energy model. We then explore using implicit-gradient based technique to learn the corresponding dynamic objectives. Our experiments show that implicitly learning a dynamic loss landscape is an effective method for improving model performance in structure prediction.

cs.LG

Neural Dependency Coding inspired Multimodal Fusion

Information integration from different modalities is an active area of research. Human beings and, in general, biological neural systems are quite adept at using a multitude of signals from different sensory perceptive fields to interact with the environment and each other. Recent work in deep fusion models via neural networks has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis, captioning and image description. However, such research has mostly focused on architectural changes allowing for fusion of different modalities while keeping the model complexity manageable. Inspired by recent neuroscience ideas about multisensory integration and processing, we investigate the effect of synergy maximizing loss functions. Experiments on multimodal sentiment analysis tasks: CMU-MOSI and CMU-MOSEI with different models show that our approach provides a consistent performance boost.

cs.NE