Searcharxiv⌕ Search

arXiv subjects

Sayan Ghosh

Publications and source records attributed to Sayan Ghosh.

At least 37 records · Page 2Linked to original sources

Martingale Characterizations of Non-Homogeneous Counting Processes and Their Fractional Variants

This paper investigates the martingale characterizations of non-homogeneous counting processes and their fractional generalizations. We show that the weighted sum of non-homogeneous Poisson processes (NPPs) is the non-homogeneous generalized counting process (NGCP). Both the compensated and exponential forms of martingale characterization for NGCP are obtained, and are shown to be equivalent. Moreover, we provide martingale characterizations for various time-changed variants of the NGCP and their Skellam versions using stable and/or inverse stable subordinators.

math.PR↗

Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search

Large Language Models (LLMs) have demonstrated impressive quality when applied to predictive tasks such as relevance ranking and semantic search. However, deployment of such LLMs remains prohibitively expensive for industry applications with strict latency and throughput requirements. In this work, we present lessons and efficiency insights from developing a purely text-based decoder-only Small Language Model (SLM) for a semantic search application at LinkedIn. Particularly, we discuss model compression techniques such as pruning that allow us to reduce the model size by up to $40\%$ while maintaining the accuracy. Additionally, we present context compression techniques that allow us to reduce the input context length by up to $10$x with minimal loss of accuracy. Finally, we present practical lessons from optimizing the serving infrastructure for deploying such a system on GPUs at scale, serving millions of requests per second. Taken together, this allows us to increase our system's throughput by $10$x in a real-world deployment, while meeting our quality bar.

cs.IR↗

Concepts for designing modern C++ interfaces for MPI

Since the C++ bindings were deleted in 2008, the Message Passing Interface (MPI) community has revived efforts in building high-level modern C++ interfaces. Such interfaces are either built to serve specific scientific application needs (with limited coverage to the underlying MPI functionalities), or as an exercise in general-purpose programming model building, with the hope that bespoke interfaces can be broadly adopted to construct a variety of distributed-memory scientific applications. However, with the advent of modern C++-based heterogeneous programming models, GPUs and widespread Machine Learning (ML) usage in contemporary scientific computing, the role of prospective community-standardized high-level C++ interfaces to MPI is evolving. The success of such an interface clearly will depend on providing robust abstractions and features adhering to the generic programming principles that underpin the C++ programming language, without compromising on either performance and portability, the core principles upon which MPI was founded. However, there is a tension between idiomatic C++ handling of types and lifetimes and MPI's loose interpretation of object lifetimes/ownership and insistence on maintaining global states. Instead of proposing "yet another" high-level C++ interface to MPI, overlooking or providing partial solutions to work around the key issues concerning the dissonance between MPI semantics and idiomatic C++, this paper focuses on the three fundamental aspects of a high-level interface: type system, object lifetimes and communication buffers, also identifying inconsistencies in the MPI specification. Presumptive solutions can be unrefined, and we hope the broader MPI and C++ communities will engage with us in productive exchange of ideas and concerns.

cs.DC↗

Anonymized Network Sensing using C++26 std::execution on GPUs

Large-scale network sensing plays a vital role in network traffic analysis and characterization. As network packet data grows increasingly large, parallel methods have become mainstream for network analytics. While effective, GPU-based implementations still face start-up challenges in host-device memory management and porting complex workloads on devices, among others. To mitigate these challenges, composable frameworks have emerged using modern C++ programming language, for efficiently deploying analytics tasks on GPUs. Specifically, the recent C++26 Senders model of asynchronous data operation chaining provides a simple interface for bulk pushing tasks to varied device execution contexts. Considering the prominence of contemporary dense-GPU platforms and vendor-leveraged software libraries, such a programming model consider GPUs as first-class execution resources (compared to traditional host-centric programming models), allowing convenient development of multi-GPU application workloads via expressive and standardized asynchronous semantics. In this paper, we discuss practical aspects of developing the Anonymized Network Sensing Graph Challenge on dense-GPU systems using the recently proposed C++26 Senders model. Adopting a generic and productive programming model does not necessarily impact the critical-path performance (as compared to low-level proprietary vendor-based programming models): our commodity library-based implementation achieves up to 55x performance improvements on 8x NVIDIA A100 GPUs as compared to the reference serial GraphBLAS baseline.

cs.DC↗

Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs

Language models can be sampled multiple times to access the distribution underlying their responses, but existing methods cannot efficiently synthesize rich epistemic signals across different long-form responses. We introduce Consensus Graphs (ConGrs), a flexible DAG-based data structure that represents shared information, as well as semantic variation in a set of sampled LM responses to the same prompt. We construct ConGrs using a light-weight lexical sequence alignment algorithm from bioinformatics, supplemented by the targeted usage of a secondary LM judge. Further, we design task-dependent decoding methods to synthesize a single, final response from our ConGr data structure. Our experiments show that synthesizing responses from ConGrs improves factual precision on two biography generation tasks by up to 31% over an average response and reduces reliance on LM judges by more than 80% compared to other methods. We also use ConGrs for three refusal-based tasks requiring abstention on unanswerable queries and find that abstention rate is increased by up to 56%. We apply our approach to the MATH and AIME reasoning tasks and find an improvement over self-verification and majority vote baselines by up to 6 points of accuracy. We show that ConGrs provide a flexible method for capturing variation in LM responses and using the epistemic signals provided by response variation to synthesize more effective responses.

cs.CL↗

Non-Homogeneous Generalized Fractional Skellam Process

This paper introduces the Non-homogeneous Generalized Skellam process (NGSP) and its fractional version NGFSP by time changing it with an independent inverse stable subordinator. We study distributional properties for NGSP and NGFSP including probability generating function, probability mass function (p.m.f.), factorial moments, mean, variance, covariance and correlation structure. Then we investigate the long and short range dependence structures for NGSP and NGFSP, and obtain the governing state differential equations of these processes along with their increment processes. We obtain recurrence relations satisfied by the state probabilities of Non-homogeneous generalized counting process (NGCP), NGSP and NGFSP. The weighted sum representations for these processes are provided. We further obtain martingale and renewal properties along with arrival time distribution for NGSP and NGFSP. An alternative version of NGFSP with a closed-form p.m.f. is introduced along with a discussion of its distributional and asymptotic properties. In addition, we study the running average processes of GCP and GSP which are of independent interest. The p.m.f. of NGSP and the simulated sample paths of GFSP, NGFSP and related processes are plotted. Finally, we discuss an application to a high-frequency financial data set pointing out the advantages of our model compared to existing ones.

math.PR↗

Subconvexity for Rankin Selberg L-Functions at Special Points

Let $f$ and $g$ be normalized Hecke-Maass cusp forms for the full modular group having spectral parameters $t_f$ and $t_g$ respectively with $t_f,t_g\asymp T\rightarrow \infty $. In this paper we show that the Rankin Selberg $L$-function associated to the pair $(f,g)$ at the special points $t=\pm(t_f+t_g)$, satisfies the subconvex bound \begin{align*} L\left(\frac{1}{2}+it,f\otimes g\right)\ll_{\eps} T^{61/84+\eps}. \end{align*} Additionally at the points $t=\pm(t_f-t_g)\asymp T^ν$ with $2/3+\eps<ν\leq 1$ we show the subconvex bound \begin{align*} L(1/2+it,f\otimes g)\ll_\eps {T^{7/12+ν/8+\eps}}, \; \text{if }\; 2/3+\eps< ν\leq 14/17, \end{align*} and \begin{align*} L(1/2+it,f\otimes g)\ll_\eps {T^{1/2+19ν/84+\eps}}, \; \text{if }\; 14/17\leq ν\leq 1. \end{align*} With the above results we are able to address the subconvexity problem in the spectral aspect for $GL(2)\times GL(2)$ Rankin Selberg $L$-functions when the parameters of both the forms vary under the additional challenge of a considerable amount conductor dropping occurring due to the special points in question.

math.NT↗

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Instructed Visual Segmentation (IVS) tasks require segmenting objects in images or videos based on natural language instructions. While recent multimodal large language models (MLLMs) have achieved strong performance on IVS, their inference cost remains a major bottleneck, particularly in video. We empirically analyze visual token sampling in MLLMs and observe a strong correlation between subset token coverage and segmentation performance. This motivates our design of a simple and effective token pruning method that selects a compact yet spatially representative subset of tokens to accelerate inference. In this paper, we introduce a novel visual token pruning method for IVS, called EVTP-IV, which builds upon the k-center by integrating spatial information to ensure better coverage. We further provide an information-theoretic analysis to support our design. Experiments on standard IVS benchmarks show that our method achieves up to 5X speed-up on video tasks and 3.5X on image tasks, while maintaining comparable accuracy using only 20% of the tokens. Our method also consistently outperforms state-of-the-art pruning baselines under varying pruning ratios.

cs.CV↗

ApproxJoin: Approximate Matching for Efficient Verification in Fuzzy Set Similarity Join

The set similarity join problem is a fundamental problem in data processing and discovery, relying on exact similarity measures between sets. In the presence of alterations, such as misspellings on string data, the fuzzy set similarity join problem instead approximately matches pairs of elements based on the maximum weighted matching of the bipartite graph representation of sets. State-of-the-art methods within this domain improve performance through efficient filtering methods within the filter-verify framework, primarily to offset high verification costs induced by the usage of the Hungarian algorithm - an optimal matching method. Instead, we directly target the verification process to assess the efficacy of more efficient matching methods within candidate pair pruning. We present ApproxJoin, the first work of its kind in applying approximate maximum weight matching algorithms for computationally expensive fuzzy set similarity join verification. We comprehensively test the performance of three approximate matching methods: the Greedy, Locally Dominant and Paz Schwartzman methods, and compare with the state-of-the-art approach using exact matching. Our experimental results show that ApproxJoin yields performance improvements of 2-19x the state-of-the-art with high accuracy (99% recall).

cs.DB↗

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection

Evaluating learned robot control policies to determine their physical task-level capabilities costs experimenter time and effort. The growing number of policies and tasks exacerbates this issue. It is impractical to test every policy on every task multiple times; each trial requires a manual environment reset, and each task change involves re-arranging objects or even changing robots. Naively selecting a random subset of tasks and policies to evaluate is a high-cost solution with unreliable, incomplete results. In this work, we formulate robot evaluation as an active testing problem. We propose to model the distribution of robot performance across all tasks and policies as we sequentially execute experiments. Tasks often share similarities that can reveal potential relationships in policy behavior, and we show that natural language is a useful prior in modeling these relationships between tasks. We then leverage this formulation to reduce the experimenter effort by using a cost-aware expected information gain heuristic to efficiently select informative trials. Our framework accommodates both continuous and discrete performance outcomes. We conduct experiments on existing evaluation data from real robots and simulations. By prioritizing informative trials, our framework reduces the cost of calculating evaluation metrics for robot policies across many tasks.

cs.RO↗

Interpretable Multi-Source Data Fusion Through Latent Variable Gaussian Process

With the advent of artificial intelligence and machine learning, various domains of science and engineering communities have leveraged data-driven surrogates to model complex systems through fusing numerous sources of information (data) from published papers, patents, open repositories, or other resources. However, not much attention has been paid to the differences in quality and comprehensiveness of the known and unknown underlying physical parameters of the information sources, which could have downstream implications during system optimization. Additionally, existing methods cannot fuse multi-source data into a single predictive model. Towards resolving this issue, a multi-source data fusion framework based on Latent Variable Gaussian Process (LVGP) is proposed. The individual data sources are tagged as a characteristic categorical variable that are mapped into a physically interpretable latent space, allowing the development of source-aware data fusion modeling. Additionally, a dissimilarity metric based on the latent variables of LVGP is introduced to study and understand the differences in the sources of data. The proposed approach is demonstrated on and analyzed through two mathematical and two materials science case studies. From the case studies, it is observed that compared to using single-source and source unaware machine learning models, the proposed multi-source data fusion framework can provide better predictions for sparse-data problems.

stat.ML↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learning on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized continuous prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15-40% improvement in end-to-end training performance on the National Energy Research Scientific Computing Center's (NERSC) Perlmutter supercomputer for various OGB datasets.

cs.DC↗

Compare without Despair: Reliable Preference Evaluation with Generation Separability

Human evaluation of generated language through pairwise preference judgments is pervasive. However, under common scenarios, such as when generations from a model pair are very similar, or when stochastic decoding results in large variations in generations, it results in inconsistent preference ratings. We address these challenges by introducing a meta-evaluation measure, separability, which estimates how suitable a test instance is for pairwise preference evaluation. For a candidate test instance, separability samples multiple generations from a pair of models, and measures how distinguishable the two sets of generations are. Our experiments show that instances with high separability values yield more consistent preference ratings from both human- and auto-raters. Further, the distribution of separability allows insights into which test benchmarks are more valuable for comparing models. Finally, we incorporate separability into ELO ratings, accounting for how suitable each test instance might be for reliably ranking LLMs. Overall, separability has implications for consistent, efficient and robust preference evaluation of LLMs with both human- and auto-raters.

cs.CL↗

Quantum optomechanical control of long-lived bulk acoustic phonons

High-fidelity quantum optomechanical control of a mechanical oscillator requires the ability to perform efficient, low-noise operations on long-lived phononic excitations. Microfabricated high-overtone bulk acoustic wave resonators ($\mathrmμ$HBARs) have been shown to support high-frequency (> 10 GHz) mechanical modes with exceptionally long coherence times (> 1.5 ms), making them a compelling resource for quantum optomechanical experiments. In this paper, we demonstrate a new optomechanical system that permits quantum optomechanical control of individual high-coherence phonon modes supported by such $\mathrmμ$HBARs for the first time. We use this system to perform laser cooling of such ultra-massive (7.5 $\mathrmμ$g) high frequency (12.6 GHz) phonon modes from an occupation of ${\sim}$22 to fewer than 0.4 phonons, corresponding to laser-based ground-state cooling of the most massive mechanical object to date. Through these laser cooling experiments, no absorption-induced heating is observed, demonstrating the resilience of the $\mathrmμ$HBAR against parasitic heating. The unique features of such $\mathrmμ$HBARs make them promising as the basis for a new class of quantum optomechanical systems that offer enhanced robustness to decoherence, necessary for efficient, low-noise photon-phonon conversion.

quant-ph↗

Final Report for CHESS: Cloud, High-Performance Computing, and Edge for Science and Security

Automating the theory-experiment cycle requires effective distributed workflows that utilize a computing continuum spanning lab instruments, edge sensors, computing resources at multiple facilities, data sets distributed across multiple information sources, and potentially cloud. Unfortunately, the obvious methods for constructing continuum platforms, orchestrating workflow tasks, and curating datasets over time fail to achieve scientific requirements for performance, energy, security, and reliability. Furthermore, achieving the best use of continuum resources depends upon the efficient composition and execution of workflow tasks, i.e., combinations of numerical solvers, data analytics, and machine learning. Pacific Northwest National Laboratory's LDRD "Cloud, High-Performance Computing (HPC), and Edge for Science and Security" (CHESS) has developed a set of interrelated capabilities for enabling distributed scientific workflows and curating datasets. This report describes the results and successes of CHESS from the perspective of open science.

cs.DC↗

Frustrated spin-1/2 Heisenberg model on a Kagome-strip chain: Dimerization and mapping to a spin-orbital Kugel-Khomskii model

We investigate the quantum phases of a frustrated antiferromagnetic Heisenberg spin-1/2 model Hamiltonian on a Kagome-strip chain (KSC), a one-dimensional analogue of the Kagome lattice, and construct its phase diagram in an extended exchange parameter space. The isolated unit cell of this lattice comprises of five spin-1/2 particles, giving rise to several types of magnetic ground states in a unit cell: a spin-$3/2$ state as well as spin-$1/2$ states with and without additional degeneracies. We explore the ground state properties of the fully connected thermodynamic system using exact diagonalization and density matrix renormalization group methods, identifying several distinct quantum phases. All but one of the phases exhibit gapless spin excitations. The exception is a dimerized spin-gapped phase that covers a large part of the phase diagram and includes the uniformly exchange coupled system. We argue that this phase can be understood by perturbing around the limit of decoupled unit cells where each unit cell has six degenerate ground states. We use degenerate perturbation theory to obtain an effective Hamiltonian, a Kugel-Khomskii model with an anisotropic spin-one orbital degree of freedom, which helps explain the origin of dimerization.

cond-mat.str-el↗

Deep Oscillatory Neural Network

We propose a novel, brain-inspired deep neural network model known as the Deep Oscillatory Neural Network (DONN). Deep neural networks like the Recurrent Neural Networks indeed possess sequence processing capabilities but the internal states of the network are not designed to exhibit brain-like oscillatory activity. With this motivation, the DONN is designed to have oscillatory internal dynamics. Neurons of the DONN are either nonlinear neural oscillators or traditional neurons with sigmoidal or ReLU activation. The neural oscillator used in the model is the Hopf oscillator, with the dynamics described in the complex domain. Input can be presented to the neural oscillator in three possible modes. The sigmoid and ReLU neurons also use complex-valued extensions. All the weight stages are also complex-valued. Training follows the general principle of weight change by minimizing the output error and therefore has an overall resemblance to complex backpropagation. A generalization of DONN to convolutional networks known as the Oscillatory Convolutional Neural Network is also proposed. The two proposed oscillatory networks are applied to a variety of benchmark problems in signal and image/video processing. The performance of the proposed models is either comparable or superior to published results on the same data sets.

cs.NE↗

Neutron and $\boldsymbolγ$-ray Discrimination by a Pressurized Helium-4 Based Scintillation Detector

Pressurized Helium-4 (PHe) based fast neutron scintillation detector offers an useful alternative to organic liquid-based scintillator due to its relatively low response to the $γ$-rays compared to the latter type of scintillator. In the present work, we have investigated the capabilities of a PHe detector for the detection of fast neutrons in a mixed radiation field where both the neutrons and the $γ$-rays are present. Discrimination between neutrons and $γ$-rays is achieved by using fast-slow charge integration method. We have also conducted systematic studies of the attenuation of fast neutrons and $γ$-rays by high-density polyethylene (HDPE). Additionally, the simulation analyses, conducted using GEANT4, provide detailed insights into the interactions of the radiation quanta with the PHe detector.

physics.ins-det↗