SearcharxivSearch

arXiv subjects

Arjun Singh

Publications and source records attributed to Arjun Singh.

18 recordsLinked to original sources

User Mobility Demands Near-Field Communications in Terahertz Band Wireless Networks Beyond 6G

Near-field propagation is often unavoidable at terahertz (THz) frequencies due to the large apertures needed for sufficient array gain, yet near-field operation complicates practical system design, especially under user mobility. This paper asks whether a mobile THz link can remain broadband, achieve the desired high rates and coverage, while operating exclusively in the radiative far field. To answer this question, we develop a proof-by-contradiction feasibility framework that jointly enforces (i) a far-field requirement based on the Fraunhofer distance and (ii) a reliability requirement specified by a target SNR at the worst-case link distance. We derive closed-form upper bounds on the far-field-feasible bandwidth for stationary and mobile links. We further incorporate practical misalignment through several UE rotation and mobility scenarios. Numerical results show that stationary THz links can remain far-field-only with physically realizable apertures while supporting extremely large bandwidths, whereas practical mobile THz systems cannot. In practically relevant mobile THz access settings, the far-field-feasible bandwidth becomes a severe limiting factor: achieving tens-of-GHz targets would require unrealistically high UE transmit power. A cross-band comparison further shows that far-field-only operation is largely attainable at sub-6~GHz and, to a significant extent, at mmWave for moderate bandwidths, while near-field-aware designs become essential for mobile THz access.

eess.SP

Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO

Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Existing GPU and TPU focused simulators are typically architecture-specific, while distributed training simulators rely on workload-specific analytical models or costly post-execution traces, limiting portability and cross-platform comparison. This work evaluates whether MLIR's StableHLO dialect can serve as a unified workload representation for cross-architecture and cross-fidelity performance modeling of distributed ML workloads. The study establishes a StableHLO-based simulation methodology that maps a single workload representation onto multiple performance models, spanning analytical, profiling-based, and simulator-driven predictors. Using this methodology, workloads are evaluated across GPUs and TPUs without requiring access to scaled-out physical systems, enabling systematic comparison across modeling fidelities. An empirical evaluation covering distributed GEMM kernels, ResNet, and large language model training workloads demonstrates that StableHLO preserves relative performance trends across architectures and fidelities, while exposing accuracy trade-offs and simulator limitations. Across evaluated scenarios, prediction errors remain within practical bounds for early-stage design exploration, and the methodology reveals fidelity-dependent limitations in existing GPU simulators. These results indicate that StableHLO provides a viable foundation for unified, distributed ML performance modeling across accelerator architectures and simulators, supporting reusable evaluation workflows and cross-validation throughout the ML system design process.

cs.DC

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth in model complexity far outpaces CMOS technology scaling, making sustainable and efficient system design a critical challenge. Addressing this requires coordinated co-design across software, hardware, and technology layers. Due to the prohibitive cost and complexity of deploying full-scale training systems, simulators play a pivotal role in enabling this design exploration. This survey reviews the landscape of distributed DNN training simulators, focusing on three major dimensions: workload representation, simulation infrastructure, and models for total cost of ownership (TCO) including carbon emissions. It covers how workloads are abstracted and used in simulation, outlines common workload representation methods, and includes comprehensive comparison tables covering both simulation frameworks and TCO/emissions models, detailing their capabilities, assumptions, and areas of focus. In addition to synthesizing existing tools, the survey highlights emerging trends, common limitations, and open research challenges across the stack. By providing a structured overview, this work supports informed decision-making in the design and evaluation of distributed training systems.

cs.DC

A Study of Optimizations for Fine-tuning Large Language Models

Fine-tuning large language models is a popular choice among users trying to adapt them for specific applications. However, fine-tuning these models is a demanding task because the user has to examine several factors, such as resource budget, runtime, model size and context length among others. A specific challenge is that fine-tuning is memory intensive, imposing constraints on the required hardware memory and context length of training data that can be handled. In this work, we share a detailed study on a variety of fine-tuning optimizations across different fine-tuning scenarios. In particular, we assess Gradient Checkpointing, Low-Rank Adaptation, DeepSpeed's Zero Redundancy Optimizer and FlashAttention. With a focus on memory and runtime, we examine the impact of different optimization combinations on GPU memory usage and execution runtime during fine-tuning phase. We provide our recommendation on the best default optimization for balancing memory and runtime across diverse model sizes. We share effective strategies for fine-tuning very large models with tens or hundreds of billions of parameters and enabling large context lengths during fine-tuning. Furthermore, we propose the appropriate optimization mixtures for fine-tuning under GPU resource limitations.

cs.LG

Impact of the Antenna on the Sub-Terahertz Indoor Channel Characteristics: An Experimental Approach

Terahertz-band (100 GHz-10 THz) communication is a promising radio technology envisioned to enable ultra-high data rate, reliable and low-latency wireless connectivity in next-generation wireless systems. However, the low transmission power of THz transmitters, the need for high gain directional antennas, and the complex interaction of THz radiation with common objects along the propagation path make crucial the understanding of the THz channel. In this paper, we conduct an extensive channel measurement campaign in an indoor setting (i.e., a conference room) through a channel sounder with 0.1 ns time resolution and 20 GHz bandwidth at 140 GHz. Particularly, the impact of different antenna directivities (and, thus, beam widths) on the channel characteristics is extensively studied. The experimentally obtained dataset is processed to develop the path loss model and, subsequently, derive key channel metrics such as the path loss exponent, delay spread, and K-factor. The results highlight the multi-faceted impact of the antenna gain on the channel and, by extension, the wireless system and, thus, show that an antenna-agnostic channel model cannot capture the propagation characteristics of the THz channel.

eess.SP

SymNoise: Advancing Language Model Fine-tuning with Symmetric Noise

In this paper, we introduce a novel fine-tuning technique for language models, which involves incorporating symmetric noise into the embedding process. This method aims to enhance the model's function by more stringently regulating its local curvature, demonstrating superior performance over the current method, NEFTune. When fine-tuning the LLaMA-2-7B model using Alpaca, standard techniques yield a 29.79% score on AlpacaEval. However, our approach, SymNoise, increases this score significantly to 69.04%, using symmetric noisy embeddings. This is a 6.7% improvement over the state-of-the-art method, NEFTune~(64.69%). Furthermore, when tested on various models and stronger baseline instruction datasets, such as Evol-Instruct, ShareGPT, OpenPlatypus, SymNoise consistently outperforms NEFTune. The current literature, including NEFTune, has underscored the importance of more in-depth research into the application of noise-based strategies in the fine-tuning of language models. Our approach, SymNoise, is another significant step towards this direction, showing notable improvement over the existing state-of-the-art method.

cs.CL

Design and Validation of a Metallic Reflectarray for Communications at True Terahertz Frequencies

Wireless communications in the terahertz band (0.1-10 THz) is a promising and key wireless technology enabling ultra-high data rate communication over multi-gigahertz-wide bandwidths, thus fulfilling the demand for denser networks. The complex propagation environment at such high frequencies introduces several challenges, such as high spreading and molecular absorption losses. As such, intelligent reflecting surfaces have been proposed as a promising solution to enable communication in the presence of blockage or to aid a resource-limited quasi-omnidirectional transmitter direct its radiated power. In this paper, we present a metallic reflectarray design achieving controlled non-specular reflection at true terahertz frequencies (i.e., 1-1.05 THz). We conduct extensive experiments to further characterize and validate its working principle using terahertz time-domain spectroscopy and demonstrate its effectiveness with information-carrying signals using a continuous-wave terahertz testbed. Our results show that the reflectarray can help facilitate robust communication links over non-specular paths and improve the reliability of terahertz communications, thereby unleashing the true potential of the terahertz band.

eess.SY

Near-field 6G Networks: Why Mobile Terahertz Communications MUST Operate in the Near Field

Near-field mobile terahertz (THz) communications is one of the candidate enablers for high-rate wireless data exchange in sixth-generation (6G) networks. However, operating in the THz near field brings both attractive opportunities and severe challenges. Hence, it becomes of interest to explore if it is possible to design a realistic mobile THz communication system without working in the THz near field. To answer this question, a mathematical framework is presented modeling a mobile THz link that works exclusively in the far field. The study leads to an interesting theoretical conclusion: while the actual frequency is of (almost) no interest, such a system must operate over a limited bandwidth not exceeding a certain threshold. It is then numerically shown that operating only in the far field imposes stringent limitations on mobile THz communications, thus making them less attractive to prospective high-rate services. In contrast, it is shown that a stationary THz link can still be broadband even when staying exclusively in the THz far field. Hence, broadband mobile THz communications MUST be near-field, while broadband stationary THz links do not have to.

cs.NI

Wavefront Engineering: Realizing Efficient Terahertz Band Communications in 6G and Beyond

Terahertz (THz) band communications is envisioned as a key technology for future wireless standards. Substantial progress has been made in this field, with advances in hardware design, channel models, and signal processing. High-rate backhaul links operating at sub-THz frequencies have been experimentally demonstrated. However, there are inherent challenges in making the next great leap for adopting the THz band in widespread communication systems, such as cellular access and wireless local area networks. Primarily, such systems have to be both: (i) wideband, to maintain desired data rate and sensing resolution; and, more importantly, (ii) operate in the massive near field of the high-gain devices required to overcome the propagation losses. In this article, it is first explained why the state-of-the-art techniques from lower frequencies, including millimeter-wave, cannot be simply repurposed to realize THz band communication systems. Then, a vision of wavefront engineering is presented to address these shortfalls. Further, it is illustrated how novel implementations of specific wavefronts, such as Bessel beams and Airy beams, offer attractive advantages in creating THz links over state-of-the-art far-field beamforming and near-field beamfocusing techniques. The paper ends by discussing novel problems and challenges in this new and exciting research area. Index Terms - Terahertz Communications; 6G; Wavefront Engineering; Bessel beams; Near field; Orbital Angular Momentum

eess.SY

Mission Apollo: Landing Optical Circuit Switching at Datacenter Scale

In this paper, we describe Apollo, to the best of our knowledge, the world's first large-scale production deployment of optical circuit switches (OCSes) for datacenter networking. We will first describe the infrastructure challenges and use cases that motivated optical switching inside datacenters. We then delve into the requirements of OCSes for datacenter applications: balancing cost, port count, switching time, and optical performance, which drive design choices and implementation details of our internally developed 3D MEMS-based OCS. To enable the Apollo optical switching layer, we employ circulators to realize bidirectional links through the OCS, effectively doubling the OCS radix. The OCS and circulator design choices were critical for meeting network bandwidth, scale, and cost targets. We review the critical co-design of WDM transceiver technology for these OCS plus circulator-based bidirectional links and their corresponding physical impairments, delivered over four generations/speeds of optical interconnect. Finally, we conclude with thoughts on future directions in hardware development and associated applications.

cs.NI

Structural investigation of Ayurveda Lauha (Iron) Bhasma

In Ayurveda, Lauha (Iron) bhasma is primarily used to cure diseases related with iron deficiency in humans. It is produced from purified raw metallic iron using a combination of multi-step traditional preparation processes described in the Ayurveda literature. Here, we present results of structural investigation performed on the medicinal grade Lauha bhasma using various X-ray based techniques. Our results indicate that after several rounds of heating and cooling in specific conditions following the Ayurvedic preparation procedure, metallic iron eventually converts to a natural iron-oxide mineral belonging to the Magnetite group. Scanning electron microscopy (SEM) and X-ray standing wave assisted fluorescence measurements carried out on powdered bhasma specimen reveal that the Magnetite micro-particles in the bhasma specimen are usually present in the form of agglomerates of nano-particles. We anticipate that the Ayurvedic Lauha Bhasma has great potential for noninvasive localized target killing of cancer cells, particularly in sensitive parts of the human body such as brain, spinal cord, and lungs, via necrosis by application of an alternating external magnetic field or photo electron generation through X-rays.

cond-mat.mtrl-sci

Numerical Investigation of Chemically Reacting Rarefied Hypersonic Flows Over A Backward-Facing Step

The present study employs the Direct Simulation Monte Carlo (DSMC) method, a widely used numerical method for studying non-continuum flows, to investigate the flow characteristics of a chemically reacting hypersonic flow over a backward-facing step. This work explores the effects of the Knudsen number on the behaviour of rarefied flow and compares the results with the non-reacting flow. The primary objective of the present thesis is to elucidate the variation of temperature in the flow domain with Knudsen number. The results of the study show that chemical reactions affect the temperature distribution in the flow domain significantly.

physics.flu-dyn

Coexistence and Spectrum Sharing Above 100 GHz

[...] This paper explores how spectrum policy and spectrum technologies can evolve to enable sharing among different stakeholders in the above 100 GHz spectrum, without introducing harmful interference or disrupting either security applications or fundamental science exploration. This portion of the spectrum presents new opportunities to design spectrum sharing schemes, based on novel antenna designs, directional ultra-high-rate communications, and active/passive user coordination. The paper provides a tutorial on current regulations above 100 GHz, and highlights how sharing is central to allowing each stakeholder to make the most out of this spectrum. It then defines - through detailed simulations based on standard International Telecommunications Union (ITU) channel and antenna models - scenarios in which active users may introduce harmful interference to passive sensing. Based on this evaluation, it reviews a number of promising techniques that can enable active/passive sharing above 100 GHz. The critical review and tutorial on policy and technologies of this paper have the potential to kickstart future research and regulations that promote safe coexistence between active and passive users above 100 GHz, further benefiting the development of digital technologies and scientific exploration.

cs.NI

Simple RGC: ImageJ plugins for counting retinal ganglion cells and determining the transduction efficiency of viral vectors in retinal wholemounts

Simple RGC consists of a collection of ImageJ plugins to assist researchers investigating retinal ganglion cell (RGC) injury models in addition to helping assess the effectiveness of treatments. The first plugin named RGC Counter accurately calculates the total number of RGCs from retinal wholemount images. The second plugin named RGC Transduction measures the co-localisation between two channels making it possible to determine the transduction efficiencies of viral vectors and transgene expression levels. The third plugin named RGC Batch is a batch image processor to deliver fast analysis of large groups of microscope images. These ImageJ plugins make analysis of RGCs in retinal wholemounts easy, quick, consistent, and less prone to unconscious bias by the investigator. The plugins are freely available from the ImageJ update site https://sites.imagej.net/Sonjoonho/.

q-bio.NC

Pretraining Federated Text Models for Next Word Prediction

Federated learning is a decentralized approach for training models on distributed devices, by summarizing local changes and sending aggregate parameters from local models to the cloud rather than the data itself. In this research we employ the idea of transfer learning to federated training for next word prediction (NWP) and conduct a number of experiments demonstrating enhancements to current baselines for which federated NWP models have been successful. Specifically, we compare federated training baselines from randomly initialized models to various combinations of pretraining approaches including pretrained word embeddings and whole model pretraining followed by federated fine tuning for NWP on a dataset of Stack Overflow posts. We realize lift in performance using pretrained embeddings without exacerbating the number of required training rounds or memory footprint. We also observe notable differences using centrally pretrained networks, especially depending on the datasets used. Our research offers effective, yet inexpensive, improvements to federated NWP and paves the way for more rigorous experimentation of transfer learning techniques for federated learning.

cs.LG

Benchmarking in Manipulation Research: The YCB Object and Model Set and Benchmarking Protocols

In this paper we present the Yale-CMU-Berkeley (YCB) Object and Model set, intended to be used to facilitate benchmarking in robotic manipulation, prosthetic design and rehabilitation research. The objects in the set are designed to cover a wide range of aspects of the manipulation problem; it includes objects of daily life with different shapes, sizes, textures, weight and rigidity, as well as some widely used manipulation tests. The associated database provides high-resolution RGBD scans, physical properties, and geometric models of the objects for easy incorporation into manipulation and planning software platforms. In addition to describing the objects and models in the set along with how they were chosen and derived, we provide a framework and a number of example task protocols, laying out how the set can be used to quantitatively evaluate a range of manipulation approaches including planning, learning, mechanical design, control, and many others. A comprehensive literature survey on existing benchmarks and object datasets is also presented and their scope and limitations are discussed. The set will be freely distributed to research groups worldwide at a series of tutorials at robotics conferences, and will be otherwise available at a reasonable purchase cost. It is our hope that the ready availability of this set along with the ground laid in terms of protocol templates will enable the community of manipulation researchers to more easily compare approaches as well as continually evolve benchmarking tests as the field matures.

cs.RO

Mapping Cloud Computing onto Useful e-Governance

Most of the services viewed in context to grid and cloud computing are mostly confined to services that are available for intellectual purposes. The grid or cloud computing are large scale distributed systems. The essence of large scale distribution can only be realized if the services are rendered to common man. The only organization which has exposure to almost every single resident is the respective governments in every country. As the size of population increases so the need for a larger purview arises. The problem of having a large purview can be solved by means of large scale grid for online services. The government services can be rendered through fully customized Service-oriented Clouds. In this paper we are presenting tight similarities between generic government functioning and the service oriented grid/cloud approach. Also, we will discuss the major issues in establishing services oriented grids for governmental organization.

cs.DC

Applying l-Diversity in anonymizing collaborative social network

To date publish of a giant social network jointly from different parties is an easier collaborative approach. Agencies and researchers who collect such social network data often have a compelling interest in allowing others to analyze the data. In many cases the data describes relationships that are private and sharing the data in full can result in unacceptable disclosures. Thus, preserving privacy without revealing sensitive information in the social network is a serious concern. Recent developments for preserving privacy using anonymization techniques are focused on relational data only. Preserving privacy in social networks against neighborhood attacks is an initiation which uses the definition of privacy called k-anonymity. k-anonymous social network still may leak privacy under the cases of homogeneity and background knowledge attacks. To overcome, we find a place to use a new practical and efficient definition of privacy called ldiversity. In this paper, we take a step further on preserving privacy in collaborative social network data with algorithms and analyze the effect on the utility of the data for social network analysis.

cs.CY