SearcharxivSearch

arXiv subjects

Muhammad Talha

Publications and source records attributed to Muhammad Talha.

13 recordsLinked to original sources

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build because object masks, semantic labels, and paired edits are rarely available at scale. We introduce SatEdit, a mask-conditioned satellite image editing framework that constructs training supervision from unlabeled imagery. SatEdit proposes object masks with a seg- mentation foundation model, assigns semantic la- bels to sampled segments with a Vision-Language Model, and applies lightweight human verification before generating paired addition and removal exam- ples through mask-guided inpainting. We fine-tune a high-resolution image editing backbone with LoRA on a SODA-A-derived dataset containing 1,014 im- ages and 852 verified object annotations across 91 classes. In controlled comparisons with open- source and proprietary image editing models, SatE- dit achieves the highest aggregate masked-region se- mantic alignment, with a CLIP score of 0.6322 and CLIP delta of 0.0726, while preserving the surround- ing scene qualitatively. These results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.

cs.CV

SplatStream: Fine Granular Scalable Gaussian Splatting for Adaptive 3D Scene Streaming

Dynamic 3D Gaussian Splatting (GS) enables high quality real-time rendering for immersive media, but its large representation size and frame-wise redundancy create significant challenges for adaptive streaming. This paper presents SplatStream, a fine granular scalable Gaussian splatting framework for dynamic 3D scene delivery. The proposed method decompose the GS scenes into quality and resolution layers, and introduces inter-layer predictive coding to achieve scalability. For temporal direction, B-frames are introduced to have temporal quality scalability. A lightweight cross-layer transformer based predictor is utilized for both cross layer and temporal predictions. In addition, a volume-opacity based importance measure is used for fine-grained Gaussian packetization, allowing visually important primitives to be transmitted earlier for progressive refinement. Finally, the scalable GS bitstream is mapped to an MPEG-DASH compatible sub-representation structure, enabling fine granular adaptive, low-latency delivery of dynamic Gaussian splatting content under bandwidth-varying conditions.

eess.IV

Quantum Reservoir Computing: Recent Advances and Future Directions

Quantum reservoir computing (QRC) uses the dynamics of a fixed or weakly tuned quantum system to transform temporal and sequential inputs into measured features, while training is typically confined to a classical readout. This separation reduces reliance on repeated quantum parameter updates and avoids the barren plateaus associated with variational circuit training. Its computational power is often attributed to the exponentially large Hilbert space of the quantum system. However, the memory, nonlinearity, and expressivity that determine what a reservoir can actually compute depend jointly on the input encoding, quantum evolution, observables, measurement, and readout, not on Hilbert space dimension alone. On hardware, these capabilities are further constrained by finite sampling, hardware noise, measurement backaction, and the cost of estimating observables, so a large state space alone does not guarantee useful computation. In this survey, we develop a common system model that connects these components and use it to organize QRC foundations, computational properties, reservoir architectures, operating protocols, and physical implementations. We examine spin, photonic, superconducting, bosonic, neutral atom, and other analog platforms, together with applications, software and high performance computing support, benchmarking, and reproducibility. The analysis distinguishes hardware demonstrations from simulations and identifies the assumptions and resources that govern comparisons across implementations. Current results do not establish a broad quantum advantage over well matched classical reservoirs. We therefore specify the resource accounting, benchmark standards, and theoretical criteria needed to evaluate claims of quantum advantage.

quant-ph

IceWatch: Forecasting Glacial Lake Outburst Floods (GLOFs) using Multimodal Deep Learning

Glacial Lake Outburst Floods (GLOFs) pose a serious threat in high mountain regions. They are hazardous to communities, infrastructure, and ecosystems further downstream. The classical methods of GLOF detection and prediction have so far mainly relied on hydrological modeling, threshold-based lake monitoring, and manual satellite image analysis. These approaches suffer from several drawbacks: slow updates, reliance on manual labor, and losses in accuracy when clouds interfere and/or lack on-site data. To tackle these challenges, we present IceWatch: a novel deep learning framework for GLOF prediction that incorporates both spatial and temporal perspectives. The vision component, RiskFlow, of IceWatch deals with Sentinel-2 multispectral satellite imagery using a CNN-based classifier and predicts GLOF events based on the spatial patterns of snow, ice, and meltwater. Its tabular counterpart confirms this prediction by considering physical dynamics. TerraFlow models glacier velocity from NASA ITS_LIVE time series while TempFlow forecasts near-surface temperature from MODIS LST records; both are trained on long-term observational archives and integrated via harmonized preprocessing and synchronization to enable multimodal, physics-informed GLOF prediction. Both together provide cross-validation, which will improve the reliability and interpretability of GLOF detection. This system ensures strong predictive performance, rapid data processing for real-time use, and robustness to noise and missing information. IceWatch paves the way for automatic, scalable GLOF warning systems. It also holds potential for integration with diverse sensor inputs and global glacier monitoring activities.

cs.LG

Optimizing ISAC MIMO Systems with Reconfigurable Pixel Antennas

The integration of sensing and communication demands architectures that can flexibly exploit spatial and electromagnetic (EM) degrees of freedom (DoF). This paper proposes an Integrated Sensing and Communication (ISAC) MIMO framework that uses Reconfigurable Pixel Antenna (RPixA), which introduces additional EM-domain DoF that are electronically controlled through binary antenna coder switch networks. We introduce a beamforming architecture combining this EM and digital precoding to jointly optimize Sensing and Communication. Based on full-wave simulation of pixel antenna, we formulate a non-convex joint optimization problem to maximize sensing rate under user-specific constraints on communication rate. We utilize an Alternating Optimization framework incorporating genetic algorithm for port states of Pixel antennas, and semi-definite relaxation (SDR) for digital beamforming. Numerical results demonstrate that the proposed EM-aware design achieves considerably higher sensing rate compared to conventional arrays and enables considerable antenna reduction for equivalent ISAC performance. These findings highlight the potential of reconfigurable pixel antennas to realize efficient and scalable EM-aware ISAC systems for future 6G networks.

eess.SP

J-SGFT: Joint Spatial and Graph Fourier Domain Learning for Point Cloud Attribute Deblocking

Point clouds (PC) are essential for AR/VR and autonomous driving but challenge compression schemes with their size, irregular sampling, and sparsity. MPEG's Geometry-based Point Cloud Compression (GPCC) methods successfully reduce bitrate; however, they introduce significant blocky artifacts in the reconstructed point cloud. We introduce a novel multi-scale postprocessing framework that fuses graph-Fourier latent attribute representations with sparse convolutions and channel-wise attention to efficiently deblock reconstructed point clouds. Against the GPCC TMC13v14 baseline, our approach achieves BD-rate reduction of 18.81\% in the Y channel and 18.14\% in the joint YUV on the 8iVFBv2 dataset, delivering markedly improved visual fidelity with minimal overhead.

eess.IV

Towards Quantum Enhanced Adversarial Robustness with Rydberg Reservoir Learning

Quantum reservoir computing (QRC) leverages the high-dimensional, nonlinear dynamics inherent in quantum many-body systems for extracting spatiotemporal patterns in sequential and time-series data with minimal training overhead. Although QRC inherits the expressive capabilities associated with quantum encodings, recent studies indicate that quantum classifiers based on variational circuits remain susceptible to adversarial perturbations. In this perspective, we investigate the first systematic evaluation of adversarial robustness in a QRC based learning model. Our reservoir comprises an array of strongly interacting Rydberg atoms governed by a fixed Hamiltonian, which naturally evolves under complex quantum dynamics, producing high-dimensional embeddings. A lightweight multilayer perceptron serves as the trainable readout layer. We utilize the balanced datasets, namely MNIST, Fashion-MNIST, and Kuzushiji-MNIST, as a benchmark for rigorously evaluating the impact of augmenting the quantum reservoir with a Multilayer perceptron (MLP) in white-box adversarial attacks to assess its robustness. We demonstrate that this approach yields significantly higher accuracy than purely classical models across all perturbation strengths tested. This hybrid approach reveals a new source of quantum advantage and provides practical guidance for the secure deployment of machine learning models on quantum-centric supercomputing with near-term hardware.

quant-ph

Full Duplex ISAC with Cluster Ray Targets: Parameter Estimation and Beamforming

This work studies a full-duplex integrated sensing and communication (ISAC) resolution framework for spatially distributed systems. Conventional high-resolution methods, such as MUSIC, fail to localize distributed targets because the signal subspace is full rank, even in the single-distributed-target setting. In an effort to resolve this, we propose a two-stage estimator, which successfully resolve multiple distributed targets and outperforms several baseline schemes without incurring any additional computational complexity. Our first-stage estimator uses the Fast Fourier transform to estimate the coarse spectrum, while in the second stage, we apply the Gauss-Newton method to fine-tune the angular estimates. Apart from this, we also propose an optimization framework for designing an adaptive beamformer capable of synthesizing both wide and directed beams to cover the full extent of the targets while also fulfilling data rate requirements of multiple users. The beamformer also meets the data-rate requirements of multiple users, maintaining quality of service. Simulation results demonstrate a threefold improvement in spread estimation under low signal-to-noise ratio (SNR) conditions and a twofold improvement for low-spread targets.

eess.SP

Lowering Barriers to CAD Adoption: A Comparative Study of Augmented Reality-Based CAD (AR-CAD) and a Traditional CAD tool

The paper presents a comparative user study between an Augmented Reality-based Computer-Aided Design (AR-CAD) system and a traditional computer-based CAD modeling software, SolidWorks. Twenty participants of varying skill levels performed 3D modeling tasks using both systems. The results showed that while the average task completion time is comparable for both groups, novice designers had a higher completion rate in AR-CAD than in the traditional CAD interface, and experienced designers had a similar completion rate in both systems. A statistical comparison of task completion rate, time, and NASA Task Load Index (TLX) showed that AR-CAD slightly reduced cognitive load while favoring a high task completion rate. Higher scores on the System Usability Scale (SUS) by novices indicated that AR-CAD was superior and worthwhile for reducing barriers to entering CAD. In contrast, the Traditional CAD interface was favored by experienced users for its advanced capabilities, while many viewed AR-CAD as a valid means for rapid concept development, education, and an initial critique of designs. This opens up the need for future research on the needed refinement of AR-CAD with a focus on high-precision input tools and its evaluation of complex design processes. This research highlights the potential for immersive interfaces to enhance design practice, bridging the gap between novice and experienced CAD users.

cs.HC

GLOFNet -- A Multimodal Dataset for GLOF Monitoring and Prediction

Glacial Lake Outburst Floods (GLOFs) are rare but destructive hazards in high mountain regions, yet predictive research is hindered by fragmented and unimodal data. Most prior efforts emphasize post-event mapping, whereas forecasting requires harmonized datasets that combine visual indicators with physical precursors. We present GLOFNet, a multimodal dataset for GLOF monitoring and prediction, focused on the Shisper Glacier in the Karakoram. It integrates three complementary sources: Sentinel-2 multispectral imagery for spatial monitoring, NASA ITS_LIVE velocity products for glacier kinematics, and MODIS Land Surface Temperature records spanning over two decades. Preprocessing included cloud masking, quality filtering, normalization, temporal interpolation, augmentation, and cyclical encoding, followed by harmonization across modalities. Exploratory analysis reveals seasonal glacier velocity cycles, long-term warming of ~0.8 K per decade, and spatial heterogeneity in cryospheric conditions. The resulting dataset, GLOFNet, is publicly available to support future research in glacial hazard prediction. By addressing challenges such as class imbalance, cloud contamination, and coarse resolution, GLOFNet provides a structured foundation for benchmarking multimodal deep learning approaches to rare hazard prediction.

cs.CV

Exploring Chalcogen Influence on Sc2BeX4 (X = S, Se) for Green Energy Applications Using DFT

We present a first-principles density functional theory study of the structural, electronic, optical, and thermoelectric properties of Sc2BeX4 (X = S, Se) chalcogenides for energy applications. Both compounds are dynamically and thermodynamically stable, exhibiting negative formation energies of -2.6 eV (Sc2BeS4) and -2.2 eV (Sc2BeSe4). They feature direct band gaps of 1.8 eV and 1.2 eV, respectively, within the TB-mBJ approximation, indicating strong visible-light absorption. Optical analysis reveals high static dielectric constants (9.0 for S and 16.5 for Se), absorption peaks near 13.5 eV, and reflectivity below 30 percent. Thermoelectric calculations predict p-type conduction with Seebeck coefficients reaching 2.5e-4 V/K and electrical conductivities of 2.45e18 and 1.91e18 (Ohm m s)^-1 at 300 K. Power factors approach 1.25e11 W/K^2 m s, with a maximum dimensionless figure of merit (ZT) of 0.80 at 800 K. Calculated Debye temperatures (420 K for Sc2BeS4 and 360 K for Sc2BeSe4) imply low lattice thermal conductivity. These findings establish Sc2BeX4 chalcogenides as promising materials for photovoltaic and thermoelectric applications.

cond-mat.mtrl-sci

Multi-Target Two-way Integrated Sensing and Communications with Full Duplex MIMO Radios

In this paper, we propose a multiple input multiple output (MIMO) Full-Duplex Integrated Sensing and Communication System consisting of multiple targets, a single downlink, and a single uplink user. We employed signal-to-interference plus noise ratio (SINR) as the performance metric for radar, downlink, and uplink communication. We use a communication-centric approach in which communication waveform is used for both communication and sensing of the environment. We develop a sensing algorithm capable of estimating the direction of arrival (DoA), range, and velocity of each target. We also propose a joint optimization framework for designing A/D transmit and receive beamformers to improve radar, downlink, and uplink SINRs while minimizing self-interference (SI) leakage. We also propose a null space projection (NSP) based approach to improve the uplink rate. Our simulation results, considering orthogonal frequency division multiplexing (OFDM) waveform, show accurate radar parameter estimation with improved downlink and uplink rate.

eess.SP

STAR-RIS-Assisted Hybrid NOMA mmWave Communication: Optimization and Performance Analysis

Simultaneously reflecting and transmitting reconfigurable intelligent surfaces (STAR-RIS) has recently emerged as prominent technology that exploits the transmissive property of RIS to mitigate the half-space coverage limitation of conventional RIS operating on millimeter-wave (mmWave). In this paper, we study a downlink STAR-RIS-based multi-user multiple-input single-output (MU-MISO) mmWave hybrid non-orthogonal multiple access (H-NOMA) wireless network, where a sum-rate maximization problem has been formulated. The design of active and passive beamforming vectors, time and power allocation for H-NOMA is a highly coupled non-convex problem. To handle the problem, we propose an optimization framework based on alternating optimization (AO) that iteratively solves active and passive beamforming sub-problems. Channel correlations and channel strength-based techniques have been proposed for a specific case of two-user optimal clustering and decoding order assignment, respectively, for which analytical solutions to joint power and time allocation for H-NOMA have also been derived. Simulation results show that: 1) the proposed framework leveraging H-NOMA outperforms conventional OMA and NOMA to maximize the achievable sum-rate; 2) using the proposed framework, the supported number of clusters for the given design constraints can be increased considerably; 3) through STAR-RIS, the number of elements can be significantly reduced as compared to conventional RIS to ensure a similar quality-of-service (QoS).

eess.SP