Searcharxiv⌕ Search

arXiv subjects

Sravan Karthick T

Publications and source records attributed to Sravan Karthick T.

3 recordsLinked to original sources

The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|ρ|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p<0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.

cs.CL↗

The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance

Retrieval-Augmented Generation (RAG) introduces a specific failure mode in discrete diffusion language models: when retrieved context contradicts parametric knowledge, the iterative denoising process becomes a visible battleground between competing knowledge sources. We identify temporal semantic divergence as an observable for detecting these conflicts and introduce the Trajectory Variance Score (TVS), a simple and interpretable measure of this divergence. TVS computes the mean pairwise cosine distance of answer embeddings across independent stochastic denoising trajectories, capturing the temporal tug of war between parametric and contextual attractors. Requiring as few as two parallel inference runs, TVS is computationally lightweight. Across four diverse datasets (Synthetic, SciQ, PopQA, and CounterFact), a simple Logistic Regression classifier using TVS achieves $70.10\%$ accuracy and $0.7647$ AUROC on LLaDA. On Dream 7B, increasing the number of trajectories from two to five improves accuracy from $63.91\%$ to $69.62\%$. More complex sequential models provide only marginal improvements over the linear classifier. Evaluation across LLaDA and Dream 7B demonstrates that conflict-induced trajectory dynamics and their key properties transfer across distinct diffusion architectures.

cs.CL↗

Expert System for Bitcoin Forecasting: Integrating Global Liquidity via TimeXer Transformers

Bitcoin price forecasting is characterized by extreme volatility and non-stationarity, often defying traditional univariate time-series models over long horizons. This paper addresses a critical gap by integrating Global M2 Liquidity, aggregated from 18 major economies, as a leading exogenous variable with a 12-week lag structure. Using the TimeXer architecture, we compare a liquidity-conditioned forecasting model (TimeXer-Exog) against state-of-the-art benchmarks including LSTM, N-BEATS, PatchTST, and a standard univariate TimeXer. Experiments conducted on daily Bitcoin price data from January 2020 to August 2025 demonstrate that explicit macroeconomic conditioning significantly stabilizes long-horizon forecasts. At a 70-day forecast horizon, the proposed TimeXer-Exog model achieves a mean squared error (MSE) 1.08e8, outperforming the univariate TimeXer baseline by over 89 percent. These results highlight that conditioning deep learning models on global liquidity provides substantial improvements in long-horizon Bitcoin price forecasting.

cs.LG↗