SearcharxivSearch

arXiv subjects

Vikram Vasudevan

Publications and source records attributed to Vikram Vasudevan.

5 recordsLinked to original sources

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver fixed within each setting: direct solving (Baseline), iterative self-correction (Single), planner-executor-reviewer collaboration (PER), and multi-agent deliberation (Broadcast). The primary benchmark comprises 4,181 competition-level math problems; paired robustness checks cover four benchmarks spanning competition math, biology, and broader science with two solver families. Across fixed policies, trained routers, and frozen LLM routers, conservative policies under-escalate, whereas higher-solve frozen routers often over-escalate. A post-answer, pre-collaboration gpt-oss-120b probe ranks Baseline failures with 0.8847 AUROC (4,151 parseable cases; 95% CI [0.8732, 0.8955]). The same score remains informative for predicting whether any collaboration helps (0.7683 AUPRC), but is much weaker for identifying PER- or Broadcast-specific value (0.1674 and 0.1041 AUPRC). Separately, the pre-answer self-confidence gate reaches 78.0% solve at 45K tokens, compared with 73.8% at 71.3K for a frozen gpt-oss-120b router and 92.4% for a retrospective fixed-order oracle. Across 10 paired model-condition settings, the oracle adds 23.2-58.3 points of retrospective coverage over Baseline, but protocol profiles vary by task. In the six settings with held-out router evaluations, oracle gaps remain 18.5-28.9 points. Confidence can therefore support initial escalation, while protocol-specific cost-aware routing remains unresolved.

cs.AI

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors. Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches higher final accuracy than a planner-executor-reviewer pipeline (PER). We ask whether this gap is explained by reviewer quality or by whether critique changes the next answer the protocol carries forward. It is not explained by reviewer precision alone: PER's reviewer is more precise than broadcast's (0.861 vs. 0.644), yet evaluator-verified useful critique is much less likely to change the next candidate and produces lower reviewer-guided repair. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy, while embedding reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Overall, reviewer-centric evaluation can overstate system quality: a protocol may spot errors well yet still fail to solve more problems if it does not act on those critiques.

cs.AI

Ultra-Wideband Double-Directional Channel Measurements and Statistical Modeling in Urban Microcellular Environments for the Upper-Midband/FR3

The upper midband, designated as Frequency Range 3 (FR3), is increasingly critical for the next-generation of wireless networks. Channel propagation measurements and their statistical analysis are essential first steps towards this direction. This paper presents a comprehensive ultra-wideband (UWB) double-directional channel measurement campaign in a large portion of FR3 (6-14 GHz) for urban microcellular environments. We analyze over 25,000 directional power delay profiles and providing key insights into line-of-sight (LoS) and obstructed line-of-sight (OLoS) conditions. This is followed by statistical modeling of path loss, shadowing, delay spread and angular spread. As the first UWB double-directional measurement campaign in this frequency range, this work offers critical insights for spectrum allocation, channel modeling, and the design of advanced communication systems, paving the way for further exploration of FR3.

eess.SY

An Ultra-Wideband Study of Vegetation Impact on Upper Midband / FR3 Communication

Growing demand for high data rates is driving interest in the upper mid-band (FR 3) spectrum (6-24 GHz). While some propagation measurements exist in literature, the impact of vegetation on link performance remains under-explored. This study examines vegetation-induced losses in an urban scenario across 6-18 GHz. A simple method for calculating vegetation depth is introduced, along with a model that quantifies additional attenuation based on vegetation depth and frequency, divided into 1 GHz sub-bands. We see that excess vegetation loss increases with vegetation depth and higher frequencies. These findings provide insights for designing reliable, foliage-aware communication networks in FR 3.

eess.SP

Ultra-wideband Double-Directionally Resolved Channel Measurements of Line-of-Sight Microcellular Scenarios in the Upper Mid-band

The growing demand for higher data rates and expanded bandwidth is driving the exploration of new frequency ranges, including the upper mid-band spectrum (6-24 GHz), which is a promising candidate for future Frequency Range 3 (FR3) applications. This paper presents ultra-wideband double-directional channel measurements in line-of-sight microcellular scenarios within the upper mid-band spectrum (6-18 GHz). Conducted in an urban street canyon environment, these measurements explore key channel characteristics such as power delay profiles, angular power spectra, path loss, delay spread, and angular spread to provide insights essential for robust communication system design. Our results reveal that path loss values for both omni-directional and best beam configurations are lower than free-space predictions due to multipath contributions from the environment. Analysis also indicates a high degree of stability in delay spread and angular spread across the entire band, with small variation between sub-bands.

eess.SY