SearcharxivSearch

arXiv subjects

Om Chabra

Publications and source records attributed to Om Chabra.

5 recordsLinked to original sources

SILC: Lookahead Caching for Short-form Video Delivery Systems

Short video platforms like TikTok, Instagram Reels, and YouTube Shorts have gained immense popularity in the last few years and are responsible for a large and growing fraction of Internet traffic. We identify two unique opportunities for improving short video delivery using their existing interactions with content delivery networks (CDNs). First, short videos use a push-based recommendation system, where the user is presented a sequence of videos recommended by the algorithm rather than user explicitly picking content to watch (e.g., in YouTube). Such push-based short video systems offer a unique opportunity for system design by providing visibility into upcoming requests. Second, the popularity of these videos follows a highly skewed Pareto distribution, leading to geographical and temporal overlap amongst videos being served. We leverage these opportunities to build SILC - a lookahead-aware caching system, aimed at (i) reducing CDN cache miss rates, as well as (ii) reducing midgress bandwidth between the CDN and the origin server. Our evaluation of SILC uses traces that we collect from real users, through (i) an in-person user study, and (ii) a data donation program involving 100 TikTok users across the world. Using a combination of these traces, we simulate traffic from 10,000 simultaneous users. Our evaluation shows that, compared to 10 state-of-the-art heuristic and learning-based cache eviction policies, SILC reduces a CDN's midgress costs by 11.1% to 111%.

cs.NI

IteRate: Autonomous AI Synthesis of In-Kernel eBPF Wi-Fi Rate Control Algorithms

Wi-Fi rate adaptation remains a persistent challenge in wireless networking. Deployed algorithms like Minstrel-HT have remained largely stagnant for over a decade, relying on hand-tuned heuristics that fail to generalize to the complexity of modern wireless environments. We present \name, an autonomous research system that closes the loop on rate control development. IteRate uses a multi-agent AI architecture to conduct the full scientific cycle: formulating hypotheses, writing eBPF programs that run inside the Linux kernel, deploying them over-the-air to Wi-Fi devices, collecting fine-grained telemetry for analysis, and iterating based on experimental evidence, all without human intervention. IteRate makes three contributions. (1) a novel kernel module that exposes per-frame hardware telemetry including modulation and coding schemes (MCS) and retry counts to eBPF programs, (2) a structured agentic AI architecture employing specialized agents for algorithm design, experiment execution, and data analysis, coordinated via a hypothesis-driven research protocol with persistent knowledge, and (3) a closed-loop pipeline that automates the cross-compilation, deployment, and evaluation of in-kernel logic onto embedded Wi-Fi targets. On a 58-node testbed running five workloads. relative to the well-known Minstrel algorithm, IteRate achieves 21% faster web-page loads, 7% higher video quality of experience (QoE), and 21% higher peak throughput. Our work demonstrates that AI agents, when equipped with appropriate kernel-level hooks and a disciplined scientific workflow, can effectively automate the research required to design Wi-Fi rate controllers.

cs.NI

KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from extraction and cleaning to integration, analysis, and modeling. This process often also demands domain knowledge and project-specific insight. While AI models have shown remarkable results in reasoning and code generation, their abilities to design and execute complex pipelines that solve these data-lake-to-insight challenges remain unclear. We introduce KramaBench which consists of 104 manually curated and solved challenges spanning 1700 files, 24 data sources, and 6 domains. KramaBench focuses on testing the end-to-end capabilities of AI systems to solve challenges which require automated orchestration of different data tasks. KramaBench also features a comprehensive evaluation framework assessing the pipeline design and individual data task implementation abilities of AI systems. We evaluate 8 LLMs using our single-agent reference framework DS-Guru, alongside both open- and closed-source single- and multi-agent systems, and find that while current agentic systems may handle isolated data-science tasks and generate plausible draft pipelines, they struggle with producing working end-to-end pipelines. On KramaBench, the best system reaches only 55% end-to-end accuracy in the full data-lake setting. Even with perfect retrieval, the accuracy tops out at 62%. Leading LLMs can identify up to 42% of important data tasks but can only fully implement 20% of individual data tasks. Our code, reference framework, and data are available at https://github.com/mitdbg/KramaBench.

cs.DB

Scalable Routing in a City-Scale Wi-Fi Network for Disaster Recovery

In this paper, we present a new city-scale decentralized mesh network system suited for disaster recovery and emergencies. When wide-area connectivity is unavailable or significantly degraded, our system, MapMesh, enables static access points and mobile devices equipped with Wi-Fi in a city to route packets via each other for intra-city connectivity and to/from any nodes that might have Internet access, e.g., via satellite. The chief contribution of our work is a new routing protocol that scales to millions of nodes, a significant improvement over prior work on wireless mesh and mobile ad hoc networks. Our approach uses detailed information about buildings from widely available maps--data that was unavailable at scale over a decade ago, but is widely available now--to compute paths in a scalable way.

cs.NI

Constraining the Regolith Composition of Asteroid (16) Psyche via Laboratory Near-infrared Spectroscopy

(16) Psyche is the largest M-type asteroid in the main belt and the target of the NASA Discovery-class Psyche mission. Despite gaining considerable interest in the scientific community, Psyche's composition and formation remain unconstrained. Originally, Psyche was considered to be almost entirely composed of metal due to its high radar albedo and spectral similarities to iron meteorites. More recent telescopic observations suggest the additional presence of low-Fe pyroxene and exogenic carbonaceous chondrites on the asteroid's surface. To better understand the abundances of these additional materials, we investigated visible near-infrared (0.35 - 2.5 micron) spectral properties of three-component laboratory mixtures of metal, low-Fe pyroxene, and carbonaceous chondrite. We compared the band depths and spectral slopes of these mixtures to the telescopic spectrum of (16) Psyche to constrain material abundances. We find that the best matching mixture to Psyche consists of 82.5% metal, 7% low-Fe pyroxene, and 10.5% carbonaceous by weight, suggesting that the asteroid is less metallic than originally estimated (~94%). The relatively high abundance of carbonaceous chondrite material estimated from our laboratory experiments implies the delivery of this exogenic material through low velocity collisions to Psyche's surface. Assuming that Psyche's surface is representative of its bulk material content, our results suggest a porosity of 35% to match recent density estimates.

astro-ph.EP