SearcharxivSearch

arXiv subjects

Aalok Thakkar

Publications and source records attributed to Aalok Thakkar.

13 recordsLinked to original sources

A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph

Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We develop two complementary lines of attack. Fixing one vertex, the conditions $λ=1$ and $μ=2$ force its neighbourhood to be a perfect matching and determine every edge between that neighbourhood and the remaining vertices. For $(99,14,1,2)$, the unresolved part is therefore a constrained $12$-regular graph on $84$ vertices. We encode this reduction in CP-SAT and validate it by recovering the unique $\mathrm{srg}(9,4,1,2)$. We also prove by exhaustive enumeration that no circulant graph on $\mathbb{Z}/99$ satisfies more than $68.0\%$ of the CAISc constraints, and we give a validated orbit formulation for prescribed automorphisms. We then study the partial-score search problem. Fourteen human-designed search configurations reached at most $69.43\%$. Separately, we supplied the scoring function to an evolutionary program-search system. It produced a degree-preserving $4$-vertex-switch tabu search whose best verified artifact scores $70.73\%$. The generated move differs from those used in our own searches and crosses a plateau that was stable under them. These results do not resolve the existence problem, but they reduce the exact search space and improve the best verified partial construction found in our experiments.

cs.AI

How Often Does Your Program Fail?

Software in production encounters inputs shaped by how it is used in practice. We study distribution-aware reliability estimation: given a program, its operational input distribution, and a condition of interest, determine how often that condition holds and certify the result with a guaranteed error bound. Symbolic and statistical methods offer two ways to answer this question. Symbolic methods reason about entire regions of the input space and can certify rates exactly, but often struggle with complex arithmetic or loops. Statistical methods instead sample inputs and apply concentration bounds. They are broadly applicable, but can require many samples when failures are rare. We bring these approaches together in a framework that includes both as special cases. Each estimator has three components: a mass estimator, a per-leaf confidence bound, and a symbolic closure rule. We prove that any instantiation satisfying three invariants returns an interval containing the true rate with confidence 1-delta, regardless of when it stops. The certified error has two parts: a statistical term, reduced by sampling, and a structural term, reduced by symbolic closure. Pure sampling and pure symbolic execution each reduce only one of these terms. Alternative choices of the components also allow the framework to support rare-event variance-reduction methods.

cs.PL

Localising Stochasticity in Weighted Automata

Weighted automata over the nonnegative reals form a fundamental model for quantitative languages. We show that, up to scaling, this model collapses to probabilistic automata. Concretely, we prove that every weighted automaton whose transition matrix has spectral radius strictly less than one can be normalised, by a semantics-preserving rescaling of transition weights, into an equivalent locally stochastic probabilistic automaton. Thus, finite-mass weighted automata and probabilistic automata coincide up to normalisation. The construction is effective and relies on Perron-Frobenius theory. We further characterise probabilistic automata by stochastic regular expressions equipped with a geometrically weighted star. Beyond the finite-mass setting, we show that the behaviour of an arbitrary weighted automaton admits a decomposition into an exponential growth rate and a normalised probabilistic component, separating quantitative growth from stochastic structure.

cs.FL

Astra: AI Safety, Trust, & Risk Assessment

This paper argues that existing global AI safety frameworks exhibit contextual blindness towards India's unique socio-technical landscape. With a population of 1.5 billion and a massive informal economy, India's AI integration faces specific challenges such as caste-based discrimination, linguistic exclusion of vernacular speakers, and infrastructure failures in low-connectivity rural zones, that are frequently overlooked by Western, market-centric narratives. We introduce ASTRA, an empirically grounded AI Safety Risk Database designed to categorize risks through a bottom-up, inductive process. Unlike general taxonomies, ASTRA defines AI Safety Risks specifically as hazards stemming from design flaws such as skewed training sets or lack of guardrails that can be mitigated through technical iteration or architectural changes. This framework employs a tripartite causal taxonomy to evaluate risks based on their implementation timing (development, deployment, or usage), the responsible entity (the system or the user), and the nature of the intent (unintentional vs. intentional). Central to the research is a domain-agnostic ontology that organizes 37 leaf-level risk classes into two primary meta-categories: Social Risks and Frontier/Socio-Structural Risks. By focusing initial efforts on the Education and Financial Lending sectors, the paper establishes a scalable foundation for a "living" regulatory utility intended to evolve alongside India's expanding AI ecosystem.

cs.CY

SocraticAI: Transforming LLMs into Guided CS Tutors Through Scaffolded Interaction

We present SocraticAI, a scaffolded AI tutoring system that integrates large language models (LLMs) into undergraduate Computer Science education through structured constraints rather than prohibition. The system enforces well-formulated questions, reflective engagement, and daily usage limits while providing Socratic dialogue scaffolds. Unlike traditional AI bans, our approach cultivates responsible and strategic AI interaction skills through technical guardrails, including authentication, query validation, structured feedback, and RAG-based course grounding. Initial deployment demonstrates that students progress from vague help-seeking to sophisticated problem decomposition within 2-3 weeks, with over 75% producing substantive reflections and displaying emergent patterns of deliberate, strategic AI use.

cs.CY

Ancient Algorithms for a Modern Curriculum

Despite ongoing calls for inclusive and culturally responsive pedagogy in computing education, the teaching of algorithms remains largely decontextualized. Foundational computer science courses often present algorithmic thinking as purely formal and ahistorical, emphasizing efficiency, correctness, and abstraction. When history is mentioned, it usually centers on the modern development of digital computers, highlighting figures such as Turing, von Neumann, and Babbage. This narrow view misrepresents the origins of algorithmic reasoning and perpetuates a Eurocentric worldview that undermines equity and representation in STEM. In contrast, algorithmic thinking predates electronic computers by millennia and has deep roots in ancient civilizations including India, China, Babylon, and Egypt. Our work responds to this gap by embedding algorithm instruction in broader historical and cultural contexts, with particular attention to classical Indian contributions.

cs.CY

BOOP: Write Right Code

Novice programmers frequently adopt a syntax-specific and test-case-driven approach, writing code first and adjusting until programs compile and test cases pass, rather than developing correct solutions through systematic reasoning. AI coding tools exacerbate this challenge by providing syntactically correct but conceptually flawed solutions. In this paper, we address the question of developing correctness-first methodologies to enhance computational thinking in introductory programming education. To this end, we introduce BOOP (Blueprint, Operations, OCaml, Proof), a structured framework requiring four mandatory phases: formal specification, language-agnostic algorithm development, implementation, and correctness proof. This shifts focus from making the code to understanding the code is correct. BOOP was implemented at Ashoka University using a VS Code extension and preprocessor that enforces constraints. Initial evaluation shows improved algorithmic reasoning and reduced trial-and-error debugging. Students reported better edge case understanding and problem decomposition, though some initially found the format verbose. Instructors observed stronger foundational skills compared to traditional approaches, suggesting that structured correctness-first approaches may significantly improve students' computational thinking abilities.

cs.SE

Stochastic Languages at Sub-stochastic Cost

When does a deterministic computational model define a probability distribution? What are its properties? This work formalises and settles this stochasticity problem for weighted automata, and its generalisation cost register automata (CRA). We show that checking stochasticity is undecidable for CRAs in general. This motivates the study of the fully linear fragment, where a complete and tractable theory is established. For this class, stochasticity becomes decidable in polynomial time via spectral methods, and every stochastic linear CRA admits an equivalent model with locally sub-stochastic update functions. This provides a local syntactic characterisation of the semantics of the quantitative model. This local characterisation allows us to provide an algebraic Kleene-Schutzenberger characterisation for stochastic languages. The class of rational stochastic languages is the smallest class containing finite support distributions, which is closed under convex combination, Cauchy product, and discounted Kleene star. We also introduce Stochastic Regular Expressions as a complete and composable grammar for this class. Our framework provides the foundations for a formal theory of probabilistic computation, with immediate consequences for approximation, sampling, and distribution testing.

cs.FL

Accessibility Beyond Accommodations: A Systematic Redesign of Introduction to Computer Science for Students with Visual Impairments

Computer science education has evolved extensively; however, systemic barriers still prevent students with visual impairments from fully participating. While existing research has developed specialized programming tools and assistive technologies, these solutions remain fragmented and often require complex technical infrastructure, which limits their classroom implementation. Current approaches treat accessibility as individual accommodations rather than integral curriculum design, creating gaps in holistic educational support. This paper presents a comprehensive framework for redesigning introductory computer science curricula to provide equitable learning experiences for students with visual impairments without requiring specialized technical infrastructure. The framework outlines five key components that together contribute a systematic approach to curriculum accessibility: accessible learning resources with pre-distributed materials and tactile diagrams, in-class learning kits with hands-on demonstrations, structured support systems with dedicated teaching assistance, an online tool repository, and psychosocial support for classroom participation. Unlike existing tool-focused solutions, this framework addresses both technical and pedagogical dimensions of inclusive education while emphasizing practical implementation in standard university settings. The design is grounded in universal design principles and validated through expert consultation with accessibility specialists and disability services professionals, establishing foundations for future empirical evaluation of learning outcomes and student engagement while serving as a template for broader institutional adoption.

cs.HC

Identity Testing for Stochastic Languages

Determining whether an unknown distribution matches a known reference is a cornerstone problem in distributional analysis. While classical results establish a rigorous framework in the case of distributions over finite domains, real-world applications in computational linguistics, bioinformatics, and program analysis demand testing over infinite combinatorial structures, particularly strings. In this paper, we initiate the theoretical study of identity testing for stochastic languages, bridging formal language theory with modern distribution property testing. We first propose a polynomial-time algorithm to verify if a finite state machine represents a stochastic language, and then prove that rational stochastic languages can approximate an arbitrary probability distribution. Building on these representations, we develop a truncation-based identity testing algorithm that distinguishes between a known and an unknown distributions with sample complexity $\widetildeΘ\left( \frac{\sqrt{n}}{\varepsilon^2} + \frac{n}{\log n} \right)$ where $n$ is the size of the truncated support. Our approach leverages the exponential decay inherent in rational stochastic languages to bound truncation error, then applies classical finite-domain testers to the restricted problem. This work establishes the first identity testing framework for infinite discrete distributions, opening new directions in probabilistic formal methods and statistical analysis of structured data.

cs.FL

Infinitude of Primes Using Formal Language Theory

Formal languages are sets of strings of symbols described by a set of rules specific to them. In this note, we discuss a certain class of formal languages, called regular languages, and put forward some elementary results. The properties of these languages are then employed to prove that there are infinitely many prime numbers.

cs.FL

Concurrency in Boolean networks

Boolean networks (BNs) are widely used to model the qualitative dynamics of biological systems. Besides the logical rules determining the evolution of each component with respect to the state of its regulators, the scheduling of component updates can have a dramatic impact on the predicted behaviours. In this paper, we explore the use of Read (contextual) Petri Nets (RPNs) to study dynamics of BNs from a concurrency theory perspective. After showing bi-directional translations between RPNs and BNs and analogies between results on synchronism sensitivity, we illustrate that usual updating modes for BNs can miss plausible behaviours, i.e., incorrectly conclude on the absence/impossibility of reaching specific configurations. We propose an encoding of BNs capitalizing on the RPN semantics enabling more behaviour than the generalized asynchronous updating mode. The proposed encoding ensures a correct abstraction of any multivalued refinement, as one may expect to achieve when modelling biological systems with no assumption on its time features.

cs.LO

Additive Collatz Trajectories

Collatz Conjecture (also known as Ulam's conjecture and 3x+1 problem) concerns the behavior of the iterates of a particular function on natural numbers. A number of generalizations of the conjecture have been subjected to extensive study. This paper explores Additive Collatz Trajectories, a particular case of a generalization of Collatz conjecture and puts forward a sufficient and necessary condition for looping of Additive Collatz Trajectories, along with two minor results. An algorithm to compute the number of equivalence classes when natural numbers are quotiented by the limiting behavior of their corresponding trajectories is also proposed.

math.NT