SearcharxivSearch

arXiv subjects

Koyar Afrasyab

Publications and source records attributed to Koyar Afrasyab.

6 recordsLinked to original sources

Edge-addition monotonicity of positive p-energy fails for every p >= 1

At a 2021 AIM workshop, Guo conjectured that the positive square energy s+ = E+_2 should inherit the familiar edge-addition monotonicity of the spectral radius, rho(G + uv) >= rho(G). That conjecture was subsequently shown to fail at p = 2. Tang, Liu, and Wang then introduced positive p-energy, proved nonmonotonicity for every 1 <= p < 3, and in version 3 of their preprint (26 March 2025) explicitly conjectured that monotonicity should hold for p >= 3. We disprove this conjectured high-exponent extension completely: for every real p > 2 there are infinitely many connected graphs G and nonedges uv such that E+_p(G + uv) < E+_p(G). Together with the Tang-Liu-Wang counterexamples below 3, this shows that no exponent p >= 1 restores the spectral-radius-style monotonicity: positive p-energy can decrease under the addition of an edge for every real p >= 1. The construction is a chain of clique blocks joined by regular bipartite graphs. An equitable quotient converges to Q = I + cA(P_k). For noninteger p, a binomial-series sign argument for a fractional power of I - cA(P_k) gives the required negative endpoint entry. At integer exponents, choosing c across the first spectral threshold leaves exactly one negative eigenvalue, and path locality forces the positive spectral contribution to have negative sign. We also give a fully rational 38-vertex certificate at p = 4 and determine the complete failure interval of a fixed 17-vertex counterexample at p = 3.

math.CO

A 50-Vertex Cubic Counterexample to the Domination-versus-Edge-Domination Conjecture

Baste, Furst, Henning, Mohr, and Rautenbach conjectured that every finite regular graph of positive degree satisfies \(γ(G) \leq γ_e(G)\), where \(γ\) is the domination number and \(γ_e\) is the edge domination number, equivalently the minimum cardinality of a maximal matching. We show that the conjecture is false already for cubic graphs. The counterexample is a previously public 50-vertex cubic graph that had been used to refute the stronger independent-domination inequality \(i(G) \leq γ_e(G)\). For this graph we prove \(γ(G) = 16 > 15 = γ_e(G)\). The equality \(γ_e(G) = 15\) has a short counting proof, and a dominating set of order 16 is displayed explicitly. For the lower bound \(γ(G) \geq 16\), we give a self-contained exact reduction: after fixing which of the 20 clause vertices lie in a putative dominating set, the remaining problem is a finite set-cover problem on the 30 literal vertices. We enumerate all \(2^{20} = 1,048,576\) clause subsets, derive two explicit lower bounds, and solve exactly the 5,931 residual cases by a recurrence stated in the paper. The complete case counts and minima are displayed, and a short standard-library Python implementation is included in an appendix. A separate 893,049-node proof-tree certificate and a direct graph search provide independent verification. Thus the regular-graph conjecture is disproved. Combined with Gupta's recent theorem that every cubic graph on at most 48 vertices satisfies the conjectured inequality, the example is order-minimal among cubic counterexamples.

math.CO

Exact Zarankiewicz Values On Two Finite Frontier Slices

The Zarankiewicz number Z(m,n,s,t) is the maximum number of edges in a bipartite graph with parts of orders m and n containing no copy of Ks,t. We give one combined, certificate-based computer-assisted proof for two finite slices and a corrected neighboring frontier: Z(12,n,3,3) = 6n (18 <= n <= 22), Z(13,22,3,3) = 137, Z(13, 18, 3, 3) = 116, Z(14, 18, 3, 3) = 124, Z(15,18,3,3) = 132, Z(14, 17, 3, 3) = 118, Z(15, 17, 3, 3) = 126, 132 <= Z(16,17,3,3) <= 133. The load-bearing new upper bounds are the exact 12 x 18 and 13 x 18 certificate packages. Their orbit certificates exclude every hypothetical matrix at the next edge count. Deletion lemmas and explicit witnesses close four neighboring cells, while the 16 x 17 entry is deliberately reported as an interval because only its 132-edge lower witness and the published 133 upper bound are certified here. Separately, the 13 x 22 proof excludes 138 ones by reducing to 83 degree profiles, rationally separating 77 of them, and eliminating the remaining six by marked-row congruences, leave enumeration, modular Gram tests, and exact Farkas certificates. All accepted claims are replayed by standard-library Python and exact integer/rational arithmetic; floating-point optimization is used only to discover certificates.

math.CO

Independence-System Realisations in Single-Source Unsplittable Flow

Additive-congestion constraints in single-source unsplittable flow can enforce stable-set structure. This note isolates and generalises that mechanism. We introduce a path-closed notion of realising an independence system by the zero-cost choices of primary terminals in a directed acyclic flow instance. The definition quantifies over every directed source-terminal path and therefore remains valid under prefix borrowing, suffix splicing, and hybrid routes.Our main result extends the triangle mechanism: every finite loopless independence system has a polynomial-size realisation, measured in the incidence size of its minimal forbidden sets. Hence every finite simple graph, and more generally every hypergraph independence system without singleton forbidden hyperedges, is representable by an acyclic single-source gadget. We then specialise the construction to odd cycles. For C_{2k+1}, a uniform rational family produces a fractional cheap-selection vector that violates the odd-cycle inequality. A potential shift converts a signed connector separator into nonnegative arc costs and gives the exact cost-preserving additive-congestion threshold tau = 1 - bq. Within the symmetric family, the supremum threshold is (k+2)/(2(k+1)), which tends to 1/2. For C5, an exact certificate independently derives all source-terminal paths and enumerates all 3^10 = 59049 unsplittable routings using rational arithmetic.

cs.DS

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.

cs.AI

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.

cs.AI