Searcharxiv⌕ Search

arXiv subjects

Binh T. Nguyen

Publications and source records attributed to Binh T. Nguyen.

At least 19 recordsLinked to original sources

Uniform High-Frequency Localization on Quantum Graphs

We study uniform high-frequency localization for the Laplacian on compact metric graphs through the least $L^2$-mass that eigenfunctions must place in a prescribed measurable observation set. We first identify this asymptotic localization constant with the minimum of a linear functional over the attainable edge-intensity set; in the generic standard-Kirchhoff setting, this set is governed by the regular Gauss image of the secular manifold. Primitive cycles and exterior-to-exterior paths consequently determine the positivity threshold, but not, in general, the positive numerical value. We introduce a boundary-aware singular-completion cone and prove that it contains all regular secular edge-energy vectors for trees, unicyclic graphs, and closed graphs of cycle rank two. We then construct a cycle-rank-two graph with two Dirichlet leaves for which this completion principle fails. An exact rational separator, combined with a validated Krawczyk enclosure, yields a nonsingular scalar secular state lying outside every boundary-compatible singular sector. A positive radial derivative identity and recurrence in the compact orbit closure convert this local separation into an exact high-frequency eigensequence for a single fixed metric. For a suitable measurable observation set, the true high-frequency localization constant $C_\infty(ω;\ell)$ and its singular-completion counterpart $C_{\mathrm{sing}}(ω;\ell)$ satisfy $$C_\infty(ω;\ell)<1/2<C_{\mathrm{sing}}(ω;\ell)$$. Thus, singular completion captures the quantitative localization geometry in several low-complexity classes but does not, in general, determine the high-frequency variational problem on a fixed quantum graph.

math.SP↗

HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rather than compositional reasoning. Fine-tuning on negation-specific data can also compromise their general purpose capabilities through catastrophic forgetting. We introduce HANCLIP (Hyperbolic, Angular, and Negation), a geometry-aware framework that improves negation sensitivity while preserving the structure of the pretrained joint embedding space. HANCLIP combines a hyperbolic contrastive objective, which models hierarchical relations and semantic asymmetries, with an angular triplet loss that separates negated descriptions from their affirmative counterparts. Using only 20,000 image-text quadruplets, HANCLIP consistently improves performance across CLIP, LongCLIP, and SmartCLIP backbones on the NegBench benchmark, while maintaining or improving zero-shot classification and image-text retrieval performance. These results show that lightweight, geometry-guided objectives can enhance negation understanding without large-scale retraining.

cs.CV↗

Uniform Eigenfunction Observability under Mixed Reflection Monodromy

We study uniform observability for Laplace eigenfunctions with mixed boundary conditions whose reflection signs do not define a scalar character. For the DDN/NND sectors of the equilateral rhombus, the resulting $\mathbb Z_2$ monodromy is resolved by a degree-two branched arithmetic translation surface of genus two. On an explicit class of admissible open sets, every exact mixed-sector eigenfunction satisfies a local $L^2$ lower bound uniform in the eigenvalue, multiplicity, and choice within the eigenspace. The proof combines exact unfolding, semiclassical defect measures, and the Veech dichotomy to exclude concentration near the saddle/conic network and inside periodic cylinders. We also obtain a reflection-overlap observability result for the full rhombus and isolate the remaining saddle-network obstruction for arbitrary open observation sets.

math.SP↗

Arithmetic Scar Accessibility and Generic Schrodinger Control on Metric Graphs

We study internal exact controllability of the free Schrödinger equation on finite compact metric graphs with Kirchhoff conditions at interior vertices and a fixed Dirichlet-Neumann assignment at exterior vertices. For a prescribed primitive cycle or exterior-to-exterior path, we identify an arithmetic criterion ensuring that the corresponding scar is accessible by exact high-frequency eigenfunctions of a fixed target metric. The criterion is formulated by intersecting the regular scar stratum of the secular set with the orbit-closure torus of the target length vector, and it can hold for resonant metrics beyond the rationally independent class. Under an intrinsic Diophantine condition on the orbit-closure flow, the construction is quantitative: along an exact eigensequence the mass outside the prescribed support decays polynomially, yielding a polynomial degeneration law for finite-frequency observation when that support is uncontrolled. Rationally independent metrics have full phase-orbit closure and therefore make every primitive obstruction accessible. Combining this spectral result with the known graph-theoretic characterization and sufficiency of the Graph Geometric Control Condition (GGCC), we prove that, for every rationally independent and hence for Lebesgue-almost every metric, GGCC is necessary and sufficient for exact controllability at every positive time in the free setting.

math.OC↗

Uniform Non-Localization under Robin Boundary Perturbations: Spectral Splitting and Bounded Eigenspace Complexity

We study the high-energy spectral effect of Robin boundary perturbations on complete Laplace eigenspaces and its consequences for eigenfunction non-localization. Highly degenerate Neumann levels generally split under Robin perturbation, but this splitting need not produce a simple spectrum: distinct modal classes may still contribute to the same Robin eigenvalue. We prove that, for every fixed positive Robin parameter, the number of modal classes contributing to any one eigenvalue is uniformly bounded over the spectrum on the equilateral triangle and on every rectangle with rational squared aspect ratio. On the equilateral triangle, for all sufficiently small Robin parameters, the Neumann complete-eigenspace observation estimate persists with a constant uniform both in the eigenvalue and in the boundary parameter. For an arbitrary fixed positive Robin parameter, a high-frequency shell-splitting analysis gives a spectrum-wide bound on the number of modal classes contributing to any one Robin eigenvalue. In the nonsquare rectangular case, the corresponding splitting is governed by a strongly convex profile on weighted quadratic shells. Since each modal class has uniformly boundedplane-wave complexity, the spectral complexity bound implies observation on every measurable set of positive measure through a multidimensional Turán--Nazarov inequality, without a frequency-separation assumption. Thus, spectral simplicity is not required for complete-eigenspace non-localization: uniformly bounded spectral coincidence complexity is sufficient.

math.SP↗

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, how good these systems are actually, especially compared to human reviewers at catching scientific gaps, remains poorly understood. In this work, we introduce PRISM (Peer Review Intelligence via Structured Multi-dimensional assessment), a benchmarking framework that evaluates review quality across four dimensions: Depth of Analysis, Novelty Assessment,Flaw Identification & Major Issues Prioritization, and Multi-dimensional Constructiveness. Unlike most existing evaluations based on surface-level metrics like ROUGE and BLEU, or unconstrained LLM-as-a-judge prompting that conflates fluency with rigor, PRISM grounds each dimension in argument mining, retrieval-augmented verification, and consensus-based scoring. We apply PRISM to benchmark five leading automated reviewer systems and human reviewers on a stratified corpus of reviews from ICLR, ICML, and NeurIPS. The results reveal that LLMs can match or beat human reviewers on individual dimensions: comparable depth of analysis, stronger novelty verification, and highly accurate critique prioritization. However, no single system consistently matches the balanced performance of the human baseline across all dimensions at once. Each exhibits a distinct specialization profile with characteristic blind spots -- failure modes that aggregate metrics miss entirely. The implication is that LLM reviewers are best understood as targeted supplements to human review, effective within specific dimensions, but unreliable as standalone replacements. Our demo and key results can be found at https://khanhthanhdev.github.io/prism-page/.

cs.CL↗

Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms

Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. Their bound contains an additional factor $\log n$, and they asked whether this factor can be removed. We answer this upper-bound question affirmatively. More specifically, let $Z=(Z_1,\ldots,Z_n)$ have independent coordinates and let $g_i(Z)$ satisfy $\mathbb E[g_i(Z)\mid Z_{-i}]=0, \ \left| \mathbb E[g_i(Z)\mid Z_i]\right|\le M, \ \text{for every } i = 1, \dots, n, $ where $Z_{-i}$ denotes all coordinates except $Z_i$. Assume additionally that changing any coordinate $Z_j$, $j\neq i$, changes $g_i$ by at most $β$, we prove that, for every $p\ge2$, for every $p\ge2$, $$ \left\| \sum_{i=1}^n g_i(Z)\right\|_p \le 16pnβ+M\sqrt{2pn}. $$ This removes the $\log n$ factor from the previous bound and matches the lower bound of Bousquet, Klochkov, and Zhivotovskiy up to universal constants in the range covered by their construction. Our proof first establishes the required estimate on the Rademacher cube, then transfers it to arbitrary product distributions by a two-copy randomization argument.

stat.ML↗

Inertia-Sensitive Kreiss Bounds for $J$-Selfadjoint Matrices

Let $A \in \mathbb{C}^{n \times n}$ have spectrum in the closed unit disk. Its maximal power growth $\operatorname{Power}(A) := \sup_{k \ge 0} \Vert{}A^k\Vert{}$ measures transient amplification, whereas the Kreiss constant $\operatorname{Kreiss}(A) := \sup_{\vert{}z\vert{}>1} (\vert{}z\vert{}-1) \Vert{}(zI-A)^{-1}\Vert{}$ measures the corresponding resolvent growth outside the disk. The classical finite-dimensional Kreiss theorem gives $\operatorname{Power}(A) \le en \operatorname{Kreiss}(A)$, and the linear dependence on $n$ is unavoidable for general matrices. We show that, for matrices selfadjoint with respect to an indefinite metric, the ambient dimension $n$ can be replaced by an effective dimension determined by the minimal polynomial and the inertia of the metric. Specifically, if $A^*J = JA$, where $J$ is a fundamental symmetry with inertia $(n-q,q)$, then $\operatorname{Power}(A) \le e \min\{d(A), 2q+1, 2(n-q)+1\} \operatorname{Kreiss}(A)$, where $d(A)$ is the degree of the minimal polynomial. Our proof requires no assumption on diagonalizability or reality on the spectrum. Instead, we associate each cyclic orbit with a finite-rank selfadjoint Hankel operator and transfer its rank and inertia to a coefficient estimate. Examples based on scaled nilpotent shifts show that the linear dependence on the smaller inertia index is asymptotically sharp, even when this index is negligible relative to the matrix size and $d(A)=n$. We also obtain scaled-disk decay estimates and weighted-norm extensions to arbitrary nonsingular Hermitian metrics.

math.FA↗

Quantitative and Uniform $L^2$ Non-Localization on Integrable Polygons

We study uniform $L^2$ non-localization of Dirichlet Laplacian eigenfunctions on planar integrable polygons, with particular emphasis on spectral degeneracy and quantitative dependence on the observation set. For a measurable set $V\subsetΩ$ of positive measure, define \[ C_2(V;Ω) := \inf_{λ\inσ(-Δ_Ω)} \inf_{0\neq u\in E_λ(Ω)} \frac{\|u\|_{L^2(V)}}{\|u\|_{L^2(Ω)}}. \] We prove that $C_2(V;Ω)>0$ for rectangles, isosceles right triangles, equilateral triangles, and hemi-equilateral triangles, uniformly over the complete eigenspaces and hence independently of spectral multiplicity. For rectangles, we obtain a quantitative refinement. If the reflected extension of $V$ has finite perimeter and $α=|V|/|Ω|$, we derive an explicit sufficient threshold $λ_*(Ω,V)$ such that every eigenfunction with $λ\geqλ_*(Ω,V)$ satisfies \[ \frac{\|u\|_{L^2(V)}}{\|u\|_{L^2(Ω)}} \geq \left[ \fracα{2} \left( 1-\frac{\sin(πα)}{πα} \right) \right]^{1/2}. \] We further establish stability under bounded real-valued potentials on the rectangular branch and under controlled spectral defects, yielding corresponding non-localization results for sufficiently accurate quasimodes and narrow spectral clusters.

math.AP↗

Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.

cs.CL↗

Decoding EEG Signals to Explore Next-Word Predictability in the Human Brain

Humans invented reading and have passed down this complex skill across generations through language. This study provides empirical evidence of the neural mechanisms underlying bottom-up (related to high-order linguistic structure) and top-down (related to next-word predictability) processes, which interact to guide comprehension during reading. While previous studies have focused on either the N400 effects of predictability or lexical categories, research on how predictability influences N400 responses across different lexical categories is limited, mainly due to constraints in publicly available datasets. Here, we examine how predictability influences brain responses, recorded at millisecond resolution using electroencephalography (EEG), with a focus on the N400 time window (300-500 ms post-stimulus) across different lexical and grammatical categories. Our results indicate that significant differences in N400 responses between high and low cloze probability levels were more pronounced for content words than function words. Among the two primary content categories, verbs exhibited greater N400 differences than nouns, while nouns carried more distinct information about their predictability than verbs. Moreover, we demonstrate that the decoding technique is more effective than the event-related potential (ERP) traditional analysis in capturing more detailed and distinct representations of cognitive processes over time.

cs.CL↗

When Generator Replay Degrades: Projected Rehearsal Orchestration for Heterogeneous Federated Class-Incremental Learning

Federated class-incremental learning (FCIL) becomes substantially harder when clients observe different label subsets, progress through tasks at different stages, and provide uneven supervision for the same semantic concepts. Existing FCIL methods often preserve old knowledge through input-space synthesis, but they can be fragile under heterogeneous task streams and difficult to transfer across modalities. To alleviate such issues, we propose PRO, a framework that replaces synthetic input replay with projected rehearsal orchestration. To remove external pretraining, we evaluate all methods under the same warmup. After this, PRO maintains compact class-level projected memories on the server and allows clients perform balanced pseudo multi-task training over current examples and old projected memories. To handle stronger representation drift, we further introduce PRO-MAX, which augments PRO with neighborhood-weighted memory alignment while preserving the same server-light principle that the server only aggregates model updates and memory statistics. Across image, text, and graph benchmarks, PRO and PRO-MAX improve retention and final utility under heterogeneous streams while remaining competitive in homogeneous FCIL. Even when baselines are given expanded replay budgets, they degrade under supervision imbalance and stage misalignment, indicating that replay quantity alone does not resolve replay-quality failures. Additional weak-task diagnostics further show that larger replay mismatch is associated with larger downstream degradation, while our method keeps projected memories better aligned with the evolving representation.

cs.LG↗

Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps

We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density. Our approach introduces Voronoi-type loss functions aligned with the gate-partition geometry and establishes finite-sample convergence rates for the maximum likelihood estimator (MLE). In over-specified models, we reveal a link between the MLE's convergence rate and the solvability of an associated system of polynomial equations characterizing near-nonidentifiable directions. For model selection, we adapt dendrograms of mixing measures to SGMoE, yielding a consistent, sweep-free selector of the number of experts that attains pointwise-optimal parameter rates under overfitting while avoiding multi-size training. Simulations on synthetic data corroborate the theory, accurately recovering the expert count and achieving the predicted rates for parameter estimation while closely approximating the regression function. Under model misspecification (e.g., $ε$-contamination), the dendrogram selection criterion is robust, recovering the true number of mixture components, while the Akaike information criterion, the Bayesian information criterion, and the integrated completed likelihood tend to overselect as sample size grows. On a maize proteomics dataset of drought-responsive traits, our dendrogram-guided SGMoE selects two experts, exposes a clear mixing-measure hierarchy, stabilizes the likelihood early, and yields interpretable genotype-phenotype maps, outperforming standard criteria without multi-size training.

stat.ML↗

High-dimensional Many-to-many-to-many Mediation Analysis

We study high-dimensional mediation analysis in which exposures, mediators, and outcomes are all multivariate, and both exposures and mediators may be high-dimensional. We formalize this as a many (exposures)-to-many (mediators)-to-many (outcomes) (MMM) mediation analysis problem. Methodologically, MMM mediation analysis simultaneously performs variable selection for high-dimensional exposures and mediators, estimates the indirect effect matrix (i.e., the coefficient matrices linking exposure-to-mediator and mediator-to-outcome pathways), and enables prediction of multivariate outcomes. Theoretically, we show that the estimated indirect effect matrices are consistent and element-wise asymptotically normal, and we derive error bounds for the estimators. To evaluate the efficacy of the MMM mediation framework, we first investigate its finite-sample performance, including convergence properties, the behavior of the asymptotic approximations, and robustness to noise, via simulation studies. We then apply MMM mediation analysis to data from the Alzheimer's Disease Neuroimaging Initiative to study how cortical thickness of 202 brain regions may mediate the effects of 688 genome-wide significant single nucleotide polymorphisms (SNPs) (selected from approximately 1.5 million SNPs) on eleven cognitive-behavioral and diagnostic outcomes. The MMM mediation framework identifies biologically interpretable, many-to-many-to-many genetic-neural-cognitive pathways and improves downstream out-of-sample classification and prediction performance. Taken together, our results demonstrate the potential of MMM mediation analysis and highlight the value of statistical methodology for investigating complex, high-dimensional multi-layer pathways in science. The MMM package is available at https://github.com/THELabTop/MMM-Mediation.

stat.ME↗

Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks

Wasserstein distributionally robust optimization (WDRO) provides a framework for adversarial robustness, yet existing methods based on global Lipschitz continuity or strong duality often yield loose upper bounds or require prohibitive computation. We address these limitations with a primal approach and adopt a notion of exact Lipschitz certificates to tighten this upper bound of WDRO. For ReLU networks, we leverage the piecewise-affine structure on activation cells to obtain an exact tractable characterization of the corresponding WDRO problem. We further extend our analysis to modern architectures with smooth activations (e.g., GELU, SiLU), such as Transformers. Additionally, we propose novel Wasserstein Distributional Attacks (WDA, WDA++) that construct candidates for the worst-case distribution. Compared to existing attacks that are restricted to point-wise perturbations, our methods offer greater flexibility in the number and location of attack points. Extensive evaluations demonstrate that our proposed framework achieves competitive robust accuracy against state-of-the-art baselines while offering tighter certificates than existing methods. Our code is available at https://github.com/OLab-Repo/WDA.

cs.LG↗

Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts

Processing high-volume, streaming data is increasingly common in modern statistics and machine learning, where batch-mode algorithms are often impractical because they require repeated passes over the full dataset. This has motivated incremental stochastic estimation methods, including the incremental stochastic Expectation-Maximization (EM) algorithm formulated via stochastic approximation. In this work, we revisit and analyze an incremental stochastic variant of the Majorization-Minimization (MM) algorithm, which generalizes incremental stochastic EM as a special case. Our approach relaxes key EM requirements, such as explicit latent-variable representations, enabling broader applicability and greater algorithmic flexibility. We establish theoretical guarantees for the incremental stochastic MM algorithm, proving consistency in the sense that the iterates converge to a stationary point characterized by a vanishing gradient of the objective. We demonstrate these advantages on a softmax-gated mixture of experts (MoE) regression problem, for which no stochastic EM algorithm is available. Empirically, our method consistently outperforms widely used stochastic optimizers, including stochastic gradient descent, root mean square propagation, adaptive moment estimation, and second-order clipped stochastic optimization. These results support the development of new incremental stochastic algorithms, given the central role of softmax-gated MoE architectures in contemporary deep neural networks for heterogeneous data modeling. Beyond synthetic experiments, we also validate practical effectiveness on two real-world datasets, including a bioinformatics study of dent maize genotypes under drought stress that integrates high-dimensional proteomics with ecophysiological traits, where incremental stochastic MM yields stable gains in predictive performance.

stat.ML↗

Fusionista2.0: Efficiency Retrieval System for Large-Scale Datasets

The Video Browser Showdown (VBS) challenges systems to deliver accurate results under strict time constraints. To meet this demand, we present Fusionista2.0, a streamlined video retrieval system optimized for speed and usability. All core modules were re-engineered for efficiency: preprocessing now relies on ffmpeg for fast keyframe extraction, optical character recognition uses Vintern-1B-v3.5 for robust multilingual text recognition, and automatic speech recognition employs faster-whisper for real-time transcription. For question answering, lightweight vision-language models provide quick responses without the heavy cost of large models. Beyond these technical upgrades, Fusionista2.0 introduces a redesigned user interface with improved responsiveness, accessibility, and workflow efficiency, enabling even non-expert users to retrieve relevant content rapidly. Evaluations demonstrate that retrieval time was reduced by up to 75% while accuracy and user satisfaction both increased, confirming Fusionista2.0 as a competitive and user-friendly system for large-scale video search.

cs.CV↗

FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text

Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Image and Bing, progress has been limited by the lack of an accessible benchmark and methods that balance performance across subtasks. Large-scale datasets such as MagicLens are comprehensive but computationally prohibitive, while existing models often favor either visual or compositional queries. We introduce FIGROTD, a lightweight yet high-quality IGROT dataset with 16,474 training triplets and 1,262 test triplets across CIR, SBIR, and CSTBIR. To reduce redundancy, we propose the Variance Guided Feature Mask (VaGFeM), which selectively enhances discriminative dimensions based on variance statistics. We further adopt a dual-loss design (InfoNCE + Triplet) to improve compositional reasoning. Trained on FIGROTD, VaGFeM achieves competitive results on nine benchmarks, reaching 34.8 mAP@10 on CIRCO and 75.7 mAP@200 on Sketchy, outperforming stronger baselines despite fewer triplets.

cs.IR↗