SearcharxivSearch

arXiv subjects

Xiao-Yu Guo

Publications and source records attributed to Xiao-Yu Guo.

At least 19 recordsLinked to original sources

Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification

Test-Time Adaptation (TTA) seeks to improve model robustness under distribution shifts by adapting parameters using unlabeled target data. However, in the absence of supervision, entropy-based adaptation is fundamentally underconstrained: multiple distinct parameter updates can achieve similarly low entropy while inducing drastically different decision boundaries. This phenomenon, known as underspecification, renders standard TTA brittle and prone to collapse into spurious modes. In this work, we reinterpret TTA through a posterior-inspired lens induced by entropy minimization, where low-entropy solutions define a pseudo-likelihood over parameters. Instead of committing to a single point estimate, we introduce a particle-based diversification framework that explores multiple plausible adaptation trajectories simultaneously. Our method can be viewed as a structured exploration of multiple plausible adaptation solutions, implemented through multi-level diversification at the output, parameter, optimizer, and input levels. Crucially, the framework acts as a plug-and-play wrapper compatible with existing TTA methods. Extensive experiments on challenging benchmarks demonstrate consistent gains in stability and robustness, achieving improvements of 3-4% under mixed shifts, 2-3% with batch size one, and 1-2.5% under label shifts, outperforming state-of-the-art baselines. Our results suggest that treating TTA as a multi-hypothesis inference problem, rather than a single-point optimization task, is key to mitigating underspecification and enabling reliable real-world deployment.

cs.CV

Bayesian Low-Rank LeArning (Bella): A Practical Approach to Bayesian Neural Networks

Computational complexity of Bayesian learning is impeding its adoption in practical, large-scale tasks. Despite demonstrations of significant merits such as improved robustness and resilience to unseen or out-of-distribution inputs over their non- Bayesian counterparts, their practical use has faded to near insignificance. In this study, we introduce an innovative framework to mitigate the computational burden of Bayesian neural networks (BNNs). Our approach follows the principle of Bayesian techniques based on deep ensembles, but significantly reduces their cost via multiple low-rank perturbations of parameters arising from a pre-trained neural network. Both vanilla version of ensembles as well as more sophisticated schemes such as Bayesian learning with Stein Variational Gradient Descent (SVGD), previously deemed impractical for large models, can be seamlessly implemented within the proposed framework, called Bayesian Low-Rank LeArning (Bella). In a nutshell, i) Bella achieves a dramatic reduction in the number of trainable parameters required to approximate a Bayesian posterior; and ii) it not only maintains, but in some instances, surpasses the performance of conventional Bayesian learning methods and non-Bayesian baselines. Our results with large-scale tasks such as ImageNet, CAMELYON17, DomainNet, VQA with CLIP, LLaVA demonstrate the effectiveness and versatility of Bella in building highly scalable and practical Bayesian deep models for real-world applications.

cs.LG

ProtDAT: A Unified Framework for Protein Sequence Design from Any Protein Text Description

Protein design has become a critical method in advancing significant potential for various applications such as drug development and enzyme engineering. However, protein design methods utilizing large language models with solely pretraining and fine-tuning struggle to capture relationships in multi-modal protein data. To address this, we propose ProtDAT, a de novo fine-grained framework capable of designing proteins from any descriptive protein text input. ProtDAT builds upon the inherent characteristics of protein data to unify sequences and text as a cohesive whole rather than separate entities. It leverages an innovative multi-modal cross-attention, integrating protein sequences and textual information for a foundational level and seamless integration. Experimental results demonstrate that ProtDAT achieves the state-of-the-art performance in protein sequence generation, excelling in rationality, functionality, structural similarity, and validity. On 20,000 text-sequence pairs from Swiss-Prot, it improves pLDDT by 6%, TM-score by 0.26, and reduces RMSD by 1.2 Å, highlighting its potential to advance protein design.

cs.AI

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. % Moreover, spatial reasoning steps are much less accurate than non-spatial ones across MLLMs. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. Spatial-MM benchmark is available at: https://github.com/FatemehShiri/Spatial-MM

cs.CV

Triangle and box diagrams in coupled-channel systems from the chiral Lagrangian

We perform an analysis of triangle- and box-loop contributions to the generalized potential in the scattering of Goldstone bosons off the J^P= 0^- and 1^- charmed mesons. Particular emphasis is put on the use of on-shell mass parameters in such contributions in terms of a renormalization scheme that ensures the absence of power-counting violating terms. This is achieved with a systematically extended set of Passarino--Veltman basis functions, that leads to manifest power-counting conserving one-loop expressions and avoids the occurrence of superficial kinematical singularities. Compact expressions to chiral order three and four are presented that are particularly useful in coding such coupled-channel systems. Our formal results are generic and prepare analogous computations for other systems, like meson-baryon scattering from the chiral Lagrangian.

hep-ph

DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding

Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions. One representative benchmark for its study is Social Intelligence Queries (Social-IQ), a dataset of multiple-choice questions on videos of complex social interactions. We define a comprehensive methodology to study the soundness of Social-IQ, as the soundness of such benchmark datasets is crucial to the investigation of the underlying research problem. Our analysis reveals that Social-IQ contains substantial biases, which can be exploited by a moderately strong language model to learn spurious correlations to achieve perfect performance without being given the context or even the question. We introduce DeSIQ, a new challenging dataset, constructed by applying simple perturbations to Social-IQ. Our empirical analysis shows DeSIQ significantly reduces the biases in the original Social-IQ dataset. Furthermore, we examine and shed light on the effect of model size, model style, learning settings, commonsense knowledge, and multi-modality on the new benchmark performance. Our new dataset, observations and findings open up important research questions for the study of social intelligence.

cs.CL

Low-energy constants in the chiral Lagrangian with baryon octet and decuplet fields from Lattice QCD data on CLS ensembles

We perform an analysis of Lattice QCD data on baryon octet and decuplet masses based on the chiral SU(3) Lagrangian. Low-energy constants (LEC) are adjusted to describe baryon masses from a large set of CLS ensembles, where finite-box and discretization effects are considered. The set is successfully compared against previous Lattice QCD data from ensembles generated with distinct QCD actions by the ETMC, QCDSF-UKQCD and HSC groups. Discretization effects are modelled by the use of action and lattice-scale dependent leading orders LEC, where uniform values are imposed in the limit of vanishing lattice scales. From the CLS data set we extract a pion-nucleon sigma term, $σ_{πN}= 58.7(1.2)$ MeV, compatible with its empirical value and a sizeable strangeness content of the nucleon with $σ_{sN} = -316(76)$ MeV.

hep-lat

Complex Reading Comprehension Through Question Decomposition

Multi-hop reading comprehension requires not only the ability to reason over raw text but also the ability to combine multiple evidence. We propose a novel learning approach that helps language models better understand difficult multi-hop questions and perform "complex, compositional" reasoning. Our model first learns to decompose each multi-hop question into several sub-questions by a trainable question decomposer. Instead of answering these sub-questions, we directly concatenate them with the original question and context, and leverage a reading comprehension model to predict the answer in a sequence-to-sequence manner. By using the same language model for these two components, our best seperate/unified t5-base variants outperform the baseline by 7.2/6.1 absolute F1 points on a hard subset of DROP dataset.

cs.CL

A coupled-channel system with anomalous thresholds and unitarity

We consider the isospin one-half example system, with $D \,π, D\,η, D_s \bar K, D^* π, D^*η, D^*\bar K$ coupled channels in the $J^P = 1^-$ partial wave, chosen such that various phenomena that come with the opening of an anomalous threshold can be illustrated in a step-wise procedure by a suitable variation of up, down and strange quark masses. We use a set of LEC in the chiral Lagrangian that were adjusted to a large set of Lattice QCD results. The six phase shifts and inelasticity parameters are presented for various choices of the pion mass. For a pion mass of 150 MeV there are no anomalous thresholds encountered. The small change from 150 MeV to 145 MeV pion mass causes a dramatic impact of the anomalous threshold on the phase shifts.

hep-ph

Coupled-channel dynamics with chiral long-range forces in the open-charm sector of QCD

We perform an analysis of Lattice QCD data in the open-charm sector based on the chiral SU(3) Lagrangian. The low-energy constants are adjusted to recover the open-charm meson masses on Lattice QCD ensembles from HPQCD, ETMC and HSC with pion and kaon masses smaller than 550 MeV. A significant set of low-energy parameters is obtainable only if the most recent information from HSC on scattering observables is included in our global fit. For the first time our analysis considers the effect of left-hand cuts as developed in terms of a generalized potential approach (GPA) previously by one of the authors. Here we use coupled-channel interaction terms at the one-loop level. The elastic s-wave and p-wave $D\,π$, $D K $ and $D \bar K $ scattering phase shifts on ensembles with nominal pion masses of about 239 MeV and 391 MeV are reproduced faithfully. Based on such low-energy parameters we predict s- and p-wave phase shifts and inelasticities at physical quark masses, where the statistical uncertainties in the phase shifts are smaller than 1 degree always. Most striking would be the exotic s-wave $D_s π$ channel, for which we predict a resonance state at about 2.287 GeV where the phase shift passes through 90 degrees.

hep-ph

Teaching Neural Module Networks to Do Arithmetic

Answering complex questions that require multi-step multi-type reasoning over raw text is challenging, especially when conducting numerical reasoning. Neural Module Networks(NMNs), follow the programmer-interpreter framework and design trainable modules to learn different reasoning skills. However, NMNs only have limited reasoning abilities, and lack numerical reasoning capability. We up-grade NMNs by: (a) bridging the gap between its interpreter and the complex questions; (b) introducing addition and subtraction modules that perform numerical reasoning over numbers. On a subset of DROP, experimental results show that our proposed methods enhance NMNs' numerical reasoning skills by 17.7% improvement of F1 score and significantly outperform previous state-of-the-art models.

cs.CL

Chiral excitations of open-beauty systems

We study the scattering of open-beauty mesons and Goldstone bosons as predicted by the chiral SU(3) Lagrangian. The impact of subleading order chiral interactions to systems with $J^P= 0^+$ and $J^P=1^+$ quantum numbers is worked out. We estimate the relevant low-energy coefficients from the open-charm sector, for which their values have been determined previously from sets of QCD lattice data. The leading order heavy-quark symmetry breaking effects are estimated by matching the $B$-meson ground-state chiral mass formula to the mass formula from the heavy-quark effective theory. We make refined predictions for the flavor anti-triplet and sextet resonances that are generated dynamically by coupled-channel interactions.

hep-ph

Improving Numerical Reasoning Skills in the Modular Approach for Complex Question Answering on Text

Numerical reasoning skills are essential for complex question answering (CQA) over text. It requires opertaions including counting, comparison, addition and subtraction. A successful approach to CQA on text, Neural Module Networks (NMNs), follows the programmer-interpreter paradigm and leverages specialised modules to perform compositional reasoning. However, the NMNs framework does not consider the relationship between numbers and entities in both questions and paragraphs. We propose effective techniques to improve NMNs' numerical reasoning capabilities by making the interpreter question-aware and capturing the relationship between entities and numbers. On the same subset of the DROP dataset for CQA on text, experimental results show that our additions outperform the original NMNs by 3.0 points for the overall F1 score.

cs.CL

From lattice QCD to predictions of scattering phase shifts at the physical point

The Hadron Spectrum Collaboration (HSC) presented new results on two of their ensembles for s-wave scattering phase shifts in the open-charm sector of QCD. For such ensembles we have made predictions that are based on the chiral Lagrangian that were published two years ago. In this talk we confront our phase shifts with those of HSC. A remarkably consistent picture emerges. In particular there is mounting evidence for the existence of a flavor-sextet state in the $D π$ and $D^*π$ channels, that show a striking quark-mass dependence.

hep-lat

Understanding Unnatural Questions Improves Reasoning over Text

Complex question answering (CQA) over raw text is a challenging task. A prominent approach to this task is based on the programmer-interpreter framework, where the programmer maps the question into a sequence of reasoning actions which is then executed on the raw text by the interpreter. Learning an effective CQA model requires large amounts of human-annotated data,consisting of the ground-truth sequence of reasoning actions, which is time-consuming and expensive to collect at scale. In this paper, we address the challenge of learning a high-quality programmer (parser) by projecting natural human-generated questions into unnatural machine-generated questions which are more convenient to parse. We firstly generate synthetic (question,action sequence) pairs by a data generator, and train a semantic parser that associates synthetic questions with their corresponding action sequences. To capture the diversity when applied tonatural questions, we learn a projection model to map natural questions into their most similar unnatural questions for which the parser can work well. Without any natural training data, our projection model provides high-quality action sequences for the CQA task. Experimental results show that the QA model trained exclusively with synthetic data generated by our method outperforms its state-of-the-art counterpart trained on human-labeled data.

cs.CL

A generalized Higgs potential with two degenerate minima for a dark QCD matter scenario

We consider the Higgs potential in generalizations of the Standard Model. The possibility of the potential to develop two almost degenerate minima is explored. This would imply that QCD matter at two distinct sets of quark masses is relevant for astrophysics and cosmology. If in the exotic minimum the QCD matter ground state is electromagnetically neutral, dark matter may consist of QCD matter and antimatter in bubbles of the Higgs field. We predict an abundance of gamma rays in the few MeV region as messengers of dark matter regions in space. In addition the ratio of dark matter to normal matter is expected to show a time dependence.

hep-ph

On a first order transition in QCD with up, down and strange quarks

We consider the quark-mass dependence of the baryon octet and decuplet ground state masses. It is predicted that QCD dynamics implies a first order transition when increasing the strange quark mass from its chiral limit towards its physical value. Our claim relies on a global fit to the available QCD lattice data on such baryon masses. Quantitative results based on an application of the chiral SU(3) Lagrangian at N$^3$LO are discussed. We predict an anomalous sector of QCD where stable baryonic matter would be composed of lambda or anti-lambda particles rather than nucleons and anti-nucleons.

hep-lat

Low-energy constants from charmed baryons on QCD lattices

We study the light quark-mass dependence of charmed baryon masses as measured by various QCD lattice collaborations. A global fit to such data based on the chiral SU(3) Lagrangian is reported on. All low-energy constants that are relevant at next-to-next-to-next-to-leading order (N$^3$LO) are determined from the lattice data sets where constraints from sum rules as they follow from large-Nc QCD at subleading order are considered. The expected hierarchy for the low-energy constants in the 1/Nc expansion is confirmed by our global fits to the lattice data. With our results the low-energy interaction of the Goldstone bosons with the charmed baryon ground states is well constrained and the path towards realistic coupled-channel computations in this sector of QCD is prepared.

hep-lat