Searcharxiv⌕ Search

arXiv subjects

Xiaohua Wu

Publications and source records attributed to Xiaohua Wu.

At least 19 recordsLinked to original sources

HIFICL: High-Fidelity In-Context Learning for Multimodal Tasks

In-Context Learning (ICL) is a significant paradigm for Large Multimodal Models (LMMs), using a few in-context demonstrations (ICDs) for new task adaptation. However, its performance is sensitive to demonstration configurations and computationally expensive. Mathematically, the influence of these demonstrations can be decomposed into a dynamic mixture of the standard attention output and the context values. Current approximation methods simplify this process by learning a "shift vector". Inspired by the exact decomposition, we introduce High-Fidelity In-Context Learning (HIFICL) to more faithfully model the ICL mechanism. HIFICL consists of three key components: 1) a set of "virtual key-value pairs" to act as a learnable context, 2) a low-rank factorization for stable and regularized training, and 3) a simple end-to-end training objective. From another perspective, this mechanism constitutes a form of context-aware Parameter-Efficient Fine-Tuning (PEFT). Extensive experiments show that HiFICL consistently outperforms existing approximation methods on several multimodal benchmarks. The code is available at https://github.com/bbbandari/HiFICL.

cs.CV↗

OMGs: A multi-agent system supporting MDT decision-making across the ovarian tumour care continuum

Ovarian tumour management has increasingly relied on multidisciplinary tumour board (MDT) deliberation to address treatment complexity and disease heterogeneity. However, most patients worldwide lack access to timely expert consensus, particularly in resource-constrained centres where MDT resources are scarce or unavailable. Here we present OMGs (Ovarian tumour Multidisciplinary intelligent aGent System), a multi-agent AI framework where domain-specific agents deliberate collaboratively to integrate multidisciplinary evidence and generate MDT-style recommendations with transparent rationales. To systematically evaluate MDT recommendation quality, we developed SPEAR (Safety, Personalization, Evidence, Actionability, Robustness) and validated OMGs across diverse clinical scenarios spanning the care continuum. In multicentre re-evaluation, OMGs achieved performance comparable to expert MDT consensus ($4.45 \pm 0.30$ versus $4.53 \pm 0.23$), with higher Evidence scores (4.57 versus 3.92). In prospective multicentre evaluation (59 patients), OMGs demonstrated high concordance with routine MDT decisions. Critically, in paired human-AI studies, OMGs most substantially enhanced clinicians' recommendations in Evidence and Robustness, the dimensions most compromised when multidisciplinary expertise is unavailable. These findings suggest that multi-agent deliberative systems can achieve performance comparable to expert MDT consensus, with potential to expand access to specialized oncology expertise in resource-limited settings.

cs.CL↗

Synthetic Spatiotemporal Plasmonic Vortices On Chip

Spatiotemporal vortices are polychromatic modes that intertwine orbital angular momentum (OAM) in space and time. Here we introduce a new class of such vortices, spatiotemporal plasmonic vortices (STPVs), carrying nontrivial topological spin textures. They are generated by chronotopic interference of temporally delayed plasmonic eigen-vortices, where a $π$-phase dislocation in the space-frequency domain maps into a 2$π$ spiraling phase in space-time, with the resulting focus-defocus dynamics emulate U(1) gauge transitions. Using interferometric time-resolved photoemission electron microscopy (ITR-PEEM), we directly image their nanometer-attosecond (nano-atto) evolution and control vortex number and position. Quantum-path analysis of coherent two-photon photoemission (2PP) processes reveals the nonlinear plasmonic polarization fields and angular-momentum conservation, establishing STPVs as a platform for probing spatiotemporally structured quantum matter.

cond-mat.mes-hall↗

BigBang-Proton Technical Report: Next-Word-Prediction is Scientific Multitask Learner

We introduce BigBang-Proton, a unified sequence-based architecture for auto-regressive language modeling pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientific multi-task learner. BigBang-Proton incorporates three fundamental innovations compared to mainstream general-purpose LLMs: Theory-Experiment Learning paradigm aligns large-scale numerical experimental data with theoretical text corpora; Binary Patch Encoding replaces byte pair encoding(BPE) tokenization; Monte Carlo Attention substitutes traditional transformer architectures. Through next-word-prediction pretraining on cross-discipline scientific datasets of real-world problems mixed with general textual corpus, followed by fine-tuning and inference on downstream tasks, BigBang-Proton demonstrates 100\% accuracy in up to 50-digit arithmetic addition operations, performance on par with leading specialized models in particle physics jet tagging, matching MAE of specialized models in inter-atomic potential simulation, performance comparable to traditional spatiotemporal models in water quality prediction, and benchmark-exceeding performance in genome modeling. These results prove that language-guided scientific computing can match or exceed the performance of task-specific scientific models while maintaining multitask learning capabilities. We further hypothesize to scale the pretraining to the universe scale as a fundamental step toward developing material world foundational model.

cs.LG↗

Magnetic order dependent photoluminescence from high energy excitons in hBN protected few-layer CrSBr

The detection and manipulation of the spin configurations in layered magnetic semiconductors hold significant interest for developing spintronic devices in two-dimensional limit. In this letter, we report a systematical study on the photoluminescence (PL) from the high energy excitons in few-layer CrSBr and its application on detecting the spin configurations. Besides the broad excitonic emission peak (Xl) at around 1.34 eV, we also observed another strong excitonic emission peak (Xh) at around 1.37 eV in hBN encapsulated 2L sample, which splits into two peaks in 3L and 4L samples. With help of the first principles calculations, we conclude that the Xh peak is associated with the transition between the top valence band and the second lowest conduction band, which is forbidden by the inversion symmetry in 1L CrSBr. Furthermore, the position and intensity of the Xh peak are strongly dependent on the interlayer magnetic order of the CrSBr samples, which provides an efficient way to probe their spin configurations. In addition, when the magnetic field is applied at the easy axis direction, we resolve an intermediate magnetic state besides the antiferromagnetic and ferromagnetic states in 3L and 4L samples. Our results reveal few-layer CrSBr as an ideal platform to study the interaction between the excitons and magnetism.

cond-mat.mes-hall↗

Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection

Advances in computer vision and deep learning have blurred the line between deepfakes and authentic media, undermining multimedia credibility through audio-visual forgery. Current multimodal detection methods remain limited by unbalanced learning between modalities. To tackle this issue, we propose an Audio-Visual Joint Learning Method (MACB-DF) to better mitigate modality conflicts and neglect by leveraging contrastive learning to assist in multi-level and cross-modal fusion, thereby fully balancing and exploiting information from each modality. Additionally, we designed an orthogonalization-multimodal pareto module that preserves unimodal information while addressing gradient conflicts in audio-video encoders caused by differing optimization targets of the loss functions. Extensive experiments and ablation studies conducted on mainstream deepfake datasets demonstrate consistent performance gains of our model across key evaluation metrics, achieving an average accuracy of 95.5% across multiple datasets. Notably, our method exhibits superior cross-dataset generalization capabilities, with absolute improvements of 8.0% and 7.7% in ACC scores over the previous best-performing approach when trained on DFDC and tested on DefakeAVMiT and FakeAVCeleb datasets.

cs.CV↗

Direct numerical simulation benchmarks for the prediction of boundary layer bypass transition in the narrow sense

We report a comprehensive set of direct numerical simulation benchmarks of bypass transition in the narrow sense with inlet freestream turbulent intensity levels of 0.75%, 1.5%, 2.25%, 3.0%, and 6.0%, respectively. Detailed descriptions of length scales and the rate of viscous dissipation are provided. We ask two key physical questions. First, how do the decay rates and length scales of freestream turbulence over a transitional and turbulent boundary layer compare to those in spatially developing isotropic turbulence without the wall? Second, what bypass mechanisms drive turbulent spot inception at the intermediate rage of freestream turbulence intensity level? We find that the boundary-layer freestream turbulence decay and length scales evolve similarly to their spatially developing isotropic turbulence flow without the wall counterparts. We also present evidence of the coexistence of two turbulent spot inception mechanisms at the inlet FST level of 2.25%: the long low-speed streak primary and secondary instabilities (only in lower inlet FST levels) and the self-amplifying process of oblique vortex filaments interacting with a Delta-shaped low-speed patch underneath (prevailing only in higher inlet FST levels).

physics.flu-dyn↗

Random Forest-of-Thoughts: Uncertainty-aware Reasoning for Computational Social Science

Social surveys in computational social science are well-designed by elaborate domain theories that can effectively reflect the interviewee's deep thoughts without concealing their true feelings. The candidate questionnaire options highly depend on the interviewee's previous answer, which results in the complexity of social survey analysis, the time, and the expertise required. The ability of large language models (LLMs) to perform complex reasoning is well-enhanced by prompting learning such as Chain-of-thought (CoT) but still confined to left-to-right decision-making processes or limited paths during inference. This means they can fall short in problems that require exploration and uncertainty searching. In response, a novel large language model prompting method, called Random Forest of Thoughts (RFoT), is proposed for generating uncertainty reasoning to fit the area of computational social science. The RFoT allows LLMs to perform deliberate decision-making by generating diverse thought space and randomly selecting the sub-thoughts to build the forest of thoughts. It can extend the exploration and prediction of overall performance, benefiting from the extensive research space of response. The method is applied to optimize computational social science analysis on two datasets covering a spectrum of social survey analysis problems. Our experiments show that RFoT significantly enhances language models' abilities on two novel social survey analysis problems requiring non-trivial reasoning.

cs.CL↗

Development of transitional Reynolds number correlation and assessment of RANS for predictions of bypass transition

We present direct numerical simulations (DNSs) of bypass transition over a flat plate with inlet freestream turbulence intensity levels of 0.75%, 1.5%, 2.25%, 3.0%, and 6.0%, respectively. A new definition of the transition intermittency is proposed based on the mean skin friction. Based on these, we develop an intermittency correlation to predict flow transition. The proposed model is consistent with the classical correlation of Abu-Ghannam and Shaw and reasonably predicts transition Reynolds number (within 10.8% error) for the experiments of Fransson & Shahinfar (2020). Accompanying Reynolds-averaged Navier-Stokes (RANS) simulations for our DNS cases simulations are performed. The RANS results are sensitive to the specification of the inlet turbulence length scale and overpredict (underpredict) the growth of the integral flow scales across the boundary layer during transitional stages when the inlet freestream turbulence is low (high), respectively.

physics.flu-dyn↗

Thermodynamic work and heat for a quantum process: Approach by Hamiltonian decomposition

The separation of internal energy into heat and work in quantum thermodynamics is a controversial issue for a long time, and we revisit and solve this problem in this work. It is shown that the Hamiltonian plays dual roles for a quantum system, and by decomposing the interaction Hamiltonian between system and environment accordingly, an ``effective Hamiltonian" for an open quantum system can be proposed. The explicit expression of the effective Hamiltonian is obtained systematically, and as a consequence, the internal energy of an open quantum system can be well defined, leading to the reasonable definitions of work and heat for a general quantum process.

quant-ph↗

Quantum steering in a star network

In this work, we will consider the star network scenario where the central party is trusted while all the edge parties (with a number of $n$) are untrusted. Network steering is defined with an $n$ local hidden state model which can be viewed as a special kind of $n$ local hidden variable model. Two different types of sufficient criteria, nonlinear steering inequality and linear steering inequality will be constructed to verify the quantum steering in a star network. Based on the linear steering inequality, how to detect the network steering with a fixed measurement will be discussed.

quant-ph↗

Primary and Secondary Factor Consistency as Domain Knowledge to Guide Happiness Computing in Online Assessment

Happiness computing based on large-scale online web data and machine learning methods is an emerging research topic that underpins a range of issues, from personal growth to social stability. Many advanced Machine Learning (ML) models with explanations are used to compute the happiness online assessment while maintaining high accuracy of results. However, domain knowledge constraints, such as the primary and secondary relations of happiness factors, are absent from these models, which limits the association between computing results and the right reasons for why they occurred. This article attempts to provide new insights into the explanation consistency from an empirical study perspective. Then we study how to represent and introduce domain knowledge constraints to make ML models more trustworthy. We achieve this through: (1) proving that multiple prediction models with additive factor attributions will have the desirable property of primary and secondary relations consistency, and (2) showing that factor relations with quantity can be represented as an importance distribution for encoding domain knowledge. Factor explanation difference is penalized by the Kullback-Leibler divergence-based loss among computing models. Experimental results using two online web datasets show that domain knowledge of stable factor relations exists. Using this knowledge not only improves happiness computing accuracy but also reveals more significative happiness factors for assisting decisions well.

cs.LG↗

On-demand single photon emission in the telecom C-band from nanowire-based quantum dots

Single photon sources operating on-demand at telecom wavelengths are required in fiber-based quantum secure communication technologies. In this work we demonstrate single photon emission from position-controlled nanowire quantum dots emitting at λ > 1530 nm. Using above-band pulsed excitation, we obtain single photon purities of g(2)(0) = 0.062. These results represent an important step towards the scalable manufacture of high efficiency, high rate single photon emitters in the telecom C-band.

quant-ph↗

Detecting the genuine multipartite two-way steerability with linear steering inequalities

According to the fundamental idea that a steering inequality can be constructed by just considering the measurements performed by Bob, and from the definitions of steering from Alice to Bob, a general scheme for designing two different kinds of linear steering inequalities (LSIs) is developed to detect the two-way steerability for bipartite system and the genuine multipartite two-way steerability for multipartite system, respectively. Besides the LSIs constructed from the known one-way criteria and the Bell operators, several other types of LSIs are also considered.

quant-ph↗

Model form uncertainty quantification of Reynolds-averaged Navier-Stokes modeling of flows over a SD7003 airfoil

It is well known that the Boussinesq turbulent viscosity hypothesis can yield inaccurate predictions when complex f low features are involved, e.g. laminar-turbulent transition. The focus of the study is to explore the capability of a physics-based uncertainty quantification (UQ) approach to quantify the model-form uncertainty in Reynolds-averaged Naiver-Stokes (RANS) simulations of laminar-turbulent transitional flows over an Selig-Donovan (SD) 7003 airfoil. This methodology perturbs the modeled Reynolds stress tensor in the momentum equations; perturbations are injected into the amplitude, eigenvalues and eigenvectors of the anisotropy Reynolds stress tensor undergone an eigen-decomposition. In this study, our analyses focus upon the amplitude perturbation. We observed a monotonic behavior of the magnitude of the predicted uncertainty bounds for different quantities of interest. High-order regressions based on the turbulence kinetic energy discrepancies are used to develop a novel switch marker function Mk to introduce perturbations in a non-uniform manner over different regions of the domain based upon prior knowledge of the limitations of the model. Importantly, the compound effect of Mk and eigenvalue perturbations show a synergy behavior, e.g., dramatically increased uncertainty bounds to account for the discrepancy in the RANS prediction; and the Mk function effectively avoids over-perturbation to the amplitude of the anisotropy Reynolds stress tensor. In this context, regression based amplitude perturbation of the anisotropy Reynolds stress tensor makes a new contribution to the RANS UQ methodology in the simulations of the airfoil transitional flows, which shows very encouraging results.

physics.flu-dyn↗

Quantification of Reynolds-averaged-Navier-Stokes model form uncertainty in transitional boundary layer and airfoil flows

It is well known that Boussinesq turbulent-viscosity hypothesis can introduce uncertainty in predictions for complex flow features such as separation, reattachment, and laminar-turbulent transition. This study adopts a recent physics-based uncertainty quantification (UQ) approach to address such model form uncertainty in Reynolds-averaged Naiver- Stokes (RANS) simulations. Thus far, almost all UQ studies have focused on quantifying the model form uncertainty in turbulent flow scenarios. The focus of the study is to advance our understanding of the performance of the UQ approach on two different transitional flow scenarios: a flat plate and a SD7003 airfoil, to close this gap. For the T3A (flat-plate flow) flow, most of the model form uncertainty is concentrated in the laminar-turbulent transition region. For the SD7003 airfoil flow, the eigenvalue perturbations reveal a decrease as well as an increase in the length of the separation bubble. As a consequence, the uncertainty bounds successfully encompass the reattachment point. Likewise, the region of reverse flow that appear in the separation bubble is either suppressed or bolstered by the eigenvalue perturbations. In this context, the UQ methodology is applied to transition and show great results. This is the first successful RANS UQ study for transitional flows.

physics.flu-dyn↗

New construction of nine-qubit error-correcting code

We report new construction of nine-qubit error-correcting code, which introduces two new nine-qubit codes and one new three-qubit code. Because both the new two nine-qubit codes have the normal logical operators, as opposed to the nine-qubit Shor code, it results in different performance when the three codes are applied in concatenated quantum error-correction. On the other hand, one of the two nine-qubit codes has the same stabilizer generators as the nine-qubit Shor code, they are more suitable for the high-wight bit-flip noise, and the other code has the different stabilizer generators, which is more suitable for the high-wight phase-flip noise. This work is enlightening to the construction of quantum error-correcting codes, and adds more options for optimizing the performance of quantum error-correction.

quant-ph↗

Classification and purification for the independent quantum channel through quantum error-correction

The essence of quantum error-correction is to use redundant Hilbert space to identify and correct errors, and the channel fidelity of the quantum channel does not affect which errors can be identified and corrected. Based on this, it is found that quantum error-correction can be used to classify the independent quantum channel into 5 types, and 4 of the 5 types can be purified. It is found in quantum error-correction, the decoherence of quantum state may be related to the degree of identification for the state under quantum noise, and the results of this work confirmed that the degree of purity of quantum channel determines its ability to retain the quantum property of the quantum state, not the fidelity. In this work, the identification of the independent Pauli channels by quantum error-correction is demonstrated.

quant-ph↗