SearcharxivSearch

arXiv subjects

Hongxuan Liu

Publications and source records attributed to Hongxuan Liu.

7 recordsLinked to original sources

Towards Trustworthy Physical AI: From Theory to Practice Across Life Cycle

Physical AI refers to AI systems that understand, reason about, and act in accordance with the physical world and its underlying laws, dynamics, and constraints. Unlike conventional AI systems, physical AI interacts continuously with uncertain physical environments, and its actions produce consequences that are physically irreversible. As existing trustworthy AI frameworks have been developed primarily for digital AI systems, they do not fully capture the distinctive challenges of physical AI, such as physical safety, cyber-physical security, and physical manufacturing process. To address this gap, we present a survey of trustworthy physical AI principles. First, we characterize the core capabilities and challenges of physical AI. Second, we examine the role of physics in AI. Third, we trace the end-to-end physical AI life cycle across five core stages and introduce Trustworthy Physical AI Operationalization (T-PAIO). Fourth, we develop the Trustworthy Physical AI (T-PAI) framework, a theoretical framework that organizes key trustworthiness principles and provides a foundation for governing trustworthy physical AI systems.

cs.AI

FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time

Tandem mass spectrometry is prominent in scientific discovery workflows for identifying unknown small molecules, yet high-throughput structural elucidation remains challenging. While recent autoregressive and graph diffusion models have shown promise in de novo elucidation, performance remains limited by poor scalability during both training and inference time. In this work, we present FRIGID, a framework with a novel diffusion language model that generates molecular structures conditioned on mass spectra via intermediate fingerprint representations and determined chemical formulae, training at the scale of hundreds of millions of unlabeled structures. We then demonstrate how forward fragmentation models enable inference-time scaling by identifying spectrum-inconsistent fragments and refining them through targeted remasking and denoising. While FRIGID already achieves strong performance with its diffusion base, inference-time scaling significantly improves its accuracy, surpassing 18% Top-1 accuracy on the challenging MassSpecGym benchmark and tripling the Top-1 accuracy of the leading methods on NPLIB1. Further empirical analyses show that FRIGID exhibits log-linear performance scaling with increasing inference-time compute, opening a promising new direction for continued improvements in de novo structural elucidation. FRIGID code is publicly available at https://github.com/coleygroup/FRIGID.

cs.LG

MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery

Reliable benchmarking is critical for developing machine learning models for tandem mass spectrometry (MS/MS) based molecule discovery. Subtle issues in experimental design and model evaluation procedures can degrade the trustworthiness of such benchmarks and lead to erroneous conclusions. We conduct a thorough review of model evaluation issues in the recent MS/MS machine learning literature, using the standard MassSpecGym benchmark suite as a case study to illustrate the impact of these issues. We find evaluation issues in at least 17 of 26 papers reporting MassSpecGym benchmark results in the first year of its adoption. We isolate three classes of failures: (i) data leakage, (ii) shortcut learning, and (iii) implementation bugs and metric divergence. Through extensive experimentation and code replication, we quantify the impact of these issues and show how they corrupt the evaluation standards MassSpecGym was designed to enforce. We distill our findings into recommendations generalizable to MS/MS challenges, benchmarks, and custom evaluation setups. We also release MassSpecGym v1.5, an implementation of our recommendations in the MassSpecGym benchmarking suite which addresses the failure modes identified in this audit. MassSpecGym v1.5 is publicly available at https://github.com/pluskal-lab/MassSpecGym.

cs.LG

PerfCoder: Large Language Models for Interpretable Code Performance Optimization

Large language models (LLMs) have achieved remarkable progress in automatic code generation, yet their ability to produce high-performance code remains limited--a critical requirement in real-world software systems. We argue that current LLMs struggle not only due to data scarcity but, more importantly, because they lack supervision that guides interpretable and effective performance improvements. In this work, we introduce PerfCoder, a family of LLMs specifically designed to generate performance-enhanced code from source code via interpretable, customized optimizations. PerfCoder is fine-tuned on a curated collection of real-world optimization trajectories with human-readable annotations, and preference-aligned by reinforcement fine-tuning using runtime measurements, enabling it to propose input-specific improvement strategies and apply them directly without relying on iterative refinement. On the PIE code performance benchmark, PerfCoder surpasses all existing models in both runtime speedup and effective optimization rate, demonstrating that performance optimization cannot be achieved by scale alone but requires optimization stratetgy awareness. In addition, PerfCoder can generate interpretable feedback about the source code, which, when provided as input to a larger LLM in a planner-and-optimizer cooperative workflow, can further improve outcomes. Specifically, we elevate the performance of 32B models and GPT-5 to new levels on code optimization, substantially surpassing their original performance.

cs.SE

Ultra-low-loss slow-light thin-film lithium-niobate optical modulator

Electro-optic modulators for next-generation optical interconnects require low loss-efficiency products, compact footprints, high modulation efficiency, broad bandwidths, and low losses. Here we propose and demonstrate a low-loss high-efficiency thin-film lithium-niobate Mach Zehnder modulator enabled by a novel ultralow-loss slow-light structure based on apodized gratings in cascade. The present loss-engineered slow-light structure achieves excess losses as low as 0.6 dB/mm experimentally, which is tens of times lower than conventional slow-light structures, and a high modulation bandwidth up to 320GHz in theory is achieved with optimally-designed capacitively-loaded traveling-wave electrodes. Experimentally, the fabricated slow-light modulator with a 2.8-mm-long modulation region has an ultra-low loss-efficiency product of 7.4 VdB and a flat electro-optic response up to 67 GHz, enabling 100-Gbps on-off keying with high ERs of 4.5 dB at a low driving voltage of 2Vpp, while 200-Gbps PAM4 and 150-Gbps PAM8 signals are also generated to show great promise for advanced modulation formats. In particular, it has also achieved the highest figure-of-merit(FOM) of 182 for high-speed optical modulation , including the bit rate, the extinction ratio normalized with respective to Vpp, the modulation efficiency. The outstanding performance of the present apodized-grating-based slow-light modulator shows great potential and paves the way for developing high-speed optical interconnects for both data-centers and high-performance computing systems.

physics.optics

Integrating Chemistry Knowledge in Large Language Models via Prompt Engineering

This paper presents a study on the integration of domain-specific knowledge in prompt engineering to enhance the performance of large language models (LLMs) in scientific domains. A benchmark dataset is curated to encapsulate the intricate physical-chemical properties of small molecules, their drugability for pharmacology, alongside the functional attributes of enzymes and crystal materials, underscoring the relevance and applicability across biological and chemical domains.The proposed domain-knowledge embedded prompt engineering method outperforms traditional prompt engineering strategies on various metrics, including capability, accuracy, F1 score, and hallucination drop. The effectiveness of the method is demonstrated through case studies on complex materials including the MacMillan catalyst, paclitaxel, and lithium cobalt oxide. The results suggest that domain-knowledge prompts can guide LLMs to generate more accurate and relevant responses, highlighting the potential of LLMs as powerful tools for scientific discovery and innovation when equipped with domain-specific prompts. The study also discusses limitations and future directions for domain-specific prompt engineering development.

cs.CL

Towards calibration-free Mach-Zehnder switches on silicon

Silicon photonic Mach-Zehnder switches (MZSs) have been extensively investigated as a promising candidate for practical optical interconnects. However, conventional 2{\times}2 MZSs are usually prone to the size variations of the arm waveguides due to imperfect fabrication, resulting in considerable random phase imbalance between the two arms, thereby imposing significant challenges for further scaling up NN MZSs. Here we propose a novel design towards calibration-free 2{\times}2 and N{\times}N MZSs, employing optimally widened arm waveguides, enabled by novel compact tapered Euler S-bends with incorporated mode filters. With standard 180-nm CMOS foundry processes, more than thirty 2{\times}2 MZSs and one 4{\times}4 Benes MZS with the new design are fabricated and characterized. Compared with their conventional counterparts with 0.45-μm-wide arm waveguides, the present 2{\times}2 MZSs exhibit ~370-fold reduction in the random phase imbalance. The measured extinction ratios of the present 2{\times}2 and 4{\times}4 MZSs operating in the all-cross state are ~30 dB and ~20 dB across the wavelength range of ~60 nm, respectively, even without any calibrations. This work paves the way towards calibration-free large-scale N{\times}N silicon photonic MZSs.

physics.app-ph