arXiv · 2312.04613
Testing LLM performance on the Physics GRE: some observations
Abstract
With the recent developments in large language models (LLMs) and their widespread availability through open source models and/or low-cost APIs, several exciting products and applications are emerging, many of which are in the field of STEM educational technology for K-12 and university students. There is a need to evaluate these powerful language models on several benchmarks, in order to understand their risks and limitations. In this short paper, we summarize and analyze the performance of Bard, a popular LLM-based conversational service made available by Google, on the standardized Physics GRE examination.
Explore related subjects
Keep this discovery
Pranav Gupta. 2023-12-07. Testing LLM performance on the Physics GRE: some observations. https://arxiv.org/abs/2312.04613
Cite the original work for its findings. Save a collection to share your selection of sources.