TY - RPRT TI - Benchmarking Llama2, Mistral, Gemma and GPT for Factuality, Toxicity, Bias and Propensity for Hallucinations AU - David Nadeau AU - Mike Kroutikov AU - Karen McNeil AU - Simon Baribeau PY - 2024 UR - https://arxiv.org/abs/2404.09785 ID - 2404.09785 ER -