TY - RPRT TI - GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching AU - Sajal Regmi AU - Chetan Phakami Pun PY - 2024 UR - https://arxiv.org/abs/2411.05276 ID - 2411.05276 ER -