Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact
延伸阅读
相关资讯
SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent
今天Codifying the Judge: Scalable Evaluation via Program Distillation
今天MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
今天DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
今天