All topics
#LLM Evaluation
Three entries on this site carry the LLM Evaluation tag: two papers and one post, dated 2026. Explore the full list of related work below.
Papers
-
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Paper · 2026 · arXiv · cited by 0 A self-audit of an LLM evaluation: eight open model variants, 293 persisted raw outputs, and a cluster bootstrap showing that a small-sample ranking identifies the worst model reliably but not the best. -
An Input-Regime Audit of Conflict Detection for Retrieval-Augmented Generation
Paper · 2026 · VecDB@VLDB 2026 submission · cited by 0 An audit of three conflict detectors for vector-database-backed RAG pipelines (a DeBERTa-MNLI cross-encoder, an LLM-as-judge baseline, and the same judge question-conditioned) across curated probes with bootstrap CIs, a 10-model LLM sweep, a prompt-robustness.
Posts
-
Diagnosing Conflicts in RAG Pipelines
Post · 2026 Why retrieval-augmented generation fails when context disagrees with parametric knowledge, and a checklist for diagnosing which failure mode you are looking at.