All topics
#Benchmarking
Two entries on this site carry the Benchmarking tag: two papers, dated 2026. Explore the full list of related work below.
Papers
-
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Paper · 2026 · arXiv · cited by 0 A self-audit of an LLM evaluation: eight open model variants, 293 persisted raw outputs, and a cluster bootstrap showing that a small-sample ranking identifies the worst model reliably but not the best. -
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
Paper · 2026 · SeT-LLM @ KDD 2026 · cited by 0 Principle-Bench: 168 scenarios mapped to two UK FCA principles with paraphrase, adversarial and boundary perturbations, evaluating LLM judges on accuracy, paraphrase robustness, adversarial robustness and calibration.