One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

arXiv:2608.183032026-08-20

LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring

Authors · Dae Lee, Mihai Delgeanu, Adel Youssef

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB