SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

Existing commonsense reasoning benchmarks lack a generative evaluation approach for model outputs and typically consider only unidirectional commonsense plausibility when assessing large language models (LLMs). This limited scope diminishes the long-term efficacy of these benchmarks and makes it challenging to fully capture the performance of generative models. To address these issues, we introduce a self-referential assessment framework for bidirectional generative commonsense reasoning SERA. We reconstruct events from existing benchmark datasets into new scenarios based on the commonsense knowledge of LLMs and evaluate performance in a self-referential manner. SERA evaluates the generative and comprehension abilities of LLMs in commonsense reasoning by analyzing their actual outputs using auxiliary measures. Our findings indicate that SERA enhances the utility of existing commonsense benchmarks and discloses previously veiled evaluation criteria. Our analysis further demonstrates that SERA addresses the shortcomings of current reference-based evaluation methods and remains robust to variations in generated samples. Furthermore, SERA shows reliable alignment with both human preference and the GPT-4o meta-evaluation framework. © 2026

키워드

Commonsense reasoning; Deep learning; Evaluation; Large language models; Natural language processing
제목
SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
저자
Seo, Jaehyung; Moon, Hyeonseok; Jang, Yoonna; Lim, Heuiseok
DOI
10.1016/j.knosys.2026.116152
발행일
2026-06
유형
Article
저널명
Knowledge-Based Systems
권
345