상세 보기
SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
- Seo, Jaehyung;
- Moon, Hyeonseok;
- Jang, Yoonna;
- Lim, Heuiseok
WEB OF SCIENCE
1SCOPUS
1초록
Existing commonsense reasoning benchmarks lack a generative evaluation approach for model outputs and typically consider only unidirectional commonsense plausibility when assessing large language models (LLMs). This limited scope diminishes the long-term efficacy of these benchmarks and makes it challenging to fully capture the performance of generative models. To address these issues, we introduce a self-referential assessment framework for bidirectional generative commonsense reasoning SERA. We reconstruct events from existing benchmark datasets into new scenarios based on the commonsense knowledge of LLMs and evaluate performance in a self-referential manner. SERA evaluates the generative and comprehension abilities of LLMs in commonsense reasoning by analyzing their actual outputs using auxiliary measures. Our findings indicate that SERA enhances the utility of existing commonsense benchmarks and discloses previously veiled evaluation criteria. Our analysis further demonstrates that SERA addresses the shortcomings of current reference-based evaluation methods and remains robust to variations in generated samples. Furthermore, SERA shows reliable alignment with both human preference and the GPT-4o meta-evaluation framework. © 2026
키워드
- 제목
- SERA: Self-referential assessment framework for bidirectional generative commonsense reasoning
- 저자
- Seo, Jaehyung; Moon, Hyeonseok; Jang, Yoonna; Lim, Heuiseok
- 발행일
- 2026-06
- 유형
- Article
- 권
- 345