SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2502.12052 · ACL · 2025

A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability

Xiaojun Wan, Li Lin, Xinyu Hu, Mingqi Gao, Zhenghan Yu

arXiv · PDF · Open in the Atlas

Code that ran

We lifted 6 functions out of this paper's own repositories and ran 6 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.

RepositoryRoleRan
PKU-ONELab/NLG-DualEval — 6 of 6
FunctionStatusWhere it lives
OC_aspect_mapping Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("6d6141715c4d337f")
calculate_mode Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("7f5722bd993450b3")
extract Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("24b1fb0c3e057797")
filter_rating Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("4babb82c84c0dbd1")
get Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("c1fdb9eccea3d1f4")
post_process Ran PKU-ONELab/NLG-DualEval/module/data_process.py
code served (permissive licence) · get_code("5b373b7f1514f260")

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine the effectiveness of meta-evaluation. In this work, we propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities, thereby providing better interpretability. In addition, we introduce a method of automatically constructing the corresponding benchmarks without requiring new human annotations. Furthermore, we conduct experiments with 16 representative LLMs as the evaluators based on our proposed framework, comprehensively analyzing their evaluation performance from different perspectives. LLM SummEval Global Topical-Chat Global Overall Coh Con Flu Rel Avg Und Nat MCtx Int UK Avg GPT-4o 0.

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2502.12052")
get_code_for_paper("2502.12052")
have("2502.12052")

Connect an agent — have() is free.