Yao Fu, Junxian He, Maosong Sun, Yikai Zhang, Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, and 1 more
We lifted 1 functions out of this paper's own repositories and ran 1 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| hkust-nlp/ceval | canonical | 1 of 1 |
| Function | Status | Where it lives |
|---|---|---|
| sample_top_p | Ran | hkust-nlp/ceval/code/evaluator_series/evaluators/llama.py code served (permissive licence) · get_code("8845976729f4c4ee") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-EVAL , the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. C-EVAL comprises multiple-choice questions across four difficulty levels: middle school, high school, college, and professional. The questions span 52 diverse disciplines, ranging from humanities to science and engineering. C-EVAL is accompanied by C-EVAL HARD, a subset of very challenging subjects in C-EVAL that requires advanced reasoning abilities to solve. We conduct a comprehensive evaluation of the most advanced LLMs on C-EVAL, including both English-and Chinese-oriented models. Results indicate that only GPT-4 could achieve an average accuracy of over 60%, suggesting that there is still significant room for improvement for current LLMs. We anticipate C-EVAL will help analyze important strengths and shortcomings of foundation models, and foster their development and growth for Chinese users. 1 * Equal Contribution. Full list of individual contributions is detailed in Appendix A.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2305.08322")
get_code_for_paper("2305.08322")
have("2305.08322")
Connect an agent — have() is free.