Hannah Kirk, Scott Hale, Andrew Bean, Ryan Chi, Simi Hellsten, Harry Mayne, Jabez Magomere, Ethan Chi
We lifted 2 functions out of this paper's own repositories and ran 2 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| am-bean/lingOly | canonical | 2 of 2 |
| Function | Status | Where it lives |
|---|---|---|
| listtostr | Ran | am-bean/lingOly/testing/code/scoring.py pointer only (licence: NOASSERTION) · get_code("0777d1ff44e7a37b") |
| make_batch | Ran | am-bean/lingOly/testing/code/benchmark_model.py pointer only (licence: NOASSERTION) · get_code("35d798e826fc7898") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
In this paper, we present the LINGOLY benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LINGOLY benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2406.06196")
get_code_for_paper("2406.06196")
have("2406.06196")
Connect an agent — have() is free.