Viktor Schlegel, Ian Pratt-Hartmann, Kamen Pavlov
We lifted 1 functions out of this paper's own repositories and ran 0 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| schlevik/nlr | canonical | 0 of 1 |
| Function | Status | Where it lives |
|---|---|---|
| generate | Not yet run | schlevik/nlr/scripts/verify_vampire.py pointer only (licence: NONE) · get_code("b525b5145e14727a") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
deep-learning-based approaches to Natural Language Processing (NLP) are credited with various capabilities that involve reasoning with natural language texts. In this paper we carry out a large-scale empirical study investigating the detection of formally valid inferences in controlled fragments of natural language for which the satisfiability problem becomes increasingly complex. We find that, while transformerbased language models perform surprisingly well in these scenarios, a deeper analysis reveals that they appear to overfit to superficial patterns in the data rather than acquiring the logical principles governing the reasoning in these fragments.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2211.05417")
get_code_for_paper("2211.05417")
have("2211.05417")
Connect an agent — have() is free.