Chaochao Lu, Sirui Chen, Junqi Chen, Jun-Qi Chen
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. The repositories linked to it are listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in this domain, their precise causal estimation capabilities are still limited, and the impact of post-training on these abilities is insufficiently explored. This paper examines the extent to which post-training can enhance LLMs' capacity for causal inference. We introduce CauGym, a comprehensive dataset comprising seven core causal tasks for training and five diverse test sets. Using this dataset, we systematically evaluate five post-training approaches: SFT, DPO, KTO, PPO, and GRPO. Across five in-domain and four existing benchmarks, our experiments demonstrate that appropriate post-training enables smaller LLMs to perform causal inference competitively, often surpassing much larger models. Our 14B parameter model achieves 93.5% accuracy on the CaLM benchmark, compared to 55.4% by OpenAI o3. Furthermore, the post-trained LLMs exhibit strong generalization and robustness under real-world conditions such as distribution shifts and noisy data. Collectively, these findings provide the first systematic evidence that targeted post-training can produce reliable and robust LLM-based causal reasoners. Our data and GRPO-model are available at https://github.com/OpenCausaLab/CauGym.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2602.06337")
get_code_for_paper("2602.06337")
have("2602.06337")
The run record, dated, one paper per request, free:
curl https://syntology.ai/api/ran/2602.06337.json
A badge for a README (the split and the date, never a ratio):
[](https://syntology.ai/paper/2602.06337)
Connect an agent — have() is free.