Wenyu Liu, Fan Zhang, Tao Wei, Zhong Li, Zaiwen Wen, Yuxuan Chen, Hongliang Lu, Yuan Lan
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Optimization modeling underpins decisionmaking in logistics, manufacturing, energy, and finance, yet translating natural-language requirements into correct optimization formulations and solver-executable code remains labor-intensive. Although large language models (LLMs) have been explored for this task, evaluation is still dominated by toy-sized or synthetic benchmarks, masking the difficulty of industrial problems with 10 3 -10 6 (or more) variables and constraints. A key bottleneck is the lack of benchmarks that align natural-language specifications with reference formulations/solver code grounded in real optimization models. To fill in this gap, we introduce MIPLIB-NL, built via a structureaware reverse construction methodology from real mixed-integer linear programs in MIPLIB 2017. Our pipeline (i) recovers compact, reusable model structure from flat solver formulations, (ii) reverse-generates natural-language specifications explicitly tied to this recovered structure under a unified model-data separation format, and (iii) performs iterative semantic validation through expert review and human-LLM interaction with independent reconstruction checks. This yields 223 one-to-one reconstructions that preserve the mathematical content of the original instances while enabling realistic natural-language-to-optimization evaluation. Experiments show substantial performance degradation on MIPLIB-NL for systems that perform strongly on existing benchmarks, exposing failure modes invisible at toy scale.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2602.10450")
get_code_for_paper("2602.10450")
have("2602.10450")
Connect an agent — have() is free.