Research

A Community Benchmark for LLMs in Optimization Modeling

INFORMS issues a call for contributions to build a shared, rigorous benchmark for evaluating how well large language models translate real-world problems into mathematical optimization models.

Editorialยท1 Sep 2026
A Community Benchmark for LLMs in Optimization Modeling

The operations research and management science community is mobilizing to create a shared, rigorous benchmark for evaluating how well large language models handle optimization modeling tasks. A new call for contributions posted on the INFORMS Open Forum invites researchers, practitioners, and developers to help design a community-driven benchmark that would provide a common yardstick for measuring LLM performance on the translation of real-world problems into mathematical optimization models. The initiative arrives at a moment when LLMs are increasingly deployed as assistants in modeling workflows, yet no standardized evaluation framework exists to separate genuine capability from marketing claims.

The absence of such a framework carries concrete costs. Optimization modeling โ€” the process of converting messy business, engineering, and logistical problems into solvable mathematical formulations โ€” underpins decisions in supply chain design, energy scheduling, transportation planning, and financial portfolio construction. When a modeler cannot reliably compare two LLMs on the same set of modeling tasks, every procurement decision, every research comparison, and every claim of progress rests on shifting sand. The INFORMS call is an explicit attempt to replace that sand with a foundation the community can build on.

Why a community benchmark now

The call arrives at a moment of rapid but uneven experimentation. A 2025 survey of optimization modeling and LLMs, published on arXiv, documents a sprawling landscape of ad hoc evaluations, each using different problem sets, prompting strategies, and success metrics. Some studies test LLMs on textbook linear programming formulations; others probe their ability to extract constraints from natural language descriptions of supply chain or scheduling problems. The result is a literature full of intriguing but incomparable results. A model that scores 80% on one researcher's test set may score 40% on another's, and neither number tells a decision-maker whether the model can handle a real production planning problem with ambiguous requirements and missing data.

Existing efforts such as CO-Bench, presented at AAAI, have begun to address this gap by benchmarking language model agents on combinatorial optimization tasks. But the INFORMS call explicitly seeks something broader: a benchmark built and maintained by the community that uses optimization modeling as its primary lens, rather than treating modeling as a side effect of general code generation or mathematical reasoning. The distinction is important. A model that writes syntactically correct Python for a solver is not necessarily a model that understands decision variables, objective functions, constraints, and the subtle art of choosing the right formulation for a given problem class. A benchmark that conflates code generation with modeling would reward surface fluency while ignoring the deeper cognitive work that makes a formulation usable, verifiable, and computationally tractable.

The organizers are soliciting contributions in several forms: problem instances drawn from real industrial and academic cases, evaluation protocols, baseline implementations, and critical feedback on what the benchmark should measure. The emphasis on community governance is deliberate. By pooling problem sets and evaluation criteria, the field can avoid the trap of a single lab defining success on its own terms. A benchmark owned by one institution, however well-intentioned, inevitably reflects that institution's research priorities, problem preferences, and blind spots. A community benchmark, by contrast, can incorporate the diversity of application domains and modeling styles that characterize operations research as a discipline.

What the benchmark is expected to measure

According to the call, the benchmark will focus on the modeling phase of the optimization pipeline โ€” the step before a solver is invoked. This includes translating natural language problem descriptions into formal mathematical models, selecting appropriate model structures, identifying decision variables and constraints, and sometimes reformulating models for computational efficiency. The organizers are particularly interested in tasks that reflect the messiness of real-world modeling: ambiguous requirements, missing data, conflicting objectives, and the need to choose among multiple valid formulations. A benchmark composed solely of clean, well-specified problems would measure an idealized version of modeling that bears little resemblance to what practitioners actually do.

The call does not prescribe a single evaluation metric. Instead, it invites proposals for metrics that capture correctness, robustness, and usefulness. Correctness might mean whether the generated model is mathematically equivalent to a reference formulation. Robustness could involve testing how models degrade when problem descriptions are perturbed or when irrelevant information is added. Usefulness is harder to quantify but could include whether a human modeler can readily understand, verify, and modify the LLM's output. A model that produces a correct but opaque formulation is less useful in practice than one that produces a slightly suboptimal but transparent formulation that a human can debug and extend.

One open question is whether the benchmark should include interactive tasks, where an LLM engages in a back-and-forth with a human modeler to refine a formulation. The call explicitly asks for input on this point, reflecting a broader debate in the field about whether LLMs are best evaluated as autonomous problem solvers or as collaborative assistants. In many real-world settings, the modeling process is iterative: a modeler proposes a formulation, a stakeholder identifies a missing constraint, the modeler revises, and the cycle repeats. A benchmark that only tests single-shot generation would miss this crucial dimension of LLM usefulness.

Challenges in building a credible benchmark

Creating a benchmark that the community trusts is not straightforward. The first challenge is problem diversity. Optimization modeling spans linear programming, integer programming, nonlinear optimization, stochastic programming, and many other subfields. A benchmark that overweights, say, small linear programs will not tell us much about LLM performance on large-scale mixed-integer problems from logistics or energy systems. The organizers acknowledge this and are asking for contributions across a wide range of problem classes and application domains. The goal is a benchmark that reflects the actual distribution of modeling work in industry and academia, not a convenience sample of problems that are easy to generate or already available in textbooks.

A second challenge is contamination. LLMs are trained on vast corpora that likely include many classic optimization problems and their solutions. If a benchmark reuses well-known textbook problems, models may score well simply because they have memorized the answer rather than because they can model a novel problem. This is not a hypothetical concern; contamination has been documented in other AI benchmarks, where models achieve high scores on test sets that overlap with their training data but fail on genuinely new instances. The call encourages contributors to submit new or significantly modified problem instances, and to document the provenance of any existing problems they adapt. Without such documentation, benchmark results become uninterpretable.

A third challenge is evaluation cost. Running a benchmark against many LLMs, especially closed-source models accessed via API, can be expensive. A single comprehensive evaluation across dozens of problem instances and multiple prompting strategies can easily run into thousands of dollars in API fees. The organizers are seeking input on how to design an evaluation protocol that is rigorous but not so costly that it excludes smaller research groups or practitioners in low-resource settings. This is a practical concern that has shaped similar efforts in other fields, such as the vLLM project's contributor guidelines, which emphasize reproducibility and clear documentation as prerequisites for community trust. A benchmark that only well-funded labs can afford to run is a benchmark that will not be widely adopted.

Reactions and early signals

The call has drawn attention from researchers who see it as a necessary corrective to the current state of evaluation. The LLM4OR initiative, which maintains a repository of LLM applications in operations research, has long argued that the field needs shared tasks and leaderboards to make progress measurable. The INFORMS call aligns with that view, though it stops short of promising a leaderboard, focusing instead on building the underlying benchmark infrastructure first. The distinction matters: a leaderboard without a credible benchmark is a race on an unmarked track, while a benchmark without a leaderboard can still provide a common language for describing model capabilities.

Some practitioners have expressed caution. In discussions on the Open Forum, a recurring theme is that optimization modeling is not a single skill but a bundle of skills โ€” abstraction, domain knowledge, mathematical fluency, and communication. A benchmark that reduces this bundle to a single score risks misleading decision-makers about what LLMs can and cannot do. A model might excel at extracting constraints from a well-specified problem but fail when asked to identify the right decision variables for a novel application. The organizers appear sensitive to this concern, noting that the benchmark should produce fine-grained results across task types rather than a single aggregate number. That approach would allow users to see, for example, that a model is strong on linear programming formulations but weak on stochastic programming, rather than receiving a single number that obscures both strengths and weaknesses.

There is also the question of how the benchmark will be maintained over time. Optimization modeling evolves as new problem classes emerge and as solvers improve. A static benchmark would quickly become obsolete, either because models learn to game its specific problem instances or because the problems themselves no longer reflect current practice. The call envisions a living benchmark, with periodic updates and versioned releases, but the governance details remain to be worked out. Who decides when a new version is released? How are deprecated problem instances handled? What happens when a contributor disputes an evaluation result? These are not trivial questions, and the answers will determine whether the benchmark remains credible over the long term.

What happens next

The immediate next step is the collection of contributions. The organizers have set no public deadline in the call, instead inviting ongoing submissions and discussion through the INFORMS Open Forum. They are particularly interested in problem instances that come with clear documentation of the underlying real-world context, reference formulations, and expected modeling pitfalls. Early contributors will likely shape the benchmark's structure, so the call is as much an invitation to participate in design as it is a request for data. A researcher who submits a well-documented problem instance from, say, a hospital scheduling application is not just providing data; they are helping to define what the benchmark measures and how it measures it.

The success of the initiative will depend on whether the community responds with the breadth and depth the organizers hope for. If it does, the result could be a benchmark that finally allows researchers, vendors, and enterprise users to compare LLMs on optimization modeling with confidence. That would mark a significant step toward understanding where these models genuinely add value in one of the most consequential areas of applied mathematics. If the effort stalls or fragments, the field will continue to rely on anecdotal evidence and incompatible one-off evaluations โ€” a state of affairs that serves no one well. The call is now open, and the next move belongs to the community.

#LLMs #optimization #benchmarking #operations research

Newsletter

Get the AI news that matters

One short brief with the day's most important AI stories โ€” written for professionals.

We send a confirmation link. No spam. Unsubscribe anytime.

WhatsApp