๐Ÿ† Text2Model Leaderboard

About

This leaderboard tracks AI models' performance in generating optimization models/solutions (e.g. MiniZinc) for the verified problems in the following dataset.

Current Rankings

gpt-4o-mini
cot_with_code_and_grammar_validation
79.09
42.73
60.91
47
110
62.5
12.5
57.47
34.48

Per-Dataset Breakdown

Execution and solution accuracy broken down by dataset source. This table shows results from our own evaluation runs.

gpt-4o-mini
agents+code_val
65
38.46
13.85
7
85.71
42.86
5
100
100
11
45.45
18.18
22
40.91
18.18
44.54
20.91