๐ Text2Model Leaderboard
About
This leaderboard tracks AI models' performance in generating optimization models/solutions (e.g. MiniZinc) for the verified problems in the following dataset.
- Dataset: skadio/text2zinc
Current Rankings
gpt-4o-mini | cot_with_code_and_grammar_validation | 79.09 | 42.73 | 60.91 | 47 | 110 | 62.5 | 12.5 | 57.47 | 34.48 |
Per-Dataset Breakdown
Execution and solution accuracy broken down by dataset source. This table shows results from our own evaluation runs.
gpt-4o-mini | agents+code_val | 65 | 38.46 | 13.85 | 7 | 85.71 | 42.86 | 5 | 100 | 100 | 11 | 45.45 | 18.18 | 22 | 40.91 | 18.18 | 44.54 | 20.91 |
gpt-4 | agents | 65 | 38.46 | 13.85 | 7 | 85.71 | 42.86 | 5 | 80 | 80 | 11 | 45.45 | 18.18 | 22 | 40.91 | 18.18 | 44.54 | 20 |
gpt-4 | agents+code_val | 65 | 41.54 | 13.85 | 7 | 28.57 | 28.57 | 5 | 80 | 80 | 11 | 54.55 | 18.18 | 22 | 40.91 | 27.27 | 43.64 | 20.91 |
gpt-4 | baseline | 65 | 32.31 | 12.31 | 7 | 42.86 | 42.86 | 5 | 20 | 20 | 11 | 45.45 | 27.27 | 22 | 27.27 | 18.18 | 32.73 | 17.27 |
gpt-4 | cot | 65 | 50.77 | 20 | 7 | 85.71 | 57.14 | 5 | 100 | 100 | 11 | 63.64 | 36.36 | 22 | 54.55 | 22.73 | 57.27 | 28.18 |
gpt-4 | cot+code+gram | 65 | 69.23 | 20 | 7 | 42.86 | 14.29 | 5 | 100 | 100 | 11 | 81.82 | 36.36 | 22 | 68.18 | 18.18 | 70 | 24.55 |
gpt-4 | cot+code_val | 65 | 50.77 | 20 | 7 | 85.71 | 57.14 | 5 | 100 | 100 | 11 | 63.64 | 36.36 | 22 | 54.55 | 22.73 | 57.27 | 28.18 |
gpt-4 | cot+gram_val | 65 | 60 | 15.38 | 7 | 42.86 | 14.29 | 5 | 80 | 80 | 11 | 81.82 | 36.36 | 22 | 63.64 | 27.27 | 62.73 | 22.72 |
gpt-4 | knowledge_graph | 65 | 44.62 | 20 | 7 | 71.43 | 42.86 | 5 | 80 | 80 | 11 | 54.55 | 36.36 | 22 | 40.91 | 18.18 | 48.19 | 25.45 |
gpt-4o | cot | 65 | 60 | 29.23 | 7 | 71.43 | 57.14 | 5 | 100 | 100 | 11 | 36.36 | 18.18 | 22 | 59.09 | 36.36 | 60 | 34.54 |
gpt-4o | cot+code+gram | 65 | 72.31 | 36.92 | 7 | 42.86 | 14.29 | 5 | 100 | 100 | 11 | 90.91 | 45.45 | 22 | 72.73 | 40.91 | 73.64 | 40 |
gpt-4o | cot+code_val | 65 | 76.92 | 35.38 | 7 | 85.71 | 42.86 | 5 | 100 | 80 | 11 | 81.82 | 36.36 | 22 | 86.36 | 54.55 | 80.91 | 41.82 |
gpt-4o-mini | gala | 65 | 32.31 | 13.85 | 7 | 57.14 | 42.86 | 5 | 40 | 40 | 11 | 36.36 | 18.18 | 22 | 27.27 | 13.64 | 33.64 | 17.28 |
gpt-5.2 | agents | 65 | 50.77 | 32.31 | 7 | 85.71 | 57.14 | 5 | 60 | 60 | 11 | 63.64 | 18.18 | 22 | 54.55 | 31.82 | 55.46 | 33.64 |
gpt-5.2 | agents+code_val | 65 | 76.92 | 29.23 | 7 | 85.71 | 57.14 | 5 | 100 | 100 | 11 | 81.82 | 36.36 | 22 | 68.18 | 36.36 | 77.27 | 36.36 |
gpt-5.2 | baseline | 65 | 63.08 | 32.31 | 7 | 100 | 57.14 | 5 | 100 | 100 | 11 | 63.64 | 36.36 | 22 | 50 | 31.82 | 64.55 | 37.27 |
gpt-5.2 | cot | 65 | 70.77 | 35.38 | 7 | 100 | 42.86 | 5 | 100 | 100 | 11 | 54.55 | 27.27 | 22 | 59.09 | 36.36 | 70 | 38.18 |
gpt-5.2 | cot+code+gram | 65 | 69.23 | 32.31 | 7 | 85.71 | 57.14 | 5 | 100 | 80 | 11 | 90.91 | 45.45 | 22 | 72.73 | 45.45 | 74.55 | 40 |
gpt-5.2 | cot+code_val | 65 | 64.62 | 30.77 | 7 | 71.43 | 42.86 | 5 | 100 | 100 | 11 | 72.73 | 36.36 | 22 | 63.64 | 36.36 | 67.28 | 36.36 |
gpt-5.2 | cot+gram_val | 65 | 76.92 | 38.46 | 7 | 85.71 | 57.14 | 5 | 100 | 100 | 11 | 90.91 | 45.45 | 22 | 72.73 | 36.36 | 79.09 | 42.73 |
gpt-5.2 | gala | 65 | 66.15 | 32.31 | 7 | 85.71 | 57.14 | 5 | 100 | 100 | 11 | 63.64 | 45.45 | 22 | 59.09 | 36.36 | 67.27 | 39.09 |
gpt-5.2 | knowledge_graph | 65 | 52.31 | 32.31 | 7 | 71.43 | 57.14 | 5 | 100 | 100 | 11 | 36.36 | 27.27 | 22 | 68.18 | 31.82 | 57.27 | 36.36 |
gpt-oss:20b | baseline | 55 | 10.91 | 5.45 | 7 | 0 | 0 | 5 | 40 | 40 | 9 | 11.11 | 0 | 22 | 18.18 | 13.64 | 13.27 | 8.16 |
gpt-oss:20b | cot | 56 | 14.29 | 5.36 | 7 | 14.29 | 14.29 | 5 | 100 | 100 | 9 | 0 | 0 | 21 | 19.05 | 9.52 | 18.37 | 11.23 |
gpt-oss:20b | gala | 65 | 15.38 | 4.62 | 7 | 0 | 0 | 5 | 40 | 40 | 11 | 27.27 | 18.18 | 22 | 18.18 | 9.09 | 17.27 | 8.18 |
o3-mini | gala | 54 | 57.41 | 27.78 | 6 | 83.33 | 50 | 5 | 80 | 80 | 8 | 37.5 | 37.5 | 22 | 54.55 | 31.82 | 57.9 | 33.69 |
Upload Evaluation Results Only
Best for: When you only have evaluation metrics and are not interested in submitting the code generated using your strategy
Required: summary.json file with these key metrics:
execution_accuracy: % of problems that executed without errorssolution_accuracy: % of problems with correct solutions
Make sure your model name/strategy name combination is unique or check overwrite if you are updating your results
Refresh leaderboard once the upload is successfull
Check this to replace existing model/strategy combination
Upload Complete Zip File
Best for: Full submissions with both code and evaluation results
Zip file should contain:
submissions/[model]/[strategy]/- For all problems inskadio/text2zinc.mzn solution filesevaluations/[model]/[strategy]/- summary.json and detailed_results.json- summary.json: Contains aggregated metrics
- detailed_results.json: Individual problem results with execution_success/solution_success flags and error outputs
Directory Structure of Submissions and Evaluations
Submissions: Contains generated MiniZinc code files generated via your strategy
submissions/
โโโ [model_name]/ # e.g., "GPT-4", "Claude-3.5", "Gemini"
โ โโโ [strategy_name]/ # e.g., "zero-shot", "few-shot", "cot", "custom_strategy_name"
โ โ โโโ problem_0.mzn
โ โ โโโ problem_1.mzn
โ โ โโโ ... (For all problems in `skadio/text2zinc`)
Evaluations: Contains execution results and performance metrics
evaluations/
โโโ [model_name]/
โ โโโ [strategy_name]/
โ โ โโโ summary.json # Aggregated metrics (required for leaderboard)
โ โ โโโ detailed_results.json # Individual problem results
โ โโโ README.md # Model description
Check this to replace existing model/strategy combinations