arXiv:2602.00887 — under review at ICML 2026
Benchmark leaderboard
10 models × 13 benchmarks × 5 frameworks — the evaluation results from the effGen paper
10
Models tested
13
Benchmarks
5
Frameworks
8×A40
GPUs (46GB)
Sort:
Model rankings
🥇
Qwen2.5-32B-Instruct
Qwen32B70.97
effGen avg
+6.0
vs Best
🥈
Gemma3-27B
Google27B69.19
effGen avg
+5.9
vs Best
🥉
GPT-OSS-20B
OpenAI20B67.82
effGen avg
+10.1
vs Best
#4
Qwen2.5-14B-Instruct
Qwen14B66.38
effGen avg
+7.3
vs Best
#5
Qwen2.5-7B-Instruct
Qwen7B63.07
effGen avg
+11.3
vs Best
#6
Gemma3-12B
Google12B60.92
effGen avg
+10.9
vs Best
#7
Qwen2.5-3B-Instruct
Qwen3B56.80
effGen avg
+13.2
vs Best
#8
Gemma3-4B
Google4B53.76
effGen avg
+12.9
vs Best
#9
Qwen2.5-1.5B-Instruct
Qwen1.5B47.44
effGen avg
+13.1
vs Best
#10
Gemma3-1B
Google1B33.38
effGen avg
+10.9
vs Best
Evaluation Setup
From arXiv:2602.00887, under review at ICML 2026
Hardware
8× NVIDIA A40 (46GB VRAM)
Software
Python 3.11|vLLM, local weights
Every model ran on this machine, through each framework in turn. The scores are the paper’s, not a re-run: they were measured at the version the paper reports, which is not necessarily the 1.0.0 on PyPI today.
13 Benchmarks
GSM8KGSM-PLUSMATH-500BB-EasyBB-MedBB-HardGAIASimpleQALoCoMoLongMemEvalARC-CARC-ECSQA
Frameworks Compared
effGenLangChainAutoGensmolagentsRaw model
Key finding
effGen scores above LangChain, AutoGen and smolagents on all 10 models and all 13 benchmarks, with the largest gains on the smaller models.
+13.15% avg gain (3B)
Up to 7.7x faster
70.97% top score (32B)