arXiv:2602.00887 — under review at ICML 2026

Benchmark leaderboard

10 models × 13 benchmarks × 5 frameworks — the evaluation results from the effGen paper

10
Models tested
13
Benchmarks
5
Frameworks
8×A40
GPUs (46GB)
Sort:

Model rankings

🥇

Qwen2.5-32B-Instruct

Qwen32B
70.97
effGen avg
+6.0
vs Best
🥈

Gemma3-27B

Google27B
69.19
effGen avg
+5.9
vs Best
🥉

GPT-OSS-20B

OpenAI20B
67.82
effGen avg
+10.1
vs Best
#4

Qwen2.5-14B-Instruct

Qwen14B
66.38
effGen avg
+7.3
vs Best
#5

Qwen2.5-7B-Instruct

Qwen7B
63.07
effGen avg
+11.3
vs Best
#6

Gemma3-12B

Google12B
60.92
effGen avg
+10.9
vs Best
#7

Qwen2.5-3B-Instruct

Qwen3B
56.80
effGen avg
+13.2
vs Best
#8

Gemma3-4B

Google4B
53.76
effGen avg
+12.9
vs Best
#9

Qwen2.5-1.5B-Instruct

Qwen1.5B
47.44
effGen avg
+13.1
vs Best
#10

Gemma3-1B

Google1B
33.38
effGen avg
+10.9
vs Best

Evaluation Setup

From arXiv:2602.00887, under review at ICML 2026

Hardware
8× NVIDIA A40 (46GB VRAM)
Software
Python 3.11|vLLM, local weights

Every model ran on this machine, through each framework in turn. The scores are the paper’s, not a re-run: they were measured at the version the paper reports, which is not necessarily the 1.0.0 on PyPI today.

13 Benchmarks
GSM8KGSM-PLUSMATH-500BB-EasyBB-MedBB-HardGAIASimpleQALoCoMoLongMemEvalARC-CARC-ECSQA
Frameworks Compared
effGenLangChainAutoGensmolagentsRaw model

Key finding

effGen scores above LangChain, AutoGen and smolagents on all 10 models and all 13 benchmarks, with the largest gains on the smaller models.

+13.15% avg gain (3B)
Up to 7.7x faster
70.97% top score (32B)