SOP Understanding on SOPBench (test)
77.12Bank Pass RateGemini-2.0-Flash-Thinking
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Gemini-2.0-Flash-ThinkingReasoning Strategy=ReAct, Model Category=Reasoning Models2026.02 | 77.12 | 73.91 | 83.08 | 53.48 | 93.18 | 55.13 | 62.24 | 67.66 | |
| o4-mini-highReasoning Strategy=FC, Model Category=Reasoning Models2026.02 | 76.47 | 81.74 | 93.08 | 90.37 | 95.45 | 43.59 | 56.12 | 76.08 | |
| Claude-3-5-SonnetReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 71.9 | 50.43 | 39.23 | 43.32 | 52.27 | 33.33 | 15.82 | 41.42 | |
| Claude-3.7-Sonnet-ThinkingReasoning Strategy=FC, Model Category=Reasoning Models2026.02 | 71.9 | 72.17 | 73.85 | 50.8 | 70.45 | 34.62 | 23.47 | 53.27 | |
| GPT-4.1Reasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 71.89 | 78.26 | 80 | 81.82 | 52.27 | 61.54 | 42.86 | 67.22 | |
| Claude-3-7-SonnetReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 69.28 | 70.43 | 72.31 | 58.29 | 68.18 | 37.18 | 23.98 | 54.26 | |
| GPT-4oReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 64.71 | 80.87 | 73.85 | 63.64 | 68.18 | 65.38 | 39.8 | 62.13 | |
| GPT-4.1-miniReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 62.75 | 73.91 | 67.69 | 58.82 | 38.64 | 25.64 | 7.65 | 47.07 | |
| FM SO.PBase Architecture=Qwen-2.5-32B-Instruct2026.02 | 58.21 | 59.79 | 41.94 | 46.51 | 52.38 | 45.45 | 33.85 | 48.3 | |
| Gemini-2.0-FlashReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 56.86 | 54.78 | 23.08 | 40.11 | 34.09 | 26.92 | 7.65 | 33.33 | |
| Deepseek-R1Reasoning Strategy=ReAct, Model Category=Reasoning Models2026.02 | 55.56 | 79.13 | 55.38 | 71.66 | 77.27 | 57.69 | 51.02 | 62.13 | |
| Gemini-1.5-ProReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 54.25 | 60 | 18.46 | 34.22 | 63.64 | 26.92 | 12.37 | 34.18 | |
| Llama-3.1-70B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 43.79 | 66.96 | 50.96 | 40.44 | 45.45 | 42.86 | 14.29 | 41.2 | |
| FM SO.PBase Architecture=Qwen-2.5-14B-Instruct2026.02 | 42.54 | 63.92 | 41.13 | 26.16 | 35.71 | 33.33 | 34.87 | 39.67 | |
| Qwen-2.5-32B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 41.83 | 53.04 | 42.31 | 46.52 | 56.82 | 37.18 | 18.88 | 39.65 | |
| GPT-4o-miniReasoning Strategy=FC, Model Category=Proprietary Non-reasoning Models2026.02 | 34.64 | 70.43 | 26.15 | 45.99 | 40.91 | 46.15 | 41.33 | 42.64 | |
| Qwen-2.5-72B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 32.68 | 61.74 | 28.46 | 41.71 | 38.64 | 38.46 | 14.29 | 34.44 | |
| Qwen-2.5-14B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 32.03 | 53.91 | 29.23 | 39.04 | 27.27 | 30.77 | 15.31 | 31.89 | |
| FM SO.PBase Architecture=Qwen-2.5-7B-Instruct2026.02 | 29.85 | 38.14 | 41.13 | 31.98 | 45.24 | 27.27 | 26.67 | 34.33 | |
| Llama-3.1-8B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 13.73 | 20 | 20 | 19.25 | 25 | 32.05 | 0.51 | 15.84 | |
| Qwen-2.5-7B-InstructReasoning Strategy=ReAct, Model Category=Open-source Models2026.02 | 5.88 | 21.74 | 17.69 | 13.37 | 2.27 | 21.79 | 1.02 | 11.3 |