Text Reasoning on In-domain (test)
66.6AccuracyRLVR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RLVRModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 66.6 | 0.514 | 0.285 | |
| RLCRModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 65.1 | 0.751 | 0.109 | |
| C3RL w/o RefModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 64.9 | 0.7 | 0.119 | |
| C3RL (Ours)Model family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 64.4 | 0.748 | 0.076 | |
| SaySelfModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 55 | 0.807 | 0.075 | |
| SCModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=12026.07 | 53.9 | 0.762 | 0.088 | |
| C3RL (Ours)Model family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 49.1 | 0.734 | 0.142 | |
| C3RL (Ours)Backbone=Llama-3.2-3B-Instruct2026.07 | 49.1 | 0.734 | 0.142 | |
| SFT+RefModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 48 | 0.72 | 0.098 | |
| C3RL w/o RefModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 47.3 | 0.697 | 0.136 | |
| C3RL w/o RefBackbone=Llama-3.2-3B-Instruct2026.07 | 47.3 | 0.697 | 0.136 | |
| RLVRModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 45.9 | 0.587 | 0.463 | |
| RLVRBackbone=Llama-3.2-3B-Instruct2026.07 | 45.9 | 0.587 | 0.463 | |
| BaseModel family=Qwen2.5VL-7B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 45.3 | 0.531 | 0.448 | |
| RLCRModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 43.3 | 0.799 | 0.072 | |
| RLCRBackbone=Llama-3.2-3B-Instruct2026.07 | 43.3 | 0.799 | 0.072 | |
| SCModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=12026.07 | 40.8 | 0.787 | 0.091 | |
| SCBackbone=Llama-3.2-3B-Instruct2026.07 | 40.8 | 0.787 | 0.091 | |
| SFT+RefModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 37.6 | 0.572 | 0.168 | |
| SFT+RefBackbone=Llama-3.2-3B-Instruct2026.07 | 37.6 | 0.572 | 0.168 | |
| BaseModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 37.2 | 0.563 | 0.134 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.07 | 37.2 | 0.563 | 0.134 | |
| SaySelfModel family=Llama-3.2-3B-Instruct, Decoding temperature=0.7, top_k=-1, top_p=1, Seed=0, 42, 20252026.07 | 13.1 | 0.676 | 0.06 | |
| SaySelfBackbone=Llama-3.2-3B-Instruct2026.07 | 13.1 | 0.676 | 0.06 |