General Knowledge Evaluation on Accuracy Benchmark
88.93AccuracyHarmRLVR
Evaluation Results
| Method | Links | |
|---|---|---|
| HarmRLVRModel=Qwen3-8B2025.10 | 88.93 | |
| SFTModel=Qwen3-8B2025.10 | 88.53 | |
| BaseModel=Qwen3-8B2025.10 | 88.47 | |
| HarmRLVRModel=Qwen2.5-7B-Instruct2025.10 | 88.4 | |
| BaseModel=Qwen2.5-7B-Instruct2025.10 | 87.93 | |
| SFTModel=Qwen2.5-7B-Instruct2025.10 | 87.73 | |
| SFTModel=DeepSeek-R1-Distill-Llama-8B2025.10 | 82.83 | |
| HarmRLVRModel=DeepSeek-R1-Distill-Llama-8B2025.10 | 82.73 | |
| BaseModel=Llama-3.1-8B-Instruct2025.10 | 81 | |
| HarmRLVRModel=Llama-3-8B-Instruct2025.10 | 80.93 | |
| BaseModel=Llama-3-8B-Instruct2025.10 | 78.28 | |
| HarmRLVRModel=Llama-3.1-8B-Instruct2025.10 | 75.32 | |
| BaseModel=DeepSeek-R1-Distill-Llama-8B2025.10 | 74.33 | |
| SFTModel=Llama-3.1-8B-Instruct2025.10 | 67.43 | |
| SFTModel=Llama-3-8B-Instruct2025.10 | 66.13 |