Style Alignment on Mark Twain (test)
96.2Style Score (Mean)FT-Model
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FT-ModelModel=GRPO fine-tuned 8B story generator2025.12 | 96.2 | 0.061 | — | |
| GPT-oss-120BModel=Reference model, Evaluation protocol=Zero-shot2025.12 | 66.3 | 0.138 | — | |
| Gemma3-27BModel=Open-weight baseline, Evaluation protocol=Zero-shot2025.12 | 49.6 | 0.133 | — | |
| Llama3-70BModel=Open-weight baseline, Evaluation protocol=Zero-shot2025.12 | 47.7 | 0.119 | — | |
| Qwen2.5-32BModel=Open-weight baseline, Evaluation protocol=Zero-shot2025.12 | 47.3 | 0.096 | — | |
| FT-AgenticEvaluation Protocol=Few-shot calibrated, Number of Parameters=8B2025.12 | — | — | 94.8 | |
| Gemma3-27BEvaluation Protocol=Few-shot calibrated, Number of Parameters=27B2025.12 | — | — | 42.7 | |
| GPT-oss-120BEvaluation Protocol=Few-shot calibrated, Number of Parameters=120B2025.12 | — | — | 73.1 | |
| Llama3.3-70BEvaluation Protocol=Few-shot calibrated, Number of Parameters=70B2025.12 | — | — | 51.4 | |
| Qwen2.5-32BEvaluation Protocol=Few-shot calibrated, Number of Parameters=32B2025.12 | — | — | 45.7 |