Natural Language Processing on BigBench II
-0.37Accuracy Degradation (%)PromptCOS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PromptCOSModel=Deepseek-d-qwen2025.09 | -0.37 | 62 | |
| PromptCOSModel=Gemma2-it2025.09 | 0 | 65 | |
| PC*Model=Gemma2-it2025.09 | 0.03 | 53 | |
| PromptCOSModel=TinyLlama-chat2025.09 | 0.08 | 68 | |
| PC*Model=TinyLlama-chat2025.09 | 0.53 | 71 | |
| PC*Model=Deepseek-d-qwen2025.09 | 0.83 | 81 | |
| PCGModel=TinyLlama-chat2025.09 | 1.41 | 45 | |
| PCGModel=Deepseek-d-qwen2025.09 | 1.6 | 52 | |
| PCGModel=Gemma2-it2025.09 | 2.33 | 46 |