Time-series reasoning on Sleep-CoT 1.0 (test)
69.88F1 ScoreOpenTSLM SoftPrompt (Llama3.2-1B)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OpenTSLM SoftPrompt (Llama3.2-1B)Category=OpenTSLM SoftPrompt, Base Model=Llama3.2-1B2025.10 | 69.88 | 81.08 | |
| OpenTSLM SoftPrompt (Llama3.2-3B)Category=OpenTSLM SoftPrompt, Base Model=Llama3.2-3B2025.10 | 54.4 | 72.04 | |
| OpenTSLM Flamingo (Gemma3-270M)Category=OpenTSLM Flamingo, Base Model=Gemma3-270M2025.10 | 51.38 | 68.49 | |
| OpenTSLM Flamingo (Llama3.2-1B)Category=OpenTSLM Flamingo, Base Model=Llama3.2-1B2025.10 | 49.33 | 67.31 | |
| OpenTSLM Flamingo (Llama3.2-3B)Category=OpenTSLM Flamingo, Base Model=Llama3.2-3B2025.10 | 45.45 | 69.14 | |
| OpenTSLM Flamingo (Gemma3-1B-pt)Category=OpenTSLM Flamingo, Base Model=Gemma3-1B-pt2025.10 | 43.69 | 60.67 | |
| OpenTSLM SoftPrompt (Gemma3-1B-pt)Category=OpenTSLM SoftPrompt, Base Model=Gemma3-1B-pt2025.10 | 30.99 | 36.56 | |
| Gemma3-4B FTCategory=Image (Plot), Protocol=Finetuned2025.10 | 18.56 | 38.28 | |
| Random BaselineCategory=Random Baseline2025.10 | 17.48 | 20 | |
| GPT-4oCategory=Tokenized Time-Series, Protocol=Zero-shot2025.10 | 15.47 | 16.02 | |
| Llama3.2-1BCategory=Tokenized Finetuned, Protocol=LoRA Finetuned2025.10 | 9.05 | 24.19 | |
| OpenTSLM SoftPrompt (Gemma3-270M)Category=OpenTSLM SoftPrompt, Base Model=Gemma3-270M2025.10 | 7.96 | 5.91 | |
| Gemma3-4B-ptCategory=Image (Plot), Protocol=Zero-shot2025.10 | 6.75 | 14.95 | |
| Llama3.2-3BCategory=Tokenized Finetuned, Protocol=LoRA Finetuned2025.10 | 5.86 | 14.3 | |
| Llama3.2-3BCategory=Tokenized Time-Series, Protocol=Zero-shot2025.10 | 5.66 | 12.15 | |
| GPT-4oCategory=Image (Plot), Protocol=Zero-shot2025.10 | 4.82 | 10.75 | |
| Llama3.2-1BCategory=Tokenized Time-Series, Protocol=Zero-shot2025.10 | 2.14 | 0.65 |