Persona classification on Reddit Truncated History Phase 2 (test)
73.7Macro F1 ScoreSNPC
Evaluation Results
| Method | Links | |
|---|---|---|
| SNPCTruncation n=30, Features=30, Evaluation cohort size (N)=1,0232026.06 | 73.7 | |
| SNPCTruncation n=20, Features=30, Evaluation cohort size (N)=9232026.06 | 69.5 | |
| SNPCTruncation n=10, Features=30, Evaluation cohort size (N)=6962026.06 | 57.4 | |
| M3: SS-only (Bayesian accumulator)Truncation n=30, Features=Ind, Evaluation cohort size (N)=5312026.06 | 56.6 | |
| M3: SS-only (Bayesian accumulator)Truncation n=20, Features=Ind, Evaluation cohort size (N)=7052026.06 | 52.4 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=30, Features=Ind, Evaluation cohort size (N)=5312026.06 | 52 | |
| SNPCTruncation n=3, Features=30, Evaluation cohort size (N)=3352026.06 | 50.6 | |
| SNPCTruncation n=5, Features=30, Evaluation cohort size (N)=4622026.06 | 49.8 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=30, Features=40fSS, Evaluation cohort size (N)=5312026.06 | 49.5 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=20, Features=Ind, Evaluation cohort size (N)=7052026.06 | 48.7 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=30, Features=40f, Evaluation cohort size (N)=5312026.06 | 47.6 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=20, Features=40fSS, Evaluation cohort size (N)=7052026.06 | 46.9 | |
| M3: SS-only (Bayesian accumulator)Truncation n=10, Features=Ind, Evaluation cohort size (N)=9442026.06 | 43.6 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=10, Features=Ind, Evaluation cohort size (N)=9442026.06 | 41.8 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=20, Features=40f, Evaluation cohort size (N)=7052026.06 | 40 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=10, Features=40fSS, Evaluation cohort size (N)=9442026.06 | 39.5 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=5, Features=Ind, Evaluation cohort size (N)=1,0982026.06 | 36.9 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=5, Features=raw text, Evaluation cohort size (N)=1,0982026.06 | 36 | |
| M3: SS-only (Bayesian accumulator)Truncation n=5, Features=Ind, Evaluation cohort size (N)=1,0982026.06 | 35.9 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=30, Features=raw text, Evaluation cohort size (N)=5312026.06 | 35.8 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=20, Features=raw text, Evaluation cohort size (N)=7052026.06 | 35 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=3, Features=raw text, Evaluation cohort size (N)=1,1642026.06 | 34.6 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=10, Features=40f, Evaluation cohort size (N)=9442026.06 | 34.4 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=5, Features=40fSS, Evaluation cohort size (N)=1,0982026.06 | 34.4 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=10, Features=raw text, Evaluation cohort size (N)=9442026.06 | 33.9 | |
| M1: uniform (Bayesian accumulator)Truncation n=30, Features=Ind, Evaluation cohort size (N)=5312026.06 | 32.4 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=5, Features=40f, Evaluation cohort size (N)=1,0982026.06 | 31.5 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=3, Features=Ind, Evaluation cohort size (N)=1,1642026.06 | 31.4 | |
| M1: uniform (Bayesian accumulator)Truncation n=20, Features=Ind, Evaluation cohort size (N)=7052026.06 | 30.9 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=3, Features=40fSS, Evaluation cohort size (N)=1,1642026.06 | 30.4 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=30, Features=raw text, Evaluation cohort size (N)=5312026.06 | 29.9 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=3, Features=40f, Evaluation cohort size (N)=1,1642026.06 | 29.7 | |
| SNPCTruncation n=1, Features=30, Evaluation cohort size (N)=1252026.06 | 29.6 | |
| M3: SS-only (Bayesian accumulator)Truncation n=3, Features=Ind, Evaluation cohort size (N)=1,1642026.06 | 29.2 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=5, Features=raw text, Evaluation cohort size (N)=1,0982026.06 | 29.1 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=20, Features=raw text, Evaluation cohort size (N)=7052026.06 | 29 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=3, Features=raw text, Evaluation cohort size (N)=1,1642026.06 | 28 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=10, Features=raw text, Evaluation cohort size (N)=9442026.06 | 27.4 | |
| Best combo (ACD+RF/ABD+SVM)Truncation n=1, Features=40f, Evaluation cohort size (N)=1,1742026.06 | 27.3 | |
| Few-shot LLM (gpt-5.4-nano, SS-only)Truncation n=1, Features=raw text, Evaluation cohort size (N)=1,1742026.06 | 27.1 | |
| Few-shot LLM (gpt-5.4-nano, All posts)Truncation n=1, Features=raw text, Evaluation cohort size (N)=1,1742026.06 | 26.3 | |
| M1: uniform (Bayesian accumulator)Truncation n=10, Features=Ind, Evaluation cohort size (N)=9442026.06 | 24.7 | |
| M2: SS-conditional (Bayesian accumulator)Truncation n=1, Features=Ind, Evaluation cohort size (N)=1,1742026.06 | 23.4 | |
| M1: uniform (Bayesian accumulator)Truncation n=5, Features=Ind, Evaluation cohort size (N)=1,0982026.06 | 23.1 | |
| M1: uniform (Bayesian accumulator)Truncation n=3, Features=Ind, Evaluation cohort size (N)=1,1642026.06 | 21.5 | |
| M1: uniform (Bayesian accumulator)Truncation n=1, Features=Ind, Evaluation cohort size (N)=1,1742026.06 | 21.4 | |
| Best combo (SS-cond.) (ABCD+SVM)Truncation n=1, Features=40fSS, Evaluation cohort size (N)=1,1742026.06 | 17.5 | |
| M3: SS-only (Bayesian accumulator)Truncation n=1, Features=Ind, Evaluation cohort size (N)=1,1742026.06 | 16.5 | |
| Majority (predict P0)Truncation n=1, Features=—2026.06 | 15.8 | |
| Majority (predict P0)Truncation n=3, Features=—2026.06 | 15.8 | |
| Majority (predict P0)Truncation n=5, Features=—2026.06 | 15.8 | |
| Majority (predict P0)Truncation n=10, Features=—2026.06 | 15.5 | |
| Majority (predict P0)Truncation n=20, Features=—2026.06 | 15 | |
| Majority (predict P0)Truncation n=30, Features=—2026.06 | 14.8 |