Preference Learning on Toy dataset 50% label noise (test)
75.7AccuracySSPO
Evaluation Results
| Method | Links | |
|---|---|---|
| SSPOprior=0.5, number of paired samples (n_L)=102025.10 | 75.7 | |
| SSPOprior=0.5, n_L=10, n_U=10002025.10 | 75.7 | |
| SSPOprior=0.5, number of paired samples (n_L)=502025.10 | 65.6 | |
| SSPOprior=0.5, n_L=50, n_U=10002025.10 | 65.6 | |
| ORPOnumber of paired samples (n_L)=102025.10 | 59.4 | |
| SimPOnumber of paired samples (n_L)=102025.10 | 59.4 | |
| ORPOn_L=10, n_U=10002025.10 | 59.4 | |
| SimPOn_L=10, n_U=10002025.10 | 59.4 | |
| ORPOnumber of paired samples (n_L)=502025.10 | 58.6 | |
| ORPOn_L=50, n_U=10002025.10 | 58.6 | |
| DPOnumber of paired samples (n_L)=102025.10 | 57.1 | |
| DPOn_L=10, n_U=10002025.10 | 57.1 | |
| DPOnumber of paired samples (n_L)=502025.10 | 56.7 | |
| DPOn_L=50, n_U=10002025.10 | 56.7 | |
| SimPOnumber of paired samples (n_L)=502025.10 | 56.3 | |
| SSPOprior=0.5, number of paired samples (n_L)=1002025.10 | 56.3 | |
| SimPOn_L=50, n_U=10002025.10 | 56.3 | |
| SSPOprior=0.5, n_L=100, n_U=10002025.10 | 56.3 | |
| DPOnumber of paired samples (n_L)=1002025.10 | 55.4 | |
| DPOn_L=100, n_U=10002025.10 | 55.4 | |
| ORPOnumber of paired samples (n_L)=1002025.10 | 54.9 | |
| ORPOn_L=100, n_U=10002025.10 | 54.9 | |
| SimPOnumber of paired samples (n_L)=1002025.10 | 53.4 | |
| SimPOn_L=100, n_U=10002025.10 | 53.4 |