Preference Learning on Toy dataset 0% label noise (test)
99.6AccuracySimPO
Evaluation Results
| Method | Links | |
|---|---|---|
| SimPOnL=100002025.10 | 99.6 | |
| SSPOthreshold=0.5, nL=100002025.10 | 99.6 | |
| DPOnL=100002025.10 | 99.5 | |
| SSPOthreshold=0.3, nL=100002025.10 | 99.3 | |
| SSPOthreshold=0.1, nL=100002025.10 | 99.2 | |
| SSPOthreshold=0.7, nL=100002025.10 | 99.1 | |
| SSPOthreshold=0.9, nL=100002025.10 | 99.1 | |
| DPOnL=10002025.10 | 98.6 | |
| SSPOthreshold=0.1, nL=10002025.10 | 98.6 | |
| SSPOthreshold=0.9, nL=10002025.10 | 98.6 | |
| SSPOthreshold=0.7, nL=10002025.10 | 98.5 | |
| SSPOthreshold=0.3, nL=10002025.10 | 98.3 | |
| SSPOthreshold=0.5, nL=10002025.10 | 98.3 | |
| SimPOnL=10002025.10 | 98.1 | |
| SSPOthreshold=0.7, nL=1002025.10 | 96.6 | |
| SSPOprior=0.7, n_L=100, n_U=10002025.10 | 96.6 | |
| SSPOthreshold=0.1, nL=1002025.10 | 96.5 | |
| SSPOthreshold=0.9, nL=1002025.10 | 96.5 | |
| SSPOprior=0.1, n_L=100, n_U=10002025.10 | 96.5 | |
| SSPOprior=0.9, n_L=100, n_U=10002025.10 | 96.5 | |
| SSPOprior=0.5, number of paired samples (n_L)=1002025.10 | 96 | |
| SSPOthreshold=0.5, nL=1002025.10 | 96 | |
| SSPOprior=0.5, n_L=100, n_U=10002025.10 | 96 | |
| SSPOthreshold=0.3, nL=1002025.10 | 95.5 | |
| SSPOprior=0.3, n_L=100, n_U=10002025.10 | 95.5 | |
| ORPOnL=100002025.10 | 93.4 | |
| SSPOthreshold=0.1, nL=502025.10 | 91 | |
| SSPOprior=0.1, n_L=50, n_U=10002025.10 | 91 | |
| SSPOthreshold=0.3, nL=502025.10 | 88.9 | |
| SSPOprior=0.3, n_L=50, n_U=10002025.10 | 88.9 | |
| SSPOprior=0.5, number of paired samples (n_L)=502025.10 | 87.9 | |
| SSPOthreshold=0.5, nL=502025.10 | 87.9 | |
| SSPOprior=0.5, n_L=50, n_U=10002025.10 | 87.9 | |
| SSPOthreshold=0.9, nL=502025.10 | 86.1 | |
| SSPOprior=0.9, n_L=50, n_U=10002025.10 | 86.1 | |
| SSPOthreshold=0.7, nL=502025.10 | 85.7 | |
| SSPOprior=0.7, n_L=50, n_U=10002025.10 | 85.7 | |
| DPOnumber of paired samples (n_L)=1002025.10 | 84.6 | |
| DPOnL=1002025.10 | 84.6 | |
| DPOn_L=100, n_U=10002025.10 | 84.6 | |
| SSPOprior=0.5, number of paired samples (n_L)=102025.10 | 84.1 | |
| SSPOthreshold=0.5, nL=102025.10 | 84.1 | |
| SSPOprior=0.5, n_L=10, n_U=10002025.10 | 84.1 | |
| SSPOthreshold=0.9, nL=102025.10 | 84 | |
| SSPOprior=0.9, n_L=10, n_U=10002025.10 | 84 | |
| ORPOnL=10002025.10 | 83.2 | |
| SSPOthreshold=0.1, nL=102025.10 | 82.2 | |
| SSPOprior=0.1, n_L=10, n_U=10002025.10 | 82.2 | |
| SimPOnumber of paired samples (n_L)=1002025.10 | 81.7 | |
| SimPOnL=1002025.10 | 81.7 | |
| SimPOn_L=100, n_U=10002025.10 | 81.7 | |
| SSPOthreshold=0.3, nL=102025.10 | 81 | |
| SSPOprior=0.3, n_L=10, n_U=10002025.10 | 81 | |
| SSPOthreshold=0.7, nL=102025.10 | 80 | |
| SSPOprior=0.7, n_L=10, n_U=10002025.10 | 80 | |
| DPOnumber of paired samples (n_L)=502025.10 | 77.7 | |
| DPOnL=502025.10 | 77.7 | |
| DPOn_L=50, n_U=10002025.10 | 77.7 | |
| SimPOnumber of paired samples (n_L)=502025.10 | 77.6 | |
| SimPOnL=502025.10 | 77.6 | |
| SimPOn_L=50, n_U=10002025.10 | 77.6 | |
| SimPOnumber of paired samples (n_L)=102025.10 | 76.2 | |
| SimPOnL=102025.10 | 76.2 | |
| SimPOn_L=10, n_U=10002025.10 | 76.2 | |
| DPOnumber of paired samples (n_L)=102025.10 | 74.3 | |
| DPOnL=102025.10 | 74.3 | |
| DPOn_L=10, n_U=10002025.10 | 74.3 | |
| ORPOnumber of paired samples (n_L)=1002025.10 | 71 | |
| ORPOnL=1002025.10 | 71 | |
| ORPOn_L=100, n_U=10002025.10 | 71 | |
| ORPOnumber of paired samples (n_L)=502025.10 | 67.9 | |
| ORPOnL=502025.10 | 67.9 | |
| ORPOn_L=50, n_U=10002025.10 | 67.9 | |
| ORPOnumber of paired samples (n_L)=102025.10 | 59 | |
| ORPOnL=102025.10 | 59 | |
| ORPOn_L=10, n_U=10002025.10 | 59 |