ResearchTasksLong-context Temporal ReasoningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedEventQAQwen3-4B RL finetuned on HanabiRewards0.856Accuracy (64K)2Feb 26, 2026