Information Extraction on FUNSD (test)
92.08F1 ScoreLayoutLMv3large
Evaluation Results
| Method | Links | |
|---|---|---|
| LayoutLMv3largeModality=Vision+Text+Layout2022.12 | 92.08 | |
| LayoutLMv3Model Category=Fine-tuned PTM, Evaluation Protocol=Fine-tuned2024.04 | 92.08 | |
| UDOPModality=Vision+Text+Layout2022.12 | 91.62 | |
| LayoutLMv3 BaseModality=Vision + Text + Layout, Pretrain Data=11M2023.09 | 90.29 | |
| LILTModality=Text+Layout2022.12 | 88.41 | |
| LiLT BaseModality=Text + Layout, Pretrain Data=11M2023.09 | 88.41 | |
| GraphDocResnetModality=Vision + Text + Layout, Pretrain Data=320k2023.09 | 87.95 | |
| UniDocModality=Vision+Text+Layout2022.12 | 87.93 | |
| UDocModality=Vision + Text + Layout, Pretrain Data=11M2023.09 | 87.93 | |
| GraphDocModality=Vision + Text + Layout, Pretrain Data=320k2023.09 | 87.77 | |
| StructuralLMlargeModality=Text+Layout2022.12 | 85.14 | |
| FormNetModality=Text+Layout2022.12 | 84.69 | |
| FormNetModality=Text + Layout, Pretrain Data=700k2023.09 | 84.69 | |
| DocFormerlargeModality=Vision+Text+Layout2022.12 | 84.55 | |
| BROSlargeModality=Text+Layout2022.12 | 84.52 | |
| LayoutLMv2largeModality=Vision+Text+Layout2022.12 | 84.2 | |
| SelfDocModality=Vision+Text+Layout2022.12 | 83.36 | |
| SelfDocModality=Vision + Text + Layout, Pretrain Data=320k2023.09 | 83.36 | |
| DocFormer BaseModality=Vision + Text + Layout, Pretrain Data=5M2023.09 | 83.34 | |
| LayoutLMv2 BaseModality=Vision + Text + Layout, Pretrain Data=11M2023.09 | 82.76 | |
| BROSModality=Text + Layout, Pretrain Data=11M2023.09 | 81.21 | |
| DocFormer BaseModality=Text + Layout, Pretrain Data=5M2023.09 | 80.54 | |
| LayoutLLM-7BModel Category=Ours, Evaluation Protocol=Zero-shot, Backbone Initialization=Vicuna-1.5-7B2024.04 | 79.98 | |
| GlobalDoc (V+T)Modality=Vision + Text, Pretrain Data=1.4M2023.09 | 79.4 | |
| LayoutLMv1 BaseModality=Vision + Text + Layout, Pretrain Data=11M2023.09 | 79.27 | |
| LayoutLMv1 BaseModality=Text + Layout, Pretrain Data=11M2023.09 | 78.66 | |
| LayoutLLM-7BModel Category=Ours, Evaluation Protocol=Zero-shot, Backbone Initialization=Llama2-7B-chat2024.04 | 78.65 | |
| LayoutLMlargeModality=Text+Layout2022.12 | 77.89 | |
| GlobalDoc (V+T)Modality=Vision + Text, Pretrain Data=320k2023.09 | 77.84 | |
| BERTlargeModality=Text2022.12 | 66.63 | |
| ROBERTabaseModality=Text2023.09 | 66.48 | |
| Qwen2-VL-7B+zero-shot=true, RIDGE fine-tuned=true2025.04 | 66.48 | |
| GlobalDoc (T)Modality=Text, Pretrain Data=1.4M2023.09 | 65.52 | |
| GlobalDoc (T)Modality=Text, Pretrain Data=320k2023.09 | 62.27 | |
| BERT BaseModality=Text2023.09 | 60.26 | |
| Qwen2-VL-7Bzero-shot=true, RIDGE fine-tuned=false2025.04 | 59.89 | |
| Vicuna-1.5-7BModel Category=LLM, Input Representation=Layout Text (Text + Box), Evaluation Protocol=Zero-shot2024.04 | 59.63 | |
| Llama2-7B-chatModel Category=LLM, Input Representation=Layout Text (Text + Box), Evaluation Protocol=Zero-shot2024.04 | 58.34 | |
| DocLLM-7BModel Size=7B, Modality=Text+Layout, Evaluation Setting=SDDS2023.12 | 51.8 | |
| Llama2-7BModel Category=LLM, Input Representation=Layout Text (Text + Box), Evaluation Protocol=Zero-shot2024.04 | 51.4 | |
| DocOwl-1.5-Chatzero-shot=true, RIDGE fine-tuned=false2025.04 | 50.88 | |
| Vicuna-7BModel Category=LLM, Input Representation=Plain Text, Evaluation Protocol=Zero-shot2024.04 | 49.79 | |
| DocLLM-1BModel Size=1B, Modality=Text+Layout, Evaluation Setting=SDDS2023.12 | 48.2 | |
| Llama2-7B-chatModel Category=LLM, Input Representation=Plain Text, Evaluation Protocol=Zero-shot2024.04 | 48.2 | |
| Vicuna-1.5-7BModel Category=LLM, Input Representation=Plain Text, Evaluation Protocol=Zero-shot2024.04 | 48.06 | |
| Qwen-VL-7BModel Category=MLLM, Evaluation Protocol=Fine-tuned, Implementation=Original Paper2024.04 | 47.09 | |
| Vicuna-7BModel Category=LLM, Input Representation=Layout Text (Text + Box), Evaluation Protocol=Zero-shot2024.04 | 42.73 | |
| Llama2-7BModel Category=LLM, Input Representation=Plain Text, Evaluation Protocol=Zero-shot2024.04 | 40.78 | |
| GPT-4+OCRModel Size=~1T, Modality=Text, Evaluation Setting=Zero-Shot (ZS)2023.12 | 37 | |
| Monkeyzero-shot=true, RIDGE fine-tuned=false2025.04 | 34.27 | |
| LLaVA-NeXT-Mistral-7B+zero-shot=true, RIDGE fine-tuned=true2025.04 | 33.41 | |
| LLaVA-NeXT-Mistral-7Bzero-shot=true, RIDGE fine-tuned=false2025.04 | 31.18 | |
| Llama2+OCRModel Size=7B, Modality=Text, Evaluation Setting=Zero-Shot (ZS)2023.12 | 17.8 | |
| LLaVA-1.5-7BModel Category=MLLM, Evaluation Protocol=Zero-shot2024.04 | 1.93 | |
| LLaVAR-7BModel Category=MLLM, Evaluation Protocol=Zero-shot, Implementation=Original Paper2024.04 | 1.71 |