Detail publikačního výsledku

Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?

JAROLÍM, A.; FAJČÍK, M.; MAKAIOVÁ, L.

Originální název

Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?

Anglický název

Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?

Druh

Stať ve sborníku v databázi WoS či Scopus

Originální abstrakt

Misinformation frequently spreads in user comments under online news articles, highlighting the need for effective methods to detect factually incorrect information. To strongly support or refute claims extracted from such comments, it is necessary to identify relevant documents and pinpoint the exact text spans that justify or contradict each claim. This paper focuses on the latter task --- fine-grained evidence extraction for Czech and Slovak claims. We create new dataset, containing two-way annotated fine-grained evidence created by paid annotators. We evaluate large language models (LLMs) on this dataset to assess their alignment with human annotations. The results reveal that LLMs often fail to copy evidence verbatim from the source text, leading to invalid outputs. Error-rate analysis shows that the llama3.1:8b model achieves a high proportion of correct outputs despite its relatively small size, while the gpt-oss-120b model underperforms despite having many more parameters. Furthermore, the models qwen3:14b, deepseek-r1:32b, and gpt-oss:20b demonstrate an effective balance between model size and alignment with human annotations.

Anglický abstrakt

Misinformation frequently spreads in user comments under online news articles, highlighting the need for effective methods to detect factually incorrect information. To strongly support or refute claims extracted from such comments, it is necessary to identify relevant documents and pinpoint the exact text spans that justify or contradict each claim. This paper focuses on the latter task --- fine-grained evidence extraction for Czech and Slovak claims. We create new dataset, containing two-way annotated fine-grained evidence created by paid annotators. We evaluate large language models (LLMs) on this dataset to assess their alignment with human annotations. The results reveal that LLMs often fail to copy evidence verbatim from the source text, leading to invalid outputs. Error-rate analysis shows that the llama3.1:8b model achieves a high proportion of correct outputs despite its relatively small size, while the gpt-oss-120b model underperforms despite having many more parameters. Furthermore, the models qwen3:14b, deepseek-r1:32b, and gpt-oss:20b demonstrate an effective balance between model size and alignment with human annotations.

Klíčová slova

Fact-checking; Fine-grained evidence; LLMs

Klíčová slova v angličtině

Fact-checking; Fine-grained evidence; LLMs

Autoři

JAROLÍM, A.; FAJČÍK, M.; MAKAIOVÁ, L.

Rok RIV

2026

Vydáno

05.12.2025

ISBN

978-80-263-1858-3

Kniha

Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025

Periodikum

Recent Advances in Slavonic Natural Language Processing

Číslo

2025

Stát

Česká republika

Strany od

25

Strany do

36

Strany počet

11

URL

BibTex

@inproceedings{BUT201605,
  author="Antonín {Jarolím} and Martin {Fajčík} and Lucia {Makaiová}",
  title="Can LLMs Extract Human-like Fine-grained Evidence for Evidence-based Fact-checking?",
  booktitle="Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025",
  year="2025",
  journal="Recent Advances in Slavonic Natural Language Processing",
  number="2025",
  pages="25--36",
  isbn="978-80-263-1858-3",
  issn="2336-4289",
  url="https://raslan2025.nlp-consulting.net/"
}