Přístupnostní navigace
E-application
Search Search Close
Master's Thesis
Author of thesis: Bc. Tomáš Ondrušek
Acad. year: 2025/2026
Supervisor: Ing. Michal Hradiš, Ph.D.
Reviewer: Ing. Jan Kohút
This thesis focuses on intelligent information retrieval from large document collections using agentic retrieval architectures. The goal is to design and evaluate an intelligent search system similar to retrieval-augmented systems without generative components, emphasizing effectiveness and interpretability. The solution integrates LangChain and LangGraph with vector databases such as Weaviate and language models from Ollama and OpenAI. A custom multilingual benchmark dataset, with a specific focus on the Czech language, among others, supports the evaluation of retrieval quality by showing how graph structure and system design influence accuracy and efficiency, and provides guidelines for building robust multilingual retrieval systems. The results demonstrate that while hybrid retrieval and cross-encoder reranking significantly improve performance on multilingual datasets, highly complex agentic pipelines and deep search architectures do not automatically guarantee better retrieval quality compared to simpler baselines, often increasing latency without proportionate gains. Additionally, the effectiveness of end-to-end agentic workflows is shown to depend heavily on the underlying language model's capability to make robust and reliable intermediate decisions.
information retrieval, large language models, agentic retrieval architectures, semantic search, vector databases, multilingual evaluation, LangChain, LangGraph
Date of defence
22.06.2026
Result of the defence
Defended (thesis was successfully defended)
Grading
B
Process of defence
Student nejprve prezentoval výsledky, kterých dosáhl v rámci své práce. Komise se poté seznámila s hodnocením vedoucího a posudkem oponenta práce. Student následně odpověděl na otázky oponenta a na další otázky přítomných. Komise se na základě posudku oponenta, hodnocení vedoucího, přednesené prezentace a odpovědí studenta na položené otázky rozhodla práci hodnotit stupněm B.
Topics for thesis defence
Language of thesis
English
Faculty
Fakulta informačních technologií
Department
Department of Computer Graphics and Multimedia
Study programme
Information Technology and Artificial Intelligence (MITAI)
Specialization
Cybersecurity (NSEC)
Composition of Committee
doc. Mgr. Kamil Malinka, Ph.D. (předseda) doc. Ing. Ondřej Ryšavý, Ph.D. (místopředseda) Ing. Zbyněk Křivka, Ph.D. (člen) doc. Ing. Ivan Homoliak, Ph.D. (člen) Ing. Libor Polčák, Ph.D. (člen) Ing. Radek Hranický, Ph.D. (člen)
Supervisor’s reportIng. Michal Hradiš, Ph.D.
Student se o téma zajímal, zapojil se do společného projektu, dobře pochopil řešené téma a provedl zajímavé experimenty.
Téma přímo vychází z potřeb projektu semANT. Student se nakonec na vývoji aplikace vytvářené v rámci projektu podílel menší měrou, než jsme původně zamýšleli i kvůli technickým a organizačním problémům, ale studentova práce minimálně poskytla informace, které dále využijeme.
Práce byla dokončená v termínu a student ji dostatečně konzultoval.
Student si aktivně vyhledal potřebné zdroje, dobře se zorientoval v řešené oblasti a získané znalosti v práci dobře využil.
Student pracoval průběžně, účastnil se koordinačních schůzek vývojářů společné aplikace. Na konzultace docházel, ale mohl trochu častěji.
Grade proposed by supervisor: B
Reviewer’s reportIng. Jan Kohút
The thesis presents numerous well-focused experiments in the area of retrieval from large document collections. The result is a dataset focused on the Czech environment, experimental findings, and the retrieval tool itself, which can be further extended.
Evaluation level: assignment fulfilled
Evaluation level: is within the usual extent
The basic structure of the thesis is appropriate; the individual chapters address relevant topics and follow one another well. In places, the text is overly structured through the use of many sections, bullet points, and highlighted passages. The graphs in Chapter 4 are stylistically too diverse, and some are not referenced in the text.
In terms of language and typography, the thesis is satisfactory.
I appreciate the high-quality overview and description of related works.
The outputs of the thesis are a tool that enables the retrieval of required information from large document collections using large language models, and experiments examining its effectiveness. The experiments are meaningfully designed and provide insights into each part of the tool. Among other things, the tool was tested on a custom dataset prepared from Czech libraries' data. The results are therefore relevant for the potential deployment of the tool within Czech libraries and archives.
The resulting tool and dataset will be further used within the semANT project.
Evaluation level: moderately difficult assignment
Grade proposed by reviewer: B
Responsibility: Mgr. et Mgr. Hana Odstrčilová