Aluno: Antonia Sophie Gemmerich
Resumo
Retrieval-augmented generation (RAG) is increasingly adopted in organisational envi-
ronments where strict data governance requirements limit the use of externally hosted
large language models. In such settings, models must often be deployed locally and op-
erate under computational and infrastructural constraints. While prior research typically
evaluates retrieval-augmented systems under benchmark conditions with carefully formu-
lated queries, far less is known about their robustness under realistic query variation and
deployment restrictions. This work investigates how structured differences in query for-
mulation influence retrieval quality and downstream generation behaviour in constrained
settings. Queries are systematically varied across expert, natural, and naive formulations
in order to examine how linguistic structure and domain knowledge shape system perfor-
mance. The analysis distinguishes between retrieval adequacy and evidence utilisation,
thereby enabling a focused assessment of how retrieved information is selected and incor-
porated into generated answers. The results show that query formulation constitutes a cen-
tral design factor. Ranking quality declines substantially from expert to naive queries in
several configurations, revealing pronounced sensitivity to input structure. Post-retrieval
refinement exhibits conditional effectiveness rather than stable improvement. Moreover,
high levels of retrieval recall do not consistently translate into correct and faithful answers.
This utilisation gap indicates that the presence of relevant evidence alone is insufficient to
guarantee reliable generation. Overall, the findings demonstrate that retrieval-augmented
systems deployed under organisational constraints are structurally sensitive to query for-
mulation. The study provides empirical evidence on robustness under realistic conditions
and highlights the need for evaluation frameworks that account for deployment constraints
and variation in user input rather than idealised benchmark scenarios.
Trabalho final de Mestrado