Search button

Query-Sensitive Retrieval-Augmented Generation under Organisational Constraints in the Regulatory Domain

Aluno: Antonia Sophie Gemmerich


Resumo
Retrieval-augmented generation (RAG) is increasingly adopted in organisational envi- ronments where strict data governance requirements limit the use of externally hosted large language models. In such settings, models must often be deployed locally and op- erate under computational and infrastructural constraints. While prior research typically evaluates retrieval-augmented systems under benchmark conditions with carefully formu- lated queries, far less is known about their robustness under realistic query variation and deployment restrictions. This work investigates how structured differences in query for- mulation influence retrieval quality and downstream generation behaviour in constrained settings. Queries are systematically varied across expert, natural, and naive formulations in order to examine how linguistic structure and domain knowledge shape system perfor- mance. The analysis distinguishes between retrieval adequacy and evidence utilisation, thereby enabling a focused assessment of how retrieved information is selected and incor- porated into generated answers. The results show that query formulation constitutes a cen- tral design factor. Ranking quality declines substantially from expert to naive queries in several configurations, revealing pronounced sensitivity to input structure. Post-retrieval refinement exhibits conditional effectiveness rather than stable improvement. Moreover, high levels of retrieval recall do not consistently translate into correct and faithful answers. This utilisation gap indicates that the presence of relevant evidence alone is insufficient to guarantee reliable generation. Overall, the findings demonstrate that retrieval-augmented systems deployed under organisational constraints are structurally sensitive to query for- mulation. The study provides empirical evidence on robustness under realistic conditions and highlights the need for evaluation frameworks that account for deployment constraints and variation in user input rather than idealised benchmark scenarios.


Trabalho final de Mestrado