Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response Planning
Minh Nguyen, Kishan KC, Toan Nguyen, Ankit Chadha, Thuy Vu
Web-scale QAECML PKDD 2023 (Research Track), LNAI 14170, pages 593–611 · Amazon Alexa AI
The paper fine-tunes a question-answering model twice, once with retrieved web text and once without, so it can fall back on what it already knows; end-to-end accuracy rises 7.40% over a ranking baseline.
Abstract
Large Language Models (LLMs) have shown impressive emergent language capabilities, especially in applications with high ambiguity, such as language reasoning and knowledge consolidation. However, previous work explores the use of LLMs for acquiring information using either parametric or external knowledge, which might lead to serious issues such as hallucination. Toward solving these issues, we present a novel approach of knowledge-aware response planning (KARP) and propose a novel framework that employs (i) a knowledge retriever to obtain relevant information from web documents or databases for a given user query, and (ii) a robust fine-tuning strategy for LLMs to exploit the retrieved external knowledge for planning a final response. Experimental results show that our proposed framework can provide natural, concise answers for open-domain questions with high accuracy.
Why it matters
Retrieval-augmented answer generators often learn to copy whatever passage the retriever hands them, so a bad retrieval means a bad answer. This work trains the same model on two input formats, question plus passage and question alone, which pushes it to keep using its own stored knowledge when the passage is thin. The training recipe is simple to copy, and the retriever piece adds an optimal-transport alignment step that improves top-1 passage selection on WikiQA.
Key results
- On WikiQA without ASNQ transfer, the paper's optimal-transport knowledge retriever raises Precision-at-1 from 63.24 to 74.16 and MAP from 75.00 to 83.29 over the TANDA baseline.
- On WikiQA with an ASNQ-adapted RoBERTa-Base encoder, Precision-at-1 goes from 78.67 to 83.77 and MAP from 86.74 to 89.28 over TANDA.
- Null or negativeOn WDRASS the retriever wins on Precision-at-1, 54.60 to 55.9, but loses on MAP, 63.50 down to 61.8, which the authors attribute to WDRASS questions having several correct answers while their model scores candidates one at a time.
- On 2,000 questions sampled from the MS MARCO QA NLG test set, BLEU goes from 14.6 for GenQA to 39.4 for the best KARP configuration, with RougeL from 0.518 to 0.632 and BERTScore from 0.698 to 0.762 for KARP config 1.
- In the end-to-end web-scale setting over roughly 100 million documents and 130 million passages, judged on 2,000 manually labeled WDRASS questions, relative accuracy over the TANDA baseline is +2.20% for GenQA, +6.20% for KARP trained on MS MARCO under data parity, and +7.40% for KARP trained on the larger OKQA collection.
- Null or negativeThe paper's own Table 3 undercuts its choice of best system: config 2 is called best-performing on BLEU 39.4 versus 38.3, yet config 1 beats it on RougeL, 0.632 versus 0.608, and on BERTScore, 0.762 versus 0.752, so no single KARP configuration wins on all three metrics.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about knowledge-aware response planning, generative question answering, retrieval-augmented generation, dense passage retrieval, answer sentence selection, optimal transport alignment, llm fine-tuning strategy, hallucination in question answering, ms marco qa nlg, wikiqa, wdrass, and open-domain question answering.
What this paper does not show
The paper never names the generative model it fine-tunes, and the three configurations in Table 3 are described only as different hyper-parameter settings with no values, so the generation results are not reproducible from the paper as written. There is no ablation separating the retriever from the two-step fine-tuning. The end-to-end table reports relative gains over TANDA with no absolute accuracy, no confidence intervals, and no annotator agreement. The method loses to TANDA on WDRASS MAP. The paper also contradicts itself about which configuration is best.
How to cite
Minh Nguyen, Kishan KC, Toan Nguyen, Ankit Chadha, and Thuy Vu. 2023. Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response Planning. In Machine Learning and Knowledge Discovery in Databases: Research Track (ECML PKDD 2023), Lecture Notes in Artificial Intelligence volume 14170, pages 593–611.
@inproceedings{nguyen2023efficient,
author = {Nguyen, Minh and K C, Kishan and Nguyen, Toan and Chadha, Ankit and Vu, Thuy},
title = {Efficient Fine-tuning Large Language Models for Knowledge-Aware Response Planning},
booktitle = {Machine Learning and Knowledge Discovery in Databases: Research Track. European Conference, ECML PKDD 2023, Turin, Italy, September 18--22, 2023, Proceedings},
series = {Lecture Notes in Computer Science (Lecture Notes in Artificial Intelligence)},
volume = {14170},
pages = {593--611},
publisher = {Springer},
address = {Cham, Switzerland},
year = {2023}
}
Related work of mine
- WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence SelectionCIKM 2022
- Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence SelectionINTERSPEECH 2023
Author copy, self-archived. Published version: Springer, LNAI 14170, pages 593-611. Cited 16 times as counted by Google Scholar on 2026-09-15.