基于少样本索引的生成式检索 (Generative Retrieval with Few-shot Indexing)

Existing generative retrieval (GR) methods rely on training-based indexing, which fine-tunes a model to memorise associations between queries and the document identifiers (docids) of relevant documents. Training-based indexing suffers from high training costs, under-utilisation of pre-trained knowledge in large language models (LLMs), and limited adaptability to dynamic document corpora. To address the issues, we propose a few-shot indexing-based GR framework (Few-Shot GR). It has a few-shot indexing process without any training, where we prompt an LLM to generate docids for all documents in a corpus, ultimately creating a docid bank for the entire corpus. During retrieval, we feed a query to the same LLM and constrain it to generate a docid within the docid bank created during indexing, and then map the generated docid back to its corresponding document. Moreover, we devise few-shot indexing with one-to-many mapping to further enhance Few-Shot GR. Experiments show that Few-Shot GR achieves superior performance to state-of-the-art GR methods requiring heavy training.

翻译：现有的生成式检索方法依赖于基于训练的索引，该方法通过微调模型来记忆查询与相关文档的文档标识符之间的关联。基于训练的索引存在训练成本高、未能充分利用大型语言模型中预训练知识以及对动态文档集适应性有限等问题。为解决这些问题，我们提出了一种基于少样本索引的生成式检索框架。该框架包含一个无需任何训练的少样本索引过程：我们通过提示大型语言模型为语料库中的所有文档生成文档标识符，最终构建整个语料库的文档标识符库。在检索阶段，我们将查询输入同一大型语言模型，并约束其生成索引阶段所构建文档标识符库中的标识符，随后将该标识符映射回对应的文档。此外，我们设计了具有一对多映射的少样本索引机制以进一步增强本框架的性能。实验表明，本方法在性能上优于需要大量训练的最先进生成式检索方法。