Retrieval-Augmented Generation for AI-Generated Content: A Survey

https://papers.cool/arxiv/2402.19473

Authors: Penghao Zhao ; Hailin Zhang ; Qinhan Yu ; Zhengren Wang ; Yunteng Geng ; Fangcheng Fu ; Ling Yang ; Wentao Zhang ; Bin Cui

Summary: The development of Artificial Intelligence Generated Content (AIGC) has been facilitated by advancements in model algorithms, scalable foundation model architectures, and the availability of ample high-quality datasets. While AIGC has achieved remarkable performance, it still faces challenges, such as the difficulty of maintaining up-to-date and long-tail knowledge, the risk of data leakage, and the high costs associated with training and inference. Retrieval-Augmented Generation (RAG) has recently emerged as a paradigm to address such challenges. In particular, RAG introduces the information retrieval process, which enhances AIGC results by retrieving relevant objects from available data stores, leading to greater accuracy and robustness. In this paper, we comprehensively review existing efforts that integrate RAG technique into AIGC scenarios. We first classify RAG foundations according to how the retriever augments the generator. We distill the fundamental abstractions of the augmentation methodologies for various retrievers and generators. This unified perspective encompasses all RAG scenarios, illuminating advancements and pivotal technologies that help with potential future progress. We also summarize additional enhancements methods for RAG, facilitating effective engineering and implementation of RAG systems. Then from another view, we survey on practical applications of RAG across different modalities and tasks, offering valuable references for researchers and practitioners. Furthermore, we introduce the benchmarks for RAG, discuss the limitations of current RAG systems, and suggest potential directions for future research. Project: https://github.com/hymie122/RAG-Survey

ca1c3836fb48e2847d5f04a2a1f9bbc.png

d4c86a60d9b5c6509a06256b357e765.png

f0ba190dac96aec392a834f6711d296.png

6bc42e23283e1e031c2ae706e20c233.png

7e2fcb40868f39d1da2e92dfff8a432.png

dd5d2933c0ce5e5cfd5cd21aff13857.png

46397f698fca3a3a504a17f0cff8516.png

a563001cc3c93ec4bf27393ce7a0bd2.png

09883bf2a667c842e33fb1158532f62.png

d23b8f5bbb3583772ab04479e1aa20e.png

45367f3e3f0b8757255545572e45b37.png

22d67bcd7117a57d7bbcb5ad0dadf36.png

fe1600dc026e0603bb743c9215f7e5c.png

6fc6798f57418de2fce8e802bbec8f1.png

5e0886957dae5d91c6f74dcd434ed76.png

79a9cbfa34f6af2942a4800b68edab1.png