paper-with-me

Finetune-RAG

홈페이지 · 논문 1편

This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning. Each line in the finetunerag_dataset.jsonl file is a JSON object: ``json { "content": "<correct content chunk retrieved>", "filename": "<original document filename>", "fictitious_filename1": "<filename of fake doc 1>", "fictitious_content1": "<misleading content chunk 1>", "fictitious_filename2": "<filename of fake doc 2>", "fictitious_content2": "<misleading content chunk 2>", "question": "<user query>", "answer": "<GPT-4o answer based only on correct content>", "content_before": "<optional preceding content>", "content_after": "<optional succeeding content>" } `` Note that the documents contain answers generated by GPT-4o. Additionally, the prompts used to generate the selected answers do not involve any ficticious data, ensuring that the answers are not contaminated when used for fine-tuning.