paper-with-me

홈 › Papers

LLaMAntino: LLaMA 2 Models for Effective Text Generation in Italian Language

2023-12-15 · Pierpaolo Basile, Elio Musacchio, Marco Polignano, Lucia Siciliani, Giuseppe Fiameni, Giovanni Semeraro

Large Language Models represent state-of-the-art linguistic models designed to equip computers with the ability to comprehend natural language. With its exceptional capacity to capture complex contextual relationships, the LLaMA (Large Language Model Meta AI) family represents a novel advancement in the field of natural language processing by releasing foundational models designed to improve the natural language understanding abilities of the transformer architecture thanks to their large amount of trainable parameters (7, 13, and 70 billion parameters). In many natural language understanding tasks, these models obtain the same performances as private company models such as OpenAI Chat-GPT with the advantage to make publicly available weights and code for research and commercial uses. In this work, we investigate the possibility of Language Adaptation for LLaMA models, explicitly focusing on addressing the challenge of Italian Language coverage. Adopting an open science approach, we explore various tuning approaches to ensure a high-quality text generated in Italian suitable for common tasks in this underrepresented language in the original models' datasets. We aim to release effective text generation models with strong linguistic properties for many tasks that seem challenging using multilingual or general-purpose LLMs. By leveraging an open science philosophy, this study contributes to Language Adaptation strategies for the Italian language by introducing the novel LLaMAntino family of Italian LLMs.

📄 PDF Abstract BibTeX arXiv:2312.09993

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingLarge Language ModelNatural Language UnderstandingPhilosophyText Generation

Similar Papers 제목 키워드 기반

Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA

2024-05-11 · Marco Polignano, Pierpaolo Basile, Giovanni Semeraro

In the pursuit of advancing natural language processing for the Italian language, we introduce a state-of-the-art Large Language Model (LLM) based on the novel Meta LLaMA-3 model: LLaMAntino-3-ANITA-8B-Inst-DPO-ITA. We f…

Computational EfficiencyLanguage ModellingLarge Language Modelzero-shot-classification+1

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

2026-05-08 · Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli 외 arxiv

This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters. Performance is investigated across a…

Addressing Hallucinations with RAG and NMISS in Italian Healthcare LLM Chatbots

2024-12-05 · Maria Paola Priola

I combine detection and mitigation techniques to addresses hallucinations in Large Language Models (LLMs). Mitigation is achieved in a question-answering Retrieval-Augmented Generation (RAG) framework while detection is …

ArticlesQuestion AnsweringRAGRetrieval-augmented Generation

Harnessing LLMs for Educational Content-Driven Italian Crossword Generation

2024-11-25 · Kamyar Zeinalipour, Achille Fusco, Asya Zanollo, Marco Maggini 외

In this work, we unveil a novel tool for generating Italian crossword puzzles from text, utilizing advanced language models such as GPT-4o, Mistral-7B-Instruct-v0.3, and Llama3-8b-Instruct. Crafted specifically for educa…

The Invalsi Benchmarks: measuring Linguistic and Mathematical understanding of Large Language Models in Italian

2024-03-27 · Giovanni Puccetti, Maria Cassese, Andrea Esuli

While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate m…

Language ModellingMath