paper-with-me

Papers

Claude 3.5 Sonnet Model Card Addendum

2024-06-24 · Preprint 2024 6 · Anthropic

This addendum to our Claude 3 Model Card describes Claude 3.5 Sonnet, a new model which outperforms our previous most capable model, Claude 3 Opus, while operating faster and at a lower cost. Claude 3.5 Sonnet offers improved capabilities, including better coding and visual processing. Since it is an evolution of the Claude 3 model family, we are providing an addendum rather than a new model card. We provide updated key evaluations and results from our safety testing.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationMMR totalmodelMulti-task Language UnderstandingQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

The Claude 3 Model Family: Opus, Sonnet, Haiku

2024-03-04 · Preprint 2024 3 · Anthropic

We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus, our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku, our fastest and least expe…

1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationCommon Sense Reasoning+8

ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography

2025-10-21 · Lara Ahrens, Wilhelm Haverkamp, Nils Strodthoff arxiv

Domain-adapted open-weight large language models (LLMs) offer promising healthcare applications, from queryable knowledge bases to multimodal assistants, with the crucial advantage of local deployment for privacy preserv…

In-Context Learning for Long-Context Sentiment Analysis on Infrastructure Project Opinions

2024-10-15 · Alireza Shamshiri, Kyeong Rok Ryu, June Young Park

Large language models (LLMs) have achieved impressive results across various tasks. However, they still struggle with long-context documents. This study evaluates the performance of three leading LLMs: GPT-4o, Claude 3.5…

In-Context LearningSentiment Analysis

Analysis of LLM Performance on AWS Bedrock: Receipt-item Categorisation Case Study

2026-04-02 · Gabby Sanchez, Sneha Oommen, Cassandra T. Britto, Di Wang 외 arxiv

This paper presents a systematic, cost-aware evaluation of large language models (LLMs) for receipt-item categorisation within a production-oriented classification framework. We compare four instruction-tuned models avai…

The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

2026-05-16 · Aleksandr Churilov arxiv

Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at rates ranging from 5.2% on commercial models to 21.7% on open-source …