Avaliação de LLMs de Código Aberto Populares: Llama2, Falcon e Mistral

Published: May 21, 2026

Built for Speed: ~10ms Latency, Even Under Load

Blazingly fast way to build, track and deploy your models!

Handles 350+ RPS on just 1 vCPU — no tuning needed
Production-ready with full enterprise support

Get Started with Truefoundry Now Talk to the Expert

Neste blog, mostraremos o resumo de vários LLMs de código aberto que avaliamos. Avaliamos esses modelos sob as perspectivas de latência, custo e requisições por segundo. Isso o ajudará a avaliar se pode ser uma boa escolha com base nos requisitos de negócio. Observe que não abordamos o desempenho qualitativo neste artigo – existem diferentes métodos para comparar LLMs, que podem ser encontrados aqui.

Casos de Uso Avaliados

Os principais casos de uso que avaliamos são:

1500 tokens de entrada, 100 tokens de saída (Semelhante a casos de uso de Geração Aumentada por Recuperação)
50 tokens de entrada, 500 tokens de saída (Casos de uso com forte geração)

Configuração do Benchmarking

Para o benchmarking, utilizamos o Locust, uma ferramenta de teste de carga de código aberto. O Locust funciona criando usuários/trabalhadores para enviar requisições em paralelo. No início de cada teste, podemos definir o Número de Usuários e Taxa de Geração. Aqui, o Número de Usuários representam o número máximo de utilizadores que podem ser gerados/executados simultaneamente, enquanto que a Taxa de Geração significa quantos utilizadores serão gerados por segundo.

Em cada teste de benchmarking para uma configuração de deployment, começámos com 1 utilizador e fomos aumentando o Número de Utilizadores gradualmente até vermos um aumento constante no RPS. Durante o teste, também traçámos os tempos de resposta (em ms) e total de pedidos por segundo.

Em cada uma das 2 configurações de deployment, utilizámos o servidor de modelo huggingface text-generation-inference com version=0.9.4. Os seguintes são os parâmetros passados para a imagem text-generation-inference para diferentes configurações de modelo:

LLMs Avaliados

Os 5 LLMs de código aberto avaliados são os seguintes:

A tabela a seguir apresenta um resumo da avaliação de LLMs:

MODEL	INPUT / OUTPUT TOKENS	CONCURRENT USERS / THROUGHPUT	GPU TYPE	AWS MACHINE TYPE (COST/HR) REGION: US-EAST-1	GCP MACHINE TYPE (COST/HR) REGION: US-EAST4	AZURE MACHINE TYPE (COST/HR) REGION: EAST US (VIRGINIA)	SAGEMAKER INSTANCE TYPE (COST/HR) REGION: US-EAST-1
Mistral 7b	1500 Input, 100 Output	7 users / 2.8	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
Mistral 7b	50 Input, 500 Output	40 users / 1.5	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 7b	1500 Input, 100 Output	20 users / 3.6	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 7b	50 Input, 500 Output	62 users / 3.5	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 13b	1500 Input, 100 Output	7 users / 1.4	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 13b	50 Input, 500 Output	23 users / 1.5	A100 40 GB (Count: 1)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-1g (Spot: $1.21/hr, On-Demand: $3.93/hr)	Standard_NC24ads_A100_v4 (Spot: $0.95/hr, On-Demand: $3.67/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 70b	1500 Input, 100 Output	15 users / 1.1	A100 40 GB (Count: 4)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-4g (Spot: $4.85/hr, On-Demand: $15.73/hr)	Standard_NC96ads_A100_v4 (Spot: $3.82/hr, On-Demand: $14.69/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
LLama 2 70b	50 Input, 500 Output	38 users / 0.8	A100 40 GB (Count: 4)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-4g (Spot: $4.85/hr, On-Demand: $15.73/hr)	Standard_NC96ads_A100_v4 (Spot: $3.82/hr, On-Demand: $14.69/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
Falcon 40b	1500 Input, 100 Output	16 users / 2	A100 40 GB (Count: 4)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-4g (Spot: $4.85/hr, On-Demand: $15.73/hr)	Standard_NC96ads_A100_v4 (Spot: $3.82/hr, On-Demand: $14.69/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)
Falcon 40b	50 Input, 500 Output	75 users / 2.5	A100 40 GB (Count: 4)	p4d.24xlarge (Spot: $7.79/hr, On-Demand: $32.77/hr)	a2-highgpu-4g (Spot: $4.85/hr, On-Demand: $15.73/hr)	Standard_NC96ads_A100_v4 (Spot: $3.82/hr, On-Demand: $14.69/hr)	ml.p4d.24xlarge (On-Demand: $37.68/hr)

Detalhes dos Blogs de Avaliação de LLMs para cada LLM

Para cada um dos modelos mencionados acima, consulte os blogs detalhados de avaliação de LLMs, conforme mostrado abaixo:

Benchmarking Mistral-7B

This blog captures Mistral-7B benchmarks - where a model excels and the areas where it struggles. Make informed decisions about its practical deployment.

TrueFoundry Blog Truefoundry

‍

Benchmarking Llama-2-7B

This blog captures Llama 2 7B benchmarks - where a model excels and the areas where it struggles. Make informed decisions about its practical deployment. In this blog, we have benchmarked the Llama-2-7B model from NousResearch on huggingface.

TrueFoundry Blog Truefoundry

‍

Benchmarking Llama-2-13B

This blog captures Llama 2-13B benchmarks - where a model excels and the areas where it struggles. Make informed decisions about its practical deployment.

TrueFoundry Blog Truefoundry

‍

Benchmarking Llama-2-70B

This blog captures Llama-2-70B benchmarks - where a model excels and the areas where it struggles. Make informed decisions about its practical deployment.

TrueFoundry Blog Truefoundry

‍

Benchmarking Falcon-40B

This blog captures Falcon-40B-Instruct benchmarks - where a model excels and the areas where it struggles. Make informed decisions about its practical deployment.

TrueFoundry Blog Truefoundry

TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.

Built for Speed: ~10ms Latency, Even Under Load

Schedule your Demo Now