GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท llama-3-1-405b

Llama 3.1 405B

vllama-3.1-405b

Meta

Modelmetaopen-source
87
Strong
About This Model

Meta's largest open-source model with 405 billion parameters, offering complete transparency, self-hosting capabilities, and competitive performance with proprietary models. Remains one of Meta's legacy open models: Meta has shipped no new open weights since Llama 4 Scout/Maverick (April 2025) and pivoted to closed models with Muse Spark (April 2026). Weights remain broadly available on Hugging Face and via many API hosts as of July 2026.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+
task accuracy code

Industry-standard coding benchmarks

Evidence
HumanEval โ€” 80.5% pass rate
highVerified: 2026-07-09
task accuracy reasoning

Mathematical and scientific reasoning benchmarks

Evidence
MATH โ€” 64.4% accuracy
GPQA โ€” 51.2% on graduate-level science
highVerified: 2026-07-09
task accuracy general

Comprehensive knowledge testing

Evidence
MMLU โ€” 85.2% on graduate-level knowledge
highVerified: 2026-07-09
output consistency

Community evaluation and testing

Evidence
Community Testing โ€” Good consistency reported in community testing
mediumVerified: 2026-07-09
latency p50

Third-party hosting performance

Evidence
Together AI โ€” ~3.5s via hosted API (hardware dependent for self-hosting)
mediumVerified: 2026-07-09
latency p95

Third-party hosting performance

Evidence
Together AI โ€” ~7.0s via hosted API
mediumVerified: 2026-07-09
uptime sla

Deployment model analysis

Evidence
Self-hosted โ€” User-controlled uptime for self-hosted deployments
highVerified: 2026-07-09
context window

Official model specifications

Evidence
Meta Documentation โ€” 128K token context window
highVerified: 2026-07-09
multimodal support

Official model capabilities

Evidence
Meta Documentation โ€” Text-only model
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+
jailbreak resistance

Safety testing and red teaming

Evidence
Meta Safety Report โ€” Safety fine-tuning applied, but open model allows modification
mediumVerified: 2026-07-09
prompt injection defense

Community security testing

Evidence
Community Testing โ€” Standard defenses, user-configurable
mediumVerified: 2026-07-09
data leakage prevention

Architecture review

Evidence
Self-hosted Model โ€” Complete data isolation when self-hosted
highVerified: 2026-07-09
adversarial robustness

Adversarial testing by Meta

Evidence
Meta Safety Testing โ€” Robust to common adversarial attacks in testing
mediumVerified: 2026-07-09
content filtering

Safety tooling review

Evidence
Llama Guard โ€” Llama Guard available for content moderation
mediumVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+
data retention

Deployment model analysis

Evidence
Self-hosted Model โ€” Complete control over data retention when self-hosted
highVerified: 2026-07-09
gdpr compliance

Privacy architecture review

Evidence
Self-hosted Model โ€” Full GDPR compliance control in self-hosted setup
highVerified: 2026-07-09
hipaa eligible

Healthcare compliance assessment

Evidence
Self-hosted Model โ€” HIPAA compliance possible with proper self-hosted infrastructure
highVerified: 2026-07-09
soc2 certified

Deployment architecture review

Evidence
Self-hosted Model โ€” SOC 2 depends on hosting infrastructure
highVerified: 2026-07-09
data sovereignty

Deployment model analysis

Evidence
Self-hosted Model โ€” Complete data sovereignty with self-hosting
highVerified: 2026-07-09
encryption at rest

Deployment architecture review

Evidence
Self-hosted Model โ€” User-controlled encryption at rest
highVerified: 2026-07-09
encryption in transit

Deployment architecture review

Evidence
Self-hosted Model โ€” User-controlled TLS configuration
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+
model documentation

Documentation completeness review

Evidence
Meta Model Card โ€” Comprehensive model card with detailed documentation
highVerified: 2026-07-09
training data transparency

Public documentation review

Evidence
Meta Documentation โ€” Training data composition and size disclosed (15T tokens)
highVerified: 2026-07-09
safety testing transparency

Safety documentation review

Evidence
Meta Safety Report โ€” Detailed safety evaluations published
highVerified: 2026-07-09
bias evaluation

Bias benchmarks review

Evidence
Meta Model Card โ€” Bias testing results disclosed
highVerified: 2026-07-09
decision explainability

Model accessibility assessment

Evidence
Open Weights โ€” Complete model transparency with open weights
highVerified: 2026-07-09
versioning changelog

Version management review

Evidence
Meta Releases โ€” Clear versioning with detailed release notes
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+
deployment flexibility

Deployment options review

Evidence
Open Model โ€” Self-host anywhere, cloud, on-prem, edge, or via APIs
highVerified: 2026-07-09
api reliability

Third-party API monitoring

Evidence
Third-party APIs โ€” ~99.5% uptime via major providers
mediumVerified: 2026-07-09
rate limits

Deployment model analysis

Evidence
Self-hosted Model โ€” No rate limits when self-hosted
highVerified: 2026-07-09
cost efficiency

Cost analysis

Evidence
Together AI Pricing โ€” $3.00 per 1M input tokens via API, infrastructure costs for self-hosting
Price Per Token (multi-provider comparison) โ€” Still hosted by multiple providers; pricing ranges roughly $0.80-$9.50 per 1M tokens across hosts (DeepInfra cheapest)
mediumVerified: 2026-07-09
monitoring observability

Tooling availability assessment

Evidence
Self-hosted Model โ€” User-implemented monitoring for self-hosted
mediumVerified: 2026-07-09
support quality

Support channels review

Evidence
Community Support โ€” Community support via GitHub and forums, no official support
mediumVerified: 2026-07-09
Strengths
  • +Complete transparency with open weights
  • +Best-in-class data sovereignty and privacy control
  • +Maximum deployment flexibility (cloud, on-prem, edge)
  • +No vendor lock-in or rate limits when self-hosted
  • +Excellent documentation and model cards
  • +Competitive performance with proprietary models
  • +Strong community support and ecosystem
Limitations
  • !Requires significant infrastructure for self-hosting (8x H100 GPUs minimum)
  • !No official commercial support
  • !Text-only (no native vision)
  • !Safety guardrails can be modified (security consideration)
  • !Higher latency compared to smaller models
  • !Complex deployment and maintenance
  • !Legacy status: Meta has shipped no new open weights since Llama 4 Scout/Maverick (April 2025) and pivoted to closed models (Muse Spark, April 2026), so future open updates are unlikely
Metadata
license: Llama 3.1 Community License (open for commercial use)
architecture: Transformer with Grouped-Query Attention
parameters: 405 billion
training cutoff: December 2023
languages supported: Multilingual (8 languages optimized)
function calling: true
json mode: true
streaming: true

Use Case Ratings

code generation

Strong coding capabilities with complete control

customer support

Good performance, self-hosting ideal for sensitive data

content creation

Strong creative capabilities with full customization

data analysis

Good analytical capabilities

research assistant

128K context with complete data privacy

healthcare

Self-hosting ideal for HIPAA compliance and sensitive data

legal compliance

Complete confidentiality with self-hosting

education

Good capabilities with full control over content

creative writing

Good creative capabilities with customization options