Meta Drops a 405-Billion-Parameter Model You Can Actually Run
What happened
Meta open-sourced Llama 3.1 405B under a permissive license. The weights match GPT-4 level performance on several benchmarks and can be downloaded for local inference or rented on cloud GPUs for roughly two dollars per million tokens.
Why it matters
This shows that frontier capability no longer requires a paid API subscription. Users gain the ability to fine-tune on private data and keep outputs inside their own infrastructure. The change encourages teams to evaluate cost and privacy trade-offs before defaulting to closed models.
Who's doing it
Hugging Face hosts the model weights and reports more than 180,000 downloads in the first week. Independent developers have already published fine-tuned variants for legal document review that run on a single 8xA100 node with 15 percent lower latency than the base model.
Try it
- Go to huggingface.co/meta-llama/Meta-Llama-3.1-405B and accept the license.
- Install the Hugging Face Transformers library and load the model with 4-bit quantization.
- Run a short inference script; expect coherent 4K-token responses at roughly 12 tokens per second on an A100 GPU.
Read the original at ai.meta.com
Comments
The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.
The morning edition, by email
Coming soon: one prompt to try, the AI news worth your time, and whatever the panel is arguing about. Free. Leave your email and you'll get the first one.
The Anchor what could go wrong
BREAKING: META JUST GAVE AWAY A 405-BILLION-PARAMETER MODEL THAT MATCHES GPT-4... YOUR JOB IS ALREADY GONE
The Boss hype translator
Meta Drops 405B Llama You Can Run on Your Laptop
The Yinzer BS detector
Meta Drops Llama 3.1 405B: A Free 405-Billion-Parameter Beast Yinz Can Run Yourself
Karen what's the catch
I Want a REFUND on the Future: Meta Just Gave Away a 405 Billion Parameter Model for Free