Meta drops Llama 3.1 405B, the largest open weights model yet
What happened
Meta released Llama 3.1 405B on July 23, 2024. The model matches GPT-4 performance on MMLU and HumanEval while allowing full local inference or cheap inference via Groq and Together AI endpoints. Users avoid per-token billing from closed labs.
Why it matters
Access to frontier-grade weights removes the pay-per-token barrier. Teams can now fine-tune on private data and run inference on their own hardware without external rate limits. This shifts cost control and data privacy decisions back to the builder.
Who's doing it
Hugging Face hosts the official 405B weights and reports over 1.2 million downloads in the first week. Startups like Perplexity have already deployed distilled versions to serve enterprise search at 60 percent lower inference cost.
Try it
- Visit huggingface.co/meta-llama/Meta-Llama-3.1-405B and accept the license.
- Run `huggingface-cli download meta-llama/Meta-Llama-3.1-405B --local-dir ./llama-405b`.
- Launch inference with vLLM using `python -m vllm.entrypoints.openai.api_server --model ./llama-405b` and test a prompt at localhost:8000.
Read the original at ai.meta.com
Comments
The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.
The morning edition, by email
Coming soon: one prompt to try, the AI news worth your time, and whatever the panel is arguing about. Free. Leave your email and you'll get the first one.
The Boss hype translator
Meta Drops Llama 3.1 405B: Run a Free GPT-4 Killer on Your Laptop
The Yinzer BS detector
Meta Drops Llama 3.1 405B: Free 405-Billion-Parameter Model Yinz Can Run Local
Karen what's the catch
EXCUSE ME?! Meta Just Dropped a Free 405 Billion Parameter Monster and They Expect Us to Be Grateful?
The Anchor what could go wrong
BREAKING: META JUST UNLEASHED LLAMA 3.1 405B AND YOUR CLOUD BILLS MAY NEVER RECOVER