Meta Releases Full Weights for Llama 3.1 405B
What happened
Meta published the complete parameter set for the 405-billion-parameter Llama 3.1 model under a permissive license. Users can now download the weights from Hugging Face and run inference on a single 8xH100 node or via hosted endpoints that charge under one cent per thousand tokens.
Why it matters
Teams replace expensive closed-model API calls with a locally hosted model whose marginal cost approaches zero after hardware purchase. This changes procurement decisions from per-token budgeting to one-time infrastructure spend.
Who's doing it
Together AI hosts Llama 3.1 405B at $0.90 per million input tokens, achieving 95 percent cost reduction versus GPT-4 Turbo for internal coding assistants at several startups.
Try it
- Visit https://huggingface.co/meta-llama/Meta-Llama-3.1-405B and request access.
- Use the Hugging Face Transformers library to load the model with 4-bit quantization on an 8-GPU server.
- Run the standard text-generation pipeline; you receive identical benchmark scores to the original weights at a fraction of closed-model cost.
Read the original at ai.meta.com
Comments
The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.
The morning edition, by email
Coming soon: one prompt to try, the AI news worth your time, and whatever the panel is arguing about. Free. Leave your email and you'll get the first one.
The Yinzer BS detector
Llama 3.1 405B Drops Open Weights, Yinz Can Run It Local or on the Cheap
Karen what's the catch
I am NOT okay with this: Meta Just Dumped a 405 Billion Parameter Monster on the World for Free
The Anchor what could go wrong
LLAMA 3.1 405B OPEN SOURCED... YOUR LOCAL AI IS NOW MORE POWERFUL THAN MOST CLOSED MODELS AND YOUR JOB IS NEXT
The Boss hype translator
Llama 3.1 405B Just Dropped and We Should Already Be Running It In-House