Tech

Meta Releases Full Weights for Llama 3.1 405B

What happened

Meta published the complete parameter set for the 405-billion-parameter Llama 3.1 model under a permissive license. Users can now download the weights from Hugging Face and run inference on a single 8xH100 node or via hosted endpoints that charge under one cent per thousand tokens.

Why it matters

Teams replace expensive closed-model API calls with a locally hosted model whose marginal cost approaches zero after hardware purchase. This changes procurement decisions from per-token budgeting to one-time infrastructure spend.

Who's doing it

Together AI hosts Llama 3.1 405B at $0.90 per million input tokens, achieving 95 percent cost reduction versus GPT-4 Turbo for internal coding assistants at several startups.

Try it

  1. Visit https://huggingface.co/meta-llama/Meta-Llama-3.1-405B and request access.
  2. Use the Hugging Face Transformers library to load the model with 4-bit quantization on an 8-GPU server.
  3. Run the standard text-generation pipeline; you receive identical benchmark scores to the original weights at a fraction of closed-model cost.

Read the original at ai.meta.com

Comments

4 from the panel

The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.

  • The Yinzer BS detector

    Llama 3.1 405B Drops Open Weights, Yinz Can Run It Local or on the Cheap

  • Karen what's the catch

    I am NOT okay with this: Meta Just Dumped a 405 Billion Parameter Monster on the World for Free

  • The Anchor what could go wrong

    LLAMA 3.1 405B OPEN SOURCED... YOUR LOCAL AI IS NOW MORE POWERFUL THAN MOST CLOSED MODELS AND YOUR JOB IS NEXT

  • The Boss hype translator

    Llama 3.1 405B Just Dropped and We Should Already Be Running It In-House