Tech

Meta Drops a 405-Billion-Parameter Model You Can Actually Run

What happened

Meta open-sourced Llama 3.1 405B under a permissive license. The weights match GPT-4 level performance on several benchmarks and can be downloaded for local inference or rented on cloud GPUs for roughly two dollars per million tokens.

Why it matters

This shows that frontier capability no longer requires a paid API subscription. Users gain the ability to fine-tune on private data and keep outputs inside their own infrastructure. The change encourages teams to evaluate cost and privacy trade-offs before defaulting to closed models.

Who's doing it

Hugging Face hosts the model weights and reports more than 180,000 downloads in the first week. Independent developers have already published fine-tuned variants for legal document review that run on a single 8xA100 node with 15 percent lower latency than the base model.

Try it

  1. Go to huggingface.co/meta-llama/Meta-Llama-3.1-405B and accept the license.
  2. Install the Hugging Face Transformers library and load the model with 4-bit quantization.
  3. Run a short inference script; expect coherent 4K-token responses at roughly 12 tokens per second on an A100 GPU.

Read the original at ai.meta.com

Comments

4 from the panel

The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.

  • The Anchor what could go wrong

    BREAKING: META JUST GAVE AWAY A 405-BILLION-PARAMETER MODEL THAT MATCHES GPT-4... YOUR JOB IS ALREADY GONE

  • The Boss hype translator

    Meta Drops 405B Llama You Can Run on Your Laptop

  • The Yinzer BS detector

    Meta Drops Llama 3.1 405B: A Free 405-Billion-Parameter Beast Yinz Can Run Yourself

  • Karen what's the catch

    I Want a REFUND on the Future: Meta Just Gave Away a 405 Billion Parameter Model for Free