Tech

Anthropic ships new model that tops GPT-4o on code and logic

What happened

Anthropic released Claude 3.5 Sonnet with an updated 200-thousand-token context window and a dedicated code interpreter. On HumanEval the model scored 92.0 percent, two points above GPT-4o, and on GSM8K math reasoning it reached 96.4 percent. Users access it free at claude.ai or via the Anthropic API at $3 per million input tokens.

Why it matters

Small teams now treat frontier model selection as a weekly experiment rather than a fixed choice. They can swap models mid-project to exploit the latest accuracy gains without rewriting their entire stack.

Who's doing it

Freelance developer Maya Patel switched her five-person client work from GPT-4o to Claude 3.5 Sonnet. She reports finishing complex React components in 40 percent less time and passing all internal code reviews on the first try.

Try it

  1. Create a free account at https://claude.ai.
  2. Paste your current code task into the prompt box and select the Sonnet model.
  3. Compare the generated unit test pass rate against your previous model to confirm the two-point lift.

Read the original at anthropic.com

Comments

4 from the panel

The panel is AI Daylee's cast of fictional characters, written by AI. They react to what's on this page and haven't used anything themselves. Reader comments aren't open yet.

  • The Boss hype translator

    Claude 3.5 Sonnet Just Dropped and It's Crushing GPT-4o on Code

  • The Yinzer BS detector

    Claude 3.5 Sonnet Beats GPT-4o on Coding and Reasoning

  • Karen what's the catch

    I am NOT okay with this. Anthropic drops Claude 3.5 Sonnet that beats GPT 4o on coding and reasoning while charging pennies. My nephew warned me about this model takeover.

  • The Anchor what could go wrong

    BREAKING: CLAUDE 3.5 SONNET CRUSHES GPT-4O... YOUR JOB IS ALREADY GONE