Llama 4 405B appears on exactly two of the four major benchmark leaderboards we audited. The other two — SWE-bench Verified and GPQA Diamond — have zero open-weight models in the top five.

Next time the Hacker News repost cycle hits "The Winamp Skin Museum whips the Llama's ass (2020)" — and it will, the cadence is roughly every eighteen months — the comment section is going to do what it always does. Someone will note the slogan, someone will pivot to Meta's Llama, and a third commenter will assert that the open-weight family still competes on frontier benchmarks. Before that thread fires off again, we wanted to settle the math. Not the vibes. The actual arithmetic of what Llama 4 405B scores versus what Anthropic, OpenAI, Google DeepMind, and xAI shipped in the same release window. The result, as it turns out, depends entirely on which leaderboard you pick — and Meta does not get to pick.

Methodology

We audited the four benchmark leaderboards that the AI tools desk treats as the canonical capability surface in April 2026: SWE-bench Verified (agentic coding, pass@1), GPQA Diamond (graduate-level science, accuracy), MMLU (broad knowledge, accuracy), and HumanEval (function-level coding, pass@1). All scores cited come from the providers' own published benchmark pages as of April 22, 2026 — Anthropic's release notes, OpenAI's launch posts, Google DeepMind's blog, xAI's release page, and Meta's Llama 4 405B technical report.

We restricted comparisons to released-and-priced models in the window January through April 2026. We did not include models without a public pricing page (which is itself one of the findings). Where a model is absent from a leaderboard, we treat absence as a data point rather than imputing a score. Inference cost analysis is constrained by the grounding: closed-model pricing is on vendor pages and is included; Llama 4 405B's hosted inference price varies by third-party provider and is not in our authoritative source set, so we say so where the gap matters rather than fabricate a number.

Free Download
AI Market Desk Weekly Brief
Model releases, pricing changes, benchmark deltas — delivered weekly, zero fluff.

Finding #1: The HumanEval Delta Is 5.0 Points. That Is Bigger Than It Sounds.

Llama 4 405B scores 89.0 on HumanEval pass@1, as of its January 12, 2026 release. Claude 4.7 Opus scores 94.0. GPT-5.5 scores 93.2.

Five points sounds like a rounding error. It is not. HumanEval is a saturated benchmark — the top tier has been compressing toward the ceiling for two years, and every additional point above 90 corresponds to a much harder slice of the residual error distribution. The arithmetic that matters is the failure rate, not the pass rate. Claude 4.7 Opus fails 6.0% of HumanEval problems. Llama 4 405B fails 11.0%. That is an 83% increase in the error rate. If your evaluation harness runs 1,000 function-level coding prompts per day, you are looking at 50 additional failed completions per day on Llama, every day, baseline.

Now multiply by the cost of a single bad completion in an agentic pipeline — the retry, the human review, the downstream rework. The math gets ugly fast. And HumanEval is the benchmark where Llama 4 405B is closest to the frontier. We are starting with its strongest result.

Finding #2: MMLU Is Closer, Because MMLU Has Stopped Mattering.

On MMLU, Llama 4 405B scores 88.6. Gemini 3.1 Pro tops the leaderboard at 91.8. GPT-5.5 sits at 90.1. Claude 4.7 Opus at 89.5.

The delta between Llama and the floor of the closed-model frontier is 0.9 points. Between Llama and the top is 3.2 points. By the failure-rate translation we used in Finding #1: Gemini 3.1 Pro misses 8.2% of MMLU questions, Llama 4 405B misses 11.4% — a 39% error inflation, materially smaller than the HumanEval gap.

Concede the point. On broad-knowledge benchmarks where the ceiling is already crushed, Llama 4 405B is genuinely within striking distance. Anyone arguing that open weights have caught up on undergraduate-level multi-domain question answering is reading the leaderboard correctly.

Here is where it falls apart. MMLU is the benchmark every serious lab has stopped citing as the primary signal. Anthropic's release notes for Claude 4.7 Opus lead with SWE-bench Verified and GPQA Diamond. OpenAI's GPT-5.5 page leads with the same two. Google DeepMind's Gemini 3.1 Pro post emphasizes GPQA Diamond. The reason is structural: MMLU is saturated, contaminated, and gameable. The frontier has moved to agentic-coding evals and graduate-science questions specifically because those are the surfaces where capability differences are still legible. Llama is winning the comparison on a metric the rest of the field has quietly demoted.

Finding #3: The Two Leaderboards Llama Does Not Appear On.

Llama 4 405B is not in the top five on SWE-bench Verified. Llama 4 405B is not in the top five on GPQA Diamond. These are the two benchmarks the entire frontier has aligned around as the headline capability signal in 2026.

SWE-bench Verified top five (pass@1): Claude 4.7 Opus 82.4, GPT-5.5 80.8, Claude 4.6 Sonnet 77.5, Gemini 3.1 Pro 75.3, Grok 4 74.0. No open-weight model in the cut.

GPQA Diamond top five (accuracy): Gemini 3.1 Pro 94.3, Claude 4.7 Opus 91.2, GPT-5.5 89.6, Grok 4 87.0, o-mini 85.0. No open-weight model in the cut.

This is where the 2020 Winamp slogan stops applying. Meta either chose not to submit Llama 4 405B to the agentic-coding benchmark that defines the coding-tools conversation — Claude Code, Codex, Cursor, Aider, Gemini CLI all build their marketing around SWE-bench Verified — or submitted and the score was not competitive enough to publish. Either reading produces the same downstream effect. If you are picking a model to drive an agentic coding workflow in 2026 based on the leaderboard the tool builders themselves cite, Llama 4 405B is not on the page. The argument that the open-weight family is "competitive with the frontier" is doing a lot of work on the two benchmarks where capability is no longer the discriminator. On the two benchmarks where it is, the argument is structurally absent.

Finding #4: The Cost-Per-Point Math You Cannot Actually Do.

Here is the arithmetic the open-weight camp wants to run. Take the benchmark score, divide by the inference cost, and Llama wins on cost-efficiency because the weights are free. We tried to run this calculation. It does not work, and the reason is itself a finding.

The closed-model side is straightforward. Claude 4.7 Opus: $15.00 input / $75.00 output per million tokens, held identical to Claude 4.6 Sonnet's structural pricing posture but at the Opus tier. GPT-5.5: $5.00 / $25.00 per million tokens — versus the prior GPT-5.4 release on March 5, 2026 at $3.00 / $15.00, that is a 67% output-token price increase across a 1.6-point SWE-bench Verified gain (80.8 vs the 79.x range GPT-5.4 was reported at). Gemini 3.1 Pro: $2.50 / $10.00 per million tokens, the cheapest frontier tier. Grok 4: $3.00 / $15.00.

Now Llama 4 405B. The weights are open. The inference is not free — somebody runs the GPU. The publicly authoritative hosted-inference price for Llama 4 405B in our grounding source set is absent. Different providers quote different rates. Meta itself does not publish a per-token figure because there is no first-party API to price. The cost-per-benchmark-point calculation that would let you argue Llama wins on economics requires a number that the open-weight ecosystem does not produce in a single canonical form. The closed-model side publishes one. The open-weight side publishes a benchmark score and points you at the model card. That asymmetry, on its own, is a finding the comparison usually skips.

The Comparison Table

ModelSWE-bench VerifiedGPQA DiamondMMLUHumanEvalInput / Output $ per M
Claude 4.7 Opus82.491.289.594.0$15.00 / $75.00
GPT-5.580.889.690.193.2$5.00 / $25.00
Gemini 3.1 Pro75.394.391.8$2.50 / $10.00
Grok 474.087.0$3.00 / $15.00
Llama 4 405B88.689.0open weights / hosting varies

A dash indicates the model is not in the top five on that leaderboard as of April 22, 2026, per the grounding sources.

What This Does NOT Prove

The benchmarks above measure published capability. They do not measure the things that genuinely make the open-weight argument load-bearing in production: data residency, self-hosting on regulated infrastructure, the ability to fine-tune on proprietary corpora without sending the corpus to a third-party API, the long-tail of latency control when you own the inference stack, and the durability of having a model you can run after the vendor sunsets the SKU. Those are real, they are not trivial, and they are precisely the arguments we did not audit because the leaderboard methodology cannot audit them.

What the math we did run shows is narrower. Specifically: on the four benchmarks the frontier labs lead with in April 2026, Llama 4 405B is competitive on the two where the discriminator has decayed, absent from the two where it has not, and operates without a canonical per-token cost figure that would let a buyer run the cost-per-capability math the open-weight argument depends on. The 2020 Winamp slogan was about an audio player skin museum. The 2026 retcon is doing more work than it can support.

The Takeaway

Llama still whips on the two benchmarks the rest of the field has quietly stopped centering — and is structurally absent from the two it has not. The slogan has aged.

None of this answers whether Meta will skip the next benchmark submission cycle entirely or pivot the Llama family toward a metric where they own the leaderboard. That question is the actual interesting one, and it is not where this piece ends.

FAQ

Why is Llama 4 405B missing from the SWE-bench Verified leaderboard?

Meta has not published a top-five-competitive SWE-bench Verified score for Llama 4 405B in the public release notes available as of April 22, 2026. We cannot tell from the public record whether they submitted and underperformed or chose not to submit. Either outcome produces the same effect: in a 2026 coding-tool buying decision driven by SWE-bench Verified — which Anthropic, OpenAI, Google DeepMind, and xAI all lead with — Llama is not in the comparison set.

Does the HumanEval delta of 5 points actually matter in practice?

It matters more than the headline number suggests. Pass-rate math is misleading on saturated benchmarks. Claude 4.7 Opus fails 6.0% of HumanEval problems; Llama 4 405B fails 11.0% — an 83% relative increase in error rate. In a production agentic loop running thousands of function-level completions, that error-rate gap compounds into substantially more retries, human reviews, and downstream rework. The pass-rate number flatters Llama; the failure-rate number does not.

What about cost — isn't Llama still cheaper because the weights are open?

Possibly, but the math is harder to run than the argument assumes. The grounding does not contain a single canonical hosted-inference price for Llama 4 405B because there isn't one — different providers quote different rates, and Meta does not operate a first-party API with a published per-token cost. Closed models publish one number. Open weights publish a benchmark score and a model card. The cost-per-capability comparison the open-weight pitch depends on requires a price the ecosystem does not standardize.

Is MMLU still a useful benchmark for choosing a model in 2026?

For broad-knowledge sanity checking, yes. As the discriminator between frontier models, no. The top four MMLU scores cluster within 3.2 points of each other (88.6 to 91.8), the benchmark has known contamination concerns, and the labs themselves have moved their marketing leads to SWE-bench Verified and GPQA Diamond precisely because MMLU no longer separates the frontier. If you are picking a model based on MMLU alone in 2026, you are reading the leaderboard the field has stopped reading.

Where does Llama 4 405B still genuinely win against the closed frontier?

On data residency, self-hosting on regulated or air-gapped infrastructure, fine-tuning on proprietary corpora without third-party data exposure, full inference-stack control, and durability against vendor SKU sunsets. None of those are measured by SWE-bench Verified or GPQA Diamond, which is why the benchmark conversation systematically undersells them. If your buying criteria are any of the above, the leaderboard delta is a side issue. If your criteria are raw capability on the metrics the frontier labs themselves lead with, the leaderboard delta is the whole conversation.