Two numbers framed the week for this desk. $5 and $25 — that's GPT-5.5's input and output per million tokens, shipped April 22, 2026, with audio folded into the modality stack. Six weeks earlier, GPT-5.4 sat at $3/$15 with no audio at all. So the price of OpenAI's flagship rose roughly 67% on the way in and 67% on the way out, and the thing you got for the premium was, among other capabilities, a model that hears.

That delta is the whole story of the news cycle that prompted this piece. When a wire headline announces that US agencies are scrambling to stop people re-creating dead pilots' voices on the open Internet, the engineering reality underneath it is mundane and specific: audio generation and audio understanding stopped being a research demo and became an API parameter you can bill against. The question every builder asked this month was not "is this scary" — the press handles that. It was "which model do I actually ground my product on, now that the audio surface is real and priced." The honest answer is *it depends*, and anyone who gives you a single name without asking what you're building is selling something. So we're going to walk three of them through it. None of these people are real. Picture them as composites — the kinds of decisions landing in this desk's inbox, not interviews we conducted.

Scenario 1: The Weekend Voice-Memo Builder

Let us say you're shipping a side project — a voice-note app that transcribes, summarizes, and answers back in audio. Imagine a developer with a day job, an OpenRouter account, and maybe forty hours a month to spend on this. The whole thing lives or dies on cost per interaction, because there's no enterprise contract behind it, just your own card.

Here's where I'd start you, and it's not where the launch-week press would. GPT-5.5 got the headlines on April 22 — multi-surface, Codex and ChatGPT, audio in the box. But look at the receipt before you fall for the noise. GPT-5.5: $5/$25 per million tokens. Gemini 3.1 Pro: $2.5/$10, shipped four days earlier on April 18, also text-plus-vision-plus-audio, and carrying a 2M-token context window against GPT-5.5's 400K. For a voice app that's stuffing whole transcripts back into the prompt, that context headroom is not a luxury — it's the difference between paginating your own conversation history and not bothering.

Now concede the strongest case for OpenAI, because it's a real one: GPT-5.5 posts 80.8 on SWE-bench Verified versus Gemini 3.1 Pro's 75.3, and if your app were a coding agent that gap would settle the argument. But your app isn't a coding agent. It's transcription and summarization, and on the benchmark that actually proxies general reasoning quality for that work — GPQA Diamond — Gemini 3.1 Pro leads the entire grounding table at 94.3, ahead of Claude 4.7 Opus at 91.2 and GPT-5.5 at 89.6. The capability you're paying GPT-5.5's premium for is the one capability your product never exercises.

So for you, weekend builder, the math points at Gemini 3.1 Pro at half the input price and a fifth less on output, with four times the SWE-bench leader's context window thrown in for the reads your app is full of. And if even that's too rich once you're handling real volume, Gemini 3 Flash sits at $0.3/$1.2 — an order of magnitude cheaper, text-plus-vision, fast tier. You'd route the heavy summarization to Pro and the cheap classification passes to Flash. That's the entire architecture. Don't overthink the rest.

Scenario 2: The Agency Engineer Running Volume

Now picture something different. Imagine you're the backend engineer at a small agency that just signed three clients wanting "AI features," and the partners promised delivery dates before anyone asked you what the inference bill looks like. You're not shipping one app. You're shipping a pipeline that processes tens of thousands of documents and audio clips a day, and the unit economics are the product. A model that's 3x better on a benchmark and 5x more expensive doesn't help you — it bankrupts the account.

This is the scenario where the launch-week leaderboard is most actively misleading, and I want you to feel why. Claude 4.7 Opus tops SWE-bench Verified at 82.4 and HumanEval at 94.0. Genuinely frontier coding. And the receipt: Claude 4.7 Opus, $15/$75 per million tokens — held flat against the prior Opus tier, which is worth noting because the category default this season was to raise prices into a capability jump, the way GPT-5.5 did. Anthropic didn't. Respect that restraint. Then never put Opus anywhere near your high-volume path, because $75 output on a pipeline doing millions of tokens a day is a number that ends client relationships.

Your workhorse is Claude 4.6 Sonnet at $3/$15, shipped March 10, prompt caching on by default — and that caching default is the line item that actually matters at your volume, because your prompts share boilerplate across every document and you stop paying full freight on the repeated prefix. Sonnet posts 77.5 on SWE-bench Verified. That's two points above Gemini 3.1 Pro's 75.3 and four below GPT-5.5, at the same nominal price as Grok 4 ($3/$15) and a fraction of Opus. For the routing layer underneath — the classify-and-triage pass before anything expensive runs — you drop to Gemini 3 Flash at $0.3/$1.2 or o-mini at $0.6/$2.4.

Here's the part the listicles skip. Your decision isn't "which model." It's "which three, and what routes between them." The agency that picks one flagship and runs everything through it is the agency that posts a margin-negative quarter and blames AI. The volume game is a routing game. The frontier model is a scalpel you keep in a drawer and bill the client separately when a task genuinely needs it.

Scenario 3: The Founder Who Has to Sign Something

Third composite. Imagine a founder building a product where the audio output goes in front of regulated users — healthcare intake, financial onboarding, anything where a hallucinated number or a mishandled voice clip becomes a liability with your name on the contract. You're not optimizing for cost per token. You're optimizing for the sentence you can say to a customer's compliance officer without lying.

And this is exactly the week that question got sharper, because the news event behind this whole piece — agencies trying to stop open-Internet voice re-creation — is a preview of the regulatory posture coming for every audio product. When the synthesized-voice capability that lets someone re-create a dead pilot is the same capability your onboarding flow uses to read back an account summary, the provider's safety story stops being marketing and becomes due diligence you'll be asked to produce.

Concede the obvious counterpoint first: on raw price you are leaving money on the table. Claude 4.7 Opus at $15/$75 is the most expensive model in the grounding by a wide margin, and Gemini 3.1 Pro would do most of your reasoning at $2.5/$10. If this were a cost decision it would already be over. But it isn't. Anthropic's entire stated focus is Constitutional AI, safety, and agents — and when your job is to hand a compliance officer a defensible answer about why you chose the model you chose, "the vendor whose public positioning is built on safety guarantees, and which leads SWE-bench at 82.4 and HumanEval at 94.0" is a stronger sentence than "the cheap one scored well on GPQA."

One caveat I won't let you skip, because it's the methodology trap. Opus tops SWE-bench Verified, but note what that benchmark measures — pass@1 on a curated coding set. It does not measure audio-handling safety, and the grounding gives Claude 4.7 Opus a text-plus-vision modality with no audio listed at all. So if your regulated product genuinely needs audio output, Opus may not even be in your candidate set, and you're back to GPT-5.5 or Gemini 3.1 Pro with a much harder compliance conversation in front of you. We could not confirm a documented audio-safety tier for any of these models from the grounding — so flag that gap to your compliance officer rather than papering over it.

What All Three Share

Strip away the personas and the same three errors were waiting at the door for every one of them.

First: they all walked in anchored to the launch-week leaderboard, and the leaderboard answers a question none of them were actually asking. SWE-bench Verified ranks coding agents. The weekend builder ships transcription. The agency runs volume. The founder signs liability. A 7-point SWE-bench gap is decisive for exactly one of those jobs and irrelevant to the other two — but the press release leads with the number that makes the launch look biggest, and that number leaks into decisions it has no business touching.

Second: none of them, left alone, would have read the pricing as a delta. GPT-5.5 at $5/$25 reads fine in isolation. Set it next to GPT-5.4's $3/$15 from March 5 and you see a 67% increase priced into the audio jump — a real number you can negotiate budget around. Set Claude 4.7 Opus's flat $15/$75 next to it and you see two labs making opposite pricing bets in the same six-week window. The absolute price is trivia. The delta is the decision input.

Third — and this is the one that costs the most — all three reached for a single model when the actual answer was a routing layer. The flagship for the hard task, the workhorse for the volume, the cheap tier for the triage. Every one of these scenarios resolves into *three* models with traffic split between them, not one. The "which model is best" framing is the wrong question wearing a confident hat.

Which Scenario Is You

So which one are you, actually. Be honest, because the wrong self-diagnosis is what burns the budget.

If your monthly inference bill is small enough that you pay it personally and you're optimizing for cost-per-interaction, you're Scenario 1 — start at Gemini 3.1 Pro, fall back to Flash, ignore the coding leaderboard entirely. If you're delivering on a deadline against client volume and your margin is the product, you're Scenario 2 — build the Sonnet-plus-Flash routing layer before you build anything else, and keep Opus in the drawer. If there's a contract with your name on it and a compliance officer in the room, you're Scenario 3 — and your model choice is a defensibility argument first and a benchmark second, with the audio-modality gap as the thing you verify before you promise anything.

Most of you are telling yourself you're Scenario 3 when your invoice says Scenario 1. That misread — buying frontier safety positioning for a hobby-scale product — is the single most common and most expensive mistake this desk sees. Price your reality, not your ambition.

FAQ

Which April 2026 models actually added audio, and which only claim multimodal?

In the grounding, two models carry text-plus-vision-plus-audio: GPT-5.5 (released April 22) and Gemini 3.1 Pro (April 18). Claude 4.7 Opus and Claude 4.6 Sonnet are listed as text-plus-vision only — no audio modality. So if your product genuinely needs to hear or speak, your candidate set narrows to OpenAI and Google immediately, regardless of how the coding leaderboard ranks the others.

Why did GPT-5.5 get more expensive than GPT-5.4?

GPT-5.4 shipped March 5, 2026 at $3/$15 per million tokens, text-plus-vision. GPT-5.5 followed April 22 at $5/$25 with audio added. That's roughly a 67% rise on both input and output, priced into the modality expansion. The category split is the interesting part: Anthropic held Claude 4.7 Opus flat at $15/$75 across its own capability jump, so the two labs made opposite pricing bets in the same window.

Is Claude 4.7 Opus worth $15/$75 over cheaper flagships?

Only for specific jobs. Opus leads SWE-bench Verified at 82.4 and HumanEval at 94.0 — frontier coding — and Anthropic's safety positioning matters when you need a defensible vendor choice. But at five-to-six times Gemini 3.1 Pro's input price, it's wrong for high-volume pipelines and wrong for audio products, since the grounding lists no audio modality for it. Pay the premium for the task that needs it, not by default.

What's the cheapest viable model for high-volume work?

Gemini 3 Flash at $0.3/$1.2 per million tokens is the lowest-cost tier in the grounding — text-plus-vision, built for the fast, high-volume lane. o-mini sits at $0.6/$2.4 and adds chain-of-thought reasoning if your triage step needs it. Neither replaces a flagship for hard reasoning, but as the classify-and-route layer underneath a Sonnet or Pro workhorse, they're where the unit economics get saved.

Does the SWE-bench leaderboard tell me which model to pick?

Only if you're building a coding agent. SWE-bench Verified measures pass@1 on a curated coding set — Claude 4.7 Opus at 82.4, GPT-5.5 at 80.8, Sonnet at 77.5, Gemini 3.1 Pro at 75.3, Grok 4 at 74.0. For transcription, summarization, or reasoning-heavy non-code work, GPQA Diamond is the better proxy, and there Gemini 3.1 Pro leads at 94.3. Match the benchmark to your actual workload before you let a ranking decide anything.

How does the voice-cloning news connect to which model I should use?

The regulatory scramble over re-created voices on the open Internet is the same audio-generation capability that GPT-5.5 and Gemini 3.1 Pro shipped this month as a billable API surface. For builders, the practical fallout is a tightening compliance posture: if your product outputs synthesized voice, expect to document your provider's safety story. The grounding doesn't specify an audio-safety tier for any model, so treat that as an open question to verify, not assume.

Can I just run everything through one model to keep it simple?

You can, and it's the most common way to torch a budget. Each scenario in this piece resolves into a routing layer — a flagship for the hard task, a workhorse like Claude 4.6 Sonnet ($3/$15, prompt caching on by default) for volume, and a cheap tier like Gemini 3 Flash for triage. One-model simplicity feels clean until the invoice arrives. Split the traffic by task difficulty from day one.

---

Three dates on the calendar will test this reading. GPT-5.5 landed April 22, 2026 at $5/$25 — watch whether the next OpenAI flagship holds that audio premium or whether competitive pressure from Gemini's $2.5/$10 pulls it back down. Gemini 3.1 Pro's 2M-context tier (April 18) is the bet that context size beats raw coding scores for real products — watch adoption among volume builders to see if it holds. And the voice-cloning regulatory response now forming will decide whether audio-capable models carry a documented safety tier by the next release cycle, or whether builders keep shipping into the gap this piece flagged. All three either confirm the routing-not-ranking thesis or break it.