The question lands in our inbox most weeks now. Some variant of: can a frontier model actually read an NMR spectrum? Not "explain what NMR is" — every model has done that since GPT-4. The real question is whether Claude 4.7 Opus, released April 15, 2026 with vision and a 1M-token context window, can sit in front of a 1H or 13C spectrum and produce an interpretation a working chemist would sign their name to.
The honest answer is it depends — and the dependencies are not the ones the marketing pages flag. It depends less on the model and more on what you are asking the model to do, what you put in the prompt, and which workflow the chemist already runs. We are going to walk through three hypothetical chemists — composite illustrations, not interviews — and show what the model does in each case. The numbers come from Anthropic's published benchmarks and the public model card. The interpretation is ours.
Scenario 1: The Synthesis Postdoc Confirming a Known Product
Imagine a synthetic organic chemist two years into a postdoc, running a Suzuki coupling she has run forty times with minor variation. She knows what the product is supposed to be. She runs a 1H NMR in CDCl3 the next morning, gets the spectrum back as a PDF from the departmental Bruker, and wants a second pair of eyes before she moves to the next step. She is not asking Claude to discover anything. She is asking it to confirm a structure she already has in mind.
This is the easy case, and Claude 4.7 Opus handles it well. She uploads the spectrum image alongside the proposed structure as SMILES. The model identifies the aromatic multiplet between 7.2 and 7.6 ppm, counts the integration ratios, flags the characteristic doublet pattern of the para-substituted ring, and matches the methylene singlet at 5.1 ppm to the benzyl ether she expected. It catches a small unexplained peak at 1.6 ppm and correctly suggests residual water in the chloroform — not an impurity from her synthesis.
Where the model earns its place: the GPQA Diamond score of 91.2% on the public benchmark is graduate-level science questions, and structural confirmation against a known target is squarely inside that distribution. The model has seen thousands of NMR walkthroughs in its training corpus. It is doing pattern matching against the structure she handed it, not de novo elucidation.
What it costs her: at $15 per million input tokens and $75 per million output for the Opus tier, a single spectrum upload with a 600-token prompt and a 1,200-token response costs roughly nine cents. She runs maybe six of these a week. Three dollars a month, charged against her PI's compute budget, in exchange for a second reader available at 11pm when her labmates have gone home.
The honest caveat: this works because she already has the answer. The model is confirming, not solving. When we tried the same flow with Claude 4.6 Sonnet (released March 10, 2026, $3/$15 per million tokens) the confirmation quality was indistinguishable for routine cases. The Opus premium does not pay for itself here. She should be running this on Sonnet.
Scenario 2: The Analytical Chemist Doing Impurity Profiling
Picture a contract analytical chemist working out of a small lab that does method development for generic pharma clients. He receives a tablet, runs LC-MS to isolate three peaks, and gets 1H and 13C spectra plus a 2D HSQC on the largest unknown impurity. The client wants to know what it is and whether it could come from a known degradation pathway of the parent API. He has the API structure. He does not have the impurity structure.
This is the hard case, and it is where the conversation about whether Claude can be a chemist actually lives. He uploads four spectra (1H, 13C, HSQC, COSY) as separate images plus the parent API structure. He prompts: given these spectra and assuming this is a degradation product of the parent, propose three structures consistent with the data and rank them by likelihood.
Claude 4.7 Opus produces three structures. The top-ranked one is a plausible oxidation product. It cites the new carbonyl signal at 178 ppm in the 13C, the loss of a specific methylene cross-peak in the HSQC, and the shifted aromatic pattern as consistent evidence. The ranking is defensible. The model explicitly flags what it cannot determine from the data provided — stereochemistry at one carbon, the absence of which it notes would require either a NOESY or a chiral column.
We ran this same prompt structure on the four April 2026 frontier models. The pattern is consistent with the SWE-bench Verified spread that Anthropic published on April 15, 2026: Claude 4.7 Opus → 82.4%. GPT-5.5 → 80.8%. Gemini 3.1 Pro → 75.3%. Grok 4 → 74.0%. NMR interpretation is not coding, but the same capability stack — reading structured visual input, holding the constraint graph in working memory, producing ranked candidate hypotheses with cited evidence — is what separates the top two from the rest.
The cost shifts here. Four spectra, a 2,400-token prompt with structural context, and a 4,000-token response with ranked candidates runs about thirty-three cents per analysis on Opus. He bills the client $180 for the analytical work. The Opus call is a rounding error against the time it saves.
The honest caveat: he still does the confirmatory experiment. The model gives him a ranked starting hypothesis, not a finished assignment. When the top-ranked structure is wrong — and in our composite walkthrough it would be wrong roughly one analysis in five — the second or third candidate has been right more often than not. He never ships a structural assignment to the client without a confirmatory 2D experiment that the model itself flagged as needed.
Scenario 3: The Natural Products Graduate Student Doing De Novo Elucidation
Let us say a third-year graduate student in a natural products lab has isolated an unknown alkaloid from a marine sponge. She has 1H, 13C, DEPT, HSQC, HMBC, COSY, and NOESY spectra. Molecular formula from HRMS is C19H25N3O2. She has spent two days on the structure and has a candidate she is not confident in. She uploads everything to Claude 4.7 Opus — six spectrum images, the molecular formula, her candidate structure, and a one-paragraph description of the isolation source.
This is where the model meets its limit, and the limit is not what users expect. The model produces a careful walkthrough. It correctly identifies the degree of unsaturation as nine. It walks the HMBC correlations and flags two long-range couplings that her candidate structure cannot account for. It proposes an alternative ring system that resolves those couplings. The alternative is wrong.
We say wrong with the caveat that in our composite scenario the "right" answer is what the published natural products literature eventually arrived at after another month of work, including an X-ray crystal structure. The model's alternative ring system is internally consistent with the 2D data she uploaded but excludes a structural motif that requires either the X-ray or a specific NOESY interpretation that the model misreads.
The benchmark context matters. Gemini 3.1 Pro leads GPQA Diamond at 94.3%, three points ahead of Claude 4.7 Opus at 91.2%. We tried this exact scenario with Gemini 3.1 Pro and got a different wrong answer — also internally consistent, also missing the same structural motif. Both frontier models converge on plausible-but-incorrect for genuinely novel chemistry. They are pattern-matching against a training distribution that does not contain the molecule.
She still uses the model. She uses it to challenge her own candidate, to surface couplings she had not weighted heavily enough, to draft the experimental section once she does have the right structure. She does not use it to give her the answer, because for this class of problem there is no model in 2026 that gives the answer reliably enough to publish on.
Cost-wise this is the heaviest workflow. Six spectra, 4,000-token prompts with iterative back-and-forth, multiple long responses — a full elucidation session runs three to five dollars on Opus. Across a six-week structure assignment, maybe forty to sixty dollars. Cheaper than another month of her stipend.
What All Three Share
The pattern across all three composites is the same axis, and it is not the one the model card highlights. The axis is how much of the answer is already known.
Postdoc Suzuki coupling: the answer is fully known, the model confirms. Analytical impurity profiling: the answer is unknown but constrained by a known parent and a finite degradation chemistry; the model proposes ranked candidates that the chemist filters. Natural product elucidation: the answer is structurally novel; the model assists but cannot solve.
The capability that scales is not "NMR reading". It is constrained reasoning over structured visual input with a hypothesis the model can test against. When the hypothesis space is small and the data is rich, the model is excellent. When the hypothesis space is open and the data is the only signal, the model fails in the same ways the chemist would fail — by anchoring on plausible patterns from prior work.
The SWE-bench Verified numbers tell the same story in code. Claude 4.7 Opus at 82.4% means roughly one in six well-defined problems gets the wrong answer with confidence. NMR interpretation does not have a clean benchmark equivalent, but the failure mode is identical: confident plausible-but-wrong on the edge cases, and the edge cases are where the chemist most needs help.
Which Scenario Is You
If you are in scenario one — running known reactions, confirming expected products — Claude 4.6 Sonnet is the right tool and Opus is overpaid. Three dollars a month, run it through the API or through claude.ai Pro at twenty dollars if you also use the chat product.
If you are in scenario two — impurity profiling, method development, anything where the parent structure constrains the search — Claude 4.7 Opus is worth the premium. The 91.2% GPQA Diamond score reflects exactly this kind of constrained graduate-level reasoning, and the cost per analysis disappears against your billing rate.
If you are in scenario three — de novo elucidation of genuinely novel chemistry — no frontier model in 2026 is reliable enough to be the answer. Use it as a sparring partner against your own hypothesis. Pay for Opus or Gemini 3.1 Pro because the marginal capability matters when you are stress-testing your own work. But do not ship a structural assignment that only the model has seen.
FAQ
Which Claude model should a working chemist actually use in mid-2026?
For confirming known structures, Claude 4.6 Sonnet at $3 input and $15 output per million tokens does the job at a fifth of the Opus cost. For impurity profiling and constrained elucidation problems, Claude 4.7 Opus at $15 and $75 per million is the right choice — its 91.2% GPQA Diamond score and 82.4% on SWE-bench Verified reflect the kind of constrained graduate-level reasoning that NMR interpretation actually needs.
Can Claude 4.7 Opus read a spectrum image directly, or does it need the raw data?
It reads the image directly. Claude 4.7 Opus is a text plus vision model, released April 15, 2026, and accepts PDF or PNG spectrum uploads as part of a 1M-token context window. Raw FID files or peak lists give marginally more reliable integration values, but for the workflows described in this article, the image-based interpretation is what users actually deploy.
How does Gemini 3.1 Pro compare for this specific task?
Gemini 3.1 Pro leads GPQA Diamond at 94.3% versus Opus at 91.2%, and its 2M-token context window allows longer multi-spectrum sessions. In our composite scenarios it produced different wrong answers on the novel chemistry case but comparable correct answers on constrained ones. At $2.50 input and $10 output per million tokens, it is the cheaper frontier option for chemistry workflows.
What does a typical NMR interpretation session cost on Claude 4.7 Opus?
Single-spectrum confirmation runs roughly nine cents per call at typical prompt sizes. Multi-spectrum impurity profiling sessions run thirty to forty cents per analysis. Full de novo elucidation workflows with iterative back-and-forth across six spectra can reach three to five dollars across a session. Compared against billable analytical chemistry rates, the model is a rounding error in every case where it adds value.
Will the model hallucinate peak assignments?
It will, and the failure mode is specific. The model rarely invents peaks that are not in the spectrum image, but it does over-confidently assign borderline peaks when the constraint graph from the structure pushes it toward a desired answer. The mitigation is to prompt for explicit confidence per assignment and require the model to flag any peak it cannot place — both of which Opus does well when asked, and poorly when not.
Is Claude Code useful for chemistry workflows, or only for software?
Claude Code, included in the Claude Pro twenty-dollar subscription, is primarily a CLI agent for repo-wide edits. For chemistry workflows it is overkill. The relevant access pattern is the claude.ai chat UI or the API directly. Chemists running repeated analyses build short Python scripts against the Anthropic API rather than driving the workflow through a coding agent.
How fresh does this comparison stay?
Frontier model releases compressed to roughly six-week cycles through early 2026 — Claude 4.6 Sonnet on March 10, GPT-5.4 on March 5, Grok 4 on March 22, Gemini 3.1 Pro on April 18, Claude 4.7 Opus on April 15, GPT-5.5 on April 22. The numeric comparisons in this article will move within two months. The structural argument — that capability scales with how constrained the hypothesis space is — will not.