Claude Opus 4.8 review
Claude Opus 4.8 review: Anthropic shipped Claude Opus 4.8 six weeks after Opus 4.7, with public pricing set at $5 per million input tokens and $25 per million output tokens. Brief benchmark testing shows that overall benchmark results and safety scores are higher than the previous Opus 4.7, with particular gains noted in math, coding, and mechanical tasks but slight declines in creative writing and imagination. The release was evaluated across six tests—creative writing, coding, math, logic, narrative reasoning, and long‑context recall—and the model’s larger token consumption was also observed.
Benchmark summary
The evaluation used six tests: creative writing, coding, math, logic, narrative reasoning, and long‑context recall. Claude Opus 4.8 performed better than Opus 4.7 on math, coding, and mechanical tasks, while it was slightly worse on imagination and creative writing. Overall benchmark results and safety scores were reported as higher than the previous version, and the update was characterized as a lateral move at best for default single‑pass performance. The testing noted that a higher‑effort thinking setting and multi‑shot prompting could push Opus 4.8 ahead but not in a single default pass.
A notable technical observation is that Opus 4.8 exhibits a token appetite described as bordering on self‑sabotage. In the coding test, Opus 4.8 produced a typing‑zombie game, Typing Dead, which was described as pretty good and as the best splash screen, zombie designs, and mechanics obtained from any Anthropic model. The six tests included logic, narrative reasoning, and long‑context recall but the primary performance differentials highlighted math, coding, mechanical tasks, and creative writing.
“`htmlClaude Opus 4.8 generated a creative writing example that is set in the Orinoco delta in the year 1000. The narrative centers on a pardo from Maracaibo named José Lanz who is sent back through eleven centuries to murder a song, and the plot incorporates paradoxical time‑travel elements tied to that mission. The piece describes the scene and characters without additional framing and concludes with the explicit closing line: “It worked perfectly. It always had.” This creative example was part of the creative writing test in the benchmark suite and was noted alongside other evaluations; the overall benchmarking reported that Opus 4.8 is slightly worse at imagination and creative writing compared to Opus 4.7. The extracted narrative content focuses on setting, character, temporal displacement, and the closing sentence provided by the model.
“`In the coding test, Claude Opus 4.8 produced a typing‑zombie game called Typing Dead as part of the benchmark exercise. Evaluators described the game’s overall quality as pretty good and noted that it implemented the typing‑zombie concept. The generated project was identified explicitly by name in the test results.
The evaluation specifically singled out the game’s splash screen, zombie designs, and mechanics for praise, listing those elements as standout components of the output. Those components were characterized as the best splash screen, the best zombie designs, and the best mechanics obtained on this coding test from any Anthropic model to date. The assessment therefore treated Typing Dead as the strongest coding‑test creative output produced by an Anthropic model within the reviewed set of examples.
Claude Opus 4.8 represents a lateral move compared to Opus 4.7, with some improvements in safety and certain benchmark categories alongside trade‑offs in creative tasks and efficiency. Evaluators noted that higher‑effort thinking settings or multi‑shot prompting could improve performance, but those gains did not appear in a single default pass.


