Forget 100 mph. Just Throw Strikes—My Experience with Opus 4.8

Whenever a new model appears, I still feel a little excitement as someone who uses Claude Code every day.

“Maybe this is the one that will finally be even smarter.”

That is the hope I have whenever I update to a new model. So naturally, I had high expectations for Opus 4.8.

My first impression was not bad at all. In fact, I thought it might be a major leap forward.

But when you use a model for work every day, you begin to notice things that are invisible in a first impression. Gradually, my assessment changed.

At First, I Thought I Could Never Go Back

What impressed me most about Opus 4.8 was the idea behind Dynamic Workflows.

Instead of simply writing code, it breaks a task down and constructs a workflow as needed. The concept itself was fascinating.

“This could change the way we work.”

It certainly had enough impact to make me think so.

When everything clicked, its performance was remarkable. It could organize a complicated specification change in one pass and anticipate issues that would normally leave a human debating one option after another.

I even found myself thinking, “There is no going back now.”

Daily Use Revealed a Different Picture

The longer I used it, however, the more uneasy I became.

Some days, it seemed exceptionally intelligent. The next day, it would make an astonishingly basic mistake. A response would stop halfway through. A connection error would suddenly appear. Assumptions it had understood only minutes earlier would be forgotten.

To use a baseball analogy:

It can throw 100 mph.

But it struggles to find the strike zone.

No one doubts its raw ability. As a manager, though, you would hesitate to put it into a real game.

Some days it is unhittable. On others, it suddenly loses command and falls apart.

In software development, we do not need 100 mph on every pitch. We need a pitcher who can throw 87 mph and consistently find the strike zone.

In the End, I Went Back to 4.7

Ultimately, I chose Opus 4.7.

It is not as spectacular as 4.8, but it puts the ball where I expect it. When a model is part of your daily work, that sense of reliability matters far more than you might imagine.

One brilliant answer is less valuable than an AI that consistently does the job you expect every day.

In business, a model that scores 80 out of 100 every time is much more dependable than one that occasionally scores 120.

Looking Back, the Timing Makes Some Sense

Reconsidering the timeline later gave me another possible perspective.

Opus 4.8 was released on May 28. Roughly two weeks later, Anthropic announced Fable 5.

At the time, Fable appeared to be conceived not as a model included in a subscription from day one, but as a premium offering billed by usage.

What follows is entirely my own speculation.

If the strategy was to charge separately for Fable, Anthropic could surely have anticipated users asking, “Then what is the benefit of paying for Max?”

Perhaps Opus 4.8 was released first to soften that reaction by showing that the everyday model had also made a substantial leap forward.

If that consideration played any part in the decision, it might explain why Opus 4.8 felt somewhat rough around the edges.

Of course, Anthropic has never publicly said that this was the case. This is nothing more than one longtime user’s personal speculation.

Still, after working with Opus 4.8, I was left with a clear impression: it was highly capable, but not yet fully polished.

If my theory is correct, Opus 4.8 did not lack ability. It may simply have been promoted to the major leagues a little too soon.

What Fable 5 Made Clear

Later, I had an opportunity to use Fable 5.

For genuinely difficult work, it was clearly a level above. It excelled at complex system design, multistep reasoning, and tasks that required sustained thinking over a long period.

But it was also expensive and computationally demanding for everyday use.

I expected the division of labor to be simple:

The problem was that the model assigned to everyday work still occasionally threw a wild pitch. That makes it difficult to trust with routine responsibilities.

What I Want from Opus 5

What I want from the next Opus is surprisingly simple.

I do not need it to be faster. I do not need it to solve harder mathematics.

What I want is command.

I want the same quality every time. I want it to avoid going off the rails halfway through a task. I want it to do today what it successfully did yesterday.

For engineers, that kind of ordinary reliability is the greatest help of all.

Whenever a new AI model is released, the industry competes over benchmark scores. In the workplace, however, a different question often matters much more:

“Will it actually finish the job properly today?”

We have seen enough 100 mph fastballs.

The next pitch does not need to be right down the middle. Low and away is fine. Inside is fine.

Just throw a strike.

That alone would make us considerably more productive.

Then again, I am partly to blame for expecting every update to be the one that finally gets everything right.

AI learns something new every month.

And every month, I repeat the same expectation.