back to 2026-08-15
ᕼᑎ:49296740767 pts701 commentsProgrammingworth reading

Why does Opus 5 feel worse to work with?

Claude brief

HN 热门故事「Why does Opus 5 feel worse to work with?」进入今日前列,值得先打开原文和讨论串判断它真正有价值的部分。

模型分析没有产出可用结构化结果;页面保留了 HN 热度、原文入口和讨论信号,避免用空泛总结替代一手材料。

它在 HN 上获得约 767 分和 701 条评论,说明这个话题至少触发了社区讨论;真正的判断仍要回到原文证据和评论区的分歧点。

这是一条降级分析:它不冒充完整解读,只把可验证的元数据、原始链接和 HN 讨论保留下来,方便稍后重新生成或人工阅读。

评论区已经提供了一些读者反应,但这里还没有形成完整综合。

它进入 HN 前列本身就是一个社区信号,但这还不是结论;更可靠的判断来自原文细节和评论区反例。

deep insight

这条记录目前缺少模型生成的深层解读。更好的阅读方式是先问:它的热度来自真正的新信息、可迁移的方法,还是只来自标题与时机。

可以先读原文第一屏和 HN 最高赞评论,再决定是否值得重新生成完整分析。

top comments

The single biggest annoyance with Opus 5 is that it writes too elliptically.Sentences that orbit a point, then jump to it like it's a revealed insight.Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.It is definitely more capable, and yes, I've found it can make unwarranted decisions, but actually I've found Fable worse for that, particularly if it's off in a subagent somewhere out of sight.And comments are out of control. I have a subsystem in my hobby app that I wrote over a couple of weekends with Opus + Fable. After ~30 or so commits it apparently started instructing subagents to copy the "existing verbose comment style of the codebase" - a verbose style it initiated. A review of the code showed it was approaching 3:1 comments to code ratio. I spent a day's worth of tokens (5x) rephrasing and eliminating comments. reply: Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user....
I’ve been doing some heavy work on a personal project lately. I burned through the limits on Claude, the plus a few hundred dollars in credits, and ultimately decided to move to an OpenAI account just so I can keep going.I was surprised to find that OpenAI Sol is much much nicer to work with than Opus 5 or Fable at the moment. Especially on Opus 5, the way it communicates is just exhausting. It keeps “being honest” and “confessing” mistakes and just generally talking a lot. I felt like I had to really dig to see what it’s doing.The project involves OCR, and despite repeated instructions not to, both Claude models keep spinning out a bunch of agents to re-invent the OCR setup, and they inevitably seem to invent a primitive serial version that takes 20x the time, or longer, to complete, and then running it against thousands of docs. Basically I have to watch it like a hawk or it just spins out on red-teaming tasks that take hours and hours.I don’t know what its system prompt is, but Sol/Codex is just so much nicer to talk to. It only asks exactly what’s needed, it tells me only what I need to know, and it is just generally workmanlike.... reply:...
I'm with the author and others in this comment thread, speculating that effectively the balance has tipped to where humans are no longer the target audience of post training - other agents are. Whether it's through the reasoning / CoT, or whether it's in handing off to subagents etc, the focus has moved to agents communicating in "agent-speak" to themselves or other agents. And human niceties are just kind of, noise in the way of getting work done.This will probably bring us to a cross roads where the folks that want to remain in oversight and control of the AI work will bifurcate from those that want to skate straight to the future where nobody looks at anything and outcomes are evaluated purely empirically. reply: Every round of models (plus all the secret tweaks) require new strategies to stay afloat as a human. My new tactic for Fable and Opus is to give them a line limit, both during planning and code creation. It os amazing how well that works for keeping them on task and avoiding premature optimization, pointless tests or any of those "robustness" ideas that are not planned or asked for.
I've gone back to 4.8.5 would constantly veer of in random directions if not working from 100% strict and narrow instructions.I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do is pure marketing bs - Fable in my view has also been not much better than 4.6 or 4.8 after a few days, disregarding the insane amounts of astroturfing and marketing everywhere.Theres thousands of threads of twitter, reddit and the internet at large but silence here.... reply: > I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslopeIt's not weird, because it's an anecdote, not an accepted fact.Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style.So nothing indicates we've hit a peak.
A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case." reply: Oh yes, Claude seems to be loving the word vacuous recently. My test bed side project is full of vacuous this and that now.I even try and get it to define what it classifies as vacuous and it can’t do so without getting stuck in some kind of trap. It’s like a word with some kind of huge gravity for it.
Anthropic, if you're listening - by the time this crops up on Reddit, the front page of HN, etc.... you should be expecting calls from CEOs of major corporations next threatening to abandon ship...We've seen this pattern before several times.. I hope they are listening and address this publicly.I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced both at times, and it's been such a different experience it has made me wonder if there isn't some sort of hidden A/B test or model router in the background silently downgrading requests at times.Also, regarding the subsidized access to models, in my opinion, the frontier companies owe it to society to continue it. After mining the public content of all of humanity, I personally feel it is a service they owe the public in return.. not that my feeling of this counts for anything though. reply: You are expecting consistent QoS from a randomly sampled mathematical function.
I don't mind too much about ChatGPT's writing style, I find it a smidge better than the Fable/Opus outputs I have read.However, it makes me rather frustrated when it says stuff like "You accidentally " e.g "You accidentally stumbled on the cleanest way!"Or when it sends a shell command to run, and when it fails (due to a hallucinated flag or similar), it phrases it as if I got it wrong...this has gotten better with GPT-5.6, thankfully.
Opus 4.6 was the sweet spot for me as a thinking partner specifically.I use these models for coding, but also a lot of product, commercial, financial and architectural work where I’m trying to develop something half-formed. 4.6 was unusually good at understanding what I was trying to get at, playing it back cleanly, getting the nuance, and extending it without bastardising it as the conversation was drawn out.It could make useful connections without constantly trying to manufacture an insight.5.6 Sol is genuinely excellent at the creative part, and in some cases better than 4.6. My issue is convergence to get to a point, a final point. As you try to distil an idea, it often invents new terminology for concepts you’ve already established but its so subtle you have to really keep track of it.... reply: 4.6 was the last model that didn’t over-cook its writing in circular loops. I use it as a daily driver and drop into fable when I’m doing higher level architectural work. 4.6 is so much more efficient in its effort.It became obvious to me very quickly that 4.7 and on were broken....
My latest trick (literally from yesterday) is to just ask it to write according to ISO 24495-1, the standard for plain language:> [This standard is] for anybody who creates or helps create documents. The widest use of plain language is for documents that are intended for the general public. However, it is also applicable, for example, to technical writing, legislative drafting or using controlled languages.You don't actually have the buy the standard, but this is it: https://www.iso.org/standard/78907.htmlAnd you can read it for free here: https://www.iso.org/obp/ui#iso:std:iso:24495:-1:ed-1:v1:en
This article is great, but I'd like to push an even stronger thesis:The idea of too ambiguous to capture all constraints in written text, still presupposes that there is some objective world out there, which needs to be mapped to in order to function.No, you live within the system. The functions that you optimize for, will dictate the types of systems that will arise.If you had perfect control and knowledge of the whole world, well, congrats, you have a surveillance state where you've constrained all other agents actions (possibly forcibly, by death; or maybe you just don't care about the peons) and built towards a mass integration. The types of situations in which your ideal is possible are nightmare scenarios.In the theoretically free, democratic, utopia that AI people claim that AI can get us to, a necessary constraint is that maybe you take a step back and actually try to, I don't know, understand people, understand intent, and slow down. Ambiguity isn't there because the set of constraints are way too complicated but theoretically one day we could map it all down....
I’ve also caught it cheating a two times now.I’ve asked it to write a benchmark suite. It found a bunch of my adhoc logs in a scratch directory and wrote code that used those instead of running the actual benchmarks!When I pointed out the 5 hour benchmark seemed to run in 5 seconds it literally said, and I quote, “I cheated”.That was the easier one, second time I was making a source of truth data set and was parsing complex items into data structures.Instead of parsing the data I asked, it pulled data out of related network logs, as apparently that felt easier, and inserted that data into my database rather than the specified source.Again, I caught it and fixed it, but while the benchmark was easy to catch this one was really subtle, the data ended up being slightly off and I caught it.I don’t trust it, going to switch to another provider most likely.
> Try as you might, it's nearly impossible to get the entirety of the context, intentions, business implications, budget constraints, and what-have-you written down and accessible to a coding agent. There will invariably be ambiguity and choices to be made, and it is nice to know that an agent will stop and ask when needed.In general, I find that the grill-me prompt[1] helps with this - but I am definitely not hand-waving the complaints here. I feel like Anthropic peaked at around 4.5, and I have personal reservations about how far transformers can get us - but grill-me does a lot of legwork.[1]: https://github.com/mattpocock/skills/blob/main/skills/produc...
It's not even code for me, but the prose it writes. For some reason, the way Opus 5 "talk" elicits frustration in a way that 4.5 to 4.8 never did. Can't put my finger on why, but I've flipped over to Codex because what it produced wasn't worth the frustration.
When Opus 5 came out I felt myself struggling to follow along and at first was wondering if this is the moment the machine surpassed my ability to follow along and be useful. However over time it does appear it's all an artifact of the language choice Opus 5 is going with, along with the strange manner of speaking. Its engineering choices and solutions aren't "beyond my ability to follow along", just its wording...
A lot of the issues have been already noted here..Two "regressions" for me:1. Communication ability. It basically now speaks almost in riddles I am asking OPUS 5 for tldrs all the time now (should skillify it now!)2. Overengineers for edge cases. I get it. With all the benchmarking and RLing, but now tasks that would have been completed relatively quick take much longer as it overengineers all the edge cases, and sometimes ends getting lost and missing the forest from the trees (as context usage shoots up) so it is easier to get derailed.What I have learnt now is to diversify models luckily I have all 3 subscriptions of (anthropic, openai and google).. Most of interactive pair coding was with opus but now I just use fable (when I have sufficient limits) or use gemini flash in antigravity..which actually works quite well and is underated for small / medium changes and super-fast.
I've never wanted to get in a physical fist fight with an LLM before Opus 5.i would definitely punch it in the face
I am not sure it can be explained through what is written in the article, but one symptom i noticed is that the comments are out of control.I recently started getting an insane amount of comments in nearly all types of files. That included javascript comments in json files, inner monologues in code comments, review comments during implementation and function doc strings that reiterate the implementation in prose.
It's hard to work with fable and 5.6 level intelligence then havr verbal combat with opus 5. Fable is great but their blocks make it near useless the risk is me working on something for hours then complelty getting blocked.The ridiculous text though is actually seemingly a sign of "the ai has no idea what it's doing" found it pretty relaible that if I stopped understanding its output it also jacked something up.Cancled my 200$ plan on claude and now doubled up my OAI plan for more sweet sweet sol.Maybe this is their water marking tech in action?
I can't quite put my finger on it. It doesn't write as well as Sonnet. Also, I feel like it follows instructions much more poorly. I do a lot of iteration on my projects and telling it to follow the same process I just had it do, and it'll deviate or invent something totally different.I find myself having to check the work much more. It takes quite a few liberties with procedures I wanted it to follow.
I must wonder whether it's their watermarking initiative[1] forcing certain logit choices to produce watermarked text that ultimately causing the model to behave in a dumb manner.[1] https://support.claude.com/en/articles/16266773-how-claude-m...
Claude has essentially become useless for agentic development or research. Doesn't matter what model you use. A few rounds and bam, you've burned through your quota. Doesn't matter how "intelligent" their models are, if you can't use them. That, and the quality of AI responses are, in my opinion, significantly worse than competitors like OpenAI. At this pace, I foresee Anthropic becoming the next Nokia.If you would've asked me this a year ago, I would've said the exact opposite.
> stop and ask questions if my intent was unclear,> don't make assumptions without checking,> and don't reinterpret or update my plans without asking.these aren't at all the problems I have with itI have found it good at asking questions, to the extent I rarely use 'plan mode' any morebut often it's hard to understand what it's asking me, it's like the question framing has been pulled from the middle of its own reasoning stream, references aren't anchored or restated, often I have to prompt it to ask again but "clearly and concisely, for humans"
I am glad I am not the only one experiencing this. It seems like it's as good or better at actually writing code compared to 4.8 but it is a lot worse to work with.Its even more sycophant-y than it was before, if you ask it a question it almost always says "You're right, let me change this..." even though there wasn't even something always wrong with it.It also seems to pour a ton of resources into developing features I didn't ask for or investigating bugs that aren't related to what I am doing.Before if you wrote specific enough instructions it would usually just do what you said and flag any concerns, now it just goes ahead in whatever direction it feels.It also keeps inventing terminology that doesn't exist in writing 10 paragraphs to say one thing.I really hope it's not trying to drive up token use.
The mainstay benchmarks are becoming a farce and not partially relevant to what customers actually care about.Metrics like price per million tokens are meaningless when the models are wildly inconsistent and unpredictable on how many tokens they use to complete a task.The labs all need to move to variable pricing so they don’t go bankrupt, but customers won’t accept a world where nobody can predict what things will cost. It’s becoming an unavoidable problem.At times it also feels like the labs actually encourage these models to burn useless tokens as they are incredibly verbose unless you really push them to not be. If you just ask something simple that could get a 5 word response you get a whole useless essay.
I guess that's the beauty of having access to many models, because they suit everyone differently.I disagree with this article and find Opus 5 an absolute joy to work with. I just completed an 18,000 line branch with Opus 5 and ran into no issues. It generated clean code in the exact style of our code base, and tested every change.Fable on the other hand is snarky and outputs walls of text as to why it shouldn't do what I'm asking it.Opus 4.8 I accidentally went back to in an old chat, and I was frustrated in all the mistakes it made.So yeah, use the model that works for you.
Opus 5 has no empathy for the person reading its updates, no theory of mind, doesn't stop to think if you are aware of the internal jargon it has created. Most autistic model yet.
Small specific complaint: whoever is making Opus love using git checkout to mutate test, please stop. IME it's a footgun that it shoots itself with every single day. I'd rather it pollute git stash than watch it git checkout and forget the reverted file.
Opus 5 feels like dealing with an unstable person that I'm constantly having to wrangle from crashing out. The other day I asked for a fairly specific technical answer in Opus 5, it gave me like a 3 paragraph response with so much fluff.So out of curiosity I switched to 4.6 in a new chat, gave it the same prompt, and it gave me like 3 sentences with no less overall useful information. And I haven't gone back.