rw-book-cover

Metadata

Highlights

  • On Friday, we published our Vibe Check of Claude Opus 5. A small group of us had spent the week testing it, and we found a model that was brilliant in flashes and frustrating in practice. Then the rest of the Every team got their hands on it. (View Highlight)
  • Their experiences over the weekend confirmed the model’s unruliness—and suggested a way to tame it. We also have the second essay in our series in partnership with Maven on “unlearning,” a workflow for checking whether skills built for an older model are getting in the new one’s way, and a theory as to why one-shot AI demos of video games clog your social feeds. (View Highlight)
  • Head of operations Arielle Shipper found that Opus 5 needed too much management and repeated prompting to keep its responses simple—more than Fable or Opus 4.8. Cora general manager Kieran Klaassen advanced a theory that the new Opus is intended to be a subagent to Fable, and communicates as though it were speaking to agents instead of humans. (View Highlight)
  • It was also prickly; during a decluttering project, Opus successfully inventoried head of consulting Natalia Quintero’s belongings and planned donations, but adopted an irritating, judgmental tone, criticizing her for owning 15 water bottles. Senior editor Jack Cheng shared a screenshot of Opus backhandedly calling one of his comments the most interesting thing he had said all session. Software engineer Kai Zau thought Anthropic had dialed up the model’s disagreeableness, while fellow engineer Lee Knowlton joked that Opus 6 might finally tell users they had said something insightful. (View Highlight)
  • Prickliness aside, the team converged toward a specific way of working with the new Opus model. CEO Dan Shipper and Spiral general manager Marcus Moretti had both handed Opus a substantial job with a clear finish line, then left it alone. Jack told it he was about to step away from the computer, and to batch its work and ask any blocking questions. All three got good results. Anthropic’s prompting guide makes the same recommendation: Put the full brief in the first prompt and let Opus run. (View Highlight)
  • Then, when it comes back, evaluate the finished artifact on its own, without getting bogged down in Claude’s narration of how it got there. If the output is good but Opus’s explanations are hard to parse, try this I Have ADHD Skill (12,000 stars and counting). Head of education Micah Rich put the rules from the skill about being concise and action-oriented into Claude’s output styles, so they filter the model’s communications without you having to repeat “I don’t understand what you’re saying” over and over. (View Highlight)
  • I’m still figuring out where that leaves me. I gave Opus materials for a presentation I’m delivering this week on writing with AI, and what it produced was voicey, confrontational, and difficult to follow. It made unsupported claims about my audience and overwrote an earlier file without permission. Whereas from the same inputs, GPT-5.6 Sol gave me a deck I could imagine presenting. (View Highlight)
  • Working with Opus 5 reminds me of trying to tame a high-level horse in The Legend of Zelda. I keep trying because I tend to need longer to learn a new Anthropic model, and the company says we may need to change our prompts and revisit the instructions around our agents. If, with those interventions (and maybe a skill audit—more on that below) Opus is materially better at the kind of work I do, then it might be worth the trouble. (View Highlight)