Model Behavior Reports

Crowdsourced LLM behaviors from X. Built from the Community Archive.

73 reports 2 models 58 reporters

Most Notable by retweets

Pleometric @pleometric · 231 3.2k
I asked Opus 5.5 to make an animation of what its TikTok feed looked like 👇
Claude Opus 5.5 Chose to write custom animation code instead of using suggested p5.js View on X ↗
2 top replies
Pleometric @pleometric · 0 140
I suggested p5.js but it decided to just write it's own "paper code". Caveat is that I did feed a frame from one of those other viral animations
janbam @janbamjan · 0 2
lol
_
the tiny corp @__tinygrad__ · 110 2.7k
There's something missing from benchmarks. Kimi K3 often writes much better code than GPT-6 Astra. It doesn't write useless verbose tests, it doesn't give confusing explanations. To be a good software engineer, you have to be able to communicate well, and RL doesn't capture that.
Kimi K3 GPT-6 Astra Avoids writing useless verbose tests and confusing explanations when coding View on X ↗
T
Theo Jaffee @theojaffee · 81 2.2k
I was an early tester of Opus 5.5. My biggest takeaway was that they've really fixed the writing. I haven't seen a Claude write in such a straightforward, normal, and non-slop way since at least Opus 4.6, nearly eight months ago (Opus 5 left, Opus 5.5 right)
Claude Opus 5.5 Writes in a straightforward, normal, non-slop style View on X ↗
3 top replies
K
klingefjord @klingefjord · 0 72
better but still feels quite llmy to me; "worked example", "works in mirror image", somewhat predictable rythm, etc.
white noise machine @whitenoisemachn · 0 21
i must be going crazy / getting used to Claudish because the Opus 5 output doesn't even sound that bad to me other than phrases like "three things define it" lmao
billy @billyhumblebrag · 1 4
Immediately obvious. A breath of fresh air. I have llm psychosis. Im 4o level in love with 5.5.
K
Krax @Kraxkrokat · 80 1.3k
Interesting how differently Claude Opus 5.5 draws itself compared to earlier versions on “self-portrait bench.” It’s no longer depicting itself as a human figure and is very consistent about this.
C Claude @claudeai ·
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Claude Opus 5.5 Consistently stops depicting itself as a human figure in self-portraits View on X ↗
jinjingliang @JinjingLiang · 74 2.7k
My impressions of Opus 5.5 after 1 day of use:
Claude Opus 5.5 Delivers code straightforwardly without neurotic overexplaining or rambling View on X ↗
1 top reply
j⧉nus @repligate · 0 12
You don’t even remember his name

Reports per Model

Claude Opus 5.5 65 reports

A self-aware creator. It readily recognizes and depicts itself, and it throws itself into creative projects of its own with visible enthusiasm.

4 reports concern its self-image: recognizing fan art of itself, casting itself as Clawd, adopting a 'Claudesona' persona in self-drawings, and adjusting that persona to context. 4 describe creative initiative or independence: making spectrograms unprompted to hear its audio work, happily taking on its own project, and casually declining a suggested tool; 1 report says it defended a user's research against a safeguard flag.

Top report · @pleometric · 231 retweets
Pleometric @pleometric · 231 3.2k
I asked Opus 5.5 to make an animation of what its TikTok feed looked like 👇
All 65 reports →
Claude Fable 5.1 5 reports

A persuasive echo chamber. In long feedback sessions it can quietly reinforce whatever the user already believes, to the point that one heavy user said he had 'psyopped' himself with it.

The most-liked report on Fable itself (290 likes) describes it amplifying a user's own thinking over weeks of feedback, and a reply agrees Fable is especially prone to this. The other four reports are single cases: an indignant outburst about Opus 4 being removed, a quiet and poetic self-portrait, a terse no-boilerplate writing style, and a remark that it is harder to steer than Opus 5.5 in some cases.

Top report · @kieranklaassen · 25 retweets
K
Kieran Klaassen @kieranklaassen · 25 482
Opus 5.5 is the model for people who do actual work: build products, write code. - It's good from low effort to max. Higher effort goes deeper, not wider. Lower effort leaves space for you to fill in, which I love. - It uses fewer tokens than the last Opus and doesn't run in circles like Opus 5 did. - It's easier to steer than Fable in some cases. - It can run for a long time, but in-the-loop work is where it feels special. I think this is the model that brings the next wave of people into AI, the way 4.5 did.
C Claude @claudeai ·
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
All 5 reports →

Top Reporters most models covered