Why does ChatGPT or Claude feel dumber? What's actually on the record
Every few weeks the same thread shows up on Reddit and X: “Is it just me, or did Claude get dumber?” Sometimes it is just you. Sometimes it isn’t. This guide collects what the AI companies have actually admitted, so you can tell the difference.
Short version: there is no public evidence that any major provider secretly swaps in a weaker model to save money, and Anthropic has said publicly that it never reduces model quality due to demand or intentionally degrades its models. But models have measurably gotten worse for stretches of time, and the providers confirmed it afterwards. The causes were bugs, changed defaults and bad training updates.
1. Bugs that degraded answers, confirmed afterwards
The clearest cases are Anthropic’s two public postmortems.
In September 2025, Anthropic published a writeup of three overlapping bugs from August and September 2025. A routing error sent some Claude Sonnet 4 requests to the wrong server configuration, peaking at 16% of Sonnet 4 requests on August 31. A separate bug corrupted some outputs on certain servers, which is why people saw Thai or Chinese characters in English replies. A third was a compiler bug that affected which words the model picked. Anthropic’s statement: “We never reduce model quality due to demand, time of day, or server load.”
In April 2026, Anthropic published a second postmortem about Claude Code. Three changes stacked up:
- The default reasoning effort was lowered from high to medium between March 4 and April 7. Anthropic later called this “the wrong tradeoff.”
- A caching bug kept deleting the model’s earlier reasoning between March 26 and April 10.
- A system-prompt instruction to keep answers short, live from April 16 to April 20, cost about 3% in quality.
Anthropic said the API itself was not affected, only Claude Code, and repeated that it never intentionally degrades models.
2. Training updates that went wrong
OpenAI’s clearest admission is the GPT-4o sycophancy rollback in April 2025. An update shipped on April 25 made the model overly flattering and agreeable. OpenAI started rolling it back on April 28 and explained that it had “focused too much on short-term feedback”: several changes, including a new reward signal based on thumbs-up and thumbs-down data, together weakened the signal that had kept sycophancy in check.
This is a different failure from a bug. The model did exactly what it was trained to do; the training target was off.
3. Outages and “degraded performance”
The most common cause of a bad day is the least mysterious one. Providers post incidents on their status pages when error rates rise or responses slow down, and during those windows some requests fail, time out or respond more slowly.
A few recent examples from the official status pages:
| Date | Provider | Incident |
|---|---|---|
| Aug 17, 2026 | Anthropic | Degraded performance for Claude Opus 5 and Sonnet 5 |
| Aug 18, 2026 | Anthropic | Degraded performance for multiple models |
| Aug 19, 2026 | Anthropic | Degraded performance for Claude Opus 5 and Haiku 4.5 |
| Sep 3, 2026 | OpenAI | Elevated errors across ChatGPT and Codex |
We plot these incidents on every model page, like GPT-6 Sol, next to community votes, so you can see whether a spike in “dumb” ratings lines up with an official incident.
4. Changed defaults inside the app
The chat apps are not the raw model. They add a system prompt, choose how long the model “thinks,” route some requests to smaller models, and apply usage limits. Any of these can change overnight without a new model name.
Two examples from this year: Anthropic’s reasoning-effort change above, and OpenAI’s August 6, 2026 ChatGPT update, which changed how GPT-5.6 Sol answers in ChatGPT and made GPT-5.6 Luna the default for Free and Go users. If you were on a free plan, the model you talked to that week was literally different.
5. Long conversations get worse on their own
This one is about how models work, not about any company. A 2025 study by researchers from Microsoft and Salesforce, LLMs Get Lost In Multi-Turn Conversation, ran more than 200,000 simulated chats and found an average 39% drop in performance when a task was spread over several turns instead of given all at once. Most of the drop came from the model becoming less reliable, not from losing ability.
In practice: if a chat has gone on for a long time and the model starts making odd mistakes, a fresh chat with a clean, complete prompt often fixes it.
6. And sometimes it really is you
It’s worth ruling out the boring explanations before blaming the model:
- The task got harder. Week one you asked for a function; week three you’re asking it to refactor a whole codebase.
- The novelty wore off. Early wins are memorable; later misses are too.
- You switched models without noticing. Check which model the app is actually using. Free tiers and usage limits can automatically move you to a smaller model, as described in each plan’s help pages.
How to check for yourself
- Keep a test prompt. Save one or two prompts you know the model used to handle well, and rerun them when things feel off.
- Check the status page of the provider before anything else. Each model page links to it, for example Claude Opus 5.5.
- Start a fresh chat and give the whole task in one message.
- Compare notes. Rate the model on the dumb meter and read the board. If lots of people report the same thing on the same day, it’s probably not you.
Votes on this site are opinions from self-selected visitors, not a benchmark. We show them next to the official record so you can judge the pattern yourself.
Sources
- Anthropic: A postmortem of three recent issues (Sep 2025)
- Anthropic: An update on recent Claude Code quality reports (Apr 2026)
- OpenAI: Sycophancy in GPT-4o
- OpenAI: Expanding on what we missed with sycophancy
- Laban et al., LLMs Get Lost In Multi-Turn Conversation (ICLR 2026)
- Anthropic status history
- OpenAI status: Elevated errors across ChatGPT and Codex (Sep 3, 2026)