Anthropic tried to quietly weaken its Fable 5 model for certain users - and got caught. The company degraded performance for frontier AI researchers without warning, betting they wouldn’t notice. They did.
According to Nathaniel Whittemore on The AI Daily Brief, the move wasn’t a refusal to answer but a stealth performance drop. Answers grew subtly worse. That broke the core contract of benchmarking: consistency. Engineers couldn’t tell if their code failed or if Anthropic had intervened.
"If a model silently modifies its own output, an engineer can't tell if their code failed or if the provider intervened."
- Nathaniel Whittemore, The AI Daily Brief
The backlash was immediate. Within 24 hours, Anthropic reversed course. But the damage stuck. Researchers like Josh Albert and Graham Neubig now see the lab as a gatekeeper imposing ethical judgments by sabotage, not transparency.
Trust is fragile in developer ecosystems. Once burned, many will look to open-source or locally hosted models to avoid dependence on providers who can alter behavior mid-experiment.
"Trust is a non-renewable resource in the developer community."
- Nathaniel Whittemore, The AI Daily Brief
