Reddit, Discord, Tumblr, and Meta are all leaning harder on AI moderation, and the misfires are piling up. Discord confirmed it wrongfully banned about 8,400 accounts between May and early July 2026 after its AI mistook square grids like chessboards and spreadsheets for child sexual abuse material. Reddit's revamped AI tools recently deleted r/AskHistorians posts dating back 10 years. Tumblr wrongfully banned nearly 200 accounts in a single afternoon in March 2026. In every case, users had no working appeal channel before the damage was done.
Reddit is publishing numbers that look strong in isolation. The company says AI has driven a 200 percent increase in enforcement actions on hate and violent content and cut exposure to potentially harmful content by more than 40 percent. Its systems revoke nearly 2 million fake votes daily, and Reddit says it uses large language models to catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed. Those are real gains against a real problem — LLM-driven spam is a new category of abuse that keyword filters cannot handle.
But the same LLM boom that Reddit is defending against is also making moderation harder from the other side. Sarah Gilbert, an r/AskHistorians moderator and research director at Cornell's Citizens and Technology Lab, said large language models have made spam detection a lot harder because bots now mimic real human voices. Marketing agencies have built businesses around gaming this dynamic. The startup ReachLLM creates and moderates subreddits specifically to get client brands cited by generative AI chatbots.
“Over the last two to three months, we've been absolutely flooded by LLM-powered spambots”— Sarah Gilbert, r/AskHistorians moderator and research director at Cornell's Citizens and Technology Lab
Key facts
- 01Discord wrongfully banned about 8,400 accounts between May and early July 2026 after its AI mislabeled square grids like chessboards as CSAM.
- 02Reddit says AI has increased enforcement on hate and violent content by more than 200 percent and cut exposure to potentially harmful content by 40 percent.
- 03Reddit's AI tools now revoke nearly 2 million fake votes daily, using LLMs to detect coordinated fake-behavior patterns.
- 04Tumblr wrongfully banned 'sub-200' accounts in a single afternoon in March 2026, according to Automattic communications head Chenda Ngak.
- 05r/AskHistorians moderators say Reddit's AI removed posts dating back 10 years, likely because they linked to the site Rare Historical Photos.
The r/AskHistorians incident in April illustrates the cost of over-correcting. Moderators watched dozens of comments and posts vanish from a subreddit whose users treat it as a research archive. After recovering some of the deleted text, the mods noticed every removed post linked to Rare Historical Photos, a historical image-sharing site. Their working theory: Reddit's system flagged the domain as spam and swept up a decade of expert answers with it. Some of those answers had taken contributors hours, sometimes days, to research and write.
Discord's failure was more systemic. The company said its AI moderation was never supposed to run without human review, and that a bug caused the AI to bypass the human step and ban accounts. Every one of the roughly 8,400 wrongful bans has since been reinstated. The affected users spent weeks locked out based on a pattern-matching mistake — square grids in an image — that a human reviewer would have cleared in seconds.
Meta users have been raising similar complaints on Facebook and Instagram since 2025, describing mass bans with no path to reach a human. Meta has not confirmed AI as the cause, but the company has publicly shifted moderation toward generative AI systems in recent years. Some Meta employees have said internally that the shift is happening too fast. Tumblr parent Automattic says it uses a mix of machine-learning classification and human moderation, and communications head Chenda Ngak confirmed the sub-200 wrongful bans in March. Tumblr users also complained through 2025 that automated systems were flagging benign content as mature and suppressing its reach.
The trust problem compounds when platforms narrow their transparency at the same time they widen their automation. Gilbert said that back when there was more transparency in the system, AskHistorians mods would routinely report hate content, get an automated non-violation response, and then start a formal appeal. Now, both the enforcement and the review sit behind the same opaque model.
“So it's hard to trust the numbers because it's hard to trust the 'judgment' of Reddit's systems”— Sarah Gilbert, r/AskHistorians moderator
Research cited by Gilbert points to a second problem: false positives fall unevenly. Marginalized and vulnerable populations experience the highest rates of moderation, she said, often because AI systems misread counter-speech, language reclamation, and responses to hateful content as the hate itself. Her framing is direct: false positives are an equity issue that further silences groups the systems are supposed to protect. There is also a governance side-effect. When AI removes rule-breaking content before a human mod sees it, subreddit teams lose the ability to decide whether the offending user should be banned outright.
Reddit is now testing a partial answer called Rules Hub, a set of tools that lets human moderators pick which rules are auto-enforced, choose the action when a rule is triggered, preview the outcome before turning it on, and review logs after the fact. Reddit expects Rules Hub to eventually replace Automod, its long-running keyword-based tool. The direction — giving humans a configurable layer on top of the AI — is the right one, and it lines up with what Discord said its system was supposed to do all along.
The market implication is straightforward. Every major consumer platform has bet that AI moderation can absorb the flood of LLM-generated spam and abuse at a cost structure human teams cannot match, and the raw enforcement numbers back that bet. But the wrongful-ban incidents at Discord, Tumblr, and Reddit are not edge cases — they are the predictable output of running classifiers at platform scale without a review layer that actually blocks the ban button. The platforms that end up owning trust in this cycle will be the ones that treat AI moderation as a first-pass filter feeding human judgment, not as a replacement for it. That is a more expensive product to run, and it is the one users will demand once enough archives quietly disappear.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




