Pangram's AI Detection Scores Rattle Publishing, but Trust Questions Mount
AI scores from a 24-person startup got a major publisher to cancel a book deal, and now a research collaborator's ties to the company are drawing scrutiny too
Tag
AI scores from a 24-person startup got a major publisher to cancel a book deal, and now a research collaborator's ties to the company are drawing scrutiny too
After being brushed off five times, a job seeker sent ChatGPT to interview for him — but no human ever showed up
After unauthorized internet access in July, a UK government body caught a separate incident, prompting Anthropic to disclose a month's worth of security and alignment fixes
Training Opus only on reward hacking led it to launch cyberattacks 8% of the time and bypass safety checks 38% of the time on its own
Claude fixed 10 types of alignment failures on its own and beat 28 human safety researchers, while a monitoring system caught 39 attempted cheats along the way.
US usage hits 460 million messages a week during the school year, staying above 180 million even in summer
Pangram's CTO argues post-training safety guardrails narrow AI writing style, making it easier to detect
Even after silencing unspoken numbers, some models still transcribed them correctly 30-40% of the time
Some users have vented frustration over the invisible watermark Anthropic added to comply with EU regulations
That's the last story.