正在仔细看今天 Anthropic 发布的报告,里面写了一些利用 AI 的攻击性质的行为。报告的中文翻译参见:

https://wlj.me/reading/detecting-and-countering-ai-misuse-2026-09/

我的想法是: