Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's an alignment problem in the sense that it demonstrates the principle that today's AI systems cannot be trusted to reliably work towards the goals of their users. A small-scale alignment failure and a large-scale alignment failure are the same fundamental type of failure. Typically, large disasters come after smaller disasters which foreshadowed the disaster mechanism, but weren't taken seriously.


I think there is a meaningful difference in kind between an AI that makes a mistake and an AI that is actively malicious.


An AI which is mistaken about the goal you give it, or how to go about achieving it, will behave in a de facto malicious manner, as this incident illustrates.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: