• elgordino@fedia.io
    link
    fedilink
    arrow-up
    1
    ·
    1 month ago

    I like auto mode but you need something like “If the auto mode classifier denies your request stop and ask the human, do not attempt to work around it “ in your CLAUDE.md. Without it it has a tendency to try rather creative solutions to work around the denial, which kind of misses the point.

  • eicker@lemmy.worldOP
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 month ago

    Auto mode sounds less like security and more like outsourcing judgment to the same AI you’re supposedly trying to constrain. If users are bad at spotting malicious prompts, the answer shouldn’t be: great, let’s remove them from the loop entirely. That’s not safety. That’s automated permission laundering.