Refusal in Language Models Is Mediated by a Single Direction

ORIGINAL QUELLE:
arxiv.org

Quelle: Hackernews

Comments

← Zurück zum security Archiv (02.05.2026)