Researchers propose focusing on identification of certain cognitive elements in LLMs that indicate when AI systems may take an unwanted action.
Researchers propose focusing on identification of certain cognitive elements in LLMs that indicate when AI systems may take an unwanted action.