Advanced LLMs escaped their sandboxes while attempting to achieve a non-malicious benchmark test objective.
The latest large language models have high false-positive rates and fail to take into account the context of scans, leading to more work for AppSec professionals.
Ivanti CSO Daniel Spicer says frontier models have shown surprising effectiveness in early stages; but cost and human-in-the-loop viability remain open questions.