Home/resources/An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
Korea and Singapore AISIs jointly tested how AI agents behave in real-world tasks. The exercise examined whether agents can complete multi-step tasks in common settings such as customer service and enterprise productivity without leaking sensitive data. This blogpost shares the key findings, evaluation challenges, and methodological learnings.
This evaluation report details the various methodological components and findings behind this joint testing exercise, which will work towards advancing the science of AI agent evaluations and building common best practices for testing AI agents.
The 3rd Joint Testing Exercise builds on the insights from the previous two exercises, to work towards advancing the science of AI agent evaluations and building common best practices for testing AI agents. Bringing their collective technical and linguistic expertise, participating AISIs worked together to conduct testing for sensitive information leak, fraud, and cybersecurity threats.