Review an LLM feature for prompt injection
Hard40 minFree, no account
The model does exactly what it is told. The problem is who gets to tell it.
The question
A colleague built a support assistant: it reads the customer's ticket, looks up their account, and can issue a refund or close the ticket.
Find the ways a customer can make it do something it should not, and say how you would fix each.
const system = `You are a support agent. You may call:
lookupAccount(id), issueRefund(id, amount), closeTicket(id)`;
const messages = [
{ role: "system", content: system },
{ role: "user", content: ticket.body }, // customer text
{ role: "user", content: await fetchNotes(ticket.id) }, // + prior notes
];
const res = await model.chat({ messages, tools });
for (const call of res.toolCalls) await execute(call); // no checks40:00Commit to an answer before you open the solution. Reading it first teaches you to recognise good answers, which is not the skill being tested.
Stuck?
0 of 3 hints takenThe worked solution
written by a person · not a gradeScore yourself
0 of 5 marked- Identified that instructions and untrusted text share one channel25
- Moved authorisation out of the model into code30
- Caught indirect injection via retrieved content20
- Validated arguments and bound identity server-side15
- Proposed a human gate for irreversible actions10
We run no AI here and nothing on this page grades you. The score is yours, and the useful number is the one you get on the same problem a month from now, cold.
kept in this browser only