The security surface of a server you exposed
You have handed a text generator a set of buttons. What happens when the text it read was written by someone else?
A tool surface is an API with an unusually credulous client. The model will be reading content from issues, emails, web pages and files, any of which may contain instructions aimed at it, and it has your buttons. The combination of broad tool access and untrusted content is the entire risk, and it is why a read-only server is a different category of thing from one that can write.
Least privilege applies exactly as it does anywhere else, with one addition: scope the tool, not just the credential. A tool called run_query with arbitrary SQL is unbounded no matter how careful the database user is; three named tools that each do one thing with bounded parameters are auditable. Narrow tools also produce better model behaviour, so this is one of the rare cases where the secure option is also the higher quality one.
Some things should not be tools at all. Anything irreversible, anything that spends money without a limit, anything that sends communication on someone behalf, and anything whose blast radius you cannot describe in a sentence. Put a human confirmation between the model and those, and log every call with its arguments, its result and who was responsible.
You should now be able to
- Enumerate what a server exposes beyond its intended use
- Apply least privilege to a tool surface
- Decide what must never be a tool
Loading…