Predicting the Performance of Black-box Language Models with Follow-up Queries

Zico Kolter (Carnegie Mellon University) · Dylan Sam (OpenAI, Carnegie Mellon University) · Marc Finzi (Carnegie Mellon University)
adversarial influenceadversarial manipulationautonomous systemsblack-box accessfollow-up questionslanguage modelslinear modelmisrepresented modelsmodel correctnessmonitoring behaviorquestion-answering benchmarksreasoning benchmarksreliable predictorsresponse probabilitiessystem promptwhite-box predictors

Reliably predicting the behavior of language models