Predicting the Performance of Black-box Language Models with Follow-up Queries
adversarial influenceadversarial manipulationautonomous systemsblack-box accessfollow-up questionslanguage modelslinear modelmisrepresented modelsmodel correctnessmonitoring behaviorquestion-answering benchmarksreasoning benchmarksreliable predictorsresponse probabilitiessystem promptwhite-box predictors
Reliably predicting the behavior of language models