llm-based agents
Intelligent agents that utilize large language models (LLMs) to perform tasks such as dialogue generation, summarization, or reasoning. These agents leverage the capabilities of LLMs to understand and respond to human-like queries.
- AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
- Can Agent Fix Agent Issues?
- EnCompass: Enhancing Agent Programming with Search Over Program Execution Paths
- MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch