LLM Evaluation
concepts · 7 notes linked
Related: Large Language Models · Langchain · GPT-4 · Fine Tuning · Claude · Prompt Engineering · AI Agents · Vicuna
Notes
- ChatGPT-4 Receives 'B' on Scott Aaronson's Quantum Information Science Final — GPT-4 scores B on honors quantum information science exam
- Evaluating RAG pipelines with Ragas + LangSmith — Reference-free RAG evaluation using Ragas and LangSmith
- GitHub - garg-ankush/scipe: SCIPE is a powerful tool for evaluating and diagnosing LLM (Large Language Model) graphs or chains. — Python tool for root-cause diagnosis of failing LLM chain nodes
- Introducing LLM-Evalkit | Google Cloud Blog — Open-source tool centralizing LLM prompt management and metric-driven evaluation
- Learning to Replicate Expert Judgment in Financial Tasks — Fine-tuned LLM beats frontier models at financial info triage
- SCIPE - Systematic Chain Improvement and Problem Evaluation — LangChain-featured tool for identifying failing nodes in LLM chains
- Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality — LLaMA fine-tune on ShareGPT data achieving near-ChatGPT quality for $300