Large Language Models
concepts · 206 notes linked
Related: Openai · Chatgpt · AI Agents · Google · Prompt Engineering · AI Safety · Anthropic · Agentic Coding
Notes
- $450 and 19 hours is all it takes to rival OpenAI's o1-preview — UC Berkeley Sky-T1-32B open-source reasoning model built for $450 in 19 hours
- 'Chemistry will no longer be an exclusive club': how AI is changing Omar Yaghi's work — AI accelerating metal-organic framework discovery for climate applications
- 'If journalism is going up in smoke, I might as well get high off the fumes': confessions of a chatbot helper — Human annotators writing gold-standard training data for LLMs
- 'The illusion of thinking': Apple research finds AI models collapse and give up with hard puzzles — Apple study showing large reasoning models collapse on hard logic puzzles
- 10 things I learned from burning myself out with AI coding agents — Practical limits of AI coding agents after 50 real projects
- 27 AI models were ranked by the public and ChatGPT came 8th — these are the models that beat it — Prolific's Humaine leaderboard ranks Gemini first in public user experience study
- 8 Google Employees Invented Modern AI. Here's the Inside Story — Origin story of "Attention Is All You Need" transformer paper at Google
- A Coder Considers the Waning Days of the Craft — Programmer's first-person reflection on GPT-4 displacing coding craft
- A Complete Breakdown of the Claude Mythos 1 Leak and Features — Leaked Claude Mythos 1 capabilities including 69% Exploit Bench score
- A Note From Ray Kurzweil on the Recent Call to Pause Work on AI More Powerful Than GPT-4 — Ray Kurzweil opposing FLI's AI pause letter citing vagueness and coordination failure
- A Wave Of Billion-Dollar Language AI Startups Is Coming — 2022 landscape survey of language AI startup ecosystem categories
- A former Google researcher behind a seminal AI paper describes how the company lost a top chatbot visionary — Google's reputational caution caused it to miss chatbot opportunity
- AI #174: You're It — Weekly AI roundup, Claude Tag, medical scanners, agent security
- AI #175: The Fable Continues — Weekly AI roundup covering Fable's return aftermath
- AI 2027 — HN debate on AI 2027 scenario: AGI timelines, LLM limits, alignment risks
- AI Index — Stanford HAI annual data-driven report tracking global AI progress and impact
- AI Is a Waste of Time — AI tools as entertainment and time-wasting before productivity gains materialize
- AI Sentience, Agency and Catastrophic Risk | TWIML - The Voice of Machine Learning & AI — Yoshua Bengio discusses AI catastrophic risk and governance
- AI critic Gary Marcus: Meta's LeCun is finally coming around to the things I said years ago — Gary Marcus argues deep learning alone cannot achieve general intelligence
- AI fears overblown? Theoretical physicist calls chatbots 'glorified tape recorders' | CNN Business — Michio Kaku argues chatbots are limited and quantum computing is the true next stage
- AI hype is built on high test scores. Those tests are flawed. — LLM benchmark scores are brittle, anthropomorphized, and often measure memorization not capability
- AI is already linked to layoffs in the industry that created it | CNN Business — AI driving tech sector layoffs and workforce skill reshuffling
- AI pair programming in your terminal — Terminal-based AI coding assistant with codebase map and git integration
- AMD Researchers Introduce Agent Laboratory: An Autonomous LLM-based Framework Capable of Completing the Entire Research Process — Autonomous LLM pipeline completing literature review, experiments, and paper writing
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support — Live human-feedback flywheel continuously improving LLM customer support system
- Amazon warns employees not to share confidential information with ChatGPT after seeing cases where its answer 'closely matches existing material' from inside the company — Amazon restricts employee ChatGPT use over corporate data leakage risk
- Artificial General Intelligence Is Already Here | NOEMA — Argument that current frontier LLMs have already achieved meaningful artificial general intelligence
- Artificial Intelligence — The Revolution Hasn't Happened Yet — Michael Jordan argues AI hype obscures the real ML/data-engineering challenge
- Beyond Code Autocomplete — AMD's holistic AI integration across the full software development lifecycle
- Billionaires including Eric Schmidt plow $300 million into a non-profit that is France's latest push to catch up in AI | Fortune — France's €300M open-source AI research nonprofit kyutai launched
- Black Hat 2023 Review: LLMs Everywhere | DNSFilter — Black Hat 2023 LLM applications in security; best use leverages embeddings not chat interfaces
- Building a C compiler with a team of parallel Claudes — 16 parallel Claude agents autonomously build Rust C compiler
- Building a Knowledge Graph for Business: The Semantic Backbone of Big AI — Building enterprise knowledge graphs with ontologies and LLM integration
- Can 1B LLM Surpass 405B LLM? Optimizing Computation for Small LLMs to Outperform Larger Models — Test-time scaling lets small LLMs outperform much larger models
- Can AI really be protected from text-based attacks? | TechCrunch — LLM prompt injection attacks are low-barrier and currently unpreventable
- ChatGPT Gets Its "Wolfram Superpowers"! — ChatGPT plugin connecting to Wolfram Alpha and Language
- ChatGPT Makes OK Clinical Decisions—Usually — Study finds ChatGPT 72% accurate across clinical decision tasks using fictional vignettes
- ChatGPT and generative AI are booming, but the costs can be extraordinary — Compute economics of training and serving large language models
- ChatGPT is 'not particularly innovative,' and 'nothing revolutionary', says Meta's chief AI scientist — Yann LeCun argues ChatGPT is solid engineering not scientific breakthrough
- ChatGPT is not all you need. A State of the Art Review of large Generative AI models — Taxonomy of 2022-2023 generative AI models across all major modalities
- ChatGPT-4 Receives 'B' on Scott Aaronson's Quantum Information Science Final — GPT-4 scores B on honors quantum information science exam
- Chegg's stock plunges on fears of competition from ChatGPT — ChatGPT disrupts edtech as Chegg stock crashes 48% in one day
- Claude Code Ralph Plugin Breaks LLM Performance (And a Simple Bash Loop Wins) — Ralph methodology critique — stateless bash loops outperform official Claude Code Ralph plugin
- Claude Sonnet 5 Is Not Frontier But Has Its Uses — Sonnet 5 system card review, cheaper faster non-frontier model
- Claude for Life Sciences — Anthropic launches Claude specialization for life sciences research workflows
- Cohere Command Models: AI-Powered Solutions for Enterprise — Cohere Command family of enterprise LLMs for agentic and RAG workflows
- Cost calculations for LLM providers — Comparative USD-per-million-token pricing table with LMSYS ELO scores
- DS-STAR: A state-of-the-art versatile data science agent — Iterative plan-verify data science agent handling heterogeneous file formats
- DSPy — Python framework replacing prompt engineering with optimizable typed signatures
- Deconstructing Geoffrey Hinton's weakest argument — Gary Marcus rebuttal of Hinton's defense of LLM understanding and hallucinations
- DeepMind Paper Provides a Mathematically Precise Overview of Transformer Architectures and Algorithms | Synced — DeepMind formal pseudocode reference for all major transformer architectures
- DeepSeek - A Wake-Up Call For US Higher Education — DeepSeek's rise reflects China's STEM education investment advantage
- Distilling step-by-step: Outperforming larger language models with less training — Rationale-based distillation trains small models to outperform 540B LLMs with far less data
- Does Prompt Caching Make RAG Obsolete? — Prompt caching economics vs RAG for LLM context management
- Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking) — LLM-assisted iterative mathematical proof via generate-referee-repair pipeline
- Elicit: AI for scientific research — AI-powered academic literature search and systematic review tool
- Employment for computer programmers in the U.S. has plummeted to its lowest level since 1980—years before the internet existed | Fortune — AI correlates with historic drop in US programming jobs
- End of an Era at Google DeepMind Hints at New Future for AI — Last "Attention Is All You Need" author departs Google, ending transformer paper era
- Even experts are surprised by AI's latest 'vibe-mathing' advance — Amateur uses GPT-5.4 Pro to solve 60-year-old Erdős primitive sets problem
- Fine-tuning ChatGPT: Surpassing GPT-4 Summarization Performance — 63% Cost Reduction and 11x Speed Enhancement using Synthetic Data and LangSmith — Fine-tuned ChatGPT beats GPT-4 summarization at 63% lower cost and 11x faster using chain-of-density synthetic data
- Fine-tuning · Hugging Face — Fine-tuning pretrained LLMs with Hugging Face Trainer API
- For Some Autistic People, ChatGPT Is a Lifeline — Autistic people using ChatGPT for social scripting and communication support
- GLM-5.2 Is The New Best Open Model — GLM-5.2 open model capabilities review, benchmarks
- Gary Sieling - Principal Engineer — Personal reflections on AI's societal impact through complementary goods theory
- Gemma: Introducing new state-of-the-art open models — Google DeepMind's Gemma open-weight model family launch announcement
- General Catalyst, Andreessen Horowitz bet on large language models for health care — VC firms co-lead $50M seed round for LLM healthcare startup
- Generative AI may be creating more work than it saves — Wharton professor argues LLM backends create more labor than they save
- Generative AI needs tools to avoid copyright infringement, Databricks' Naveen Rao says — or more companies could meet Napster's fate — Generative AI copyright risk parallels Napster; open-source training on proprietary data as solution
- Generative AI: A Creative New World — Sequoia's 2022 thesis on generative AI market opportunity and waves
- GitHub - HandsOnLLM/Hands-On-Large-Language-Models: Official code repo for the O'Reilly Book - "Hands-On Large Language Models — Official code repo for illustrated O'Reilly LLM book
- GitHub - badaramoni/wave-field-llm: Wave Field AI — a efficient attention architecture for language models — O(N log N) FFT-based attention replacing quadratic dot-product attention
- GitHub - darrenburns/elia: A snappy, keyboard-centric terminal user interface for interacting with large language models. — Keyboard-focused terminal app for chatting with multiple LLM providers
- GitHub - garg-ankush/scipe: SCIPE is a powerful tool for evaluating and diagnosing LLM (Large Language Model) graphs or chains. — Python tool for root-cause diagnosis of failing LLM chain nodes
- GitHub - tatsu-lab/stanford_alpaca: Code and documentation to train Stanford's Alpaca models, and generate the data. — Instruction-following LLaMA model fine-tuned via self-instruct data
- Github Copilot and ChatGPT alternatives — Survey of AI coding assistants and privacy-safe ChatGPT alternatives in 2023
- Godfather of Artificial Intelligence" Geoffrey Hinton on the promise, risks of advanced AI — Geoffrey Hinton 60 Minutes interview warning of AI existential and societal risks
- Google AI Introduces DS STAR: A Multi Agent Data Science System That Plans, Codes And Verifies End To End Analytics — Multi-agent system converting natural language to Python for heterogeneous data analytics
- Google Publishes Scaling Principles for Agentic Architectures — Predictive regression framework for selecting optimal multi-agent coordination strategies
- Google Replaces BERT Self-Attention with Fourier Transform: 92% Accuracy, 7 Times Faster on GPUs — FNet replaces transformer self-attention with Fourier Transform for faster training
- Google Researchers Can Create an AI That Thinks a Lot Like You After Just a Two-Hour Interview — Stanford/Google study simulates 1,000 people as LLM agents from interview transcripts
- Google publishes document on more notable ranking systems — Google's official guide listing active and retired Search ranking systems
- Google takes aim at Duolingo with new English tutoring tool | TechCrunch — Google Search adds AI-powered English speaking practice feature for language learners
- Google's medical AI chatbot is already being tested in hospitals — Google Med-PaLM 2 hospital pilot tests medical LLM accuracy and safety
- Gradient Descent Models Are Kernel Machines (Deep Learning) — Deep networks trained by gradient descent are kernel machines
- Harvard '21 grad says Gen Z just uses A.I. to do their homework — Gen Z uses ChatGPT for homework completion rather than genuine learning
- Here are the top 10 generative-AI startups founded by ex-Googlers that are taking on ChatGPT — 2023 roundup of ex-Google generative-AI startups
- How AI can keep disappearing languages alive — Training AI on low-resource African languages to preserve cultural diversity
- How ChatGPT is changing the way cybersecurity practitioners look at the potential of AI — ChatGPT dual-use cybersecurity capabilities surprise skeptical security researchers
- How Claude Helps Me Manage My Calendar (but ChatGPT Stumbles!) — Claude processes ICS calendar files reliably where ChatGPT fails context limits
- How I Automate Grading With LangChain And GPT-4 — LangChain GPT-4 pipeline automating bootcamp homework grading
- How I use Google's new Antigravity IDE without hitting rate limits — Managing Google Antigravity IDE quota via model and mode selection
- How IBM's Watson Went From the Future of Health Care to Sold Off for Parts — IBM Watson Health $5B AI healthcare failure sold for parts
- How Much of the World Is It Possible to Model? — Limits and nature of mathematical models from climate to LLMs
- How Smart is ChatGPT? — GPT-4 vs GPT-3.5 exam-percentile benchmark comparison
- How the Foundation Model Transparency Index Distorts Transparency — EleutherAI critique of Stanford FMTI biasing transparency toward corporate services
- How to Build An AI Agent with Function Calling and GPT-5 | Towards Data Science — Building a web-search AI agent via GPT-5 function calling
- How to Design a Production-Grade Multi-Agent Communication System Using LangGraph Structured Message Bus, ACP Logging, and Persistent Shared State Architecture — LangGraph tutorial building ACP message bus with Planner-Executor-Validator agents
- How to Use RLMs in Deep Agents — Recursive language models combat context rot via programmatic subagent orchestration
- Hugging Face Introduces StackLLaMA: A 7B Parameter Language Model Based on LLaMA and Trained on Data from Stack Exchange Using RLHF — Hugging Face RLHF fine-tuning of LLaMA 7B on Stack Exchange Q&A data
- Hugging Face Just Released SmolAgents: A Smol Library that Enables to Run Powerful AI Agents in a Few Lines of Code — Lightweight Hugging Face library for building AI agents in three lines
- Hugging Face clones OpenAI's Deep Research in 24 hours — Open-source research agent replicates OpenAI Deep Research in one day
- Human intuition fuels AI-driven quantum materials discovery — Expert-curated ML model encodes human intuition for quantum materials discovery
- I created over a dozen personal apps using AI in 60 days, here's what I learned — Non-programmer builds 15+ apps using Claude Sonnet and Bolt.new
- I used Gemini 2.0 to create an AI shopping assistant — it's surprisingly good at saving me time and money — Gemini 2.0 Flash used to build reusable agentic shopping assistant prompt
- I worked on Google's AI. My fears are coming true — Fired Google engineer warns LaMDA and Bing chatbots show sentient behavior
- I've been using Claude Code for a couple of days — Hacker News debate on Claude Code coding experience
- Improving language models by retrieving from trillions of tokens — RETRO retrieval-augmented LM outperforms models 25x larger on language modeling
- Inside the Music Industry's High-Stakes A.I. Experiments — Universal Music Group navigating generative AI threats and partnerships
- Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations — Co-STORM multi-agent system for serendipitous unknown-unknown discovery
- Introducing Claude for education — Anthropic launches Claude specialized version for higher education
- Introducing LLM-Evalkit | Google Cloud Blog — Open-source tool centralizing LLM prompt management and metric-driven evaluation
- Is AI a danger to humanity or our salvation? — Hinton, LeCun, and Bengio split on AI existential risk after ChatGPT
- Jailbreaking Black Box Large Language Models in Twenty Queries — PAIR algorithm uses attacker LLM to jailbreak target LLM in under 20 queries
- Keynote at NVIDIA GTC San Jose 2026 — Jensen Huang's 2026 GTC keynote covering full AI stack advances
- LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others — Artificial Analysis leaderboard ranking 100+ LLMs
- Language, Statistics, & Category Theory, Part 1 — Category theory framework unifying algebraic and statistical structure in language
- Large Language Model Performance Doubles Every 7 Months — LLM task-completion capability doubles every seven months exponentially
- Large Language Models as Optimizers — OPRO uses LLMs as gradient-free optimizers via natural language prompts
- Leaked org chart reveals the 58 top leaders and engineers at Google DeepMind — Google DeepMind org structure after 2023 Brain-DeepMind merger
- Learning to Replicate Expert Judgment in Financial Tasks — Fine-tuned LLM beats frontier models at financial info triage
- Let's Build the GPT Tokenizer — A Complete Guide to Tokenization in LLMs — Karpathy tokenizer video as book chapter
- LightPROF: A Lightweight AI Framework that Enables Small-Scale Language Models to Perform Complex Reasoning Over Knowledge Graphs (KGs) Using Structured Prompts — Lightweight Retrieve-Embed-Reason framework for small LLMs on KGs
- Llamar.ai: A deep dive into the (in)feasibility of RAG with LLMs — RAG product prototype built and shut down due to GPT-4 API cost infeasibility
- Look behind the curtain: Don't be dazzled by claims of 'artificial intelligence' | Op-Ed — Emily Bender op-ed demystifying AI as probabilistic pattern matching
- Machine Learning Course Series — Three-course ML curriculum from Python basics to LLMs
- Machine Learning Course Series — Three-course ML curriculum from Python basics to LLMs
- Machine Learning Crash Course — Google for Developers — Google's free practical ML course with interactive modules
- Mark Zuckerberg says AI could soon do the work of Meta's midlevel engineers — Zuckerberg predicts AI replacing midlevel software engineers by 2025
- Meet 'Stack,' A 3TB of Permissively Licensed Source Code for LLMs (Large Language Models) — BigCode releases 3TB permissively licensed code dataset for LLMs
- Meet AnythingLLM: A Full-Stack Application That Transforms Your Content into Rich Data for Enhanced Large Language Models LLMs Interactions — Open-source full-stack app for chatting with documents using LLMs
- Meet DrugAgent: A Multi-Agent Framework for Automating Machine Learning in Drug Discovery — LLM multi-agent system automating end-to-end drug discovery ML pipelines
- Meet TinyLlama: A Small AI Model that Aims to Pretrain a 1.1B Llama Model on 3 Trillion Tokens — Small 1.1B LLM pre-trained on 3 trillion tokens
- Meta is playing the AI game with house money — Meta Q2 2025 earnings signal all-in superintelligence bet
- Meta's Yann LeCun predicts 'new paradigm of AI architectures' within 5 years and 'decade of robotics' | TechCrunch — LeCun forecasts LLM paradigm obsolescence and rise of world models
- Microsoft's relationship with OpenAI cracked when it hired Mustafa Suleyman, rival Marc Benioff says | TechCrunch — Microsoft-OpenAI partnership fracturing over competing AI ambitions
- Must-Have Prompt Engineering Skills for 2024 — In-demand prompt engineering skills, platforms, and tools
- NOUS RESEARCH — Open-source AI lab training and distributing open language models
- Navigating the Jagged Technological Frontier | Harvard Business School AI Institute — Field experiment showing AI boosts consultant productivity unevenly across tasks
- NeurIPS 2023: Key Takeaways From Invited Talks — NeurIPS 2023 invited talk summaries on LLM efficiency, generative AI, and responsible data
- New MIT Research Shows Spectacular Increase In White Collar Productivity From ChatGPT — MIT RCT finds ChatGPT makes white-collar workers 37% faster at writing tasks
- Noam Chomsky says A.I. is far from 'true intelligence' and ChatGPT is the 'banality of evil' | Fortune — Chomsky argues LLMs lack true reasoning and exhibit indifference to truth
- NotebookLM's automatically generated podcasts are surprisingly effective — Google NotebookLM generates convincing AI podcast episodes from user documents
- Nvidia CEO: We bet the farm on AI and no one knew it | TechCrunch — Jensen Huang recounts Nvidia's 2018 AI pivot from rasterization to GPU-accelerated ML
- Nvidia Wants to Replace Nurses With AI for $9 an Hour — Nvidia-Hippocratic AI partnership deploys generative AI nurses at $9/hr
- OmniThink: A Cognitive Framework for Enhanced Long-Form Article Generation Through Iterative Reflection and Expansion — Iterative reflection framework improving knowledge density in LLM long-form writing
- Once "too scary" to release, GPT-2 gets squeezed into an Excel spreadsheet — GPT-2 fully implemented in Excel spreadsheet for LLM education
- One AI Tutor Per Child: Personalized learning is finally here — LLMs enable scalable one-on-one personalized tutoring for every child
- Open-source AI matches top proprietary model in solving tough medical cases — Open-source Llama 3.1 405B matches GPT-4 on clinical diagnostic reasoning
- OpenAI CEO Sam Altman says ChatGPT would have passed for an AGI 10 years ago — Sam Altman on shifting AGI definition, hallucinations, and AI training data consent
- OpenAI Publishes GPT Prompt Engineering Guide — OpenAI's six-strategy guide for eliciting better GPT-4 responses
- OpenAI executives say releasing ChatGPT for public use was a last resort after running into multiple hurdles — and they're shocked by its popularity — OpenAI's surprise at ChatGPT's viral adoption and CEO's AI risk warnings
- Paper page - Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models — GPT-4 achieves near-perfect scores on MIT Math and EECS curriculum
- Practical Guide to Task Automation using ChatGPT & Python — Using ChatGPT and Python APIs to automate common knowledge-work tasks
- Prompt Engineering Guide — Primer on LLM-powered agents: capabilities, design patterns, use cases
- Prompt Engineering Urges 'Hermeneutic Prompting' As A Powerful Technique Unlocking The True Value Of Generative AI — Hermeneutic circle prompting technique for richer LLM responses
- ReAct: Synergizing Reasoning and Acting in Language Models — Interleaved reasoning traces and external actions improve LLM agent reliability
- Read the internal memo Alphabet sent in merging A.I.-focused groups DeepMind and Google Brain — Alphabet merges DeepMind and Google Brain into Google DeepMind unit
- Reddit to charge for API access; CEO blames A.I. | Fortune — Reddit monetizing API access to prevent free AI training data extraction
- Replit CEO on AI breakthroughs: 'We don't care about professional coders anymore' — Replit Agent targets non-coders after Claude 3.5 Sonnet SWE-bench breakthrough
- Researchers From Stanford And DeepMind Come Up With The Idea of Using Large Language Models LLMs as a Proxy Reward Function — LLMs as proxy reward functions for RL agent alignment via natural language
- Researchers at Boston University Release the Platypus Family of Fine-Tuned LLMs — Cheap fast LLM fine-tuning via curated Open-Platypus dataset and LoRA merging
- SCIPE - Systematic Chain Improvement and Problem Evaluation — LangChain-featured tool for identifying failing nodes in LLM chains
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Language-independent subword tokenizer trained from raw sentences
- Sharing new breakthroughs and artifacts supporting molecular property prediction, language processing, and neuroscience — Meta FAIR releases OMol25 dataset, UMA model, Adjoint Sampling, and brain language study
- Simon Willison (@simonw) on X — Simon Willison asks for real-world LLM fine-tuning commercial success stories
- Simulating scientists: A new tool for AI-powered scientific discovery — LLM tool mimics scientists for molecular property prediction
- Space Force Gets Scared, Pauses All Use of Generative AI — US Space Force bans generative AI tools on government devices over data security concerns
- State of AI, September 2025: A Broad Primer from an AI Team Lead in Education — Education-focused overview of AI history, capabilities, agents, and AGI trajectory
- Stop Writing Code, Start Writing Docs — Best practices for agentic coding tools emphasizing documentation over single-shot prompting
- Technology Innovation Institute Open-Sourced Falcon LLMs: A New AI Model That Uses Only 75 Percent of GPT-3's Training Compute, 40 Percent of Chinchilla's, and 80 Percent of PaLM-62B's — TII open-sources Falcon-40B and Falcon-7B compute-efficient decoder LLMs
- Test-Time Training (TTT): A New Approach to Sequence Modeling — Hidden state as a learnable model updated at inference
- The 'Godfather of AI' Has a Hopeful Plan for Keeping Future AI Friendly — Geoffrey Hinton's views on LLM risks and analog computing as AI safety mitigation
- The Batch | DeepLearning.AI | AI News & Insights — DeepLearning.AI weekly AI news and insights newsletter homepage
- The Capacity for Moral Self-Correction in Large Language Models — RLHF-trained LLMs can self-correct harmful outputs when instructed
- The Era of Agentic Organization: Learning to Organize with Language Models — AsyncThink paradigm: concurrent LLM reasoning optimized via reinforcement learning
- The False Promise of Imitating Proprietary LLMs — Finetuning open models on ChatGPT outputs mimics style but not factuality or capability
- The Foundation Model Development Cheatsheet — Quick-start guide covering full pipeline for responsible open foundation model development
- The Goopification of AI — AI chatbots colonizing the self-help genre via probabilistic text assembly
- The Pentagon says AI is speeding up its 'kill chain' | TechCrunch — AI accelerating Pentagon kill chain planning amid usage policy tensions
- The Trick to Make LLaMa Fit into Your Pocket: Meet OmniQuant, an AI Method that Bridges the Efficiency and Performance of LLMs — OmniQuant learnable post-training quantization achieves low-bit LLM compression efficiently
- The second wave of AI coding is here — Next-gen AI coding agents targeting functional correctness via process data
- These ex-Apple employees are bringing AI to the desktop — Ex-Apple founders launch startup for LLM-powered desktop OS reimagining
- This AI Paper Demonstrates An End-to-End Training Flow on An Large Language Model LLM-13 Billion GPT-Using Sparsity And Dataflow — End-to-end sparse LLM training on SambaNova RDU hardware
- This startup says it's made ChatGPT for construction sites. Read the pitch deck it used to raise $40 million. — Trunk Tools construction-specific AI agent raises $40M Series B
- Timeline of Open and Proprietary Large Language Models | NextBigFuture.com — Chronological overview of open and proprietary LLM releases through 2023
- Tiny DeepSeek 1.5B Models Run on $249 NVIDIA Jetson Nano | NextBigFuture.com — DeepSeek R1 1.5B distilled model runs locally on sub-$250 edge hardware
- Toolformer: Language Models Can Teach Themselves to Use Tools — Self-supervised LLM fine-tuning to autonomously call and integrate external APIs
- TradingAgents: Multi-Agents LLM Financial Trading Framework — Multi-agent LLM framework simulating specialized trading firm roles
- Training Your Own LLM using privateGPT — Running a private local LLM on sensitive data without cloud exposure
- Transformer Explainer: LLM Transformer Model Visually Explained — Interactive visual walkthrough of Transformer/GPT-2 architecture
- Transformers Explained Visually (Part 3): Multi-head Attention, deep dive — Deep dive into multi-head attention mechanics and data dimensions
- Tx-LLM: Supporting therapeutic development with large language models — PaLM-2 fine-tuned on 66 drug discovery tasks across full pipeline
- Unbabel says its new AI model has dethroned OpenAI's GPT-4 as the tech industry's best language translator — TowerLLM 7B/13B model edges GPT-4o on multilingual translation benchmarks
- Using ChatGPT as a technical writing assistant — ChatGPT as iterative technical writing draft tool
- Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality — LLaMA fine-tune on ShareGPT data achieving near-ChatGPT quality for $300
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension — Hybrid convolutional-transformer model via knowledge distillation for fast NLU inference
- What Really Made Geoffrey Hinton Into an AI Doomer — Hinton's reasons for leaving Google to warn about accelerating AI risk
- What was 60 Minutes thinking, in that interview with Geoff Hinton? — Gary Marcus rebuttal annotating 60 Minutes Hinton AI interview
- Why LLaMa Is A Big Deal — LLaMA enables GPT-3-class LLM inference on consumer hardware
- Why machine learning struggles with causality - TechTalks — ML systems lack causal reasoning needed for robust generalization
- Why some college professors are adopting ChatGPT AI as quickly as students — ChatGPT disruption of higher education from professor perspective
- Will Transformers Take Over Artificial Intelligence? | Quanta Magazine — Transformers expanding from NLP to vision, generative, and multimodal AI tasks
- Word embeddings visualization from 'Build a Large Language Model (From Scratch)' — Animated 2D visualization of LLM token-to-vector embedding training
- Worldcoin, co-founded by Sam Altman, is betting the next big thing in AI is proving you are human | TechCrunch — Iris-scanning proof-of-personhood system as defense against AI-indistinguishable bots
- deeplearning.ai — DeepLearning.AI GitHub organization hosting course materials
- elvis (@omarsar0) on X — Stanford CME295 new course on Transformers and LLMs announced