interpretability
Interpretability in AI refers to the degree to which a human can understand the model's decisions or behavior, allowing for insights into how and why decisions are made. This is crucial for validating and trusting AI systems, especially in high-stakes applications.
- A Geometry-Aware Metric for Mode Collapse in Time Series Generative Models
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
- AdaptGrad: Adaptive Sampling to Reduce Noise
- Additive Models Explained: A Computational Complexity Approach
- Advancing Interpretability of CLIP Representations with Concept Surrogate Model
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
- An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
- Approximating Shapley Explanations in Reinforcement Learning
- Association-Focused Path Aggregation for Graph Fraud Detection
- Backdoor Mitigation via Invertible Pruning Masks
- Bayesian Concept Bottleneck Models with LLM Priors
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- Beyond Pairwise Connections: Extracting High-Order Functional Brain Network Structures under Global Constraints
- BioOSS: A Bio-Inspired Oscillatory State System with Spatio-Temporal Dynamics
- CF-VLM:CounterFactual Vision-Language Fine-tuning
- CHiQPM: Calibrated Hierarchical Interpretable Image Classification
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- Causal Discovery over Clusters of Variables in Markovian Systems
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- Conformal Information Pursuit for Interactively Guiding Large Language Models
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Curvature Tuning: Provable Training-free Model Steering From a Single Parameter
- DeepHalo: A Neural Choice Model with Controllable Context Effects
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- Dynamic Algorithm for Explainable $k$-medians Clustering under $\ell_p$ Norm
- Dynamic Diffusion Schrödinger Bridge in Astrophysical Observational Inversions
- Emergent Risk Awareness in Rational Agents under Resource Constraints
- Empowering Decision Trees via Shape Function Branching
- Enhancing Interpretability in Deep Reinforcement Learning through Semantic Clustering
- Evaluating LLMs in Open-Source Games
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- Fast Rate Bounds for Multi-Task and Meta-Learning with Different Sample Sizes
- FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
- GMM-based VAE model with Normalising Flow for effective stochastic segmentation
- GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
- Geometry-Aware Edge Pooling for Graph Neural Networks
- Global Minimizers of $\ell^p$-Regularized Objectives Yield the Sparsest ReLU Neural Networks
- Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain
- Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
- High-order Interactions Modeling for Interpretable Multi-Agent Q-Learning
- How do Transformers Learn Implicit Reasoning?
- Identifiability of Deep Polynomial Neural Networks
- Identifiability of Deep Polynomial Neural Networks
- Improved Representation Steering for Language Models
- Influence Functions for Edge Edits in Non-Convex Graph Neural Networks
- Integrating Drug Substructures and Longitudinal Electronic Health Records for Personalized Drug Recommendation
- Interpretable Next-token Prediction via the Generalized Induction Head
- Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
- LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
- Latent Retrieval Augmented Generation of Cross-Domain Protein Binders
- LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
- Less is More: Local Intrinsic Dimensions of Contextual Language Models
- Localizing Knowledge in Diffusion Transformers
- MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion Prediction
- MIX: A Multi-view Time-Frequency Interactive Explanation Framework for Time Series Classification
- MODEL SHAPLEY: Find Your Ideal Parameter Player via One Gradient Backpropagation
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural Networks
- Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
- Metritocracy: Representative Metrics for Lite Benchmarks
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition
- MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
- Multi-step Visual Reasoning with Visual Tokens Scaling and Verification
- NeuSymEA: Neuro-symbolic Entity Alignment via Variational Inference
- Online Feedback Efficient Active Target Discovery in Partially Observable Environments
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions
- PhysDiff: A Physically-Guided Diffusion Model for Multivariate Time Series Anomaly Detection
- Pinpointing Attention-Causal Communication in Language Models
- Preserving Task-Relevant Information Under Linear Concept Removal
- Probabilistic Stability Guarantees for Feature Attributions
- Probabilistic Token Alignment for Large Language Model Fusion
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
- ProtoPairNet: Interpretable Regression through Prototypical Pair Reasoning
- Quantifying Generalisation in Imitation Learning
- Quantifying Statistical Significance of Deep Nearest Neighbor Anomaly Detection via Selective Inference
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- Random Forest Autoencoders for Guided Representation Learning
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- Rethinking Gradient Step Denoiser: Towards Truly Pseudo-Contractive Operator
- Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
- SORTeD Rashomon Sets of Sparse Decision Trees: Anytime Enumeration
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
- Sample-efficient Learning of Concepts with Theoretical Guarantees: from Data to Concepts without Interventions
- Semantic Representation Attack against Aligned Large Language Models
- SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
- Steering Information Utility in Key-Value Memory for Language Model Post-Training
- SymMaP: Improving Computational Efficiency in Linear Solvers through Symbolic Preconditioning
- TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
- The Rashomon Set Has It All: Analyzing Trustworthiness of Trees under Multiplicity
- TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop
- Timely Clinical Diagnosis through Active Test Selection
- TopER: Topological Embeddings in Graph Representation Learning
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Towards Understanding Transformers in Learning Random Walks
- Towards Unsupervised Training of Matching-based Graph Edit Distance Solver via Preference-aware GAN
- Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
- Transformer brain encoders explain human high-level visual responses
- Uncertainty Quantification for Physics-Informed Neural Networks with Extended Fiducial Inference
- Unveiling Concept Attribution in Diffusion Models
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
- Variational Polya Tree
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive Hypotheses
- Who Reasons in the Large Language Models?