concept bottleneck models

These models propose a structured way of learning where a model is forced to learn interpretable concepts as intermediate representations, advocating for improved interpretability and error analysis by aligning learned concepts with human-understandable categories.

6 papers