vision transformer
A vision transformer is a type of model architecture that applies transformer techniques—particularly self-attention—to image data, enhancing the model's ability to capture spatial relationships and context within images.
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
- GSPN-2: Efficient Parallel Sequence Modeling
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
- VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution