transformer architecture
The structural design of transformers, which relies on self-attention mechanisms and feed-forward layers, enabling parallel processing of sequences and providing significant improvements over recurrent neural networks for sequence tasks.
- ALINE: Joint Amortization for Bayesian Inference and Active Data Acquisition
- ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative Verification
- CausalPFN: Amortized Causal Effect Estimation via In-Context Learning
- Chain-of-Model Learning for Language Model
- DeltaFormer: Unlock the state space of Transformer
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification Policy
- E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products
- FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
- Improving Formal Reasoning of Transformer with State Stack
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- Kinaema: a recurrent sequence model for memory and pose in motion
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Learning Urban Climate Dynamics via Physics-Guided Urban Surface–Atmosphere Interactions
- Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- Overcoming Long Context Limitations of State Space Models via Context Dependent Sparse Attention
- Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-training
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- SAS: Simulated Attention Score
- Time-Masked Transformers with Lightweight Test-Time Adaptation for Neural Speech Decoding
- Towards Provable Emergence of In-Context Reinforcement Learning
- Transformer brain encoders explain human high-level visual responses
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Wavy Transformer
- When No Paths Lead to Rome: Benchmarking Systematic Neural Relational Reasoning