transformer blocks
The basic units of transformer architectures that consist of self-attention mechanisms and feedforward neural networks, essential for building deeper networks that capture complex patterns.
- FFN Fusion: Rethinking Sequential Computation in Large Language Models
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
- Large Language Models Think Too Fast To Explore Effectively
- One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
- ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
- Towards Fully FP8 GEMM LLM Training at Scale