Song Han
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Radial Attention: $\mathcal O(n \log n)$ Sparse Attention for Long Video Generation
- Scaling RL to Long Videos
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
- Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
- WorldModelBench: Judging Video Generation Models As World Models