Daniel Soudry
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
- FP4 All the Way: Fully Quantized Training of Large Language Models
- Optimal Rates in Continual Linear Regression via Increasing Regularization
- Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
- Tensor-Parallelism with Partially Synchronized Activations