model size
Model size refers to the complexity and capacity of an AI model, typically measured by the number of parameters it contains. Larger models are often more capable of capturing intricate patterns in data, but they also require more computational resources and can be prone to overfitting.
- Chain-of-Model Learning for Language Model
- Co-PatcheR: Collaborative Software Patching with Component-specific Small Reasoning Models
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- Dimension-adapted Momentum Outscales SGD
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- Fast Training of Large Kernel Models with Delayed Projections
- Grids Often Outperform Implicit Neural Representation at Compressing Dense Signals
- Mellow: a small audio language model for reasoning
- PINN Balls: Scaling Second-Order Methods for PINNs with Domain Decomposition and Adaptive Sampling
- Power Lines: Scaling laws for weight decay and batch size in LLM pre-training
- Progressive Data Dropout: An Embarrassingly Simple Approach to Train Faster
- SAS: Simulated Attention Score
- Scaling and context steer LLMs along the same computational path as the human brain
- Scaling can lead to compositional generalization
- Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families