685B MoE model for general intelligence
DeepSeek V3 is a 685-billion parameter mixture-of-experts model that activates only 37B parameters per forward pass, making it highly efficient. It outperforms all previous open-source models on coding, math, and reasoning.
Try DeepSeek V3Sourced from official model cards and academic papers. Higher is better.