Reasoning capabilities distilled into 70B Llama
DeepSeek R1 Distill 70B transfers the chain-of-thought reasoning abilities of the full R1 model into a 70B Llama architecture via knowledge distillation. It offers strong reasoning performance at a fraction of the full model's cost.
Try DeepSeek R1 Distill 70BSourced from official model cards and academic papers. Higher is better.