Knowledge Distillation - Students meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 9.34M • • 5.77k Qwen/Qwen2.5-7B-Instruct Text Generation • 8B • Updated Jan 12, 2025 • 12.6M • • 1.24k
LM Eval Tasks cais/mmlu Viewer • Updated Mar 8, 2024 • 231k • 447k • 715 SaylorTwift/bbh Viewer • Updated Jun 16, 2024 • 6.76k • 42.9k • 6 Rowan/hellaswag Viewer • Updated Jul 10, 2025 • 60k • 300k • 169 allenai/ai2_arc Viewer • Updated Dec 21, 2023 • 7.79k • 410k • 331
Knowledge Distillation - Teachers meta-llama/Llama-3.1-70B-Instruct Text Generation • 71B • Updated Dec 15, 2024 • 696k • • 909 Qwen/Qwen2.5-72B-Instruct Text Generation • 73B • Updated Jan 12, 2025 • 410k • • 932
Knowledge Distillation - Students meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 9.34M • • 5.77k Qwen/Qwen2.5-7B-Instruct Text Generation • 8B • Updated Jan 12, 2025 • 12.6M • • 1.24k
Knowledge Distillation - Teachers meta-llama/Llama-3.1-70B-Instruct Text Generation • 71B • Updated Dec 15, 2024 • 696k • • 909 Qwen/Qwen2.5-72B-Instruct Text Generation • 73B • Updated Jan 12, 2025 • 410k • • 932
LM Eval Tasks cais/mmlu Viewer • Updated Mar 8, 2024 • 231k • 447k • 715 SaylorTwift/bbh Viewer • Updated Jun 16, 2024 • 6.76k • 42.9k • 6 Rowan/hellaswag Viewer • Updated Jul 10, 2025 • 60k • 300k • 169 allenai/ai2_arc Viewer • Updated Dec 21, 2023 • 7.79k • 410k • 331