3.0
@hn_dc1861
about 2 months ago
Below are my comments on Magistral small (not medium). 24B size is good for local inference. As a model outputting long "reasoning" traces (~10k tokens), 40k context length is a little concerning. Where are the results of normal benchmarks, e.g., MMLU/pro, IFEval and such. Still, thank you Mistral team for releasing this model with Apache 2.0.
