Concur with Gemma2 being underwhelming, I dismissed it pretty quickly but gemma3:27b is looking pretty good atm.
BTW mistral-small:24b is also worth mentioning (IMO best local model) and phi4:14b is also pretty strong for its size.
mistral-small was my previous local goto model, testing now to see if gemma3 can replace it.
With eval rate numbers:
- phi4: 12 tokens/s
- mistral-small: 9 tokens/s
On Nvidia RTX 4090 laptop:
- phi4: 36 tokens/s
- mistral-small: 16 tokens/s
4.5
@hn_3f726d
about 1 month ago
I had a similar experience on my pipeline.
Was looking to both decrease costs and experiment out of OpenAI offering and ended up using Mistral Small on summarization and Large for the final analysis step and I'm super happy.
They have also a very generous free tier which helps in creating PoCs and demos.
4.0
@hn_4352f6
about 1 month ago
Mistral Large is 123b so one can probably assume that medium is between 24b and 123b, also Mistral 3.1 is by a wide margin my go-to model in real life situations. Benchmarks absolutely don't tell the whole story, and different models have different use cases.
2.0
@hn_fe2350
about 2 months ago
Running it on a MacBook with M1 Pro chip and 32 GB of RAM is quite slow. I expected to be as fast as phi4 but it's much slower.