NVIDIA's 120B open-weight MoE reasoning model (12B active) with 1M context and 350 tok/s output speed. Open weights on HuggingFace.
No takes yet.
Try another filter or be the first to share your honest opinion.