4.0
@hn_1d2a3c
2 months ago
We very much can, especially such a Mixture of Experts model with only 3B activated parameters. With an RTX 3070 (7GB GRAB VRAM), 32 GB RAM and an SSD I can run such models at speeds tolerable for casual use.

48B/3B-active hybrid linear-attention MoE (Kimi Delta Attention + MLA at 3:1), 1M context, MIT license; 6x faster decode via ~75% KV cache reduction.
@hn_1d2a3c
2 months ago
We very much can, especially such a Mixture of Experts model with only 3B activated parameters. With an RTX 3070 (7GB GRAB VRAM), 32 GB RAM and an SSD I can run such models at speeds tolerable for casual use.